Reward distribution device, question answering system, reward distribution program, and reward distribution method
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NEC CORP
- Filing Date
- 2025-12-17
- Publication Date
- 2026-08-06
Smart Images

Figure JP2025044161_06082026_PF_FP_ABST
Abstract
Description
Reward Distribution Device, Question-Answering System, Reward Distribution Program, and Reward Distribution Method
[0001] The present disclosure relates to a reward distribution device, a question-answering system, a reward distribution program, and a reward distribution method.
[0002] Conventionally, various techniques have been proposed to facilitate obtaining data necessary for improving the performance of a learned model (to enhance the motivation of data holders to provide data). For example, Patent Document 1 describes a reward distribution system that calculates the distribution of reward amounts to a plurality of operators who provide data or processing services constituting an application. The distribution management device of this reward distribution system includes an estimation processing unit, a contribution estimation unit, and a distribution amount calculation unit. The estimation processing unit estimates the effect obtained as a result of application execution as a quantitative value. The contribution estimation unit combines information on the history of the application (number of uses / number of selections) and input / output relationship information between data and processing services to estimate the contribution of each data and service processing unit to the quantitative value. The distribution amount calculation unit calculates the reward amount for the provider of the data or service processing unit based on the contribution.
[0003] Japanese Patent Application Laid-Open No. 2015-099492
[0004] For changing the version of a learned model, a large amount of change data is required. However, in the prior art as described in Patent Document 1, rewards are paid only to those with particularly large contributions. That is, in the prior art, the providers of change data who are the targets of reward payment are very limited. Therefore, conventionally, it has been difficult to motivate many people who hold change data to provide the change data.
[0005] The present disclosure has been made in view of the above problems, and aims to enhance the motivation of holders of change data that can be used for changing the version of a learned model to provide the change data to the developer of the learned model.
[0006] An exemplary reward distribution device relating to this disclosure includes: a first evaluation means that, when a question is input, evaluates the contribution of modification data used to change the version of a trained model to improving the quality of the answer to the question based on the quality of the answer to the question, based on the quality of the answer to the question, based on the quality of the answer to the question; a second evaluation means that evaluates the degree of relevance between the modification data and the content of the multiple questions that have been input to the trained model to date; a calculation means that calculates a reward to be paid to the provider of the modification data based on the contribution and the degree of relevance; and an output means that outputs the content of the reward.
[0007] An illustrative aspect of the Question Answering System of this Disclosure comprises a trained model constructed to output an answer to a question when a question is input, a modification means for changing the version of the trained model, and the reward distribution device described above.
[0008] An exemplary aspect of the present disclosure relates to a reward distribution program that causes at least one processor to perform: a first evaluation process that evaluates the contribution of modification data used to change the version of a trained model to improving the quality of answers to a question, based on the quality of answers to the trained model, which is constructed to output answers to the question when a question is input; a second evaluation process that evaluates the degree of relevance between the modification data and the content of the multiple questions that have been input to the trained model so far; a calculation process that calculates a reward to be paid to the provider of the modification data based on the contribution and the degree of relevance; and an output process that causes the content of the reward to be output to an output means.
[0009] An exemplary aspect of the present disclosure of a reward distribution method includes: a first evaluation step in which at least one processor evaluates the contribution of modification data used to change the version of a trained model to improving the quality of answers to a question, based on the quality of answers to the trained model, which is constructed to output answers to a question when a question is input; a second evaluation step in which at least one processor evaluates the degree of relevance between a plurality of questions and the modification data, based on the content of a plurality of questions that have been input to the trained model to date; a calculation step in which at least one processor calculates a reward to be paid to the provider of the modification data based on the contribution and the degree of relevance; and an output step in which at least one processor causes an output means to output the content of the reward.
[0010] According to an illustrative aspect of this disclosure, it is possible to increase the motivation of holders of modification data that can be used to change the version of a trained model to provide that modification data to the developers of the trained model.
[0011] This is a block diagram showing an example of the functional configuration of a reward distribution device according to the first exemplary embodiment of this disclosure. This is a flowchart showing an example of the flow of a reward distribution method according to the first exemplary embodiment of this disclosure. This is a block diagram showing an example of the functional configuration of a question answering system equipped with a reward distribution device according to the second exemplary embodiment of this disclosure. This is a diagram showing an example of output by an output means provided in the reward distribution device according to the second exemplary embodiment of this disclosure. This is a flowchart showing an example of the flow of a reward distribution method according to the second exemplary embodiment of this disclosure. This is a block diagram showing the hardware configuration of a computer functioning as a reward distribution device according to this disclosure.
[0012] The following are examples of embodiments of the present invention. However, the present invention is not limited to the exemplary embodiments shown below, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining the technical means employed in each of the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, embodiments obtained by appropriately omitting some of the technical means employed in each of the exemplary embodiments shown below may also be included in the scope of the present invention. In addition, the effects mentioned in each of the exemplary embodiments shown below are examples of effects that can be expected in that exemplary embodiment and do not define the scope of the present invention. That is, embodiments that do not produce the effects mentioned in each of the exemplary embodiments shown below may also be included in the scope of the present invention.
[0013] [First Exemplary Embodiment] First, a first exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. This exemplary embodiment is the basic form for each of the exemplary embodiments described later. The scope of application of each technical means adopted in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means adopted in this exemplary embodiment can also be adopted in other exemplary embodiments included in this disclosure, to the extent that no particular technical problems occur. Furthermore, each technical means shown in the drawings referenced to explain this exemplary embodiment can also be adopted in other exemplary embodiments included in this disclosure, to the extent that no particular technical problems occur.
[0014] (Pre-trained model) A pre-trained model is constructed to output an answer to a given question when a question is input. A "pre-trained model" may consist of a single model or multiple models.
[0015] (Configuration of Reward Distribution Device 1) Next, the configuration of the reward distribution device 1 will be explained with reference to Figure 1. Figure 1 is a block diagram showing the configuration of the reward distribution device 1. As shown in Figure 1, the reward distribution device 1 includes a first evaluation means 11, a second evaluation means 12, a calculation means 13, and an output means 14.
[0016] - First evaluation means 11 The first evaluation means 11 evaluates the contribution of the modification data used to change the version of the trained model to improving the response quality, based on the response quality of the trained model. "Version change" includes version upgrades and version downgrades. "Modification data" includes at least one of a submodel to be added to the trained model and the training data.
[0017] - Second evaluation means 12 The second evaluation means 12 evaluates the degree of relevance between multiple questions and the data to be modified, based on the content of multiple questions that have been input into the trained model so far.
[0018] Calculation means 13 The calculation means 13 calculates the compensation to be paid to the provider of the modification data based on the degree of contribution and the degree of relevance.
[0019] Output means 14 Output means 14 outputs the content of the reward calculated by calculation means 13. Output means 14 includes a monitor that displays the content of the reward, a communication module that transmits data of the content of the reward, and terminals that connect to other devices (monitors, etc.) that receive data of the content of the reward, etc.
[0020] (Effects of the Reward Distribution Device 1) The reward distribution device 1 described above employs a configuration that includes a calculation means 13 that calculates the reward to be paid to the provider of the modification data based on the degree of contribution and the degree of relevance. In other words, the reward distribution device 1 calculates the reward based on multiple evaluation axes. Therefore, according to the reward distribution device 1 of this embodiment, not only those who provide modification data that contributes to improving the quality of answers, but also those who provide modification data that makes it easier for users (users of the trained model) to output answers on information of high interest to them will be paid. In other words, the reward distribution device 1 of this embodiment calculates rewards for a wider range of target individuals than before. As a result, more people will think that the modification data they possess may lead to the acquisition of rewards. Therefore, according to the reward distribution device 1 of this embodiment, it is possible to increase the motivation of holders of modification data that can be used to change the version of the trained model to provide modification data to the developers of the trained model. As a result, it becomes easier for developers of trained models to obtain high-quality modification data that leads to improvements in the performance of the trained model.
[0021] (Flow of Reward Distribution Method S1) Next, the flow of the reward distribution method S1 will be explained with reference to Figure 2. Figure 2 is a flowchart showing the flow of the reward distribution method S1. As shown in Figure 2, the reward distribution method S1 includes a first evaluation step S11, a second evaluation step S12, a calculation step S13, and an output step S14.
[0022] - First evaluation step S11 In the first evaluation step S11, at least one processor evaluates the contribution of the modification data used to change the version of the trained model to improving the response quality, based on the response quality of the trained model. The processor that evaluates the contribution may be provided by the reward distribution device 1 or by another device.
[0023] - Second evaluation step S12 After the first evaluation step S11, the process moves to the second evaluation step S12. In the second evaluation step S12, at least one processor evaluates the degree of relevance between a plurality of questions and the modification data based on the content of the plurality of questions that have been input to the trained model so far. The processor that evaluates the degree of relevance may be provided by the reward distribution device 1 or by another device. Note that the second evaluation step S12 may be performed before the first evaluation step S11 or in parallel with the first evaluation step S11.
[0024] Calculation process S13 After the first evaluation process S11 and the second evaluation process S12, the process moves to calculation process S13. In calculation process S13, at least one processor calculates the reward to be paid to the provider of the modification data based on the degree of contribution and the degree of relevance. The processor that calculates the reward may be provided by the reward distribution device 1 or by another device.
[0025] - Output process S14 After the calculation process S13, the process moves to the output process S14. In the output process S14, at least one processor causes the content of the reward to be output to the output means 14. The processor that causes the content of the reward to be output may be provided by the reward distribution device 1, or it may be provided by another device.
[0026] (Effects of Reward Distribution Method S1) As described above, the reward distribution method S1 employs a configuration in the calculation step S13 in which the reward to be paid to the provider of modification data is calculated based on the degree of contribution and the degree of relevance. Therefore, according to the reward distribution method S1 of this embodiment, not only those who provide modification data that contributes to improving the quality of answers, but also those who provide modification data that makes it easier for users (users of the trained model) to output answers on information of high interest to them will receive a reward. In other words, the reward distribution method S1 of this embodiment calculates rewards for a wider range of target individuals than before. As a result, more people will think that the modification data they possess may lead to receiving a reward. Therefore, according to the reward distribution method S1 of this embodiment, the motivation of holders of modification data that can be used to change the version of the trained model to provide modification data to the developers of the trained model can be increased. And as a result, it becomes easier for developers of trained models to obtain high-quality modification data that leads to improvements in the performance of the trained model.
[0027] [Second Exemplary Embodiment] Next, a second exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same function as those described in the above-described exemplary embodiment are denoted by the same reference numerals, and their descriptions are omitted as appropriate. The scope of application of each technical means adopted in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means adopted in this exemplary embodiment can also be adopted in other exemplary embodiments included in this disclosure, to the extent that no particular technical hindrance occurs. Furthermore, each technical means shown in each drawing referenced to describe this exemplary embodiment can also be adopted in other exemplary embodiments included in this disclosure, to the extent that no particular technical hindrance occurs.
[0028] (Configuration of the Question Answering System 100) First, the configuration of the question answering system 100 will be described with reference to Figure 3. Figure 3 is a block diagram showing the configuration of the question answering system 100. The question answering system 100 according to the second exemplary embodiment includes a reward distribution device 1A and a question answering device 2, as shown in Figure 3. The question answering system 100 according to the second exemplary embodiment further includes a user interface 3, a first database 4, and a second database 5. Each device 1A, 2-5 is connected to each other via a communication network.
[0029] • Question answering device 2 The question answering device 2 comprises a trained model 21 and a modification means 22.
[0030] The pre-trained model 21 according to the second exemplary embodiment is constructed to output an answer to a question when a question is input, similar to the pre-trained model 21 according to the first exemplary embodiment. The pre-trained model 21 according to the second exemplary embodiment may also consist of multiple submodels. "Composed of multiple models" includes the merging of multiple submodels, or the composition of multiple submodels and a gateway model that selects a submodel to input a question. Furthermore, the pre-trained model 21 according to the second exemplary embodiment is updated by adding new submodels or training new training data. For this reason, the pre-trained model 21 also includes models that are still being trained. The submodels constituting the pre-trained model 21 according to the second exemplary embodiment may include at least one image recognition model. For this reason, the pre-trained model 21 according to the second exemplary embodiment can respond to text questions (prompts) or questions that combine text and images.
[0031] Modification means 22 Modification means 22 modifies the version of the trained model 21 to be modified. In the second exemplary embodiment, modification means 22 modifies the version of the trained model 21 when modification data used to modify the version of the trained model 21 is input, when a predetermined modification start operation is received, etc.
[0032] ・User Interface 3 User Interface 3 accepts input of questions from the user. Upon receiving input of a question, User Interface 3 transmits the received question to the question answering device 2. When User Interface 3 receives an answer to the question from the question answering device 2, it displays the received answer. User Interface 3 can be configured as, for example, a PC, mobile phone, tablet terminal, etc., that communicates with the question answering device 2 via a network. Note that User Interface 3 may also be configured as the question answering device 2. Furthermore, User Interface 3 does not have to be configured as the question answering system 100. In addition, User Interface 3 may accept input of evaluation results from the user regarding the answers output by the trained model 21. The evaluation results may be a binary choice such as "good" or "bad," or a multi-level scale such as "1," "2," "3," "4," or "5." Furthermore, the evaluation results may be subjective, such as "thank you" or "completely wrong."
[0033] - First Database 4 The first database 4 stores questions previously input to the trained model 21 and answers to those questions output by the trained model 21. The first database 4 also stores words included in questions previously input to the trained model 21 and the frequency of those words appearing. Furthermore, the first database 4 accumulates questions and answers each time a question is input to the trained model 21 and the trained model 21 outputs an answer. Note that the first database 4 may be configured as the reward distribution device 1A. Also, the first database 4 does not have to be configured as the question answering system 100.
[0034] - Second Database 5 The second database 5 stores various modification data used / used to change the version of the trained model 21. In addition, the second database 5 stores modification data each time it is provided. The second database 5 may also be configured as the reward distribution device 1A. Furthermore, the second database 5 does not have to be configured as the question answering system 100.
[0035] - Reward distribution device 1A The reward distribution device 1A according to the second exemplary embodiment comprises a first evaluation means 11A, a second evaluation means 12A, a calculation means 13A, and an output means 14A. The reward distribution device 1A according to the second exemplary embodiment further comprises an acquisition means 15, a identification means 16, and a third evaluation means 17.
[0036] - Acquisition means 15: Each time the trained model 21 outputs an answer, the acquisition means 15 acquires the content of the outputted answer from the trained model 21 along with the question. The acquisition means 15 also acquires the content of multiple questions that have been input to the trained model 21 so far from the first database 4. Note that the questions are not limited to text. That is, the questions may consist of images, audio, etc.
[0037] - First evaluation means 11A The first evaluation means 11A in the second exemplary embodiment evaluates the contribution of the modification data used to change the version of the trained model to improving the response quality, based on the response quality of the trained model, similar to the first evaluation means 11 in the first exemplary embodiment. The first evaluation means 11A in the second exemplary embodiment evaluates the contribution when the version of the trained model 21 is changed. The first evaluation means 11A in the second exemplary embodiment includes an improvement calculation means 111 and a contribution calculation means 112.
[0038] The improvement calculation means 111 calculates the degree of improvement in the response quality of the modified trained model 21 by performing benchmark tests on the trained model 21 before the version change and the trained model 21 after the version change. Specifically, the improvement calculation means 111 quantifies the degree of deviation from a pre-prepared ideal answer for a given question, for both the answer output by the trained model 21 before the version change and the answer output by the trained model 21 after the version change. The "ideal answer" includes information obtained a certain time after the question is initially input. Specifically, for example, for a question about predicting the sales of a certain company before the financial results are announced, the sales figures that become public at the financial results announcement (for example, three months after the question) (this is compared with the answer (prediction) that the trained model 21 had output before the financial results announcement) would be relevant. The improvement calculation means 111 then calculates the improvement as the difference between the degree of deviation of the answer after the version change and the degree of deviation of the answer before the version change. Specifically, the improvement calculation means 111 calculates a higher improvement if the deviation of the responses after the version change is smaller than the deviation of the responses before the version change. More specifically, the improvement calculation means 111 may calculate the improvement as the difference between the accuracy rate of the trained model 21 after the version change and the accuracy rate of the trained model 21 before the version change. Alternatively, the improvement calculation means 111 may calculate the improvement as the accuracy rate of the trained model 21 after the version change.
[0039] The improvement calculation means 111 may be configured to calculate the degree of improvement in the response quality of the modified trained model 21 based on the user's evaluation results for the responses output by the trained model 21 before the version change and the user's evaluation results for the responses output by the modified trained model 21. Specifically, the improvement calculation means 111 may calculate a higher degree of improvement if there are many positive evaluation results. More specifically, the improvement calculation means 111 may calculate the degree of improvement as the difference between the number of positive evaluation results and the number of negative evaluation results. Alternatively, the improvement calculation means 111 may calculate the degree of improvement as the number of positive evaluation results.
[0040] Furthermore, the improvement calculation means 111 may be configured to calculate the improvement based on the user's evaluation results and the results of benchmark tests.
[0041] The contribution calculation means 112 calculates the contribution based on the calculated improvement level. Specifically, the contribution calculation means 112 calculates a higher improvement level the higher the improvement level. If multiple types of modification data are used for version changes, the contribution calculation means 112 may be configured to calculate a provisional contribution level for each modification data. The obtained non-provisional contribution levels may then be apportioned based on the ratio of the provisional contribution levels. Furthermore, the contribution calculation means 112 may be configured to calculate the contribution level based on the improvement level calculated using a benchmark and the improvement level calculated based on user evaluation. In this case, the contribution calculation means 112 can calculate the contribution level using, for example, the following formula (1). In the following formula (1), α and β are proportionality coefficients set by the service provider that provides services using the reward distribution device 1A. Contribution level = α (improvement level based on user evaluation results) + β (improvement level based on benchmark test) ... (1)
[0042] - Second evaluation means 12A The second evaluation means 12A according to the second exemplary embodiment evaluates the degree of relevance between a plurality of questions and the change data based on the content of the plurality of questions input to the learned model so far, similar to the second evaluation means 12 according to the first exemplary embodiment. The second evaluation means 12A according to the second exemplary embodiment classifies each question into one of a plurality of categories (category A, category B,...), and calculates the ratio of the number of questions classified into each category to the total number of questions as the degree of relevance to each category of the change data.
[0043] The second evaluation means 12A according to the second exemplary embodiment extracts each of the plurality of words included in each question from the plurality of questions acquired by the acquisition means 15, and calculates the frequency of appearance of each word. Then, for the learning data, the second evaluation means 12A calculates, as the degree of relevance, how much the word with the highest frequency or the words with relatively high frequencies (top 〇%) are included. On the other hand, for the small model, the second evaluation means 12A calculates, as the degree of relevance, how much the weighting regarding the word with the highest frequency or the words with relatively high frequencies is.
[0044] - Third evaluation means 17 When the learned model 21 is composed of a plurality of small models, the third evaluation means 17 evaluates, based on the answer of the learned model 21, the second contribution degree, which is the contribution degree of each small model constituting the learned model to the answer.
[0045] ・Calculation means 13A In the second exemplary embodiment, calculation means 13A calculates the remuneration to be paid to the provider of the modification data based on the degree of contribution and the degree of relevance, similar to calculation means 13 in the first exemplary embodiment.Calculation means 13A in the second exemplary embodiment calculates the remuneration using, for example, the following formula (2).In formula (2) below, γ is a proportionality coefficient that determines the remuneration, and is determined by the provider, etc., based on the amount of the source of the remuneration to be distributed.Remuneration = γ × degree of contribution × (degree of relevance of the modification data to category A + degree of relevance to category B + degree of relevance to category C)...(2)Calculation means 13A in the second exemplary embodiment calculates the remuneration to be paid to the provider of the modification data based on the degree of contribution, the degree of relevance, and the second degree of contribution.Calculation means 13A may be configured to calculate the remuneration to be paid to the provider of the modification data with a high degree of contribution, and separately to calculate the remuneration to be paid to the provider of the modification data with a high degree of relevance.
[0046] Incidentally, while contribution is evaluated based on the change data used in the most recent version change, relevance may also be evaluated based on the change data used in past version changes. For this reason, the calculation means 13A in the second exemplary embodiment also calculates the compensation to be paid to those who previously provided change data with a high degree of relevance.
[0047] Furthermore, the trained model 21 may experience a decrease in response quality after a version change. In such cases, it becomes necessary to further modify the version of the trained model 21 by deleting at least one of the submodels and training data used in the version change from the trained model 21 after the version change. For this reason, the calculation means 13A according to the second exemplary embodiment calculates a reward to be paid to the provider of the specific information that led to the identification of at least one of the deleted submodels and training data when the evaluation result of at least one of the contribution and relevance improves as a result of deleting at least one of the submodels and training data used in the version change from the trained model 21.
[0048] - Specific means 16 The specific means 16 identifies improvement information necessary for the learned model 21 to output an answer with higher quality than the answer output by the learned model 21, based on the answer output by the learned model 21. The data serving as improvement information is not particularly limited. That is, the improvement information may be, for example, text, images, audio, etc. These may be acquired from devices such as cameras and microphones, or may be downloaded from web pages on the Internet or social network services. In addition, examples of images include images with higher resolution than the previously input image, images taken at the same location at multiple times with time intervals, etc.
[0049] - Output means 14A The output means 14A according to the second exemplary embodiment outputs the content of the reward calculated by the calculation means 13A, similar to the output means 14 according to the first exemplary embodiment. The content of the reward includes, in addition to the reward amount, information identifying the change data that is the source of the obtained reward. When the specific means 16 identifies improvement information, the output means 14A according to the second exemplary embodiment also outputs information regarding the solicitation of improvement information. Specifically, the output means 14A displays a screen showing, for example, as shown in FIG. 4, the type of information required as improvement information (monitoring camera images, aerial images, satellite images, etc.), as well as that a reward corresponding to the contribution of the provided improvement information can be obtained if improvement information is provided, the acquisition location of the information, the acquisition time, etc. The developer of the learned model 21, financial institutions, etc. who have confirmed the content of the reward will pay the reward to the provider of the change data based on the output content of the reward.
[0050] (Effect of the reward distribution device 1A) According to the reward distribution device 1A described above, the same effects as those of the reward distribution device 1 according to the first exemplary embodiment can be obtained. That is, according to the reward distribution device 1A, an effect can be obtained that the holder of the change data that can be used for version change of the learned model 21 can be motivated to provide the change data to the developer of the learned model 21.
[0051] Furthermore, in the reward distribution device 1A described above, the calculation means 13A is configured to calculate a reward to be paid to the provider of the specific information that triggered the identification of the deleted modification data when the evaluation result of at least one of the contribution and relevance improves as a result of deleting the modification data from the trained model 21. Therefore, with the reward distribution device 1A, a reward is also paid to the provider of the specific information (the number of people eligible for payment is expanded), which has the effect of further increasing the motivation of holders of modification data to provide the modification data.
[0052] Furthermore, the reward distribution device 1A is configured to include identification means 16 that identify improvement information based on the responses output by the trained model 21. As a result, with the reward distribution device 1A, a person who has confirmed the improvement information will either provide modification data that matches the improvement information from the modification data they possess, or create and provide new modification data that matches the improvement information. As a result, the response quality of the trained model 21 improves, and the calculation means 13A calculates a reward for the provider of the modification data (creating a virtuous cycle of provision of modification data and increased rewards).
[0053] (Flow of reward distribution method S1A) Next, the flow of the reward distribution method S1A will be explained with reference to Figure 5. Figure 5 is a flowchart showing the flow of the reward distribution method S1A. The reward distribution method S1A according to the second exemplary embodiment includes a first evaluation step S11A, a second evaluation step S12A, a calculation step S13A, and an output step S14A, as shown in Figure 5. The reward distribution method S1A according to the second exemplary embodiment further includes an acquisition step S15 and a identification step S16.
[0054] - Acquisition process S15 In the initial acquisition process S15, at least one processor acquires the content of the outputted answers from the trained model 21 along with the questions. Also in the acquisition process S15, at least one processor acquires the content of multiple questions that have been input to the trained model 21 so far from the first database 4. The processor that acquires various information may be provided in the reward distribution device 1A or in another device.
[0055] - First evaluation step S11A After acquisition step S15, the process moves to the first evaluation step S11A. In the first evaluation step S11A according to the second exemplary embodiment, at least one processor evaluates the contribution of the modification data used to change the version of the trained model to improving the response quality, based on the response quality of the trained model, similar to the first evaluation step S11 according to the first exemplary embodiment. In the first evaluation step S11A according to the second exemplary embodiment, at least one processor evaluates the contribution when the version of the trained model 21 is changed. The processor that evaluates the contribution may be provided by the reward distribution device 1A or by another device. The first evaluation step S11A according to the second exemplary embodiment includes an improvement calculation step S111 and a contribution calculation step S112.
[0056] In the improvement calculation step S111, at least one processor performs benchmark tests on the trained model 21 before the version change and the trained model 21 after the change, thereby calculating the degree of improvement in the response quality of the trained model 21 after the change.
[0057] In the contribution calculation step S112, at least one processor calculates the contribution based on the calculated improvement.
[0058] - Second evaluation step S12A After the first evaluation step S11A, the process moves to the second evaluation step S12A. In the second evaluation step S12A according to the second exemplary embodiment, at least one processor evaluates the degree of relevance between a plurality of questions and the modification data based on the content of a plurality of questions that have been input to the trained model so far, similar to the second evaluation step S12 according to the first exemplary embodiment. The processor that evaluates the degree of relevance may be provided by the reward distribution device 1A or by another device. The second evaluation step S12A may be performed before the first evaluation step S11A or in parallel with the first evaluation step S11A.
[0059] - Third evaluation step S17 After the first evaluation step S11A and the second evaluation step S12A, the process moves to the third evaluation step S17. The third evaluation step S17 includes the judgment step S171 and the evaluation step S172.
[0060] In the initial decision step S171, at least one processor determines whether the trained model 21 is composed of multiple submodels.
[0061] If the decision step S171 determines that the trained model 21 is composed of multiple submodels (S171: YES), the process moves to the evaluation step S172. In the evaluation step S172, at least one processor evaluates the second contribution of each submodel constituting the trained model, which is the degree of contribution to the answer, based on the answer of the trained model 21. If the decision step S171 determines that the trained model 21 is not composed of multiple submodels (S171: NO), the process moves to the calculation step S13A.
[0062] Calculation process S13A After the first evaluation process S11A and the second evaluation process S12A, the process moves to calculation process S13A. Calculation process S13A according to the second exemplary embodiment includes a determination process S131, a first calculation process S132, and a second calculation process S133.
[0063] In the initial decision step S131, at least one processor determines whether the improvement in response quality is due to the deletion of data to be modified.
[0064] If the determination step S131 determines that the improvement in response quality is not due to the deletion of the data to be modified (S131: NO), the process proceeds to the first calculation step S132. In the first calculation step S132, at least one processor calculates the reward to be paid to the provider of the data to be modified, based on the degree of contribution and the degree of relevance, similar to the calculation step S13 in the first exemplary embodiment described above. The processor that calculates the reward may be provided by the reward distribution device 1A, or it may be provided by another device.
[0065] If, in the judgment step S131, it is determined that the improvement in response quality is due to the deletion of the modified data (that the evaluation result of at least one of the contribution and relevance improved by deleting at least one of the submodels and training data used to change the version from the trained model 21) (S131: YES), then the process moves to the second calculation step S133. In the second calculation step S133, at least one processor calculates the reward to be paid to the provider of the specific information that led to the identification of at least one of the deleted submodels and training data.
[0066] Furthermore, in the calculation step S13A of the second exemplary embodiment, at least one processor also calculates the compensation to be paid to those who have previously provided highly relevant modification data.
[0067] - Specific step S16 In specific step S16, at least one processor identifies improvement information necessary for the trained model 21 to output a higher quality response than the response output by the trained model 21, based on the response output by the trained model 21. The processor that identifies the improvement information may be provided by the reward distribution device 1A or by another device.
[0068] - Output process S14A In the output process S14 of the second exemplary embodiment, at least one processor causes the output means 14A to output the content of the reward calculated in the calculation process S13A, similar to the output process S14 of the first exemplary embodiment. If improvement information is identified in the identification process S16, in the output process S14A of the second exemplary embodiment, at least one processor also causes the output means 14A to output information related to soliciting improvement information.
[0069] (Effects of reward distribution method S1A) The reward distribution method S1A described above can be used to obtain the same effects as the reward distribution method S1 in the first exemplary embodiment described above. In other words, the reward distribution method S1A can be used to increase the motivation of holders of modification data that can be used to change the version of the trained model to provide the modification data to the developer of the trained model.
[0070] Furthermore, in the reward distribution method S1A described above, if the evaluation result of at least one of the contribution and relevance improves as a result of deleting the modification data from the trained model 21, then in the calculation step S13A, at least one processor calculates the reward to be paid to the provider of the specific information that triggered the identification of the deleted modification data. Therefore, according to the reward distribution method S1A, a reward is also paid to the provider of the specific information (the number of payees is expanded), which has the effect of further increasing the motivation of holders of modification data to provide the modification data.
[0071] Furthermore, the reward distribution method S1A includes a specific step S16 that identifies improvement information based on the responses output by the trained model 21. Therefore, according to the reward distribution method S1A, the person who confirms the improvement information will either provide modification data that matches the improvement information from the modification data they possess, or create and provide new modification data that matches the improvement information. As a result, the response quality of the trained model 21 improves, and in the calculation step S13A, a reward is calculated for the provider of the modification data (creating a positive cycle of provision of modification data and increased rewards).
[0072] [Example of implementation by software] Some or all of the functions of the reward distribution devices 1 and 1A (hereinafter also referred to as "each of the above devices") may be implemented by hardware such as integrated circuits (IC chips) or by software.
[0073] In the latter case, each of the above devices is implemented, for example, by a computer that executes instructions for a program, which is software that realizes each function. An example of such a computer (hereinafter referred to as computer C) is shown in Figure 6. Figure 6 is a block diagram showing the hardware configuration of computer C, which functions as each of the above devices.
[0074] Computer C comprises at least one processor C1 and at least one memory C2. Memory C2 stores a reward distribution program P for operating Computer C as described above. In Computer C, the processor C1 reads the reward distribution program P from memory C2 and executes it, thereby realizing each of the functions described above.
[0075] For processor C1, for example, a CPU (Central Processing Unit), GPU (Graphic Processing Unit), DSP (Digital Signal Processor), MPU (Micro Processing Unit), FPU (Floating Point Number Processing Unit), PPU (Physics Processing Unit), TPU (Tensor Processing Unit), quantum processor, microcontroller, or a combination thereof can be used. For memory C2, for example, flash memory, HDD (Hard Disk Drive), SSD (Solid State Drive), or a combination thereof can be used.
[0076] Furthermore, computer C may also be equipped with RAM (Random Access Memory) for deploying the reward distribution program P at runtime and for temporarily storing various data. Computer C may also be equipped with a communication interface for sending and receiving data with other devices. Furthermore, computer C may also be equipped with an input / output interface for connecting input / output devices such as a keyboard, mouse, display, and printer.
[0077] Furthermore, the reward distribution program P can be recorded on a tangible, non-temporary recording medium M that is readable by the computer C. Such a recording medium M could be, for example, a tape, disk, card, semiconductor memory, or a programmable logic circuit. The computer C can acquire the reward distribution program P via such a recording medium M. The reward distribution program P can also be transmitted via a transmission medium. Such a transmission medium could be, for example, a communication network or broadcast waves. The computer C can also acquire the reward distribution program P via such a transmission medium.
[0078] [Addendum 1] This disclosure includes the technologies described in the following addendums. However, the present invention is not limited to the technologies described in the following addendums, and various modifications are possible within the scope of the claims.
[0079] (Note 1) A reward distribution device comprising: a first evaluation means that evaluates the contribution of modification data used to change the version of a trained model to improving the quality of answers to a question, based on the quality of answers of a trained model that has been constructed to output answers to a question when a question is input; a second evaluation means that evaluates the degree of relevance between a plurality of questions and the modification data based on the content of a plurality of questions that have been input to the trained model so far; a calculation means that calculates a reward to be paid to the provider of the modification data based on the degree of contribution and the degree of relevance; and an output means that outputs the content of the reward.
[0080] (Note 2) The reward distribution device according to Note 1, wherein the first evaluation means comprises: an improvement calculation means for calculating the degree of improvement in the response quality of the modified trained model by performing benchmark tests on the trained model before the version change and the trained model after the change; and a contribution calculation means for calculating the contribution based on the calculated improvement.
[0081] (Note 3) The reward distribution device according to Note 1 or 2, wherein the first evaluation means comprises: an improvement calculation means for calculating the degree of improvement in the response quality of the modified trained model based on the user's evaluation results for the responses output by the trained model before the version was changed and the user's evaluation results for the responses output by the trained model after the version was changed; and a contribution calculation means for calculating the contribution based on the calculated degree of improvement.
[0082] (Note 4) The reward distribution device according to any one of Notes 1 to 3, further comprising a third evaluation means for evaluating a second contribution, which is the degree of contribution of each sub-model constituting the trained model to the response, based on the response of the trained model, wherein the calculation means calculates a reward to be paid to the provider of the modification data based on the contribution, the relevance, and the second contribution.
[0083] (Note 5) The trained model is composed of multiple submodels, and its version is changed by adding a new submodel or training new training data, and the calculation means calculates a reward to be paid to the provider of specific information that led to the identification of at least one of the deleted submodel and training data when the evaluation result of at least one of the contribution and relevance improves as a result of deleting at least one of the submodel and training data used to change the version of the trained model, according to any one of Notes 1 to 4.
[0084] (Note 6) The reward distribution device according to any one of Notes 1 to 5, wherein the trained model is composed of multiple submodels, and its version is changed by adding a new submodel or training new training data, and if the evaluation result of at least one of the contribution and relevance decreases as a result of changing the version of the trained model, the device further comprises a version change means for deleting at least one of the submodels and training data added at the time of the change from the trained model after the version change.
[0085] (Note 7) The reward distribution device according to any one of Notes 1 to 6, wherein the calculation means also calculates the reward to be paid to persons who have previously provided the relevant modification data.
[0086] (Note 8) The reward distribution device according to any one of Notes 1 to 7, further comprising identification means for identifying improvement information necessary for the trained model to output a higher quality response than the response output by the trained model, wherein the output means also outputs information regarding the solicitation of the improvement information.
[0087] (Note 9) A question answering system comprising: a trained model constructed to output an answer to a question when a question is input; a means for changing the version of the trained model; and a reward distribution device as described in any one of Notes 1 to 8.
[0088] (Note 10) A reward distribution program that causes at least one processor to execute: a first evaluation process that evaluates the contribution of modification data used to change the version of a trained model to improving the quality of answers to a question, based on the quality of answers of a trained model that is built to output answers to a question when a question is input; a second evaluation process that evaluates the degree of relevance between a plurality of questions and the modification data, based on the content of a plurality of questions that have been input to the trained model so far; a calculation process that calculates a reward to be paid to the provider of the modification data based on the contribution and the degree of relevance; and an output process that outputs the content of the reward.
[0089] (Note 11) A reward distribution method comprising: a first evaluation step in which at least one processor evaluates the contribution of modification data used to change the version of a trained model to improving the quality of answers, based on the quality of answers of a trained model that is constructed to output answers to questions when a question is input; a second evaluation step in which at least one processor evaluates the degree of relevance between a plurality of questions and the modification data, based on the content of a plurality of questions that have been input to the trained model so far; a calculation step in which at least one processor calculates a reward to be paid to the provider of the modification data based on the contribution and the degree of relevance; and an output step in which at least one processor causes an output means to output the content of the reward.
[0090] (Note 12) The reward distribution program according to Note 10, wherein in the first evaluation process, at least one processor is made to perform an improvement calculation process to calculate the degree of improvement in the response quality of the modified trained model by performing benchmark tests on the trained model before the version change and the trained model after the change, and a contribution calculation process to calculate the contribution based on the calculated improvement.
[0091] (Note 13) The reward distribution program according to Note 10 or 12, wherein in the first evaluation process, at least one processor is made to perform: an improvement calculation process that calculates the degree of improvement in the response quality of the modified trained model based on the user's evaluation results for the responses output by the trained model before the version change and the user's evaluation results for the responses output by the modified trained model; and a contribution calculation process that calculates the contribution based on the calculated improvement.
[0092] (Note 14) A reward distribution device according to Note 10, 12, or 13, wherein at least one processor further executes a third evaluation means that evaluates a second contribution, which is the degree of contribution of each sub-model constituting the trained model to the answer, based on the answer of the trained model, when the trained model is composed of a plurality of sub-models, and in the calculation process, calculates a reward to be paid to the provider of the modification data based on the contribution, the degree of relevance, and the second contribution.
[0093] (Note 15) The pre-trained model is composed of multiple sub-models, and its version is changed by adding new sub-models or training new training data, and in the calculation process, at least one processor is instructed to calculate a reward to be paid to the provider of specific information that led to the identification of at least one of the sub-models and training data that were deleted from the pre-trained model, if the evaluation result of at least one of the contribution and relevance improves as a result of deleting at least one of the sub-models and training data used to change the version of the pre-trained model. This is the reward distribution program described in any one of Notes 10, 12 to 14.
[0094] (Note 16) The pre-trained model is composed of multiple sub-models, and its version is changed by adding a new sub-model or training new training data, and the reward distribution program according to any one of Notes 10, 12 to 15 further causes at least one processor to perform a version change process to delete at least one of the sub-models and training data added at the time of the change from the pre-trained model after the version change if the evaluation result of at least one of the contribution and the relevance decreases as a result of changing the version of the pre-trained model.
[0095] (Note 17) The reward distribution program according to any one of Notes 10, 12 to 16, wherein in the calculation process, at least one processor is also made to calculate the reward to be paid to persons who have previously provided the relevant modification data.
[0096] (Note 18) A reward distribution program according to any one of Notes 10, 12 to 17, wherein at least one processor further performs a specific processing to identify improvement information necessary for the trained model to output a higher quality response than the response output by the trained model, and in the output processing, the output means also outputs information regarding the solicitation of the improvement information.
[0097] (Note 19) The reward distribution method according to Note 11, wherein the first evaluation step includes: an improvement calculation step in which at least one processor performs benchmark tests on the trained model before the version change and the trained model after the change to calculate the degree of improvement in the response quality of the trained model after the change; and a contribution calculation step in which at least one processor calculates the contribution based on the calculated improvement.
[0098] (Note 20) The reward distribution method according to Note 11 or 19, wherein the first evaluation step includes: an improvement calculation step in which at least one processor calculates the degree of improvement in the response quality of the modified trained model based on the user's evaluation results for the responses output by the trained model before the version change and the user's evaluation results for the responses output by the modified trained model; and a contribution calculation step in which at least one processor calculates the contribution based on the calculated improvement.
[0099] (Note 21) The reward distribution method according to Note 11, 19, or 20, further comprising a third evaluation step in which at least one processor evaluates a second contribution, which is the degree of contribution of each sub-model constituting the trained model to the answer, based on the answer of the trained model, wherein in the calculation step, at least one processor calculates a reward to be paid to the provider of the modification data based on the contribution, the relevance, and the second contribution.
[0100] (Note 22) The reward distribution method according to any one of Notes 11, 19 to 21, wherein the trained model is composed of multiple submodels, and its version is changed by adding a new submodel or training new training data, and in the calculation step, at least one processor calculates a reward to be paid to the provider of the specific information that led to the identification of at least one of the deleted submodel and training data when the evaluation result of at least one of the contribution and relevance improves as a result of deleting at least one of the submodel and training data used to change the version of the trained model.
[0101] (Note 23) The reward distribution method according to any one of Notes 11, 19 to 22, wherein the trained model is composed of multiple submodels, and its version is changed by adding a new submodel or training new training data, and at least one processor further includes a version change step of deleting at least one of the submodels and training data added at the time of the change from the trained model after the version change if the evaluation result of at least one of the contribution and the relevance decreases as a result of changing the version of the trained model.
[0102] (Note 24) The reward distribution method according to any one of Notes 11, 19 to 23, wherein in the calculation step, at least one processor also calculates the reward to be paid to persons who have previously provided the relevant modification data.
[0103] (Note 25) The reward distribution method according to any one of Notes 11, 19 to 24, further comprising a selection step in which at least one processor identifies improvement information necessary for the trained model to output a higher quality response than the response output by the trained model, wherein in the output step, at least one processor also causes the output means to output information regarding the solicitation of the improvement information.
[0104] [Addendum 2] This disclosure includes the technologies described in the following addendums. However, the present invention is not limited to the technologies described in the following addendums, and various modifications are possible within the scope of the claims.
[0105] (Note 1) A reward distribution device comprising at least one processor, the at least one processor performing: a first evaluation process that, when a question is input, evaluates the contribution of modification data used to change the version of a trained model to improving the quality of the answer based on the quality of the answer of a trained model that has been constructed to output an answer to the question; a second evaluation process that evaluates the degree of relevance between a plurality of questions and the modification data based on the content of a plurality of questions that have been input to the trained model so far; a calculation process that calculates a reward to be paid to the provider of the modification data based on the contribution and the degree of relevance; and an output process that causes the content of the reward to be output to an output means.
[0106] (Note 2) The reward distribution device according to Note 1, wherein the at least one processor performs, in the first evaluation process, an improvement calculation process that calculates the degree of improvement in the response quality of the modified trained model by performing benchmark tests on the trained model before the version change and the trained model after the change, and a contribution calculation process that calculates the contribution based on the calculated improvement.
[0107] (Note 3) The reward distribution device according to Note 1, wherein the at least one processor performs, in the first evaluation process, an improvement calculation process that calculates the degree of improvement in the response quality of the modified trained model based on the user's evaluation result for the response output by the trained model before the version change and the user's evaluation result for the response output by the modified trained model, and a contribution calculation process that calculates the contribution based on the calculated improvement.
[0108] (Note 4) When the trained model is composed of a plurality of submodels, the reward distribution device according to Note 1, wherein the at least one processor further performs a third evaluation process to evaluate a second contribution, which is the degree of contribution of each submodel constituting the trained model to the answer, based on the answer of the trained model, and in the calculation process, the at least one processor calculates a reward to be paid to the provider of the modification data based on the contribution, the relevance, and the second contribution.
[0109] (Note 5) The trained model is composed of multiple submodels, and its version is changed by adding a new submodel or training new training data, and in the calculation process, the at least one processor calculates a reward to be paid to the provider of specific information that led to the identification of the deleted submodel and training data when the evaluation result of at least one of the contribution and relevance improves as a result of deleting at least one of the submodel and training data used to change the version of the trained model. (Note 5) The reward distribution device according to Note 1.
[0110] (Note 6) The reward distribution device according to Note 1, wherein the trained model is composed of multiple submodels, and its version is changed by adding a new submodel or training new training data, and the at least one processor further performs a version change process to delete at least one of the submodels and training data added at the time of the change from the trained model after the version change if the evaluation result of at least one of the contribution and the relevance decreases as a result of changing the version of the trained model.
[0111] (Note 7) The reward distribution device according to Note 1, wherein in the calculation process, at least one processor also calculates the reward to be paid to persons who have previously provided the relevant modification data.
[0112] (Note 8) The reward distribution device according to Note 1, wherein at least one processor further performs a specific processing to identify improvement information necessary for the trained model to output a higher quality response than the response output by the trained model, and in the output processing, also causes the output means to output information regarding the solicitation of the improvement information.
[0113] (Note 9) A non-temporary recording medium that records a reward distribution program, which causes at least one processor to execute: a first evaluation process that evaluates the contribution of modification data used to change the version of a trained model to improving the quality of answers to a question, based on the quality of answers of a trained model that is constructed to output an answer to a question when a question is input; a second evaluation process that evaluates the degree of relevance between a plurality of questions and the modification data, based on the content of a plurality of questions that have been input to the trained model so far; a calculation process that calculates a reward to be paid to the provider of the modification data based on the contribution and the degree of relevance; and an output process that causes the content of the reward to be output to an output means.
[0114] Although the present disclosure has been described above with reference to embodiments, the present disclosure is not limited to the embodiments described above. Various modifications to the structure and details of the present disclosure can be made as can be understood by those skilled in the art within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.
[0115] Each drawing is merely illustrative to illustrate one or more embodiments. Each drawing may be associated with one or more other embodiments, rather than being associated with only one specific embodiment. As those skilled in the art will understand, various features or steps described with reference to any one drawing can be combined with features or steps shown in one or more other drawings, for example, to create embodiments not explicitly shown or described. Not all features or steps shown in any one drawing to illustrate an exemplary embodiment are necessarily required, and some features or steps may be omitted. The order of steps described in any of the drawings may be changed as appropriate.
[0116] This application claims priority based on Japanese Patent Application No. 2025-015409, filed on 31 January 2025, and incorporates all of its disclosures herein.
[0117] 100 Question Answering System 1, 1A Reward Distribution Device 11, 11A First Evaluation Means 111 Improvement Calculation Means 112 Contribution Calculation Means 12, 12A Second Evaluation Means 13, 13A Calculation Means 14, 14A Output Means 15 Acquisition Means 16 Identification Means 17 Third Evaluation Means 2 Question Answering Device 21 Trained Model 22 Modification Means 3 User Interface 4 First Database 5 Second Database C1 Processor C2 Memory S1, S1A Reward Distribution Method S11, S11A First Evaluation Process S111 Improvement Calculation Process S112 Contribution Calculation Process S12, S12A Second Evaluation Process S13, S13A Calculation Process S131 Judgment Process S132 First Calculation Process S133 Second Calculation Process S14, S14A Output process S15 Acquisition process S16 Identification process S17 Third evaluation process S171 Judgment process S172 Evaluation process
Claims
1. A reward distribution device comprising: a first evaluation means that evaluates the contribution of modification data used to change the version of a trained model to improving the quality of answers to a question, based on the quality of answers to the trained model, which is constructed to output answers to the question when a question is input; a second evaluation means that evaluates the degree of relevance between the modification data and the content of the multiple questions that have been input to the trained model to date; a calculation means that calculates a reward to be paid to the provider of the modification data based on the contribution and the degree of relevance; and an output means that outputs the content of the reward.
2. The reward distribution device according to claim 1, wherein the first evaluation means comprises: an improvement calculation means for calculating the degree of improvement in the response quality of the modified trained model by performing benchmark tests on the trained model before the version change and the trained model after the change; and a contribution calculation means for calculating the contribution based on the calculated improvement.
3. The reward distribution device according to claim 1 or 2, wherein the first evaluation means comprises: an improvement calculation means for calculating the degree of improvement in the response quality of the modified trained model based on the user's evaluation results for the responses output by the trained model before the version change and the user's evaluation results for the responses output by the modified trained model; and a contribution calculation means for calculating the contribution based on the calculated improvement.
4. When the trained model is composed of a plurality of submodels, the reward distribution device according to any one of claims 1 to 3 further comprises a third evaluation means for evaluating a second contribution, which is the degree of contribution of each submodel constituting the trained model to the answer, based on the answer of the trained model, wherein the calculation means calculates a reward to be paid to the provider of the modification data based on the contribution, the relevance, and the second contribution.
5. The reward distribution device according to any one of claims 1 to 4, wherein the trained model is composed of multiple submodels, and its version is changed by adding a new submodel or training new training data, and the calculation means calculates a reward to be paid to the provider of specific information that led to the identification of at least one of the deleted submodel and training data when the evaluation result of at least one of the contribution and relevance improves as a result of deleting at least one of the submodel and training data used to change the version of the trained model.
6. The reward distribution device according to any one of claims 1 to 5, wherein the trained model is composed of a plurality of submodels, and its version is changed by adding a new submodel or training new training data, and the device further comprises a version change means for deleting at least one of the submodels and training data added at the time of the change from the trained model after the version change if the evaluation result of at least one of the contribution and the relevance decreases as a result of the version change of the trained model.
7. The reward distribution device according to any one of claims 1 to 6, wherein the calculation means also calculates the reward to be paid to persons who have previously provided the relevant modification data.
8. The reward distribution device according to any one of claims 1 to 7, further comprising identification means for identifying improvement information necessary for the trained model to output a higher quality response than the response output by the trained model, wherein the output means also outputs information related to soliciting the improvement information.
9. A question answering system comprising: a trained model constructed to output an answer to a question when a question is input; a means for changing the version of the trained model; and a reward distribution device according to any one of claims 1 to 8.
10. A reward distribution method comprising: a first evaluation step in which the processor evaluates the contribution of modification data used to change the version of a trained model to improving the quality of answers to a question, based on the quality of answers of the trained model, which is constructed to output answers to a question when a question is input; a second evaluation step in which the processor evaluates the degree of relevance between a plurality of questions and the modification data, based on the content of a plurality of questions that have been input to the trained model so far; a calculation step in which the processor calculates a reward to be paid to the provider of the modification data based on the contribution and the degree of relevance; and an output step in which the processor causes an output means to output the content of the reward.
11. The reward distribution method according to claim 10, wherein the first evaluation step includes: an improvement calculation step in which at least one processor performs benchmark tests on the trained model before the version change and the trained model after the change to calculate the degree of improvement in the response quality of the trained model after the change; and a contribution calculation step in which at least one processor calculates the contribution based on the calculated degree of improvement.
12. The reward distribution method according to claim 10 or 11, wherein the first evaluation step includes: an improvement calculation step in which at least one processor calculates the degree of improvement in the response quality of the modified trained model based on the user's evaluation results for the responses output by the trained model before the version change and the user's evaluation results for the responses output by the modified trained model; and a contribution calculation step in which at least one processor calculates the contribution based on the calculated degree of improvement.
13. The reward distribution method according to any one of claims 10 to 12, wherein, if the trained model is composed of a plurality of submodels, at least one processor further includes a third evaluation step in which it evaluates a second contribution, which is the degree of contribution of each submodel constituting the trained model to the answer, based on the answer of the trained model, and in the calculation step, at least one processor calculates a reward to be paid to the provider of the modification data based on the contribution, the relevance, and the second contribution.
14. The reward distribution method according to any one of claims 10 to 13, wherein the trained model is composed of a plurality of submodels, and its version is changed by adding a new submodel or training new training data, and in the calculation step, at least one processor calculates a reward to be paid to the provider of specific information that led to the identification of at least one of the deleted submodel and training data, when the evaluation result of at least one of the contribution and relevance improves as a result of deleting at least one of the submodel and training data used to change the version of the trained model.
15. The reward distribution method according to any one of claims 10 to 14, wherein the trained model is composed of a plurality of submodels, and its version is changed by adding a new submodel or training new training data, and the method further includes a version change step in which at least one processor deletes at least one of the submodels and training data added at the time of the change from the trained model after the version change if the evaluation result of at least one of the contribution and the relevance decreases as a result of the version change of the trained model.
16. The reward distribution method according to any one of claims 10 to 15, wherein in the calculation step, at least one processor also calculates the reward to be paid to persons who have previously provided the modification data with a high degree of relevance.
17. The reward distribution method according to any one of claims 10 to 16, further comprising a selection step in which at least one processor identifies improvement information necessary for the trained model to output a higher quality response than the response output by the trained model, wherein in the output step, at least one processor also causes the output means to output information regarding the solicitation of the improvement information.
18. A reward distribution program that causes at least one processor to execute: a first evaluation process that evaluates the contribution of modification data used to change the version of a trained model to improving the quality of answers to a question, based on the quality of answers of a trained model that has been constructed to output an answer to a question when a question is input; a second evaluation process that evaluates the degree of relevance between a plurality of questions and the modification data, based on the content of a plurality of questions that have been input to the trained model so far; a calculation process that calculates a reward to be paid to the provider of the modification data based on the contribution and the degree of relevance; and an output process that causes the content of the reward to be output to an output means.
19. The reward distribution program according to claim 18, wherein in the first evaluation process, at least one processor is made to perform: an improvement calculation process which calculates the degree of improvement in the response quality of the modified trained model by performing benchmark tests on the trained model before the version change and the trained model after the change; and a contribution calculation process which calculates the contribution based on the calculated improvement.
20. The reward distribution program according to claim 18 or 19, wherein in the first evaluation process, at least one processor is made to perform: an improvement calculation process that calculates the degree of improvement in the response quality of the modified trained model based on the user's evaluation results for the responses output by the trained model before the version change and the user's evaluation results for the responses output by the modified trained model; and a contribution calculation process that calculates the contribution based on the calculated improvement.