Information processing device, information processing method, and information processing program
Patent Information
- Application Number
- PCT/JP2025/012341
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2026-10-01
Smart Images

Figure JP2025012341_01102026_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and information processing program
[0001] This invention relates to an information processing device, an information processing method, and an information processing program.
[0002] Fine-tuning (FT), a technique that further trains the parameters of a given base model using task-specific data, is widely used as a method to specialize a given base model for a specific task or domain (hereinafter referred to as "task, etc.").
[0003] The general knowledge and capabilities of a fully modified (FT) model depend on the knowledge and capabilities of the underlying model from which it was created. To update an FT model, a new underlying model must be prepared and the FT process performed again.
[0004] Performing feasibility studies (FT) on a new base model incurs not only training costs but also maintenance costs for storing task-specific data for reuse. Therefore, as the frequency of base model updates and the number of specialized tasks increase, these costs increase linearly. Two approaches are known to address this challenge: learning transfer and inference-time transfer techniques.
[0005] Learning transfer is a technique that produces the same effect as updating the base model by transferring the learning process (parameter sequence) of FT from the old base model to the new base model within the parameter space.
[0006] For example, a disadvantage of transfer learning is that it requires the parameter sizes to be the same in the models before and after the transfer. The transfer itself is a difficult operation, and additional training may be necessary to achieve high accuracy.
[0007] Inference-time transfer technology is a technique that performs a transfer during inference by adding the difference between the outputs of the old FT model and the old base model to the output of the new base model. With inference-time transfer technology, the effect of FT on the new base model can be achieved simply by calculating the output during inference, without additional training. Furthermore, in inference-time transfer technology, the parameter sizes of the models before and after the transfer do not need to be the same.
[0008] Known inference time transition techniques include Emulated Fine-Tuning (see, for example, Non-Patent Document 1) and Proxy-Tuning (see, for example, Non-Patent Document 2).
[0009] Eric Mitchell, Rafael Rafailov, Archit Sharma, Chelsea Finn, Christopher D. Manning, "An Emulator for Fine-Tuning Large Language Models using Small Language Models", [online], [Retrieved October 9, 2024], Internet <URL: https: / / arxiv.org / abs / 2310.12962> Alisa Liu, Xiaochuang Han, Yizhong Wang, Yulia Tsvetkov, Yejin Choi, Noah A. Smith, "Tuning Language Models by Proxy", [online], [Retrieved October 9, 2024], Internet <URL: https: / / arxiv.org / abs / 2401.08565>
[0010] However, conventional inference time transfer techniques have the problem of high computational cost. For example, the methods described in Non-Patent Documents 1 and 2, while allowing for the omission of fine-tuning, require the computation of three models of similar size during inference. For instance, to obtain the output of the new FT model using conventional techniques, computation of three models is required: the old base model, the old FT model, and the new base model.
[0011] Therefore, the present invention aims to eliminate the need to compute the old base model during the inference-time transfer in the inference-time transfer technology.
[0012] To solve the aforementioned problems, the information processing device of the present invention is characterized by having a correction unit that calculates a correction value which is calculated by correction using a first output value calculated by inputting an input value into a first model, and the correction value is obtained by correcting a second output value calculated by inputting the input value into a pre-trained second model using the first output value.
[0013] According to the present invention, in the inference-time transfer technology, it is possible to eliminate the need to calculate the old base model during the inference-time transfer.
[0014] Figure 1 is a diagram showing an example configuration of the information processing device of the first embodiment. Figure 2 is a flowchart showing the processing procedure of the information processing device of the first embodiment. Figure 3 is a flowchart showing the processing procedure for learning the reward model. Figure 4 is a flowchart showing the processing procedure for inference using the reward model. Figure 5 is a diagram illustrating the data flow of the learning process. Figure 6 is a diagram illustrating the data flow of the inference process. Figure 7 is a diagram illustrating the data flow of the inference process. Figure 8 is a diagram illustrating Example 1. Figure 9 is a diagram illustrating Example 2. Figure 10 is a diagram illustrating Example 3. Figure 11 is a diagram showing an example configuration of a computer that executes an information processing program.
[0015] The embodiments for carrying out the present invention will be described below with reference to the drawings. The present invention is not limited to these embodiments.
[0016] [First Embodiment] The configuration of the information processing device that performs processing related to the inference time transfer technology of this embodiment will be explained using Figure 1. Figure 1 is a diagram showing an example of the configuration of the information processing device of the first embodiment. As shown in Figure 1, the information processing device 10 includes an input / output unit 11, a storage unit 12, and a control unit 13.
[0017] The input / output unit 11 is an interface that handles the input and output of various types of data. For example, the input / output unit 11 accepts inputs of model parameters and data to be input into the model.
[0018] The storage unit 12 stores data, programs, etc., that are referenced when the control unit 13 performs various processes. The storage unit 12 is implemented by semiconductor memory elements such as RAM (Random Access Memory) and flash memory, or by storage devices such as hard disks and optical discs.
[0019] The memory unit 12 stores the reward model information 121. Here, the model information in the following explanation refers to the model parameters. For example, if the model is a neural network, the model information would be weights, biases, etc.
[0020] Here, the reward model is a model that calculates values to correct the output of the base model. The information processing device 10 can train the reward model. The information processing device 10 can also perform inference using the trained reward model. Details of the reward model will be described later.
[0021] The control unit 13 is responsible for controlling the entire information processing device 10. The functions of the control unit 13 are realized, for example, by the CPU (Central Processing Unit) executing a program stored in the memory unit 12. The control unit 13 includes a learning unit 131 and an inference unit 132.
[0022] The learning unit 131 includes an input receiving unit 1311, a calculation unit 1312, a correction unit 1313, a reward model optimization unit 1314, and an output processing unit 1315.
[0023] The input reception unit 1311 accepts data input. The calculation unit 1312 performs calculations using the base model and the reward model. The correction unit 1313 corrects the calculation results of the calculation unit 1312.
[0024] The reward model optimization unit 1314 optimizes the reward model. That is, the reward model optimization unit 1314 updates the parameters of the reward model. The output processing unit 1315 outputs the updated reward model parameters. The output processing unit 1315 may store the reward model parameters as reward model information 121 in the storage unit 12, or it may output them to the outside of the information processing device 10. Note that the reward model optimization unit 1314 is an example of an update unit.
[0025] The inference unit 132 includes an input receiving unit 1321, a calculation unit 1322, a correction unit 1323, and an output processing unit 1324.
[0026] The input reception unit 1321 accepts data input. The calculation unit 1322 performs calculations using the base model and the reward model. The correction unit 1323 corrects the calculation results of the calculation unit 1322.
[0027] Although the input receiving unit 1321, calculation unit 1322, and correction unit 1323 are described separately here from the input receiving unit 1311, calculation unit 1312, and correction unit 1313 of the learning unit 131, they may be the same thing (for example, a module or library that can be called from both the learning unit 131 and the inference unit 132).
[0028] The output processing unit 1324 outputs the inference result based on the values corrected by the correction unit 1323. For example, if the base model is for recognizing objects in an image and the corrected values are the probability values for each object, the output processing unit 1324 considers the object with the largest probability value to be the recognized object and outputs a label to identify that object.
[0029] Figure 2 is a flowchart showing the processing procedure of the information processing device of the first embodiment. As shown in Figure 2, first, the information processing device 10 learns the reward model (step S1). Next, the information processing device 10 performs inference using the reward model (step S2).
[0030] The information processing apparatus 10 may perform learning of the reward model and inference using the reward model at different timings. For example, when an old base model is introduced, the information processing apparatus 10 performs both learning of the reward model and inference using the reward model. Thereafter, when the old base model is updated to a new base model, the information processing apparatus 10 performs inference using the reward model without performing learning of the reward model.
[0031] Here, the reward model is a model for correcting the output of the base model such that the output is specialized for a task or the like. That is, the correction by the reward model can replace FT.
[0032] Furthermore, if the reward model has been learned using the old base model, the information processing apparatus 10 does not need to further learn the reward model during inference by the new base model. Accordingly, the information processing apparatus 10 can reduce the computational cost during inference. For example, the information processing apparatus 10 can eliminate the need for computation of the old base model during inference by the new base model.
[0033] The information processing apparatus 10 of the present embodiment aims to obtain an output after FT of the new base model, similarly to conventional inference-time transfer techniques.
[0034] Here, the output in the inference-time transfer techniques described in Non-Patent Document 1 and Non-Patent Document 2 is defined as shown in Formula (1).
[0035]
[0036] θ 1 pt is a parameter of the old base model before FT. θ 1 ft is a parameter of the old base model after FT. θ 2 pt is a parameter of the new base model before FT.
[0037] On the other hand, the output in the present embodiment (the result of correction by the correction unit 1313 and the correction unit 1323) is defined as shown in Formula (2).
[0038]
[0039] θ pt represents the parameters of the old base model before fine-tuning or the new base model. Q θ is a reward model.
[0040] π θ (y|x) and Q θ (y|x) notations mean the output of a model constructed based on the parameter θ. For example, π θ (y|x) can be regarded as a probability distribution for an output y∈Y conditioned on an input x. For example, inference is performed by inputting a specific input value to x.
[0041] For example, in a model for image classification, x is an image and y is a label. Also, for example, in an autoregressive language model, x is a sequence of words up to the (t-1)th position, and y is the t-th word.
[0042] Formulas (1) and (2) mean that the output of the model on the left side (the new base model after fine-tuning) is pseudo-calculated by the right side.
[0043] Here, the reward model in formula (2) outputs a real-valued vector obtained by arranging reward values q y obtained when outputting y for an input x, as shown in formula (3).
[0044]
[0045] In addition, Z is a normalization factor, defined as shown in formula (4).
[0046]
[0047] Furthermore, Q θ (y|x) is input to the exponential function exp(). exp(Q θ (y|x)) is Euler's number e raised to the power of a value calculated using the reward model.
[0048] When training the reward model, the information processing apparatus 10 updates the parameter θ of the reward model such that a loss function as shown in formula (5) is minimized. Note that the parameter θ of the base model pt is fixed.
[0049]
[0050] For example, if the reward model structure is the same as that of the base model, the information processing device 10 will use the parameters θ of the base model. pt Reward model Q with initial value of θ θ By updating this data, the reward model can be trained.
[0051] As shown in equation (1), conventional technology requires calculations for three base models to obtain the output after FT of the base model. On the other hand, as shown in equation (2), the information processing device 10 of this embodiment can obtain the output after FT of the base model with calculations for two base models (base model and reward model).
[0052] Figure 3 will be used to specifically explain the processing procedure for learning the reward model and the processing content of each part of the learning unit 131. Figure 3 is a flowchart showing the processing procedure for learning the reward model.
[0053] First, the input receiving unit 1311 receives input for the parameters of the old base model and the training dataset (step S101). The training dataset is a set of input-output pairs. The output values of the training dataset are, for example, the correct labels.
[0054] Furthermore, the input receiving unit 1311 sets the initial values of the reward model parameters (step S102). The input receiving unit 1311 may also set the initial values of the reward model parameters by copying the parameters of the old base model. In this case, if the old base model is a deep learning model, the reward model will be a deep learning model with the same configuration as the old base model.
[0055] Next, the input receiving unit 1311 samples pairs of input and output values from the training dataset (step S103).
[0056] Next, the calculation unit 1312 inputs the sampled input values into the old base model and calculates the base value (step S104). The base value here is the output of the old base model.
[0057] The correction unit 1313 inputs the input value into the reward model and calculates a correction value (step S105). The correction unit 1313 inputs the same input value as the input value in step S104 into the reward model. The correction value is the output of the reward model.
[0058] Then, the correction unit 1313 corrects the base value with the correction value and calculates the probability value (correction value) (step S106). Note that, as shown in equation (2), etc., the correction unit 1313 can calculate the correction value using a value based on the correction value, rather than the correction value itself. The parameters of the old base model are θ 1 pt If we assume this, the corrected probability value can be expressed as shown in equation (6).
[0059]
[0060] The reward model optimization unit 1314 updates the parameters of the reward model so that the probability values approach the sampled output values (step S107). The output values are the values associated with the input values in step S104, that is, the values that form a pair with the input values.
[0061] If the output value is the correct label, it may be represented as 1 or 0. On the other hand, the probability value obtained by the correction can take on a continuous value in the range of 0 to 1. Step S107 may be a process of updating the parameters so that the probability value approaches 1. For example, the reward model optimization unit 1314 updates the parameters by gradient descent using cross-entropy loss.
[0062] If further learning is required (step S108; Yes), the information processing device 10 returns to step S103 and repeats the process. For example, the information processing device 10 determines that further learning is not required based on conditions such as the parameters being updated a predetermined number of times, a predetermined amount of time elapsed since the start of processing, and the amount of parameter updates converging.
[0063] If further learning is not required (step S108; No), the output processing unit 1315 stores the updated reward model parameters as reward model information 121 in the storage unit 12 (step S109). The output processing unit 1315 may also output the updated reward model parameters to the outside of the information processing device 10.
[0064] Figure 4 will be used to specifically explain the processing procedure for inference using the reward model and the processing content of each part of the inference unit 132. Figure 4 is a flowchart showing the processing procedure for inference using the reward model.
[0065] First, the input receiving unit 1321 receives input of parameters for the new base model and data to be inferred (input values) (step S201). The data to be inferred may be input values whose output values constituting a pair are unknown.
[0066] Next, the input receiving unit 1321 reads the parameters of the reward model, i.e., the updated reward model information 121 (step S202).
[0067] Next, the calculation unit 1322 inputs the input values of the data to be inferred into the new base model and calculates the base values (step S203). The base values here are the output of the new base model.
[0068] The correction unit 1323 inputs the input value into the reward model and calculates the correction value (step S204). The correction unit 1323 inputs the same input value as the input value in step S203 into the reward model.
[0069] Then, the correction unit 1323 corrects the base value with the correction value and calculates the probability value (step S205). The parameters of the new base model are θ 2 pt If we assume this, the corrected probability value can be expressed as shown in equation (7).
[0070]
[0071] The output processing unit 1324 outputs the inference result based on the probability value (step S206). For example, the output processing unit 1324 outputs the label corresponding to the largest probability value among multiple probability values. The output processing unit 1324 may also output the probability value as is.
[0072] Note that the order of steps S201 and S202 may be reversed from that shown in Figure 4. Also, steps S201 and S202 may be executed simultaneously. The order of the other steps may also be changed or executed simultaneously, as long as it does not create inconsistencies.
[0073] The effects of the embodiment will now be described. Note that the calculation unit 1312 and the correction unit 1313 may be replaced with calculation unit 1322 and correction unit 1323 as appropriate.
[0074] The correction unit 1313 calculates a correction value, which is a value obtained by correcting a base value calculated by inputting an input value into the reward model, using the correction value. The base value is a value obtained by correcting a base value calculated by inputting an input value into a pre-trained base model using the correction value. The reward model is an example of a first model. The base model is an example of a second model. The correction value is an example of a first output value. The base value is an example of a second output value.
[0075] As a result, the information processing device 10 can obtain the output of the base model after FT by calculating the outputs of two models: the base model before FT and the reward model. Compared to the conventional technology, which required the calculation of three models, the information processing device 10 can eliminate the need to calculate the old base model during the transfer in inference. For example, a part of equation (1) in the conventional technology (the part shown in equation (8)) is expressed as exp(Q) in this embodiment. θ The computational complexity is reduced because it is replaced by (y | x) (see equation (2)).
[0076]
[0077] The reward model optimization unit 1314 updates the parameters of the reward model so that the value calculated by the correction unit 1313 approaches the output value associated with the input value. In this way, the information processing device 10 can efficiently learn the reward model by utilizing the base model.
[0078] The correction unit 1313 performs correction by multiplying a value based on the correction value by a base value calculated by inputting the input value into the base model. For example, the correction unit 1313 performs correction by multiplying the base value by a value that is a power of e with the correction value as the exponent, and then normalizing it. As a result, the information processing device 10 can adjust the scale of the corrected output by a function.
[0079] Furthermore, the correction performed by the correction unit 1313 is not limited to multiplying the base value by a value based on the correction value; it may be performed using both the correction value and the base value. The correction unit 1313 may also add a value based on the correction value to the base value.
[0080] Furthermore, the output dimensions of the reward model and the base model may be the same. That is, the correction unit 1313 calculates a correction value, which is a value obtained by correcting the base value calculated by inputting the input value into the base model, which has the same output dimension as the reward model. This allows sufficient information to correct the output value of the base model from the reward model with minimal computational effort.
[0081] Here, the data flow in the first embodiment will be explained using Figures 5, 6, and 7.
[0082] Figure 5 illustrates the data flow during the learning process. As shown in Figure 5, the calculation unit 1312 calculates a first output value (correction value) by inputting the input values of the training dataset into the first model (reward model). The calculation unit 1312 also calculates a second output value (base value) by inputting the input values of the training dataset into the second model (old base model).
[0083] The correction unit 1313 calculates a corrected value by correcting the second output value using the first output value. The reward model optimization unit 1314 updates the parameters of the first model based on the corrected value and the ground truth value. The ground truth value is the value that pairs with the input value in the training dataset.
[0084] Figure 6 illustrates the data flow of the inference process. As shown in Figure 6, the calculation unit 1322 calculates a first output value by inputting input values into the first model. The calculation unit 1322 also calculates a second output value by inputting input values into the second model. The second model here may be a new base model. For example, the input values are values entered by the user. The output values that correspond to the input values may be unknown.
[0085] Figure 7 shows other variations of the inference process. Figure 7 is a diagram illustrating the data flow of the inference process. The calculation unit 2322 in Figure 7 is provided in an external device to the information processing device 10.
[0086] In the example shown in Figure 7, the calculation unit 2322 receives an input value from the input receiving unit 1321. The calculation unit 2322 calculates a second output value by inputting the input value into the second model. The calculation unit 2322 passes the second output value to the input receiving unit 1321. The correction unit 1313 calculates a corrected value by correcting the second output value using the first output value.
[0087] [Example 1] The following describes an example based on the first embodiment. Figure 8 is a diagram illustrating Example 1. The starting point of the arrow between the two information providing servers 10a shows the configuration during learning. The ending point of the arrow between the two information providing servers 10a shows the configuration during inference.
[0088] In Embodiment 1, the base model providing server 20a provides the parameters of the base model to the information providing server 10a. The information providing server 10a also provides the inference results from the base model after FT to the user U10a's terminal 30a. In this case, the information providing server 10a uses the reward model created by the method of the first embodiment.
[0089] Users of the information provision server 10a receive the base model from the operator of the base model provision server 20a and provide the inference results to end-users, user U10a.
[0090] As shown in Figure 8, the base model providing server 20a provides the old base model information 31a to the information providing server 10a. The calculation unit 1312a, the correction unit 1313a, and the reward model optimization unit 1314a perform the same processing as the calculation unit 1312, the correction unit 1313, and the reward model optimization unit 1314, respectively.
[0091] During training, the information provision server 10a trains the reward model using the training dataset 41a in accordance with the method of the first embodiment. The initial value of the reward model information 121a is the old base model information 31a.
[0092] The calculation unit 1312a calculates a first output value by inputting the input values of the training dataset 41a into a model based on the reward model information 121a. The calculation unit 1312a also calculates a second output value by inputting the input values of the training dataset 41a into a model based on the old base model information 31a. The correction unit 1313a calculates a corrected value by correcting the second output value using the first output value. The reward model optimization unit 1314a updates the reward model information 121a based on the corrected value.
[0093] Subsequently, suppose the base model providing server 20a begins providing a new base model that is an updated version of the old base model. At this time, the information providing server 10a receives new base model information 32a, which is different from the old base model information 31a. The information providing server 10a can then provide inference results by correcting the output of the new base model with the output of the trained reward model. In this case, the information providing server 10a does not need to retrain the reward model. Of course, the information providing server 10a can still perform inference using the old base model before it was updated.
[0094] During inference, the input receiving unit 1321a receives input values from the user U10a's terminal 30a. The calculation unit 1322a calculates a first output value by inputting the input values into a model based on the reward model information 121a. The calculation unit 1322a also calculates a second output value by inputting the input values into a model based on the new base model information 32a. The correction unit 1323a calculates a corrected value by correcting the second output value using the first output value. The information provision server 10a transmits the corrected value to the terminal 30a as an inference result.
[0095] [Example 2] Figure 9 is a diagram illustrating Example 2. The starting point of the arrow between the two information providing servers 10b shows the configuration during training. The ending point of the arrow between the two information providing servers 10b shows the configuration during inference.
[0096] As shown in Figure 9, in Embodiment 2, the base model providing server 20b provides the output of the base model to the information providing server 10b. The calculation unit 1312b, the correction unit 1313b, and the reward model optimization unit 1314b perform the same processing as the calculation unit 1312, the correction unit 1313, and the reward model optimization unit 1314, respectively.
[0097] In other words, in Embodiment 2, the information provision server 10b does not receive the parameters of the base model. The information provision server 10b also provides the inference results from the base model after FT to the user U10b's terminal 30b. In doing so, the information provision server 10b uses the reward model created using the method of the first embodiment.
[0098] Users of the information provision server 10b receive the output of the base model from the operator of the base model provision server 20b and provide the inference results to end-user user U10b.
[0099] During training, the calculation unit 1312b calculates a first output value by inputting the input values of the inference dataset 41b into the model based on the reward model information 121b. At this time, the information provision server 10b may receive the old base model information 31b as the initial value of the reward model information 121b, or it may set an arbitrary initial value. The calculation unit 1312b also receives a second output value from the base model provision server 20b. The calculation unit 201b of the base model provision server 20b calculates the second output value by inputting the input values of the inference dataset 41b into the model based on the old base model information 31b. The correction unit 1313b calculates a corrected value by correcting the second output value using the first output value. The reward model optimization unit 1314b updates the reward model information 121b based on the corrected value.
[0100] Suppose the old platform model is then updated to the new platform model. At this time, the information provision server 10b receives the output of the new platform model information 32b. The information provision server 10b can, of course, perform inference using the old platform model before the update.
[0101] During inference, the calculation unit 1312b calculates a first output value by inputting the input values of the inference dataset 41b into a model based on the reward model information 121b. The calculation unit 1312b also receives a second output value from the base model provision server 20b. The calculation unit 201b of the base model provision server 20b calculates a second output value by inputting the input values of the inference dataset 41b into a model based on the new base model information 32b. The correction unit 1313b calculates a corrected value by correcting the second output value using the first output value. The information provision server 10b transmits the corrected value as an inference result to the terminal 30b.
[0102] The information provision server 10b can provide inference results by correcting the output of the new base model with the output of the trained reward model. In this case, the information provision server 10b does not need to retrain the reward model.
[0103] In Example 1, the information provision server 10a performed calculations using the base model itself, whereas in Example 2, the information provision server 10b receives the results of calculations using the base model from an external server (base model provision server 20b). This reduces the computational cost for the information provision server 10b.
[0104] [Example 3] In Example 3, the information provision server 10c calculates a correction value by inputting the corresponding input value for each of the multiple users into the reward model corresponding to each of the multiple users, and then calculates a value that is corrected using each of the correction values by inputting the input value into the base model. The information provision server 10c updates the parameters of the reward model so that the corrected value approaches the output value associated with the input value of the corresponding user. This provides a dedicated reward model for each user.
[0105] Figure 10 illustrates Embodiment 3. The starting point of the arrow between the two information-providing servers 10c shows the configuration during training. The ending point of the arrow between the two information-providing servers 10c shows the configuration during inference.
[0106] The information provision server 10c provides inference results to user servers 20_1c, 20_2c, and 20_3c. The information provision server 10c can provide each of the user servers 20_1c, 20_2c, and 20_3c with inference results using a dedicated reward model.
[0107] The learning unit 131c and the inference unit 132c perform the same processing as the learning unit 131 and the inference unit 132, respectively.
[0108] During training, the learning unit 131c receives the training dataset 41_1c from the user server 20_1c. Using the old base model information 31c and the received training dataset 41_1c, the learning unit 131c trains the reward model according to the method of the first embodiment and obtains the reward model information 121_1c. The reward model information 121_1c is the parameters of the reward model specific to the user server 20_1c.
[0109] The learning unit 131c receives the training dataset 41_2c from the user server 20_2c. The learning unit 131c uses the old base model information 31c and the received training dataset 41_2c to train the reward model according to the method of the first embodiment and obtain the reward model information 121_2c. The reward model information 121_2c is the parameters of the reward model specific to the user server 20_2c.
[0110] The learning unit 131c receives the training dataset 41_3c from the user server 20_3c. The learning unit 131c uses the old base model information 31c and the received training dataset 41_3c to train the reward model according to the method of the first embodiment and obtain the reward model information 121_3c. The reward model information 121_3c is the parameters of the reward model specific to the user server 20_3c.
[0111] Suppose the information provision server 10c then updates the old base model to the new base model. At this time, the inference unit 132c can provide inference results by correcting the output of the new base model based on the new base model information 32c with the output of the trained reward model. In this case, the information provision server 10c does not need to retrain the reward model. The information provision server 10b can, of course, perform inference using the old base model before the update.
[0112] Furthermore, multiple new platform models may be created in parallel from a single old platform model. Also, there may be multiple new platform models with different versions, each updated from a single old platform model. In such cases, a user may combine any of these new platform models to create a reward model corresponding to their user server. Alternatively, a user may select one of several old platform models existing on the system and use that selected old platform model to create a reward model corresponding to their user server.
[0113] [System Configuration, etc.] Furthermore, the components of each part shown in the diagram are functional concepts and do not necessarily need to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown in the diagram, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions. In addition, all or any part of the processing functions performed by each device can be realized by a CPU and the program executed on that CPU, or by hardware using wired logic.
[0114] Furthermore, among the processes described in the embodiments described above, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically by known methods. In addition, the processing procedures, control procedures, specific names, and information including various data and parameters shown in the above document and drawings can be arbitrarily changed unless otherwise specified.
[0115] [Program] The information processing device 10 described above can be implemented by installing a program (information processing program) as packaged software or online software on a desired computer. For example, by having the information processing device execute the above program, the information processing device can be made to function as the information processing device 10. The information processing device referred to here includes mobile communication terminals such as smartphones, mobile phones and PHS (Personal Handyphone System), and terminals such as PDA (Personal Digital Assistant).
[0116] Figure 11 shows an example configuration of a computer that executes an information processing program. Computer 1000 has, for example, memory 1010 and CPU 1020. Computer 1000 also has a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.
[0117] Memory 1010 includes ROM (Read Only Memory) 1011 and RAM (Random Access Memory) 1012. ROM 1011 stores, for example, a boot program such as BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to the hard disk drive 1090. The disk drive interface 1040 is connected to the disk drive 1100. For example, a removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to, for example, a mouse 1110 and a keyboard 1120. The video adapter 1060 is connected to, for example, a display 1130.
[0118] The hard disk drive 1090 stores, for example, the OS 1091, application programs 1092, program modules 1093, and program data 1094. That is, the programs that define each process executed by the information processing device 10 are implemented as program modules 1093 in which executable code for a computer is written. The program modules 1093 are stored, for example, in the hard disk drive 1090. For example, a program module 1093 for executing processes similar to the functional configuration of the information processing device 10 is stored in the hard disk drive 1090. Note that the hard disk drive 1090 may be replaced by an SSD (Solid State Drive).
[0119] Furthermore, the data used in the processing of the above-described embodiment is stored as program data 1094 in, for example, memory 1010 or hard disk drive 1090. The CPU 1020 then reads the program module 1093 and program data 1094 stored in memory 1010 or hard disk drive 1090 into RAM 1012 as needed and executes them.
[0120] Furthermore, the program module 1093 and program data 1094 are not limited to being stored in the hard disk drive 1090; for example, they may be stored in a removable storage medium and read by the CPU 1020 via a disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (LAN (Local Area Network), WAN (Wide Area Network), etc.). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via a network interface 1070.
[0121] The following additional information is disclosed regarding the embodiments described above.
[0122] (Note 1) An information processing device comprising: a memory; and at least one processor connected to the memory, wherein the processor calculates a correction value which is a correction value calculated by using a first output value calculated by inputting an input value to a first model, and the correction value is a value obtained by correcting a second output value calculated by inputting the input value to a pre-trained second model using the first output value. (Note 2) An information processing device according to Note 1, wherein the processor updates the parameters of the first model so that the correction value approaches the correct output value associated with the input value. (Note 3) An information processing device according to Note 1 or 2, wherein the processor performs a correction by multiplying the second output value by a value based on the first output value. (Note 4) An information processing device according to Note 3, wherein the processor performs a correction by multiplying the second output value by a value which is a power of e with the first output value as the exponent, and then normalizing it. (Appendix 5) An information processing device according to Appendix 1 or 2, wherein the processor calculates a first output value by inputting the input value to the first model, calculates a second output value by inputting the input value to the second model, and calculates a value obtained by correcting the second output value using the first output value. (Appendix 6) An information processing device according to Appendix 1 or 2, wherein the processor calculates a value obtained by correcting the second output value received from an external server using the first output value.(Appendix 7) An information processing device according to Appendix 2, wherein the processor calculates a correction value calculated by a correction using the first output value calculated by inputting the input value corresponding to each of the plurality of datasets into the first model corresponding to each of the plurality of datasets, the value calculated by inputting the input value into the second model, corrected using each of the first output values, and updates the parameters of the first model so that the calculated value approaches the correct output value associated with the input value of the corresponding dataset. (Appendix 8) An information processing device according to Appendix 1, wherein the processor calculates the correction value which is the value obtained by correcting the second output value calculated by inputting the input value into the second model having the same output dimension as the first model. (Note 9) An information processing device as described in Note 1, wherein the processor calculates the corrected value by correcting the value calculated by inputting input values into a pre-trained model so that the value calculated by inputting input values into a pre-trained model approaches the correct output value associated with the input values, using the first output value calculated by inputting input values into a first model which is a correction model that has been pre-trained to correct the value calculated by inputting training input values into a pre-trained model so that it approaches the correct output value associated with the training input values, and the processor calculates a third output value by inputting inference input values into a correction model whose parameters have been updated so that the value calculated by inputting inference input values into a pre-trained model approaches the correct output value associated with the training input values, and the information processing device calculates a value calculated by inputting inference input values into a pre-trained model using the third output value.(Appendix 11) An information processing device as described in Appendix 10, wherein the processor calculates the third output value by inputting the inference input value into the correction model, whose parameters have been updated so that the value obtained by correcting the value obtained by inputting the training input value into the correction model approaches the correct output value associated with the training input value, and calculates the value obtained by correcting the value obtained by inputting the inference input value into a second pre-trained model different from the first pre-trained model using the third output value. (Appendix 12) A non-temporary storage medium storing a program executable by a computer, wherein the program causes the computer to perform a process of calculating a correction value which is a correction value calculated by correction using a first output value calculated by inputting the input value into the first model, and is a correction value which is a value obtained by correcting the second output value calculated by inputting the input value into a second pre-trained model using the first output value.
[0123] The memory stores the model parameters corresponding to the first and second models. The processor then reads the model parameters from memory.
[0124] 10 Information processing device 11 Input / output unit 12 Storage unit 13 Control unit 121 Reward model information 131 Learning unit 132 Inference unit 1311, 1321 Input receiving unit 1312, 1322 Calculation unit 1313, 1323 Correction unit 1314 Reward model optimization unit 1315, 1324 Output processing unit
Claims
1. An information processing device characterized by having a correction unit that calculates a correction value, which is calculated by a correction using a first output value calculated by inputting an input value into a first model, wherein the correction value is obtained by correcting a second output value calculated by inputting the input value into a pre-trained second model using the first output value.
2. The information processing apparatus according to claim 1, further comprising an update unit that updates the parameters of the first model so that the correction value approaches the correct output value associated with the input value.
3. The information processing apparatus according to claim 1 or 2, characterized in that the correction unit performs correction by multiplying the second output value by a value based on the first output value.
4. The information processing apparatus according to claim 3, characterized in that the correction unit performs the correction by multiplying the second output value by a value which is a power of e with the first output value as the exponent, and then normalizing it.
5. The information processing apparatus according to claim 1 or 2, further comprising a calculation unit that calculates a first output value by inputting the input value into the first model and calculates a second output value by inputting the input value into the second model, wherein the correction unit calculates a value obtained by correcting the second output value using the first output value.
6. The information processing apparatus according to claim 1 or 2, characterized in that the correction unit calculates a value obtained by correcting the second output value received from an external server using the first output value.
7. The information processing apparatus according to claim 2, wherein the correction unit calculates a correction value calculated by using the first output value calculated by inputting the input value corresponding to each of the plurality of datasets into the first model corresponding to each of the plurality of datasets, and the correction value is obtained by correcting the value calculated by inputting the input value into the second model using each of the first output values, and the update unit updates the parameters of the first model so that the value calculated by the correction unit approaches the correct output value associated with the input value of the corresponding dataset.
8. The information processing apparatus according to claim 1, characterized in that the correction unit calculates a correction value which is a value obtained by correcting the second output value calculated by inputting the input value into the second model having the same output dimension as the first model.
9. The information processing apparatus according to claim 1, characterized in that the correction unit calculates the correction value by correcting the value calculated by inputting the input value into a pre-trained model so that it approaches the correct output value associated with the input value, using the first output value calculated by inputting the input value into the first model, which is a correction model that has been pre-trained to correct the value calculated by inputting the input value into a pre-trained model.
10. An information processing apparatus characterized by comprising: a calculation unit that calculates a value by inputting inference input values into a correction model whose parameters have been updated so that the value obtained by correcting the value obtained by inputting the training input values into a correction model approaches the correct output value associated with the training input values; and a correction unit that calculates a value obtained by correcting the value obtained by inputting the inference input values into a pre-trained model using the value calculated by the calculation unit.
11. The information processing apparatus according to claim 10, characterized in that the calculation unit calculates a value by inputting the inference input value into the correction model, whose parameters have been updated so that the value obtained by inputting the training input value into the first pre-trained model and correcting it using the value obtained by inputting the training input value into the correction model approaches the correct output value associated with the training input value, and the correction unit calculates a value obtained by inputting the inference input value into a second pre-trained model different from the first pre-trained model and correcting it using the value calculated by the calculation unit.
12. An information processing method performed by a computer, comprising a correction step of calculating a correction value which is calculated by a correction using a first output value calculated by inputting an input value into a first model, wherein the correction value is a value obtained by correcting a second output value calculated by inputting the input value into a pre-trained second model using the first output value.
13. An information processing program characterized by causing a computer to perform a correction step to calculate a correction value, which is a correction value calculated by using a first output value calculated by inputting an input value into a first model, wherein the correction value is a value obtained by correcting a second output value calculated by inputting the input value into a pre-trained second model using the first output value.