Information processing device, information processing method, and information processing program
The information processing apparatus and method enhance LLM performance by dynamically selecting and executing actions, addressing the performance plateau issue in existing techniques, enabling continuous model improvement.
Patent Information
- Application Number
- JP2025022624
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2026-08-26
AI Technical Summary
Existing techniques for improving Large Language Models (LLMs) face a performance plateau where further enhancements are not achieved through repeated model improvement measures.
An information processing apparatus and method that includes an acquisition unit for models, a selection unit for candidate actions, and a generation unit to generate improved models by executing selected actions, utilizing a trained language model to dynamically select and execute actions for model enhancement.
This approach enables further improvement in LLM performance by dynamically selecting and executing actions, preventing performance plateauing during autonomous learning and enhancing model performance through actions like fine-tuning and merging.
Smart Images

Figure 2026136843000001_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to an information processing apparatus, an information processing method, and an information processing program.
Background Art
[0002] Techniques for the autonomous improvement of models such as Large Language Models (LLMs) are known. As an example of the technique for the self-improvement of LLMs, for example, the technique described in Non-Patent Document 1 can be cited.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the technique described in Non-Patent Document 1, there is a problem that the performance of the model may not improve when the improvement of the model is repeated.
[0005] This disclosure has been made in view of the above problems, and an exemplary object thereof is to provide a technique capable of further improving the performance of a model in the autonomous learning of the model.
Means for Solving the Problems
[0006] An information processing apparatus according to an exemplary aspect of the present disclosure includes an acquisition unit that acquires one or more models, a selection unit that selects any one of a plurality of candidate actions each of which defines a method for improving the model, and a generation unit that generates an improved model from the one or more models by executing the selected action.
[0007] An information processing method relating to an illustrative aspect of this disclosure includes: an acquisition process in which at least one processor acquires one or more models; a selection process in which the at least one processor selects any action from a plurality of candidate actions each defining a method for improving a model; and a generation process in which the at least one processor generates an improved model from the one or more models by executing the selected action.
[0008] An illustrative aspect of the present disclosure relates to an information processing program, which is a program that causes a computer to function as an information processing device, wherein the computer functions as an acquisition means for acquiring one or more models, a selection means for selecting one of a plurality of candidate actions, each of which defines a method for improving a model, and a generation means for generating an improved model from the one or more models by executing the selected action. [Effects of the Invention]
[0009] One exemplary aspect of this disclosure is that it provides a technique that can further improve the performance of a model during autonomous learning. [Brief explanation of the drawing]
[0010] [Figure 1] This is a block diagram showing the configuration of the information processing device related to this disclosure. [Figure 2] This is a flowchart showing the flow of the information processing method related to this disclosure. [Figure 3] This is a block diagram showing the configuration of the information processing device related to this disclosure. [Figure 4] This figure shows an example of the functional configuration of the information processing device related to this disclosure. [Figure 5] This is a flowchart showing the flow of the information processing method related to this disclosure. [Figure 6] This is a block diagram showing the configuration of a computer that functions as an information processing device related to this disclosure. [Modes for carrying out the invention]
[0011] The following are examples of embodiments of the present invention. However, the present invention is not limited to the exemplary embodiments shown below, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining some or all of the technologies (things or methods) employed in each of the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, embodiments obtained by appropriately omitting some of the technologies employed in each of the exemplary embodiments shown below may also be included in the scope of the present invention. In addition, the effects mentioned in each of the exemplary embodiments shown below are examples of effects that can be expected in that exemplary embodiment and do not define the scope of the present invention. That is, embodiments that do not produce the effects mentioned in each of the exemplary embodiments shown below may also be included in the scope of the present invention.
[0012] [First Exemplary Embodiment] A first exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. This exemplary embodiment is the basic form for each of the exemplary embodiments described later. The scope of application of each technology adopted in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technology adopted in this exemplary embodiment can also be adopted in other exemplary embodiments included in this disclosure, to the extent that no particular technical problems occur. Furthermore, each technology shown in the drawings referenced to explain this exemplary embodiment can also be adopted in other exemplary embodiments included in this disclosure, to the extent that no particular technical problems occur.
[0013] (Configuration of information processing device) The configuration of the information processing device 1 will be explained with reference to Figure 1. Figure 1 is a block diagram showing the configuration of the information processing device 1. As shown in Figure 1, the information processing device 1 comprises an acquisition unit 11, a selection unit 12, and a generation unit 13. The acquisition unit 11 acquires one or more models. The selection unit 12 selects one of several candidate actions, each defining a method for improving the model. The generation unit 13 generates an improved model from the one or more models by executing the action selected by the selection unit 12.
[0014] (Effects of information processing equipment) As described above, the information processing device 1 employs a configuration comprising an acquisition unit 11 that acquires one or more models, a selection unit 12 that selects one of several candidate actions, each defining a method for improving the model, and a generation unit 13 that generates an improved model from the one or more models by executing the selected action. Therefore, the information processing device 1 has the effect of being able to further improve the performance of the model during autonomous learning.
[0015] (Information processing flow) The flow of the information processing method S1 will be explained with reference to Figure 2. Figure 2 is a flowchart showing the flow of the information processing method S1. As shown in Figure 2, the information processing method S1 includes an acquisition process S11, a selection process S12, and a generation process S13. In the acquisition process S11, at least one processor acquires one or more models. In the selection process S12, the at least one processor selects one of several candidate actions, each defining a method for improving the model. In the generation process S13, the at least one processor generates an improved model from the one or more models by executing the selected action.
[0016] (Effects of information processing methods) As described above, the information processing method S1 includes: an acquisition process S11 in which at least one processor acquires one or more models; a selection process S12 in which the at least one processor selects any one action from a plurality of candidate actions that each define a model improvement method; and a generation process S13 in which the at least one processor generates an improved model from the one or more models by executing the selected action. Therefore, according to the information processing method S1, an effect can be obtained that the performance of the model can be further improved in the self-learning of the model.
[0017] 〔Second Exemplary Embodiment〕 A second exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same functions as those described in the above exemplary embodiments are denoted by the same reference numerals, and the description thereof will be omitted as appropriate. Note that the scope of application of each technology adopted in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technology adopted in this exemplary embodiment can be adopted in other exemplary embodiments included in the present disclosure as long as there are no particular technical obstacles. In addition, each technology shown in each drawing referred to for explaining this exemplary embodiment can be adopted in other exemplary embodiments included in the present disclosure as long as there are no particular technical obstacles.
[0018] (Overview of Information Processing Apparatus) The information processing apparatus 1A is an apparatus that realizes self-learning of a model. Examples of the model include, but are not limited to, a neural network, a decision tree model, etc. As an example, the model may be a language model such as an LLM.
[0019] The information processing device 1A selects one of several candidate actions and executes the selected action to generate an improved model from one or more models. An action is a measure to improve a model. An action includes, for example, an action type and input data. The action type is information indicating the type of action, and for example, it indicates the learning method (fine-tuning, etc.), the type of loss function, the model merging method, etc. Examples of action types include "Supervised Fine-Tuning" and "TIES Merging".
[0020] Input data is the input to an action and varies depending on the action type. For example, if the action type is "Supervised Fine-Tuning," the input may include the base model, training data, and hyperparameters. Similarly, if the action type is "TIES Merging," the input may include the first model, the second model, and hyperparameters.
[0021] (Configuration of information processing device) The configuration of the information processing device 1A will be described with reference to Figure 3. Figure 3 is a block diagram showing the configuration of the information processing device 1A. The information processing device 1A comprises a control unit 10A, a storage unit 20A, a communication unit 30A, an input unit 40A, and an output unit 50A.
[0022] (Communications Department) The communication unit 30A communicates with external devices of the information processing device 1A via a communication line N. The specific configuration of the communication line N is not limited to this exemplary embodiment, but examples of communication line N include a wireless LAN (Local Area Network), a wired LAN, a WAN (Wide Area Network), a public telephone network, a mobile data communication network, or a combination thereof. The communication unit 30A transmits data supplied from the control unit 10A to other devices and supplies data received from other devices to the control unit 10A.
[0023] (Input section) The input unit 40A is configured to receive input to the information processing device 1A, and may include, for example, an input device such as a keyboard, mouse, touch panel, camera, or microphone. Alternatively, the input unit 40A may be configured to receive data from the input device via an interface such as USB (Universal Serial Bus).
[0024] (Output section) The output unit 50A is configured to output from the information processing device 1A, and may include, for example, an output device such as a display, printer, touch panel, or speaker. The output unit 50A may also be configured to have an interface such as USB, and to output data to the output device via this interface.
[0025] (Storage part) The memory unit 20A stores various types of information referenced by the control unit 10A. The memory unit 20A specifically includes an initial model memory unit 201, an initial object memory unit 202, and an initial policy memory unit 203. The initial model memory unit 201 stores one or more models. One example of one or more models is a language model. The initial object memory unit 202 stores objects. Objects are data used to execute actions described later, and include, for example, training data, hyperparameters, etc.
[0026] The initial policy memory unit 203 stores the policy. The policy is referenced when the selection unit 15A, described later, selects an action from candidate actions. The policy takes multiple candidate actions as input and outputs a single action. As an example, the policy may include a trained language model.
[0027] If the policy includes a language model, for example, an action might be expressed as text such as "TIES Merging using {Model M1} and {Model M2}", and the prompt to be input to the language model might be text such as "Select the best action from {Action A1}, {Action A2}, ... {Action Ak}". In other words, if the policy includes a language model, each of the multiple candidate actions is expressed as text, and the selection unit 15A, described later, inputs a prompt to the language model indicating that one of the multiple candidate actions should be selected, thereby selecting an action.
[0028] Furthermore, the policy may include a hypothesis. Examples of hypotheses include statements such as, "Supervised Fine-Tuning is effective," or "The order of Supervised Fine-Tuning followed by TIES Merging is effective."
[0029] Furthermore, the policy is not limited to a language model. The policy may, for example, be to select an action according to a predetermined rule. More specifically, the policy may be to select multiple candidate actions in a predetermined order, or to randomly select an action from multiple candidate actions. Alternatively, the policy may be to select an action by referring to the evaluation value for each action included in the evaluation result of the evaluation unit 17A, which will be described later.
[0030] (Control Unit) Figure 4 shows an example of the functional configuration of the control unit 10A. The control unit 10A comprises a model holding unit 11A, an object holding unit 12A, a policy holding unit 13A, a candidate action enumeration unit 14A, a selection unit 15A, a generation unit 16A, an evaluation unit 17A, and an update unit 18A. The model holding unit 11A, the selection unit 15A, the generation unit 16A, the evaluation unit 17A, and the update unit 18A are examples of acquisition means, selection means, generation means, evaluation means, and update means according to this disclosure, respectively.
[0031] (Model holding unit, object holding unit, policy holding unit) The model holding unit 11A retrieves one or more models stored in the initial model storage unit 201. One or more models include, for example, a language model. Here, "retrieving a model" includes obtaining information that identifies the model. It is not necessary to obtain all of the multiple parameters that define the model; it is sufficient to identify the model that is the target of improvement. The object holding unit 12A retrieves objects stored in the initial object storage unit 202. The policy holding unit 13A retrieves policies stored in the initial policy storage unit 203.
[0032] (Candidate Action Enumeration Section) The candidate action enumeration unit 14A enumerates candidate actions from combinations of models held by the model holding unit 11A and objects held by the object holding unit 12A. Candidate actions are measures that define how to improve a model. As an example, the candidate action enumeration unit 14A refers to a table that describes the action type and the types of objects and models required for that action type, and enumerates all possible combinations of objects and models from the models held by the model holding unit 11A and the objects held by the object holding unit 12A to create candidate actions. However, the method of enumerating candidate actions is not limited to the example described above.
[0033] (Selection section) The selection unit 15A selects one of the multiple candidate actions listed by the candidate action enumeration unit 14A.
[0034] (Example 1 of action selection process) As an example, the selection unit 15A selects one of the above-mentioned candidate actions by referring to the policy or the information obtained by the policy. If the policy includes a trained language model, the selection unit 15A selects one of the above-mentioned candidate actions by referring to the output of the language model as an example.
[0035] Furthermore, if the strategy includes a hypothesis, the selection unit 15A may, as an example, select an action based on that hypothesis. The hypothesis may also be obtained using a language model.
[0036] (Example 2 of action selection process) Furthermore, the selection unit 15A may, for example, select one of several candidate actions according to predetermined rules. More specifically, the selection unit 15A may, for example, select several candidate actions in a predetermined order, or it may randomly select one of several candidate actions.
[0037] Alternatively, the strategy may involve selecting an action by referring to an evaluation value for each action. In this case, the selection unit 15A selects an action by referring to an evaluation value for each action calculated by the evaluation unit 17A, which will be described later. Here, the evaluation value may be, for example, a value included in the evaluation result calculated by the evaluation unit 17A, which will be described later, or it may be the cumulative value of the evaluation values for each action calculated by the evaluation unit 17A. The selection unit 15A may, for example, select the action with the highest evaluation value from among multiple candidate actions.
[0038] (Generation part) The generation unit 16A generates an improved model from the one or more models by executing the action selected by the selection unit 15A. The generation unit 16A adds the improved model and improved objects to the model holding unit 11A and the object holding unit 12A, respectively. The actions executed by the generation unit 16A include, as an example, the process of merging the one or more models.
[0039] Furthermore, the actions performed by the generation unit 16A include, as an example, data generation and fine tuning. In other words, the actions performed by the generation unit 16A include, as an example, the process of generating data using at least one of the one or more models by inputting a prompt to at least one of the aforementioned models, and the process of updating at least one of the one or more models by referring to the generated data. Here, the process of updating the model is the process of training the model.
[0040] (Evaluation Department) The evaluation unit 17A evaluates the improved model generated by the generation unit 16A. For example, the evaluation unit 17A evaluates the performance of the improved model using a pre-prepared benchmark. The evaluation results from the evaluation unit 17A include, for example, a reward corresponding to at least one of the following: an indicator of the accuracy of the improved model, and the computational cost performed by the improved model. The reward is an indicator of the quality of the action. For example, the reward may be a value indicating the accuracy of the improved model, or it may be a monotonically increasing function of accuracy and a monotonically decreasing function of computational cost. Specifically, for example, the reward may be the value obtained by dividing accuracy by computational cost.
[0041] (Update section) The update unit 18A updates the above-mentioned policy or the information referenced by the policy by referring to the evaluation results from the evaluation unit 17A. The information referenced by the policy includes, as an example, prompts, which are inputs to the language model included in the policy. Hereinafter, the policy updated by the update unit 18A will also be referred to as the "policy model".
[0042] The update unit 18A updates the policy or the information (memory) referenced by the policy using reinforcement learning with an indicator of the quality of the action as a reward. In other words, the update unit 18A can also update the policy or the information referenced by the policy by referring to the reward included in the evaluation result by the evaluation unit 17A. Any reinforcement learning method can be applied. The update unit 18A may update the policy by directly updating the parameters of the policy model, or it may update the information referenced by the policy rather than the policy itself. Examples of methods for directly updating the parameters of the policy model include PPO (Proximal Policy Optimization), DQN (Deep Q-Networks), and REINFORCE. An example of a method for updating the information referenced by the policy is the Reflexion method, which updates memory.
[0043] (Information processing flow) Figure 5 is a flowchart showing the flow of the information processing method S1A. In step S101, the model holding unit 11A, the object holding unit 12A, and the policy holding unit 13A each acquire the initial model, initial object, and initial policy, respectively. In step S102, the candidate action enumeration unit 14A enumerates multiple candidate actions using the model held by the model holding unit 11A and the object held by the object holding unit 12A.
[0044] In step S103, the selection unit 15A selects one of the multiple candidate actions enumerated in step S102. In step S104, the generation unit 16A generates the improved model by executing the action selected by the selection unit 15A. In step S105, the generation unit 16A adds the improved model and the improved object to the model holding unit 11A and the object holding unit 12A, respectively.
[0045] In step S106, the evaluation unit 17A evaluates the improved model. In step S107, the update unit 18A determines whether to terminate the update. For example, the update unit 18A determines to terminate the update if the number of loops exceeds a certain number, or if a desired evaluation result is obtained. The desired evaluation result may be, for example, that an index indicating the accuracy of the improved model meets a predetermined condition, or, for example, that the computation cost performed by the improved model meets a predetermined condition. If the update is not terminated (NO in step S107), the update unit 18A proceeds to the process in step S108. On the other hand, if the update is terminated (YES in step S107), the update unit 18A terminates the process.
[0046] In step S108, the update unit 18A updates the policy. For example, the update unit 18A updates the policy using reinforcement learning. In this case, for example, the update unit 18A uses the accuracy of the model obtained as a result of the action as a reward to learn the policy so that it outputs actions with high accuracy and avoids outputting actions with low accuracy.
[0047] If the policy consists of memory (which contains past action-reward pairs) and a language model used to create prompts for selecting actions, the update unit 18A may update the policy by updating the prompts, by updating both the language model and the prompts, or by updating the language model. Furthermore, if a hypothesis is included in memory, the update unit 18A may verify the hypothesis based on the evaluation results of the evaluation unit 17A and update the hypothesis based on the verification results.
[0048] The process from steps S102 to S108 is repeated, enabling the model to learn autonomously.
[0049] (Examples of application) Information processing device 1A can be applied to improve models used in various industries and sectors. For example, information processing device 1A can be used in the medical field to support decision-making by medical professionals such as doctors regarding medical treatment. In this case, the model to be improved may be, for example, a generative model that takes patient information as input and outputs the results of differential diagnosis of the patient's condition. Another example is a generative model that takes text representing the content of a meeting as input and creates a meeting summary and minutes.
[0050] (Effects of information processing equipment) Incidentally, in the development of models such as LLMs, research into autonomous improvement of LLMs is attracting attention in order to reduce human resource costs. However, in the technology described in Non-Patent Document 1, for example, fixed model improvement measures are iteratively repeated, so while performance improves up to a certain point, it may stop improving beyond that point.
[0051] In contrast, the information processing device 1A employs a configuration in which the selection unit 15A selects one of the above-mentioned candidate actions by referring to the policy or the information obtained by the policy. Therefore, with the information processing device 1A, instead of executing actions uniformly, the action selected by referring to the policy or the information obtained by the policy is executed, thereby preventing the model's performance from plateauing during autonomous learning.
[0052] Furthermore, in the information processing device 1A, the above strategy includes a trained language model, and the selection unit 15A refers to the output of the language model and selects one of the multiple candidate actions. Therefore, the information processing device 1A enables autonomous learning of the model using a trained language model.
[0053] Furthermore, in the information processing device 1A, the actions performed by the generation unit 16A include a process of generating data using at least one of the one or more models by inputting a prompt to that model, and a process of updating at least one of the one or more models by referring to the generated data. Therefore, with the information processing device 1A, the performance of the model can be further improved in autonomous learning by selecting one of several candidate actions, including fine tuning, to generate an improved model.
[0054] Furthermore, the information processing device 1A employs a configuration in which the actions executed by the generation unit 16A include the process of merging one or more of the above-mentioned models. Therefore, with the information processing device 1A, by selecting one of several candidate actions, including the process of merging models, to generate an improved model, the performance of the model can be further improved during autonomous learning of the model.
[0055] Furthermore, the information processing device 1A employs a configuration in which one or more of the above models include a language model. Therefore, the information processing device 1A can further improve the performance of the language model during autonomous learning.
[0056] Furthermore, the information processing device 1A employs a configuration that includes an evaluation unit 17A that evaluates the improved model generated by the generation unit 16A, and an update unit 18A that updates the policy or the information referenced by the policy by referring to the evaluation results of the evaluation unit 17A. Therefore, with the information processing device 1A, the selected action can be dynamically changed by selecting an action based on the policy updated based on the evaluation results of the evaluation unit 17A, thereby further improving the performance of the model during autonomous learning.
[0057] Furthermore, in the information processing device 1A, the evaluation result from the evaluation unit 17A includes a reward corresponding to at least one of (i) an index indicating the accuracy of the improved model, and (ii) the cost of the calculations performed by the improved model. The update unit 18A updates the policy or the information referenced by the policy by referring to the reward. Therefore, the information processing device 1A can further improve the performance of the model during autonomous learning by updating the policy or the information referenced by the policy by referring to the reward included in the evaluation result.
[0058] [Examples of implementation using software] Some or all of the functions of the information processing devices 1 and 1A (hereinafter also referred to as "the above devices") may be implemented by hardware such as integrated circuits (IC chips) or by software.
[0059] In the latter case, each of the above devices is implemented, for example, by a computer that executes instructions for a program, which is software that realizes each function. An example of such a computer (hereinafter referred to as Computer C) is shown in Figure 6. Figure 6 is a block diagram showing the hardware configuration of Computer C, which functions as each of the above devices.
[0060] Computer C comprises at least one processor C1 and at least one memory C2. Memory C2 stores a program P that causes computer C to operate as each of the above-mentioned devices. In computer C, processor C1 reads program P from memory C2 and executes it, thereby realizing each of the above-mentioned devices.
[0061] For processor C1, for example, a CPU (Central Processing Unit), GPU (Graphic Processing Unit), DSP (Digital Signal Processor), MPU (Micro Processing Unit), FPU (Floating Point Number Processing Unit), PPU (Physics Processing Unit), TPU (Tensor Processing Unit), quantum processor, microcontroller, or a combination thereof can be used. For memory C2, for example, flash memory, HDD (Hard Disk Drive), SSD (Solid State Drive), or a combination thereof can be used.
[0062] Computer C may also be equipped with RAM (Random Access Memory) for loading program P at runtime and for temporarily storing various data. Furthermore, computer C may be equipped with communication interfaces for sending and receiving data with other devices. Additionally, computer C may be equipped with input / output interfaces for connecting input / output devices such as keyboards, mice, displays, and printers.
[0063] Furthermore, program P can be recorded on a non-temporary, tangible recording medium M that is readable by computer C. Such a recording medium M could be, for example, tape, disk, card, semiconductor memory, or programmable logic circuitry. Computer C can acquire program P via such a recording medium M. Program P can also be transmitted via a transmission medium. Such a transmission medium could be, for example, a communication network or broadcast waves. Computer C can also acquire program P via such a transmission medium.
[0064] Furthermore, each of the above functions of each of the above devices may be implemented by a single processor in a single computer, by multiple processors in a single computer working together, or by multiple processors in each of multiple computers working together. In addition, the programs for implementing each of the above functions in each of the above devices may be stored in a single memory in a single computer, distributed and stored in multiple memories in a single computer, or distributed and stored in multiple memories in each of multiple computers.
[0065] [Additional Note A] This disclosure includes the technologies described in the following appendices. However, the present invention is not limited to the technologies described in the following appendices, and various modifications are possible within the scope of the claims.
[0066] (Note A1) Acquisition means for acquiring one or more models, A selection method in which each person chooses one of several candidate actions that define how to improve the model, A generation means that generates an improved model from the one or more models by performing a selected action, An information processing device equipped with the following features.
[0067] (Appendix A2) The selection means selects one of the multiple candidate actions by referring to the policy or the information obtained by said policy. The information processing device described in Appendix A1.
[0068] (Note A3) The aforementioned policy includes a trained language model, The selection means selects one of the multiple candidate actions by referring to the output of the language model. The information processing device described in Appendix A2.
[0069] (Note A4) The actions performed by the generation means include: A process of generating data using at least one of the aforementioned models by inputting a prompt to at least one of the aforementioned one or more models, A process of updating at least one of the one or more models by referring to the generated data, Includes, An information processing device as described in any one of the appendices A1 to A3.
[0070] (Note A5) The actions performed by the generation means include: This process includes merging one or more of the aforementioned models. An information processing device as described in any one of the appendices A1 to A4.
[0071] (Note A6) The aforementioned one or more models include a language model. An information processing device as described in any one of the appendices A1 to A5.
[0072] (Note A7) An evaluation means for evaluating the improved model generated by the generation means, The system includes an update means that updates the strategy or the information referenced by the strategy by referring to the evaluation results obtained by the evaluation means. The information processing device described in Appendix A2 or A3.
[0073] (Note A8) The evaluation results obtained by the aforementioned evaluation means include: An indicator showing the accuracy of the improved model, and The computational cost of the improved model The reward includes at least one of the following: The update means updates the policy or the information referenced by the policy by referring to the reward. The information processing device described in Appendix A7.
[0074] [Additional Notes B] This disclosure includes the technologies described in the following appendices. However, the present invention is not limited to the technologies described in the following appendices, and various modifications are possible within the scope of the claims.
[0075] (Note B1) At least one processor performs an acquisition process to acquire one or more models, The aforementioned at least one processor performs a selection process in which it selects one of several candidate actions that define how to improve the model, The at least one processor performs a generation process that generates an improved model from the one or more models by executing a selected action, An information processing method that includes this.
[0076] (Note B2) In the selection process, the at least one processor selects one of the multiple candidate actions by referring to a policy or information obtained by the policy. The information processing method described in Appendix B1.
[0077] (Note B3) The aforementioned policy includes a trained language model, In the selection process, the at least one processor selects one of the multiple candidate actions by referring to the output of the language model. The information processing method described in Appendix B2.
[0078] (Note B4) The actions performed by the generation process include: The process of generating data using at least one of the models by prompting at least one of the one or more models, The process of updating at least one of the one or more models by referring to the generated data of the at least one processor, Includes, The information processing method described in any one of the appendices B1 to B3.
[0079] (Note B5) The actions performed by the generation process include: The at least one processor performs the process of merging the one or more models. Includes, The information processing method described in any one of the appendices B1 to B4.
[0080] (Note B6) The aforementioned one or more models include a language model. The information processing method described in any one of the appendices B1 through B5.
[0081] (Note B7) The at least one processor performs an evaluation process to evaluate the improved model generated by the generation process, The at least one processor performs an update process that updates the policy or the information referenced by the policy by referring to the evaluation result of the evaluation process, Includes, The information processing method described in Appendix B2 or B3.
[0082] (Note B8) The evaluation results obtained through the aforementioned evaluation process include: An indicator showing the accuracy of the improved model, and The computational cost of the improved model The reward includes at least one of the following: In the update process, the at least one processor updates the policy or the information referenced by the policy by referring to the reward. The information processing method described in Appendix B7.
[0083] [Additional Note C] This disclosure includes the technologies described in the following appendices. However, the present invention is not limited to the technologies described in the following appendices, and various modifications are possible within the scope of the claims.
[0084] (Note C1) A program that makes a computer function as an information processing device. The aforementioned computer, Acquisition means for acquiring one or more models, A selection method in which each person chooses one of several candidate actions that define how to improve the model, A generation means that generates an improved model from the one or more models by performing a selected action, An information processing program that functions as such.
[0085] (Note C2) The selection means selects one of the multiple candidate actions by referring to the policy or the information obtained by said policy. The information processing program described in Appendix C1.
[0086] (Note C3) The aforementioned policy includes a trained language model, The selection means selects one of the multiple candidate actions by referring to the output of the language model. The information processing program described in Appendix C2.
[0087] (Note C4) The actions performed by the generation means include: A process of generating data using at least one of the aforementioned models by inputting a prompt to at least one of the aforementioned one or more models, A process of updating at least one of the one or more models by referring to the generated data, Includes, An information processing program described in any one of the appendices C1 to C3.
[0088] (Note C5) The actions performed by the generation means include: A process for merging one or more of the aforementioned models, Includes, An information processing program described in any one of the appendices C1 to C4.
[0089] (Appendix C6) The aforementioned one or more models include a language model. An information processing program described in any one of the appendices C1 to C5.
[0090] (Note C7) The aforementioned computer, An evaluation means for evaluating the improved model generated by the generation means, An update process that updates the strategy or the information referenced by the strategy, based on the evaluation results obtained by the evaluation means, To make it function as The information processing program described in Appendix C2 or C3.
[0091] (Note C8) The evaluation results obtained by the aforementioned evaluation means include: An indicator showing the accuracy of the improved model, and The computational cost performed by the improved model, The reward includes at least one of the following: The update means updates the policy or the information referenced by the policy by referring to the reward. The information processing program described in Appendix C7.
[0092] [Additional Note D] This disclosure includes the technologies described in the following appendices. However, the present invention is not limited to the technologies described in the following appendices, and various modifications are possible within the scope of the claims.
[0093] (Note D1) It comprises at least one processor, and the at least one processor is A retrieval process to obtain one or more models, A selection process in which each person chooses one of several candidate actions that define how to improve the model, A generation process that generates an improved model from the one or more models by executing the selected action, An information processing device that performs the following actions.
[0094] The information processing device may also include memory. Furthermore, the memory may store a program that causes at least one processor to execute each of the aforementioned processes.
[0095] (Note D2) In the selection process, the at least one processor selects one of the multiple candidate actions by referring to a policy or information obtained by the policy. The information processing device described in Appendix D1.
[0096] (Note D3) The aforementioned policy includes a trained language model, In the selection process, the at least one processor selects one of the multiple candidate actions by referring to the output of the language model. The information processing device described in Appendix D2.
[0097] (Note D4) The actions performed by the generation process include: A process of generating data using at least one of the aforementioned models by inputting a prompt to at least one of the aforementioned one or more models, A process of updating at least one of the one or more models by referring to the generated data, Includes, An information processing device as described in any one of the appendices D1 to D3.
[0098] (Note D5) The actions performed by the generation process include: A process for merging one or more of the aforementioned models, Includes, An information processing device as described in any one of the appendices D1 to D4.
[0099] (Note D6) The aforementioned one or more models include a language model. An information processing device as described in any one of the appendices D1 to D5.
[0100] (Note D7) The aforementioned at least one processor is The generation process includes an evaluation process for evaluating the improved model generated by the generation process, An update process that updates the policy or the information referenced by the policy, based on the evaluation results obtained by the evaluation process, Execute The information processing device described in Appendix D2 or D3.
[0101] (Note D8) The evaluation results obtained through the aforementioned evaluation process include: An indicator showing the accuracy of the improved model, and The computational cost performed by the improved model, The reward includes at least one of the following: In the update process, the at least one processor updates the policy or the information referenced by the policy by referring to the reward. The information processing device described in Appendix D7.
[0102] [Additional Note E] This disclosure includes the technologies described in the following appendices. However, the present invention is not limited to the technologies described in the following appendices, and various modifications are possible within the scope of the claims.
[0103] (Note E1) A program for causing a computer to function as an information processing device, wherein the computer, A retrieval process to obtain one or more models, A selection process in which each person chooses one of several candidate actions that define how to improve the model, A generation process that generates an improved model from the one or more models by executing the selected action, A non-temporary recording medium that stores an information processing program for executing [the specified action]. [Explanation of Symbols]
[0104] 1. 1A Information Processing Device 11 Acquisition Department 11A Model holding section 12, 15A Selection Section 12A Object holding section 13, 16A generation section 13A Steering mechanism holding section 14A Candidate Action Enumeration Section 17A Evaluation Department 18A Update Department
Claims
1. Acquisition means for acquiring one or more models, A selection method in which each person chooses one of several candidate actions that define how to improve the model, A generation means that generates an improved model from the one or more models by performing the selected action. An information processing device equipped with the following features.
2. The selection means selects one of the multiple candidate actions by referring to the policy or the information obtained by the policy. The information processing apparatus according to claim 1.
3. The aforementioned policy includes a trained language model, The selection means selects one of the multiple candidate actions by referring to the output of the language model. The information processing apparatus according to claim 2.
4. The actions performed by the generation means include: A process of generating data using at least one of the aforementioned models by inputting a prompt to at least one of the one or more models, A process to update at least one of the one or more models by referring to the generated data. Includes An information processing apparatus according to any one of claims 1 to 3.
5. The actions performed by the generation means include: Process to merge one or more of the aforementioned models Includes An information processing apparatus according to any one of claims 1 to 3.
6. The one or more models mentioned above include a language model. An information processing apparatus according to any one of claims 1 to 3.
7. An evaluation means for evaluating the improved model generated by the generation means, An update means that updates the strategy or the information referenced by the strategy by referring to the evaluation results of the evaluation means. It is equipped with The information processing apparatus according to claim 2 or 3.
8. The evaluation results obtained by the aforementioned evaluation means include: An indicator showing the accuracy of the improved model, and The computational cost of the improved model The reward includes at least one of the following: The update means updates the policy or the information referenced by the policy by referring to the reward. The information processing apparatus according to claim 7.
9. At least one processor performs an acquisition process to acquire one or more models, The aforementioned at least one processor performs a selection process in which it selects one of several candidate actions that define how to improve the model, The at least one processor performs a generation process that generates an improved model from the one or more models by executing a selected action. An information processing method that includes this.
10. An information processing program for causing a computer to function as an information processing device, wherein the computer, Acquisition means for acquiring one or more models, A selection method in which each person chooses one of several candidate actions that define how to improve the model, A generation means that generates an improved model from the one or more models by performing the selected action. An information processing program designed to function as such.