Learning device, learning method, and learning program

JP7927449B2Active Publication Date: 2026-10-01MITSUBISHI HEAVY IND LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022076197
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-05-02
Publication Date
2026-10-01
Estimated Expiration
2042-05-02

AI Technical Summary

Benefits of technology

【0010】 本開示によれば、多様な対戦相手であっても、ハイパーパラメータを含む学習モデルの学習を、適切に効率よく実行することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007927449000001
    Figure 0007927449000001
  • Figure 0007927449000002
    Figure 0007927449000002
  • Figure 0007927449000003
    Figure 0007927449000003
Patent Text Reader

Abstract

To appropriately and efficiently execute training of learning models including hyper parameters even if various opponents are set against an agent.SOLUTION: A learning device comprises a processing unit for causing reinforcement learning of agents' learning models to be performed in a competitive environment where the agents compete against each other. The learning model includes hyper parameters, and the processing unit executes a step of evaluating the strength of the plurality of agents to be opponents of the agent to be learned, a step of setting the competition probability according to the strength of the agents to be the opponents to the agent to be learned, a step of setting the agents to be the opponents based on the competition probability, and a step of causing the agents to be the opponents after the setting to compete against each other to execute reinforcement learning of the agent to be learned.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a learning device, a learning method, and a learning program. Background Art

[0002] Conventionally, it is known to perform reinforcement learning for agents based on battle results between agents (see, for example, Patent Document 1). Prior Art Literature Patent Literature

[0003] Patent Document 1 Japanese Unexamined Patent Publication No. 2019-197592 Summary of Invention Problem to be Solved by Invention

[0004] Here, when the opponent of the agent 5 to be learned is fixed, as a result of performing reinforcement learning, there is a possibility that a learning model that becomes a policy specialized for that opponent is generated. For this reason, it is conceivable to obtain a versatile learning model by setting a plurality of opponents with different policies, replacing each opponent according to a certain learning step (also referred to as a swap step), and causing learning through battles.

[0005] However, when opponents are replaced according to a fixed learning step, there is a high possibility that the learning model will learn to acquire a large amount of reward in battles against specific opponents that are easy to win against. For this reason, there is a problem that it is difficult to obtain a general-purpose model that can acquire large amounts of reward against a variety of opponents.

[0006] Accordingly, an object of the present disclosure is to provide a learning device, a learning method, and a learning program capable of appropriately and efficiently executing learning of a learning model including hyperparameters even against a variety of opponents. Means for Solving the Problem

[0007] The learning device of this disclosure is a learning device comprising a processing unit for performing reinforcement learning on a learning model of an agent in a competitive environment in which agents compete against each other, wherein the learning model includes hyperparameters, and the processing unit performs the steps of: evaluating the strength of a plurality of agents that will be opponents of the agent to be learned; setting a match probability for the agent to be learned according to the strength of the opponent agents; setting the opponent agents based on the match probability; and performing reinforcement learning on the agent to be learned by having the set opponent agents compete against each other.

[0008] The learning method of this disclosure is a learning method for performing reinforcement learning on a learning device in a competitive environment in which agents compete against each other, wherein the learning model includes hyperparameters, and the learning device is instructed to perform the following steps: evaluate the strength of a plurality of agents that will be opponents of the agent to be learned; set a match probability for the agent to be learned according to the strength of the opponent agents; set the opponent agents based on the match probability; and perform reinforcement learning on the agent to be learned by having the set opponent agents compete against each other.

[0009] The learning program of this disclosure is a learning program for performing reinforcement learning on a learning device in a competitive environment in which agents compete against each other, wherein the learning model includes hyperparameters, and the learning device is caused to perform the following steps: evaluate the strength of a plurality of agents that will be opponents of the agent to be learned; set a probability of a match for the agent to be learned according to the strength of the opponent agents; set the opponent agents based on the probability of a match; and perform reinforcement learning on the agent to be learned by having the set opponent agents compete against each other. [Effects of the Invention]

[0010] According to this disclosure, it is possible to appropriately and efficiently train a learning model, including hyperparameters, even against a variety of opponents. [Brief explanation of the drawing]

[0011] [Figure 1] Figure 1 is an explanatory diagram of learning using the learning model according to this embodiment. [Figure 2] Figure 2 is a schematic representation of the learning device according to this embodiment. [Figure 3] Figure 3 is a diagram showing the flow of the learning method according to this embodiment. [Figure 4] Figure 4 is an explanatory diagram of the learning method according to this embodiment. [Modes for carrying out the invention]

[0012] Embodiments of the present invention will be described in detail below with reference to the drawings. However, the present invention is not limited by these embodiments. Furthermore, some of the components in the embodiments described below are easily substituted or substantially identical to those that are easily substituted by those skilled in the art. Moreover, the components described below can be combined as appropriate, and if there are multiple embodiments, each embodiment can be combined.

[0013] [Embodiment] The learning device 10 and learning method according to this embodiment are a device and method for training a learning model that includes hyperparameters. Figure 1 is an explanatory diagram of learning using the learning model according to this embodiment. Figure 2 is a schematic representation of the learning device according to this embodiment. Figure 3 is a diagram showing the flow of the learning method according to this embodiment. Figure 4 is an explanatory diagram of the learning method according to this embodiment.

[0014] (Learning using a learning model) First, referring to Figure 1, we will explain the learning process using the learning model M. The learning model M is installed on agent 5, which performs an action At. Agent 5 can be any machine capable of performing actions, such as a robot, vehicle, ship, or aircraft. Agent 5 uses the learning model M to perform a predetermined action At under a predetermined environment 6.

[0015] As shown in Figure 1, the learning model M is a neural network having multiple nodes. The neural network is a network formed by connecting multiple nodes, and has multiple layers, with multiple nodes in each layer. Parameters of the neural network include weights and biases between nodes. In addition, parameters of the neural network include hyperparameters such as the number of layers, the number of nodes, and the learning rate. In this embodiment, the learning of the learning model M is performed, including learning of the learning model M and its hyperparameters.

[0016] Next, learning using the learning model M will be described. Examples of the learning include reinforcement learning. Reinforcement learning is a type of unsupervised learning in which the weights and biases between nodes in the learning model M are learned such that the reward Rt given to the agent 5 under a predetermined environment 6 is maximized.

[0017] In reinforcement learning, the agent 5 acquires a state St from the environment 6 (the environment unit 22 described later), and acquires a reward Rt from the environment 6. Then, the agent 5 selects an action At from the learning model M based on the acquired state St and reward Rt. When the action At selected by the agent 5 is executed, the state St of the agent 5 in the environment 6 transitions to a state St+1. Further, a reward Rt+1 based on the executed action At, the state St before transition and the state St+1 after transition is given to the agent 5. Then, in reinforcement learning, the above learning is repeated for a predetermined number of evaluable steps such that the reward Rt given to the agent 5 is maximized.

[0018] (Learning Device) Next, the learning device 10 will be described with reference to FIG. 2. The learning device 10 executes reinforcement learning of the agent 5 under an environment that is a virtual space. The learning device 10 includes an environment unit 22 and a learning unit 23. The environment unit 22 and the learning unit 23 function as a processing unit for causing the learning model of the agent to learn, and a storage unit that stores various data used in learning. The hardware configuration of the learning device 10 is not particularly limited. In the present embodiment, in FIG. 2, a block indicated by a rectangular shape functions as the processing unit, and a block indicated by a cylindrical shape functions as the storage unit.

[0019] The environment unit 22 provides the agents 5 with a battle environment in which the agents 5 compete against each other. Specifically, the environment unit 22 gives a reward Rt to the agent 5, and derives the state St of the agent 5 that transitions according to an action At. Various models such as a motion model Ma, an environment model Mb, and a battle model Mc are stored in the environment unit 22. The battle model Mc serves as a learning model for the opponent agent 5 with respect to the agent 5 to be learned, and battle models Mc from opponent A to opponent N are prepared. The environment unit 22 uses the motion model Ma, the environment model Mb, and the battle model Mc to receive, as an input, the action At taken by the agent 5, and calculates the state St of the agent 5 as an output. The calculated state St is output to the learning unit 23. Additionally, a reward model Md for calculating a reward is stored in the environment unit 22. The reward model Md is a model that receives, as inputs, the action At taken by the agent 5, the state St, and the transition-destination state St+1, and calculates the reward Rt to be given to the agent 5 as an output. The calculated reward Rt is output to the learning unit 23.

[0020] The learning unit 23 executes learning of the learning model M. The learning unit 23 performs reinforcement learning as the learning. The learning unit 23 includes a model comparison unit 31 that compares the strength of the battle model Mc, a battle probability calculation unit 32 that calculates a battle probability according to the strength of the battle model Mc, and a reinforcement learning unit 33 that performs reinforcement learning. The learning unit 23 also includes a database 35 that stores a reinforcement learning model M (hereinafter, also simply referred to as learning model M) as a learning result obtained by reinforcement learning.

[0021] The model comparison unit 31, as an example, compares the strength of the learning model M of agent 5 to be trained with the strength of the opponent model Mc of agent 5, and evaluates whether the opponent model Mc is strong or weak against the learning model M based on the comparison results. Specifically, the model comparison unit 31 has the learning model M (after training has stopped) compete against the opponent model Mc, and uses the win rate or rating against the opponent model Mc as an indicator of strength. The model comparison unit 31 determines that the opponent model Mc is strong if its win rate or rating is high against the learning model M, and determines that the opponent model Mc is weak if its win rate or rating is low against the learning model M.

[0022] The model comparison unit 31 may also compare the strengths of the opposing model Mcs and evaluate whether an opposing model Mc is strong or not based on the comparison results. Specifically, the model comparison unit 31 uses a predetermined opposing model Mc as a reference, has the reference opposing model Mc compete against other opposing model Mcs, and uses the win rate or ELO rating against the opposing model Mc as an indicator of strength. The model comparison unit 31 determines that an opposing model Mc is strong if its win rate or rating is high compared to the reference opposing model Mc, while determining that an opposing model Mc is weak if its win rate or rating is low compared to the reference opposing model Mc.

[0023] Furthermore, the model comparison unit 31 may use KL divergence (KL distance) instead of win rate or rating. KL divergence is an indicator that shows the similarity of the probability distributions of state St and action At between models, and the models can be a learning model M and a battle model Mc, or two battle models Mc. A small KL divergence indicates high similarity, and a large KL divergence indicates low similarity. The model comparison unit 31 determines that the battle model Mc is strong if the KL divergence is above a preset threshold, and determines that the battle model Mc is weak if the KL divergence is below the threshold.

[0024] The match probability calculation unit 32 calculates and sets the match probability according to the strength of the match model Mc. Specifically, the match probability calculation unit 32 calculates the match probability so that the weaker the strength of the opponent agent 5 evaluated by the model comparison unit 31, the lower the match probability. In other words, the match probability calculation unit 32 calculates the match probability so that the stronger the strength of the opponent agent 5 evaluated by the model comparison unit 31, the higher the match probability. Here, the match probability is the proportion of times each of the multiple opponent agents 5 will play against the agent 5 to be learned, and it is calculated so that the sum of all match probabilities of the multiple match models Mc equals 100%. For example, as shown in Figure 4, three match models Mc are prepared as opponents, with a match probability of 10% for the weak match model Mc (opponent A), a match probability of 60% for the strong match model Mc (opponent B), and a match probability of 30% for the even match model Mc (opponent C).

[0025] The reinforcement learning unit 33 performs learning based on the reward Rt provided by the environment unit 22, and executes reinforcement learning of the learning model M. Specifically, the reinforcement learning unit 33 performs reinforcement learning of the learning model M for a predetermined number of learning steps T, updating various parameters to maximize the reward Rt provided to each agent 5. Here, the predetermined learning steps T include a swap step set for each opponent, a fixed step marking the end of opponent changes, and a maximum learning step marking the end of learning. In addition, by executing reinforcement learning of the learning model M, the reinforcement learning unit 33 obtains the reinforcement learning model M, which is the learning result of reinforcement learning, and stores the reinforcement learning model M obtained each time the weights and biases between nodes are updated in the database 35. If the initial values ​​of the weights and biases between nodes are 0 and the update values ​​are N, and the initial step of learning step T is 0 and the final step is S, then the database 35 will contain reinforcement learning models M0 to M N The learning models up to M0 are memorized, and each reinforcement learning model M0~M N In this case, from learning step T0 to learning step T SThe reinforcement learning model M up to this point is stored in memory.

[0026] (Learning Methods) Next, with reference to Figures 3 and 4, the learning method performed by the learning device 10 will be described. In the learning method, first, the learning device 10 performs the step of setting the parameter values ​​of the hyperparameters of the learning model M (step S1). In step S1, the parameter values ​​of the hyperparameters are set arbitrarily.

[0027] Next, in the learning method, the model comparison unit 31 of the learning device 10 performs a step of evaluating the strength of the opponent (step S2). Specifically, in step S2, the model comparison unit 31 calculates evaluation indicators of the strength of the opponent model Mc, such as the opponent's win rate, rating, or KL divergence.

[0028] Next, in the learning method, the battle probability calculation unit 32 of the learning device 10 performs the step of setting the battle probability according to the strength of the battle model Mc calculated in step S2 (step S3). In step S3, the battle probability calculation unit 32 calculates the battle probability to be set for each battle model Mc from the evaluation index of the strength of the battle model Mc calculated in step S2.

[0029] After step S3 is executed, in the learning method, the learning device 10 performs the step of setting up an opponent agent 5 (battle model Mc) based on the battle probability calculated in step S3 (step S4). In step S4, the learning device 10 sets up an opponent agent 5 by random drawing based on the battle probability.

[0030] Then, in the learning method, the learning model M of agent 5, which is the target of learning, is made to compete against the battle model Mc set in step S4, and reinforcement learning of the learning model M is performed (step S5). In step S5, the reinforcement learning unit 33 of the learning device 10 performs reinforcement learning of the learning model M in order to maximize the reward Rt given to agent 5. Also in step S5, the reinforcement learning unit 33 performs reinforcement learning on the learning model M obtained by performing reinforcement learning. 0~N This is stored in database 35.

[0031] Next, in the learning method, the reinforcement learning unit 33 of the learning device 10 determines whether the learning step T has reached the swap step (step S6). In step S6, if the reinforcement learning unit 33 determines that the learning step T has reached the swap step (step S6: Yes), it executes a step to change the opponent by random drawing based on the match probability (step S7). In step S7, similar to step S4, the learning device 10 changes the opponent agent 5 (match model Mc) based on the match probability calculated in step S3.

[0032] On the other hand, in step S6, if the reinforcement learning unit 33 determines that the learning step T has not reached the swap step (step S6: No), it proceeds back to step S5 and repeatedly executes steps S5 through S6 until the learning step T reaches the swap step.

[0033] After step S7 is executed, in the learning method, the reinforcement learning unit 33 of the learning device 10 determines whether the learning step T has reached a certain step (step S8). In step S8, if the reinforcement learning unit 33 determines that the learning step T has reached a certain step (step S8: Yes), the model comparison unit 31 calculates and evaluates the strength of the opponent against the reinforced learning model M (step S9). In step S9, the model comparison unit 31 of the learning device 10 performs the step of evaluating the strength of the opponent using the same evaluation method as in step S2, with the reinforced learning model M as the reference.

[0034] On the other hand, in step S8, if the reinforcement learning unit 33 determines that the learning step T has not reached a certain step (step S8: No), it proceeds back to step S5 and repeatedly executes steps S5 through S8 until the learning step T reaches a certain step.

[0035] After step S9 is executed, in the learning method, the reinforcement learning unit 33 of the learning device 10 determines whether the learning step T has reached the maximum learning step S (step S10). In step S10, if the reinforcement learning unit 33 determines that the learning step T has reached the maximum learning step S (step S10: Yes), reinforcement learning is terminated and the series of learning methods is terminated. On the other hand, in step S10, if the reinforcement learning unit 33 determines that the learning step T has not reached the maximum learning step S (step S10: No), the process proceeds to step S3, and steps S3 to S10 are repeatedly executed until the maximum learning step S is reached.

[0036] Thus, the learning unit 23, which executes steps S1 to S10 above, functions as a processing unit for reinforcement learning of agent 5. The learning device 10 stores a learning program P for executing the above learning method in its memory unit.

[0037] Figure 4 is an explanatory diagram of the learning method described above. As shown in Figure 4, the learning model M performs reinforcement learning by playing against an opponent of a predetermined strength (strong opponent B in Figure 4) that is set based on the match probability, until the learning step T reaches the swap step. After this, the opponent is changed based on the match probability, and the learning model M again performs reinforcement learning by playing against an opponent of a predetermined strength (even opponent C in Figure 4) until the learning step T reaches the swap step. When changing opponents based on match probability, the opportunities to play against weaker opponents decrease compared to strong opponents because the probability of playing against weaker opponents is low.

[0038] As described above, the learning device 10, learning method, and learning program P described in this embodiment can be understood, for example, as follows.

[0039] The learning device 10 according to the first embodiment is a learning device 10 that includes a processing unit for performing reinforcement learning on a learning model M of an agent 5 in a battle environment in which agents 5 compete against each other, wherein the learning model M includes hyperparameters, and the processing unit performs the following steps: step S2 for evaluating the strength of a plurality of agents 5 that will be opponents of the agent 5 to be learned; step S3 for setting a battle probability for the agent 5 to be learned according to the strength of the opponent agents 5; step S4 for setting the opponent agents 5 based on the battle probability; and step S5 for performing reinforcement learning on the agent 5 to be learned by having the set opponent agents 5 compete against each other.

[0040] This configuration allows for learning opportunities tailored to the opponent's strength, enabling reinforcement learning of the learning model M that is appropriate to the opponent's strength. Therefore, even with diverse opponents in a competitive environment, the learning of the learning model M, including its hyperparameters, can be performed appropriately and efficiently.

[0041] In a second embodiment, in step S3, where the match probability is set, the weaker the strength of the opponent agent 5 among the multiple agents 5 that will be opponents, the lower the match probability is set.

[0042] This configuration allows for more opportunities for reinforcement learning of the learning model M when the opponent is strong, and fewer opportunities when the opponent is weak. Therefore, it is possible to perform appropriate reinforcement learning according to the opponent.

[0043] In a third embodiment, step S2, which evaluates the strength of agent 5, includes at least one of the following as an indicator of the strength of agent 5: match win rate, rating, and KL divergence.

[0044] This configuration allows for an accurate assessment of the opponent's strength.

[0045] The learning method according to the fourth aspect is a learning method for performing reinforcement learning on a learning device 10 in a competitive environment in which agents 5 compete against each other, wherein the learning model M includes hyperparameters, and the learning device 10 is instructed to perform the following steps: step S2 evaluate the strength of a plurality of agents 5 that will be opponents of the agent 5 to be learned; step S3 set a match probability for the agent 5 to be learned according to the strength of the opponent agents 5; step S4 set the opponent agents 5 based on the match probability; and step S5 have the set opponent agents 5 compete and perform reinforcement learning on the agent 5 to be learned.

[0046] This configuration allows for learning opportunities tailored to the opponent's strength, enabling reinforcement learning of the learning model M that is appropriate to the opponent's strength. Therefore, even with diverse opponents in a competitive environment, the learning of the learning model M, including its hyperparameters, can be performed appropriately and efficiently.

[0047] The fifth learning program P is a learning program for performing reinforcement learning on a learning device 10 in a competitive environment where agents 5 compete against each other, wherein the learning model M includes hyperparameters, and the learning device 10 is instructed to perform the following steps: step S2 evaluate the strength of a plurality of agents 5 that will be opponents of the agent 5 to be learned; step S3 set a match probability for the agent 5 to be learned according to the strength of the opponent agents 5; step S4 set the opponent agents 5 based on the match probability; and step S5 have the set opponent agents 5 compete and perform reinforcement learning on the agent 5 to be learned.

[0048] This configuration allows for learning opportunities tailored to the opponent's strength, enabling reinforcement learning of the learning model M that is appropriate to the opponent's strength. Therefore, even with diverse opponents in a competitive environment, the learning of the learning model M, including its hyperparameters, can be performed appropriately and efficiently. [Explanation of Symbols]

[0049] 5 Agents 10 Learning device 22 Environment Department 23 Learning Department 31 Model Comparison Section 32. Match Probability Calculation Unit 33 Reinforcement Learning Department 35 Databases M Learning Model Ma motion model Mb Environment Model MC Battle Model Md reward model P Learning Program

Claims

1. A learning device comprising a processing unit for reinforcing the learning model of an agent in a competitive environment where agents compete against each other, The aforementioned learning model includes hyperparameters, The aforementioned processing unit, A step of evaluating the strength of multiple agents that will be opponents of the agent to be learned, The steps include setting a match probability for the agent to be trained, according to the strength of the opponent agent, The steps include setting the agent that will be the opponent based on the aforementioned match probability, The process involves having the configured opponent agent compete against another agent, and then performing reinforcement learning on the agent that will be the target of the learning process. The predetermined learning steps performed in the step of performing the reinforcement learning are: The swap steps are set for each opponent, There are certain steps that mark the end of the opponent change, This includes the maximum number of learning steps required to complete the learning process. When it is determined that the learning step of the agent to be learned has reached the swap step, the step of setting the agent is executed to change the agent that will be the opponent, When it is determined that the learning step of the agent to be learned has reached a certain step, the step of evaluating the strength of the agent is performed to evaluate the strength of the opponent against the reinforcement-learned learning model. A learning device that repeatedly performs the steps from evaluating the strength of the agent to executing the reinforcement learning until the learning step of the agent to be learned reaches the maximum learning step.

2. The learning device according to claim 1, wherein in the step of setting the probability of a match, the weaker the strength of the opponent agent among the multiple agents that will be opponents, the lower the probability of a match.

3. The learning device according to claim 1 or 2, wherein the step of evaluating the strength of the agent includes at least one of the match win rate, rating, and KL divergence as an indicator of the strength of the agent.

4. A learning method for reinforcement learning an agent's learning model using a learning device in a competitive environment where agents compete against each other, The aforementioned learning model includes hyperparameters, The learning device, A step of evaluating the strength of multiple agents that will be opponents of the agent to be learned, The steps include setting a match probability for the agent to be trained, according to the strength of the opponent agent, The steps include setting the agent that will be the opponent based on the aforementioned match probability, The process involves having the agent that will be the opponent after setup compete against another agent, and then performing reinforcement learning on the agent that will be the target of learning. The predetermined learning steps performed in the step of performing the reinforcement learning are: The swap steps are set for each opponent, There are certain steps that mark the end of the opponent change, This includes the maximum number of learning steps required to complete the learning process. When it is determined that the learning step of the agent to be learned has reached the swap step, the step of setting the agent is executed to change the agent that will be the opponent, When it is determined that the learning step of the agent to be learned has reached a certain step, the step of evaluating the strength of the agent is performed to evaluate the strength of the opponent against the reinforcement-learned learning model. A learning method that repeatedly performs the steps from evaluating the strength of the agent to executing the reinforcement learning until the learning step of the agent to be learned reaches the maximum learning step.

5. A learning program for reinforcement learning the learning model of an agent using a learning device in a competitive environment where agents compete against each other, The aforementioned learning model includes hyperparameters, The learning device, A step of evaluating the strength of multiple agents that will be opponents of the agent to be learned, The steps include setting a match probability for the agent to be trained, according to the strength of the opponent agent, The steps include setting the agent that will be the opponent based on the aforementioned match probability, The process involves having the agent that will be the opponent after setup compete against another agent, and then performing reinforcement learning on the agent that will be the target of learning. The predetermined learning steps performed in the step of performing the reinforcement learning are: The swap steps are set for each opponent, There are certain steps that mark the end of the opponent change, This includes the maximum number of learning steps required to complete the learning process. When it is determined that the learning step of the agent to be learned has reached the swap step, the step of setting the agent is executed to change the agent that will be the opponent, When it is determined that the learning step of the agent to be learned has reached a certain step, the step of evaluating the strength of the agent is performed to evaluate the strength of the opponent against the reinforcement-learned learning model. A learning program that repeatedly performs the steps from evaluating the strength of the agent to executing the reinforcement learning until the learning step of the agent to be learned reaches the maximum learning step.

Citation Information

Patent Citations

  • Strategy model training method, device and equipment

    CN114330754A

  • Information processor and information processing program

    JP2019197592A

  • Model selection method, model selection program, and information processing device

    WO2021038759A1