Reinforcement learning device, reinforcement learning method, and reinforcement learning program
The reinforcement learning device dynamically adjusts the search space based on predicted rewards to address convergence time and local solution issues in conventional reinforcement learning methods, achieving efficient exploration and optimal control in continuous action spaces.
Patent Information
- Application Number
- JP2024505854
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-11
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-03-11
AI Technical Summary
Conventional reinforcement learning methods face challenges such as prolonged convergence time and the risk of falling into local solutions due to incomplete policies, especially when dealing with continuous action spaces.
A reinforcement learning device and method that dynamically adjusts the search space based on predicted rewards, allowing for efficient exploration and policy updates in continuous action spaces.
This approach reduces the time to convergence by adapting the search space according to reward predictions, preventing local solutions and enabling optimal control in reinforcement learning for continuous action spaces.
Smart Images

Figure 0007687520000004 
Figure 0007687520000005 
Figure 0007687520000006
Abstract
Description
Technical Field
[0001] The disclosed technology relates to a reinforcement learning device, a reinforcement learning method, and a reinforcement learning program.
Background Art
[0002] Reinforcement learning is a technique that can learn better actions for an unknown environment. It is also possible to use continuous values for actions. When dealing with continuous actions, the policy establishment density function can be treated as a normal distribution with mean μ and variance σ 2 (see, for example, Non-Patent Document 1). At this time, the larger σ is, the greater the variation in the calculated actions, and a wide range of exploration is performed.
[0003] Also, since reinforcement learning learns by trial and error, it has the drawback of slow learning, and studies have been conducted to shorten the calculation time such as parallelization (see, for example, Non-Patent Document 2).
Prior Art Documents
Non-Patent Documents
[0004]
Non-Patent Document 1
Non-Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0005] Conventional reinforcement learning methods have the following first and second problems. The first problem is that it takes a great deal of time for reinforcement learning to converge. Therefore, if it is possible to learn with as few trial numbers as possible and prevent the calculation time for a single trial from increasing excessively for efficient exploration, the calculation time can be reduced and learning can be converged.
[0006] The second problem is that in reinforcement learning that performs exploration and trial and error based on a policy, if the policy is incomplete, exploration may not be performed well, and the system may fall into a local solution and optimal control may not be achieved.
[0007] The disclosed technology has been made in view of the above points, and an object thereof is to provide a reinforcement learning device, a reinforcement learning method, and a reinforcement learning program that can dynamically adjust a search space according to a predicted reward in reinforcement learning for a continuous action space.
Means for Solving the Problem
[0008] A first aspect of the present disclosure is a reinforcement learning apparatus that performs reinforcement learning for a continuous action space, in which preset settings for simulation and an agent model are stored. In the simulation based on the settings in the reinforcement learning, a defined action is used as an input to obtain a state in the next trial, a reward corresponding to the state, and a flag indicating whether the simulation execution has ended. The agent model estimation unit inputs the state obtained by the simulation into the agent model to obtain a policy, an action determination unit that calculates the action based on the policy and a predefined exploration amount, and an exploration amount estimation unit for estimating the exploration amount. The agent model estimation unit updates the agent model according to the settings of the agent model based on the state, the reward, the flag, and the action. The exploration amount estimation unit updates the exploration amount based on a predicted reward obtained for the reward and the exploration amount in the previous trial. The calculation of the action, the update of the agent model, and the update of the exploration amount are repeated until a predetermined condition according to the flag and the settings is satisfied.
[0009] A second aspect of the present disclosure is a reinforcement learning method for performing reinforcement learning on a continuous action space, in which preset settings for a simulation and an agent model are stored. In the simulation based on the settings in the reinforcement learning, with a predefined action as input, a state in the next trial, a reward corresponding to the state, and a flag indicating whether the simulation execution has ended are acquired. The state acquired by the simulation is input into the agent model to obtain a policy, and based on the policy and a predefined exploration amount, the action is calculated. Further, based on the state, the reward, the flag, and the action, the agent model is updated according to the settings of the agent model. The exploration amount is updated based on a predicted reward obtained for the reward and the exploration amount in the previous trial. The calculation of the action, the update of the agent model, and the update of the exploration amount are repeated until a predetermined condition according to the flag and the settings is satisfied, and the computer is caused to execute the process.
[0010] A third aspect of the present disclosure is a reinforcement learning program for performing reinforcement learning on a continuous action space, in which preset settings for a simulation and an agent model are stored. In the simulation based on the settings in the reinforcement learning, with a predefined action as input, a state in the next trial, a reward corresponding to the state, and a flag indicating whether the simulation execution has ended are acquired. The state acquired by the simulation is input into the agent model to obtain a policy, and based on the policy and a predefined exploration amount, the action is calculated. Further, based on the state, the reward, the flag, and the action, the agent model is updated according to the settings of the agent model. The exploration amount is updated based on a predicted reward obtained for the reward and the exploration amount in the previous trial. The calculation of the action, the update of the agent model, and the update of the exploration amount are repeated until a predetermined condition according to the flag and the settings is satisfied, and the computer is caused to execute the process.
Advantages of the Invention
[0011] According to the disclosed technology, in reinforcement learning for a continuous action space, the search space can be dynamically adjusted according to the predicted reward.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Embodiments for Carrying Out the Invention
[0013] Hereinafter, an example of an embodiment of the disclosed technology will be described with reference to the drawings. In each drawing, the same or equivalent components and parts are given the same reference numerals. Also, the dimensional ratios in the drawings are exaggerated for convenience of explanation and may be different from the actual ratios.
[0014] FIG. 1 is a block diagram showing the hardware configuration of the reinforcement learning device 100.
[0015] As shown in FIG. 1, the reinforcement learning device 100 includes a CPU (Central Processing Unit) 11, a ROM (Read Only Memory) 12, a RAM (Random Access Memory) 13, a storage 14, an input unit 15, a display unit 16, and a communication interface (I / F) 17. Each component is connected to be communicable with each other via a bus 19.
[0016] The CPU 11 is a central processing unit that executes various programs and controls each unit. That is, the CPU 11 reads a program from the ROM 12 or the storage 14 and executes the program using the RAM 13 as a work area. The CPU 11 performs control of the above-described components and various arithmetic processes according to the programs stored in the ROM 12 or the storage 14. In the present embodiment, a reinforcement learning program is stored in the ROM 12 or the storage 14.
[0017] The ROM 12 stores various programs and various data. The RAM 13 temporarily stores a program or data as a work area. The storage 14 is composed of a storage device such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive), and stores various programs including an operating system and various data.
[0018] The input unit 15 includes a pointing device such as a mouse and a keyboard, and is used to perform various inputs.
[0019] The display unit 16 is, for example, a liquid crystal display, and displays various information. The display unit 16 may adopt a touch panel method and function as the input unit 15.
[0020] The communication interface 17 is an interface for communicating with other devices such as terminals. For this communication, for example, a wired communication standard such as Ethernet (registered trademark) or FDDI, or a wireless communication standard such as 4G, 5G, or Wi-Fi (registered trademark) is used.
[0021] Next, each functional configuration of the reinforcement learning device 100 will be described. FIG. 2 is a block diagram showing the functional configuration of the reinforcement learning device of the present embodiment. Each functional configuration is realized by the CPU 11 reading out the reinforcement learning program stored in the ROM 12 or the storage 14, expanding it in the RAM 13, and executing it. The reinforcement learning device 100 performs reinforcement learning for a continuous action space.
[0022] As shown in FIG. 2, the reinforcement learning device 100 includes a learning setting storage unit 110, an agent model estimation unit 111, an exploration amount estimation unit 112, and an action determination unit 113. This configuration is the main configuration 100A of the reinforcement learning device 100. Further, the reinforcement learning device 100 includes a setting input unit 101, a simulation execution unit 102, a model storage unit 103, an action storage unit 104, and an operation output unit 105 as processing units responsible for input / output functions.
[0023] The setting input unit 101 stores the data received by the input from the user in the learning setting storage unit 110. Note that the setting input unit 101 corresponds to the input unit 15 as hardware.
[0024] In the learning setting storage unit 110, the data received from the user by the setting input unit 101 is stored as settings. FIG. 3 shows an example of the data stored in the learning setting storage unit 110. Information is stored for each column of "setting item", "setting content", and "setting target". The "setting content" is the setting or set value for the "setting item". The "setting target" is the target processing unit of the reinforcement learning device 100. In the first and second rows, the settings for the "setting items" {parameter of exploration amount estimator} and {initial exploration amount} are stored and used by the exploration amount estimator 112. As the {parameter of exploration amount estimator}, α, λ, and C are defined. In the third row, the setting for the "setting item" {name of reinforcement learning algorithm} is stored and used by the agent model estimator 111. The reinforcement learning algorithm (hereinafter, simply referred to as the algorithm when referring to the reinforcement learning algorithm) that determines the processing content in the agent model estimator 111 is selected according to the {name of reinforcement learning algorithm}. In the fourth and eighth rows, the settings for the "setting items" {maximum number of steps} and {frequency of saving agent model}, which are required for each algorithm, are stored. In the fifth to seventh rows, the settings for the "setting items" {simulation type name}, {simulation initialization parameter}, and {initial action value} indicating the action value at the start of execution, which select the processing content of the simulation execution unit 102, are stored. Note that these settings are only examples, and the learning setting storage unit 110 can appropriately store the settings required for reinforcement learning. The learning setting storage unit 110 transmits each stored set value to the simulation execution unit 102, the agent model estimator 111, and the exploration amount estimator 112, which are the setting targets of each set value, respectively.
[0025] The simulation execution unit 102 executes a simulation for the input action a. The input of the action a that triggers the simulation execution receives the initial action value from the learning setting storage unit 110 during the initial operation, and sets the initial action value as the action a. When it is not during the initial operation, it receives the action a from the action determination unit 113. The simulation execution unit 102 outputs the state s (next state s) observed as a result of the simulation, the reward r defined for the state, and the flag d by executing the simulation, and transmits these outputs to the agent model estimation unit 111. The flag d is a boolean value indicating whether the simulation has ended and the simulation environment should be reset.
[0026] The internal algorithm of the simulation execution unit 102 is set according to the {simulation type name} stored in the learning setting storage unit 110. The internal algorithm corresponds to, for example, a video game or a board game characterized by screen or state transitions for specific operations, or a simulator can be used. The simulator reproduces the state change when the device is operated for a specific state prepared by the user in advance. For example, a simulator that reproduces the indoor temperature and humidity change when the air conditioner is controlled can be used. In addition, if there is a real environment having the same input and output as the simulation execution unit 102, the real environment may be used. The real environment is, for example, an environment where there is a building that can control the air conditioner, and the indoor temperature and humidity change can be measured by a sensor or the like and data can be collected.
[0027] The agent model estimation unit 111 receives the outputs (state s, reward r, and flag d) transmitted from the simulation execution unit 102, and receives the action a of the previous trial transmitted from the action determination unit 113. In addition, the agent model estimation unit 111 reads various setting values stored in the learning setting storage unit 110, and extracts the agent model stored in the model storage unit 103.
[0028] The agent model estimation unit 111 inputs the state s obtained from the simulation execution unit 102 into the agent model, and acquires a policy π as part of the output of the agent model. The agent model estimation unit 111 transmits the acquired policy π to the action determination unit 113. Also, the agent model estimation unit 111 inputs the state s, reward r, flag d, and action a extracted into the agent model, and updates the agent model. Here, the state s used for updating the agent model is the next state s for the next time (trial) as described later, and the reward r and flag d corresponding to the next state s. Also, the action a is the one updated after the acquisition of the policy π. The internal algorithm (reinforcement learning algorithm) used for the calculation of the agent model is defined by the {reinforcement learning algorithm name} stored in the learning setting storage unit 110. The reinforcement learning algorithm may use existing techniques, and an algorithm targeting continuous value actions may be used. The agent model defined by the algorithm is in the form of a function or a neural network, and their hyperparameters and neural network weight coefficients are updated by the method defined by each algorithm. Depending on the algorithm, in the agent model estimation unit, the history of the state s, reward r, and action a may be stored and used for model update. When the update of the agent model defined by the algorithm is executed or when the storage frequency of the agent model is described in the learning setting storage unit 110, the agent model estimation unit 111 transmits the updated agent model to the model storage unit 103 based on the set value.
[0029] The exploration amount estimation unit 112 receives the reward r transmitted from the simulation execution unit 102, and based on the reward r, updates the predicted reward r, for example, from Equation (1). pred The parameter α in Equation (1) is the learning rate of the predicted reward, and the set value of the {exploration amount estimation parameter} stored in the learning setting storage unit 110 is used. Also, r on the right side pred is the predicted reward before update. Note that an arbitrary value such as 0 is used as the initial value of r on the right side pred .
Number
[0030] Next, the exploration amount estimation unit 112 determines the exploration amount σ from Equation (2) based on the predicted reward r and transmits it to the action decision unit 113. Note that the parameters λ and C in Equation (2) use the set values of the {exploration amount estimation parameters} stored in the learning setting storage unit 110. pred
Number
[0031] Equations (1) and (2) are simplified models for dynamically adjusting the variation of movements in the motor learning of animals and humans.
[0032] The action decision unit 113 calculates and determines the action a for the next trial based on the policy π transmitted from the agent model estimation unit 111 and the exploration amount σ output from the exploration amount estimation unit 112, and transmits the action a to the simulation execution unit 102.
[0033] Here, when the policy π represents a normal distribution with a mean μ and a variance σ 2 , the probability density function of the action a can be expressed as follows in Equation (3) by the exploration amount σ output from the exploration amount estimation unit 112. x is a random variable. The action a is probabilistically determined according to the probability density function. As a result, reinforcement learning for a continuous action space can be performed.
Number
[0034] In the model storage unit 103, the agent model updated in the agent model estimation unit 111 is stored. FIG. 4 shows an example of the data of the agent model stored in the model storage unit 103. It is assumed that the model is stored in principle every time it is updated. When {agent model storage frequency} is described in the learning setting storage unit 110, the setting is followed.
[0035] Also, when the learning of the agent model estimation unit 111 is interrupted halfway and executed again, if a model with the same {reinforcement learning algorithm name} and the same {simulation type name} is stored in the model storage unit 103 in advance, the model with the larger number of steps can be read and used.
[0036] In the action storage unit 104, the actions at each time transmitted from the action decision unit 113 are stored. FIG. 5 shows an example of the data of the actions stored in the action storage unit 104.
[0037] The operation output unit 105 extracts the actions for a specific period stored in the action storage unit 104 and outputs the control content to the target controller.
[0038] (Processing flow of the reinforcement learning device 100) Next, the operation of the reinforcement learning device 100 will be described. FIG. 6 is a flowchart showing the flow of the reinforcement learning process by the reinforcement learning device 100. The CPU 11 reads the reinforcement learning program from the ROM 12 or the storage 14, expands it in the RAM 13, and executes it, thereby performing the reinforcement learning process.
[0039] In step S100, the CPU 11 initializes the agent model and the exploration amount. The initialization of the agent model is performed by the agent model estimation unit 111, and the exploration amount is performed by the exploration amount estimation unit 112.
[0040] In the agent model estimation unit 111, the agent model is initialized based on the setting items such as the {reinforcement learning algorithm name} stored in the learning setting storage unit 110. In the model storage unit 103, if a model corresponding to the combination of the {reinforcement learning algorithm name} and the {simulation type name} stored in the learning setting storage unit 110 already exists, the one with the larger number of steps among the corresponding models is read out as the weight of the agent model. Here, when reading out the model stored in the model storage unit 103, the current number of steps is defined by the number of steps of the stored model. When not reading out the model stored in the model storage unit 103, the current number of steps is set to 0. The current number of steps is stored inside the agent model estimation unit 111.
[0041] Also, in the exploration amount estimation unit 112, the {exploration amount estimation parameter} and the {initial exploration amount} stored in the learning setting storage unit 110 are extracted. Initialization is performed so that the value of the initial exploration amount is used as the exploration amount.
[0042] In step S102, the CPU 11 initializes the simulator in the simulation execution unit 102 and acquires the state s.
[0043] In the simulation execution unit 102, the {simulation type name} and the {simulation initialization parameter} stored in the learning setting storage unit 110 are read out, and the simulation environment corresponding to the simulation type name is initialized using the simulation initialization parameter. The simulation execution unit 102 outputs the initial state s by the initialization and outputs it to the agent model estimation unit. Also, the state s is stored inside the simulation execution unit 102.
[0044] In step S104, the CPU 11 inputs the state s acquired from the simulation execution unit 102 into the agent model as the agent model estimation unit 111 and acquires the policy π as the output.
[0045] In step S106, the CPU 11, acting as the action determination unit 113, calculates and determines the action a based on the policy π output from the agent model estimation unit 111 and the exploration amount σ defined by the exploration amount estimation unit 112. The action a is output to the simulation execution unit 102 and the action storage unit 104 and is also stored inside the agent model estimation unit 111.
[0046] In step S108, the CPU 11 adds 1 to the current step stored in the agent model estimation unit 111.
[0047] In step S110, the CPU 11, acting as the simulation execution unit 102, acquires the next state s, the reward r, and the flag d. In the simulation execution unit 102, based on the action a acquired from the action determination unit 113 and the state s stored inside the simulation execution unit 102, the next state s for the next time (next trial) is acquired. Also, in the simulation execution unit 102, along with the acquisition of the next state s, the reward r and the flag d indicating whether the simulation execution has ended are calculated according to the next state s.
[0048] In step S112, the CPU 11, acting as the agent model estimation unit 111, updates the agent model. The update is executed according to {reinforcement learning algorithm name} based on the state s, the reward r, and the flag d acquired from the simulation execution unit 102 and the action a stored inside the agent model estimation unit 111. The updated agent model is stored in the model storage unit 103 when it corresponds to the frequency described in {agent model storage frequency}.
[0049] Note that depending on the type of reinforcement learning algorithm, there are cases where the agent model is updated every time a simulation is executed, and cases where the model is updated collectively at intervals without updating the model each time. If the algorithm name registered in {Reinforcement learning algorithm name} is an algorithm that performs collective updates at intervals, it is saved inside the agent model estimator 111 without being updated. That is, except for the update timing defined by the algorithm, the update process of the agent model is not performed. Instead, the state s, reward r, and flag d obtained from the simulation execution unit 102 are saved inside the agent model estimator 111. At the update timing defined by the algorithm, the agent model is updated based on the history of the state s, reward r, flag d, and action a saved inside the agent model estimator 111.
[0050] In step S114, the CPU 11 updates the exploration amount as the exploration amount estimator 112 based on the predicted reward obtained from the reward r acquired from the simulation execution unit 102 and the exploration amount σ at the previous time saved inside the exploration amount estimator 112. The update of the exploration amount is performed by calculating the predicted reward in the above formula (1) and the exploration amount σ in formula (2). The updated exploration amount is saved inside the exploration amount estimator 112 and used during the operation of the action decision unit 113. By updating the exploration amount in this way, the update can be performed so as to expand the exploration space when the reward amount is small.
[0051] In step S116, the CPU 11 determines whether the flag d in the simulation execution unit 102 is True or False. If it is True, the initialization of the simulation execution unit 102 in step S102 is executed, and the subsequent processing is executed again. If the flag d is False, the process proceeds to step S118. The fact that it is False is an example of satisfying the predetermined condition regarding the flag of the present disclosure.
[0052] In step S118, the CPU 11 determines whether the current step, which is a variable held in the agent model estimator 111, exceeds the maximum number of steps stored in the learning setting storage unit 110. If it does not exceed, the processes after step S104 are executed again, and if it exceeds, all processes are terminated. Exceeding the maximum number of steps is an example of satisfying a predetermined condition regarding the setting of the present disclosure.
[0053] As described above, according to the reinforcement learning device 100 of the present embodiment, in reinforcement learning targeting a continuous action space, the search space can be dynamically adjusted according to the predicted reward. Thereby, in a situation where a large reward can be obtained, the time until learning convergence is shortened without expanding the search space, while in a situation where no reward can be obtained, optimal control can be realized without falling into a local solution by expanding the search space.
[0054] Generally, in a situation where a large reward can be obtained, since good control can be executed with the current policy, the need for wide-ranging exploration is low. On the contrary, in a situation where no reward can be obtained, since good control cannot be executed, it is necessary to explore widely. In the method of the present disclosure, by dynamically adjusting the search space according to the predicted reward, when falling into a policy where no reward is obtained, the search space is expanded, and by trying a wide range of actions, it is possible to escape from the local solution and search for the optimal solution.
[0055] According to the method of the present disclosure, in reinforcement learning targeting a continuous action space, by efficiently performing exploration, the time until learning convergence can be shortened, and the first problem can be solved. Also, by expanding the search space when the amount of reward is small, it is possible to learn a policy that can obtain more rewards without falling into a local solution, and the second problem can be solved.
[0056] (Utilization in various industrial fields) The method using the reinforcement learning device 100 in the present disclosure can be used in various industrial fields, and each case will be described by giving utilization examples.
[0057] <When used for air conditioning control> In this embodiment, as the simulation execution unit 102, a simulator that predicts future room temperature changes and heat consumption using weather data, the number of visitors, past room temperature, air conditioning control data, etc. as inputs is used, and the set value of air conditioning control is treated as an action. As a result, an agent model that learns optimal air conditioning control to achieve energy savings while maintaining comfort can be created.
[0058] Regarding the temperature prediction in the simulation execution unit 102, it can be realized by using a neural network or a regression model that takes various data as inputs and outputs the room temperature. Also, regarding the heat consumption prediction, it can be realized by using a regression model that predicts the required heat amount with weather data, the number of visitors, and the set value of the air conditioner as inputs. Also, these can be used in combination.
[0059] At this time, the simulation execution unit 102 internally holds data acquired from various sensors such as weather data, the number of visitors, past room temperature, and air conditioning control history (this is defined as environmental data). Also, the simulation execution unit 102 is assumed to have previously learned a model that reproduces environmental changes at a future time using these environmental data and a model that estimates the amount of heat consumed by air conditioning equipment (heat consumption) accompanying air conditioning control. Also, it is assumed that rules for evaluating whether these values are comfortable and energy-saving based on the estimated temperature and humidity and heat consumption are defined in advance.
[0060] Regarding step S100, it is as described in the explanation of the above processing flow. Regarding step S102, in the simulation execution unit 102, the {simulation type name} and {simulation initialization parameters} registered in the learning setting storage unit 110 are read. Here, when used for air conditioning control, the name of the simulator that predicts future room temperature changes according to weather, the number of visitors, past room temperature, and air conditioning control (e.g., indoor temperature and humidity reproduction env) is specified. The simulation execution unit 102 is initialized according to the {simulation initialization parameters}. For example, one day is randomly selected from the dates with existing environmental data and where simulation is possible, and according to the time t specified by the simulation initialization parameters, the environmental data required for indoor temperature and humidity reproduction from time t of that date is loaded and held in the simulation execution unit 102. Also, the environmental data required for heat consumption estimation from time t of that date is similarly loaded and held in the simulation execution unit 102. As the initial state, the indoor temperature and humidity data at time t is acquired and output to the agent model estimation unit 111.
[0061] Regarding steps S104, S108, S112 and subsequent steps, it is as described in the explanation of the above processing flow.
[0062] Regarding step S106, it is as described in the above process flow. Here, action a indicates the air conditioning control method at a certain time, and as shown in FIG. 4, it indicates the set values for each air conditioning device. Regarding step S110, in the simulation execution unit 102, the indoor temperature and humidity at the next time (for example, 10 minutes later) are predicted. The indoor temperature and humidity are predicted based on the action a (that is, the air conditioning control method) acquired from the action determination unit 113, the state s stored in the simulation execution unit 102, and the environmental data loaded in advance. Also, regarding the amount of heat consumed by the air conditioning device along with the air conditioning control, it is estimated using the state s and the environmental data loaded in advance. Based on a rule for evaluating whether it is a good state from the viewpoints of pre-defined comfort and energy saving, a reward is determined. If the time when the simulation is performed is the last time when data exists in the date, the flag d indicating whether the simulation has ended is set to True, and in other cases, it is set to False. The state s, the reward r, and the flag d are output to the agent model estimation unit 111 and the exploration amount estimation unit 112.
[0063] <When used for controlling devices such as robots> In the case of this usage form, as the simulation execution unit 102, a simulator that predicts the future state of the device using information indicating the state of the device and device operations as inputs is used, and device operation commands (motor operations and device movement instructions) are handled as actions. The information indicating the state of the device is the joint angle, speed, robot position information, etc. Thereby, an agent model for learning the optimal device control for realizing the target operation can be created. At this time, it is assumed that the simulation execution unit 102 has learned in advance so that it can predict changes in the device state from the previously measured data, or can predict changes in the device state by a physical simulator. Also, it is assumed that a rule for evaluating that it is the target operation is defined in advance.
[0064] Regarding step S100, it is as described in the above process flow. Regarding step S102, in the simulation execution unit 102, the {simulation type name} and {simulation initialization parameters} registered in the learning setting storage unit 110 are read. Here, when used for robot control, the name of the simulator that predicts the next state from the previous state and device operations (e.g., robot arm env) is specified. The simulation execution unit 102 is initialized according to the {simulation initialization parameters}.
[0065] Regarding steps S104, S108, S112 and later, it is as described in the above process flow.
[0066] Regarding step S106, it is as described in the above process flow. Here, the action a indicates the device control method at a certain time. Regarding step S110, in the simulation execution unit 102, the state change at the next time (e.g., 1 second later) is predicted. The state change is predicted based on the action a (i.e., the device control method) obtained from the action decision unit 113, the state s stored in the simulation execution unit, and the environment data loaded in advance. Also, the reward is determined based on the rule for evaluating that it is a predetermined target operation. As a result of the simulation, if the simulation ends due to the failure of the operation, the flag d indicating whether the simulation has ended is set to True, and otherwise it is set to False. The state s, the reward r, and the end flag d are output to the agent model estimation unit and the exploration amount estimation unit. The failure of the operation means, for example, dropping an object when carrying it with a robot arm, or a moving robot going outside the operation target area.
[0067] <When used for game operations> In this application example, as the simulation execution unit 102, a game in which the state transitions with information indicating the state (such as a game screen) and game operations as inputs is used as a simulator, and the game operations are treated as actions. As a result, an agent model that learns game operations that can achieve high scores can be created. At this time, the rules of the game are determined in advance, assuming that they can be obtained as rewards.
[0068] Regarding step S100, it is as described in the above process flow. In step S102, in the simulation execution unit 102, the {simulation type name} and {simulation initialization parameters} registered in the learning setting storage unit 110 are read out. Here, when using for game operations, the name of the simulator (game) (e.g., block-breaking env) is specified. The simulation execution unit is initialized according to the {simulation initialization parameters}.
[0069] Regarding steps S104, S108, S112 and later, it is as described in the above process flow.
[0070] Regarding step S106, it is as described in the above process flow. Here, the action a indicates a device control method at a certain time. Regarding step S110, in the simulation execution unit 102, the game is executed using the action a (that is, the game operation) acquired from the action determination unit 113, and the state change at the next time (for example, after 1 frame) is obtained. Also, a reward is acquired based on the predetermined game rules. As a result of the simulation, when the simulation (game execution) ends due to game over or the like, the flag d indicating whether the simulation has ended is set to True, and False otherwise. The state s, the reward r, and the end flag d are output to the agent model estimation unit and the exploration amount estimation unit.
[0071] The above is the description of the application example.
[0072] Note that, in the above-described embodiment, the reinforcement learning process in which the CPU reads and executes software (program) may be executed by various processors other than the CPU. Examples of the processor in this case include PLDs (Programmable Logic Devices) whose circuit configuration can be changed after manufacture, such as FPGAs (Field-Programmable Gate Arrays), GPUs (Graphics Processing Units), and dedicated electric circuits such as ASICs (Application Specific Integrated Circuits) having a circuit configuration designed specifically to execute specific processes. Further, the reinforcement learning process may be executed by one of these various processors, or may be executed by a combination of two or more processors of the same type or different types (for example, a plurality of FPGAs, and a combination of a CPU and an FPGA, etc.). Further, the hardware structure of these various processors is, more specifically, an electric circuit combining circuit elements such as semiconductor elements.
[0073] Also, in the above-described embodiment, the mode in which the reinforcement learning program is stored (installed) in the storage 14 in advance has been described, but it is not limited to this. The program may be provided in a form stored in a non-transitory storage medium such as a CD-ROM (Compact Disk Read Only Memory), a DVD-ROM (Digital Versatile Disk Read Only Memory), and a USB (Universal Serial Bus) memory. Further, the program may be in a form downloaded from an external device via a network.
[0074] Regarding the above embodiments, the following additional remarks are further disclosed.
[0075] (Additional Clause 1) A memory, At least one processor connected to the memory, Comprising, The processor is, A reinforcement learning device that performs reinforcement learning for a continuous action space, settings defined in advance for the simulation and the agent model are stored, in the simulation based on the settings in the reinforcement learning, using a predefined action as an input, a state in the next trial, a reward corresponding to the state, and a flag indicating whether the simulation execution has ended are to be acquired, inputting the state acquired by the simulation into the agent model to obtain a policy, calculating the action based on the policy and a predefined exploration amount, further, updating the agent model according to the settings of the agent model based on the state, the reward, the flag, and the action, updating the exploration amount based on a predicted reward obtained for the reward and the exploration amount in the previous trial, repeating the calculation of the action, the update of the agent model, and the update of the exploration amount until a predetermined condition according to the flag and the settings is satisfied, a reinforcement learning device configured as described above.
[0076] (Additional clause 2) A non-transitory storage medium storing a program executable by a computer to execute reinforcement learning processing, the program is a reinforcement learning program that performs reinforcement learning for a continuous action space, settings defined in advance for the simulation and the agent model are stored, in the simulation based on the settings in the reinforcement learning, using a predefined action as an input, a state in the next trial, a reward corresponding to the state, and a flag indicating whether the simulation execution has ended are to be acquired, inputting the state acquired by the simulation into the agent model to obtain a policy, Calculate the action based on the above strategy and a predefined exploration amount. Furthermore, update the agent model according to the settings of the agent model based on the state, the reward, the flag, and the action. Update the exploration amount based on the predicted reward obtained for the reward and the exploration amount in the previous trial. Repeat the calculation of the action, the update of the agent model, and the update of the exploration amount until a predetermined condition according to the flag and the settings is satisfied. Non-transitory storage medium.
Explanation of Signs
[0077] 100 Reinforcement learning device 100 Learning device 101 Setting input unit 102 Simulation execution unit 103 Model storage unit 104 Action storage unit 105 Operation output unit 110 Learning setting storage unit 111 Agent model estimation unit 112 Exploration amount estimation unit 113 Action decision unit
Claims
1. A reinforcement learning device that performs reinforcement learning for a continuous action space, wherein settings defined in advance for simulation and an agent model are stored, in the simulation based on the settings in the reinforcement learning, using a predefined action as input, a state in the next trial, a reward corresponding to the state, and a flag indicating whether the simulation execution has ended are acquired, an agent model estimation unit that inputs the state acquired by the simulation into the agent model and acquires a policy, an action determination unit that calculates the action based on the policy and a predefined exploration amount, and an exploration amount estimation unit for estimating the exploration amount, wherein the agent model estimation unit updates the agent model according to the settings of the agent model based on the state, the reward, the flag, and the action, the exploration amount estimation unit updates the exploration amount based on a predicted reward obtained for the reward and the exploration amount in the previous trial, and repeats the calculation of the action, the update of the agent model, and the update of the exploration amount until a predetermined condition according to the flag and the settings is satisfied, a reinforcement learning device.
2. The reinforcement learning device according to claim 1, wherein the exploration amount estimation unit calculates the predicted reward based on a parameter of a learning rate of the predicted reward defined in the settings and the reward, and updates the exploration amount based on the calculated predicted reward and a parameter for exploration amount estimation in the settings.
3. The reinforcement learning device according to claim 1 or claim 2, wherein the action determined by the action determination unit is probabilistically determined according to a probability density function using a random variable, the mean and variance of a normal distribution represented by the policy.
4. The reinforcement learning device according to claim 1 or claim 2, wherein the action is an air conditioning control method.
5. A reinforcement learning method for performing reinforcement learning for a continuous action space, wherein settings defined in advance for simulation and an agent model are stored, in the simulation based on the settings in the reinforcement learning, using a predefined action as input, a state in the next trial, a reward corresponding to the state, and a flag indicating whether the simulation execution has ended are acquired, Input the state obtained by the simulation into the agent model to obtain a policy, Calculate the action based on the policy and a predefined exploration amount, Furthermore, update the agent model according to the settings of the agent model based on the state, the reward, the flag, and the action, Update the exploration amount based on the predicted reward obtained for the reward and the exploration amount in the previous trial, Repeat the calculation of the action, the update of the agent model, and the update of the exploration amount until a predetermined condition according to the flag and the settings is satisfied, A reinforcement learning method for causing a computer to execute a process.
6. A reinforcement learning program for performing reinforcement learning on a continuous action space, Predefined settings for the simulation and the agent model are stored, In the simulation based on the settings in the reinforcement learning, a predefined action is used as an input, and in the next trial, a state, a reward corresponding to the state, and a flag indicating whether the simulation execution has ended are obtained, Input the state obtained by the simulation into the agent model to obtain a policy, Calculate the action based on the policy and a predefined exploration amount, Furthermore, update the agent model according to the settings of the agent model based on the state, the reward, the flag, and the action, Update the exploration amount based on the predicted reward obtained for the reward and the exploration amount in the previous trial, Repeat the calculation of the action, the update of the agent model, and the update of the exploration amount until a predetermined condition according to the flag and the settings is satisfied, A reinforcement learning program for causing a computer to execute a process.
Citation Information
Patent Citations
Cross-modal video moment positioning method based on space-time reinforcement learning
CN111782871A
Reward function estimation apparatus, reward function estimation method and program
JP2013225192A
Method and device for reinforcement learning using novel centering operation based on probability distribution
US20200110964A1
Apparatus and method for configuring a communication link
US20200113017A1