Learning device, control system, learning method, and program
Patent Information
- Application Number
- JP2024557258
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2025-05-01
- Publication Date
- 2025-07-15
AI Technical Summary
Defining an internal evaluation function in reinforcement learning to achieve effective learning results is challenging due to the difficulty in specifying a high evaluation for desired states and actions.
A learning device and system that uses an internal evaluation function to calculate evaluation values for operations in a controlled environment, where the evaluation value is higher for states with higher frequency of desired operations, allowing for offline reinforcement learning and determining optimal actions based on these values.
This approach enables the definition of an internal evaluation function that enhances reinforcement learning outcomes by prioritizing states with higher desired operation frequencies, reducing computational complexity and improving control system performance.
Abstract
Description
Learning device, control system, learning method, and recording medium
[0001] The present disclosure relates to a learning device, a control system, a learning method, and a recording medium.
[0002] Reinforcement learning determines the next action to be taken in a certain state based on an evaluation function, and by repeating this process, it searches for a better state or action. The evaluation function in reinforcement learning is specified, for example, to highly evaluate the closeness of the distance between the target state and the detected state, and the effectiveness of learning is improved by repeating trial-and-error calculations based on the results of this evaluation function. Value iteration and policy iteration, known as reinforcement learning methods, are examples of optimization methods for searching for the optimal solution.
[0003] Japanese Patent Application Publication No. 2021-030359
[0004] However, it can be difficult to define an internal evaluation function to obtain good results from such reinforcement learning.
[0005] Therefore, an example of an object of this disclosure is to provide a learning device, a control system, a learning method, and a recording medium that can solve the above-mentioned problems.
[0006] According to a first aspect of the present disclosure, a learning device is provided with a calculation processing unit that uses an internal evaluation function in reinforcement learning to calculate at least one of an evaluation value for an operation in a state in an environment including a controlled object and an evaluation value for a next state of the state, and determines an operation for the controlled object based on the calculated evaluation value, and the internal evaluation function is defined to generate the evaluation value that indicates that the higher the frequency of a state in the desired operation of the controlled object, the higher the evaluation.
[0007] According to a second aspect of the present disclosure, there is provided a control system comprising the above-described learning device and a control unit that has undergone offline reinforcement learning by the learning device based on the frequency of the state, wherein the control unit controls the control target using the results of the offline reinforcement learning.
[0008] According to a third aspect of the present disclosure, there is provided a learning method including: using an internal evaluation function in reinforcement learning to determine at least one of an evaluation value for an operation in a state in an environment including a controlled object and an evaluation value for a next state of the state; and determining an operation for the controlled object based on the obtained evaluation value; and the internal evaluation function is defined to generate the evaluation value indicating that the higher the frequency of a state in a desired operation of the controlled object, the higher the evaluation.
[0009] According to a fourth aspect of the present disclosure, there is provided a recording medium storing a program for causing an apparatus that executes a program for a learning process related to reinforcement learning to use an internal evaluation function in reinforcement learning to determine at least one of an evaluation value for an operation in a state in an environment including a controlled object and an evaluation value for a next state of the state, and to determine an operation for the controlled object based on the obtained evaluation value, and to generate, using the internal evaluation function, the evaluation value that indicates that the higher the frequency of a state in a desired operation of the controlled object, the higher the evaluation.
[0010] According to the present disclosure, this defines an internal evaluation function for achieving good reinforcement learning results.
[0011] FIG. 1 is a configuration diagram of a control system 1 according to some embodiments of the present disclosure. FIG. 1 is a configuration diagram of a calculation processing device 10 according to some embodiments of the present disclosure. FIG. 2 is a diagram for describing application examples of some embodiments of the present disclosure. FIG. 3 is a diagram for describing another application example of some embodiments of the present disclosure. FIG. 4 is a diagram for describing using a value function according to some embodiments of the present disclosure as a learning target. FIG. 5 is a diagram for describing the definition of a "value function" according to some embodiments of the present disclosure. FIG. 6 is a diagram for describing an example of estimating value from a limited number of state-action data according to some embodiments of the present disclosure. FIG. 7 is a diagram for describing a method for approximating a state distribution according to some embodiments of the present disclosure. FIG. 8 is a diagram for describing approximation of a density function according to some embodiments of the present disclosure. FIG. 9 is a diagram for describing an adjustment technique for balancing parameters according to some embodiments of the present disclosure. FIG. 10 is a diagram for describing an adjustment technique for balancing parameters according to some embodiments of the present disclosure. FIG. 11 is a diagram for describing an optimization process according to some embodiments of the present disclosure. FIG. 12 is a diagram for describing the gradient of a value function in a state space according to some embodiments of the present disclosure. FIG. 13 is a diagram for describing the gradient of a value function in a state space showing the relationship between a parameter θv and a state s according to some embodiments of the present disclosure. FIG. 14 is a flowchart for describing processing of a calculation processing device according to some embodiments of the present disclosure. FIG. 15 is a diagram illustrating an example of the configuration of a calculation processing device according to some embodiments of the present disclosure. 1 is a flowchart illustrating an example of a processing procedure in a calculation method according to some embodiments of the present disclosure.
[0012] Below, computers according to several embodiments of the present disclosure will be described with reference to the drawings. The following description describes embodiments of the present disclosure, and the present disclosure is not limited to the following embodiments. For clarity of explanation, the following description has been omitted or simplified as appropriate. Furthermore, each element of the following embodiments can be easily modified, added, or converted within the scope of the present disclosure. Note that the internal evaluation function in reinforcement learning is an example of a function estimated by reinforcement learning, such as a policy function or value function in reinforcement learning. In the following description, the policy function or value function may be referred to as the internal evaluation function. In reinforcement learning, the state frequency refers to the frequency with which a state corresponding to the desired behavior of the control object occurs in the data used in reinforcement learning when the results of reinforcement learning are used to control the control object. In the following description, this may be referred to as an empirical value. Trajectory data indicating a frequency distribution refers to data indicating the frequency of samples or a representative value obtained based on the frequency of samples. For example, the representative value obtained based on the frequency of samples may be a result approximating a Gaussian distribution based on the frequency of samples. When a state in reinforcement learning is used as a sample, trajectory data indicating the frequency distribution of the state may be data indicating the frequency of the state or a representative value obtained based on the frequency of the state. More specifically, the policy distribution in reinforcement learning may be the probability that action a will be selected in state s when the state s is at a certain stage as the state in reinforcement learning transitions through multiple stages. Trajectory data based on conditional probability is data indicating a probability under specified conditions or a representative value obtained based on the probability under specified conditions. For example, the representative value obtained based on the probability under specified conditions may be a result approximating a Gaussian distribution based on the probability under specified conditions. The policy function in reinforcement learning may be set based on trajectory data based on conditional probability depending on the state of the action determined in reinforcement learning. An expert in the control of a control object is an entity that performs an operation related to the control of the control object or determines the operation and determines a command related to the control, and is an example of an entity that has sufficient knowledge about the operation related to the control of the control object.
[0013] 1 is a configuration diagram of a control system 1 according to some embodiments of the present disclosure. The control system 1 controls a plant in real space. For example, the illustrated plant is an example of a chemical plant.
[0014] The control system 1 includes a calculation processing device 10 (learning device) and a plant P. The calculation processing device 10 includes, for example, a regulator 11, a sensor 12, a control unit 13, a storage unit 15, and an input / output unit 16. Note that either the regulator 11 or the sensor 12 may be provided as a component of the plant P.
[0015] The adjuster 11 adjusts the state of the plant P, which is the control target, in accordance with control from the control unit 13. The sensor 12 detects the state of the plant P and outputs the detection result to the control unit 13. The memory unit 15 includes a memory area such as a semiconductor memory element and holds written data. The input / output unit 16 includes a display unit that displays information and an input device such as a touch panel. The communication unit 17 provides an interface for communicating with an external higher-level device.
[0016] Next, an overview of the operation of the control system 1 according to some embodiments of the present disclosure will be described.
[0017] The control unit 13 controls the plant P as a control target, and performs feedback control based on the state value (state s) output by the sensor 12. The control system 1 receives an instruction to switch control tasks from a higher-level device (scheduler), which changes the control target accordingly. Even in such a case, the control system 1 aims to stably and quickly transition to the state indicated by the new control target. When the control target changes in response to the control task switch, the control system 1 can adjust the manipulated variable and generation timing for the plant P.
[0018] For example, the control unit 13 adjusts the state of the plant P by adjusting, for example, the opening degree of a valve, opening and closing of a valve, etc., based on the detection results of the sensor 12, thereby controlling the injection amount of the target material.
[0019] Incidentally, when an instruction to switch control tasks is given, the control system 1 can adjust the magnitude and occurrence timing of state changes acting on the plant P when the control target changes accordingly.
[0020] Note that the control unit 13 in some embodiments of the present disclosure includes a neural network NN (referred to as the NN of the control unit 13) that is trained to ensure the safety of the plant P that is the control target. The NN of the control unit 13 generates a control amount (action a) for the adjuster 11 based on a state s of the plant P. The state of the plant P includes various data such as the liquid volume, pressure, and concentration detected by the sensor 12. The control unit 13 controls the adjuster 11 based on the detection results of the sensor 12.
[0021] Next, the configuration of the control unit 13 during learning of the NN and the learning will be described.
[0022] The NN of the control unit 13 in some embodiments of the present disclosure acquires, through offline learning, a control rule determined through learning, before controlling the actual plant P. This adjusts the NN to a state where it can function as the control unit 13. This learning of the NN of the control unit 13 is performed by reinforcement learning by the calculation processing device 10. Note that the NN of the control unit 13 does not restrict the acquisition of the vibration control rule, for example, readjustment of its response characteristics, after being applied to the actual plant P. The calculation processing device 10 when applied to such reinforcement learning is an example of a learning device.
[0023] Some embodiments of the present disclosure will be described with reference to Fig. 2. Fig. 2 is a configuration diagram of a processing device 10 according to some embodiments of the present disclosure.
[0024] The control unit 13 includes, for example, a data acquisition unit 131, a mode control unit 132, a control execution unit 133, a probability distribution estimation unit 134, a probability distribution approximation unit 135, a balance adjustment unit 136, and a value function adjustment unit 137. The data acquisition unit 131 acquires various data from outside and adds the acquired data to the storage unit 15.
[0025] The mode control unit 132 switches the operation mode, which specifies the processing to be performed by the control unit 13, to a selected one from a learning preparation mode, a learning mode, and a control execution mode. The learning preparation mode provides a state in which sample data for offline learning is collected and added to dataset D. The learning mode provides a state in which a probability distribution is generated by performing estimation processing based on the data, mainly using the collected and accumulated data of dataset D. This learning mode may further perform processing such as approximating the generated probability distribution and adjusting the value function. The control execution mode provides a state in which actual control is executed after the learning processing in the learning mode is completed.
[0026] When the control execution mode is selected, the control execution unit 133 executes actual control of the plant P, which is the control target. By this control, the control execution unit 133 acquires a control target from a higher-level device for the plant P, which is the control target, and adjusts each unit so as to execute control according to the control target. The control execution unit 133 executes sequential actions a t The control amount based on this result is output. Details of these will be described later.
[0027] The probability distribution estimation unit 134 performs a process of estimating a probability distribution based on the data of the accumulated data set D.
[0028] The probability distribution approximation unit 135 generates a probability distribution (second probability distribution) by converting the probability distribution (first probability distribution) generated by the estimation process by the probability distribution estimation unit 134. The probability distribution approximation unit 135 may include processing such as nonlinear mapping transformation in the conversion to the second probability distribution.
[0029] The balance adjustment unit 136 performs a balancing process to prevent bias in the probability distribution estimated by the probability distribution estimation unit 134. More specifically, when multiple types of processing are performed, the balance adjustment unit 136 adjusts the processing so that the same processing is not performed consecutively.
[0030] Value function adjuster 137 adjusts the gradient of value function V in the state space. For example, probability distribution estimator 134, probability distribution approximator 135, balance adjuster 136, and value function adjuster 137 are an example of a calculation processor that executes learning processing in the learning mode.
[0031] FIG. 3A is a diagram illustrating an application example of some embodiments of the present disclosure. The time chart shown in this figure includes control of a chemical plant (plant P) transitioning from the start of operation (startup) to a first steady state, control of transitioning from the first steady state through an unsteady state associated with a load change to a second steady state, control of transitioning from the second steady state through an unsteady state associated with a change in concentration designation to a third steady state, and control associated with shutdown. In the first steady state, production is planned with a target production volume of 100% and a product concentration of 95%. In the second steady state, production is planned with a target production volume of 80% and a product concentration of 95%. In the third steady state, production is planned with a target production volume of 80% and a product concentration of 99%. The target ratios and switchover counts described above are merely examples and are not limited thereto and can be changed as appropriate. As such, during each steady state, a large amount of data distributed around the target value is generated and accumulated as history data. By repeating such operations, more data is generated, and the density of data around the control target in the state space and the action space increases.
[0032] In this way, by utilizing the distinction between steady states and non-steady states, it is possible to classify data based on task switching requests from a higher-level device. In some embodiments of the present disclosure, when there are discontinuous sections due to changes in the operating state of the control system 1, it is preferable to divide the data into several tasks, each of which is a unit. In some embodiments of the present disclosure, when there are discontinuous sections due to changes in the operating state of the control system 1, it is preferable to divide the data into several tasks, each of which is a unit.
[0033] (Regarding Learning Value Functions) Referring to FIG. 4, learning a value function in some embodiments of the present disclosure will be described. FIG. 4 is a diagram for explaining learning a value function in some embodiments of the present disclosure. This diagram shows a model of a general reinforcement learning procedure. It is assumed that events progress from the top to the bottom of this diagram. At the top, "state s k-1 " shows the state with the oldest time history in this diagram. "State s" in this diagram k-1 If you check the " column from top to bottom, you will see that the top row "State s k-1 " followed by "state s k ", "State s k+1 ", "State s k+2 " are lined up. "State s k From the left end of the same row, "Policy function π", "Action a t ", environment f, state s, reward function γ, reward R t " are arranged in order. The top row and "state s k Between the two columns, "Value Function V" and to the right of this, "Value V (s k-1 ) are arranged in a row. k-1 " is related to "value function V", and "value V(s)" is calculated by "value function V". k-1 ) is generated. k-1 " is related to "policy function π" and "environment f". t " is an example of behavioral cloning. The same explanation as above can be applied to the other stages.
[0034] (Regarding pre-training for value function) In general reinforcement learning, learning the value function is performed before estimating the reward function. This will be explained for comparison. In the reinforcement learning of the comparative example, learning progresses as follows: state => reward function => value function => policy function (=> action). The value function V(s) of the comparative example is shown in the following equation (1). The value function V(s) is an example of a state value function (state).
[0035]
[0036] In the above formula (1), E (s,a)~π [·] indicates the value of state s expected under policy π when action a is performed in state s. The value in [·] is the cumulative reward. In this case, the reward function r(s, a) for the combination of state s and action a at time t and its discount rate γ are used to obtain the state trajectory from the accumulated reward at each time t. The discount rate γ is a coefficient that estimates the influence of reward r at each time t, and is, for example, a positive real number less than 1 (0<γ≦1). By determining this discount rate γ, the influence of reward expected in the distant future can be adjusted. For example, trajectory data for the entire range of time t is obtained from the trajectory data at each time t, and a value function V(s) is defined based on the distance to this trajectory data. According to this formula (1), the value function V(s) is defined using the reward function r(s, a). Also, according to FIG. 4, when action a t Going back to the value function V(s) from t It is easier to find the value function V(s) than to find the reward function r in order to define such a relationship as an inverse function.
[0037] Database D records information on "state => action => next state" for each stage. The elements of database D are shown in formula (2). Each stage may be identified using the identification index "k."
[0038]
[0039] According to this formula (2), a pair of state s and action a, or a pair of action a and state s, at each stage is associated with each other and recorded as time history information of the stage identified by the identification index "k." If a stage (or time point) is arbitrarily determined from among the stages of this database D, the state probability at that stage (or time point), the transition probability from state s to state s, the probability that action a will be selected in state s, and the like can be derived from the elements of database D. The state probability, the transition probability from state s to state s, and the probability that action a will be selected in state s are respectively shown in the following formula (3).
[0040]
[0041] However, it is not easy to comprehensively set the elements of database D, and there are cases where only a small portion of the entire selectable space is included. The relationship between state set S indicating the selectable state space and state set S bar of state s bar actually stored in database D is shown in the following formula (4). Note that characters with a line above them are indicated by adding a "bar" to the character. The relationship between action set A indicating the selectable action space and action set A bar of action a bar actually stored in database D is shown in the following formula (5).
[0042]
[0043]
[0044] In this way, state set S and action set A contain state s and action a, respectively, other than the elements recorded as elements of database D. Note that if the number of recorded elements is small compared to the number of elements that can be selected from the entire space, the selection of state s and action a will be limited.
[0045] Therefore, the range of the selected state s and action a is expanded to the range of state set S and action set A. As a result, unrecorded state s (i.e., s bar) and action a (i.e., a bar) can also be used as long as they are included in the range of state set S and action set A.
[0046] Incidentally, if the data to be analyzed is "business data" used in business processing, it is unlikely that "poor operation results" will exist in the operation history for business processing. According to such "business data," as shown in formula (6), the probability P(s) of state s in which bad action a is selected is bad ) is the probability P(s) of state s in which good action a is selected. good ) is significantly less than
[0047]
[0048] However, the state s required for calculating the value function V is not necessarily recorded. For example, in the case of a task with a production volume of 95% or 100%, the numerical value of the target state, such as "0.95" or "1.00", is often not recorded.
[0049] In contrast, air traffic control systems that expect no accidents to occur record almost no collisions, meaning that only ideal situations are recorded.
[0050] (Definition Using "Policy Function" and Probability) Hereinafter, the definition of the "value function" using the "policy function" and probability will be described. FIG. 5 is a diagram for explaining the definition of the "value function" according to some embodiments of the present disclosure. In FIG. 5A, when the state is state s 0 , s 1 , s 2 , s 3 In this case, the selected action a 0 , a 1 , a 2、 a 3 and the reward obtained in each state r 0 , r 1 , r 2 , r 3 State s k In this case, the "value function V" when the policy π is shown is shown in equation (7).
[0051]
[0052] The right side of this formula (7) differs from the above formula (1) in the following respects. For example, action a k:∞ indicates the action from the stage identified by the identification information "k" to the future. π (a t |s t-1 ) is the state s at time t t-1 In action a t In other words, this policy distribution indicates a frequency distribution based on the frequency of the states. k+1:∞ indicates a state from the stage identified by the identification information "k+1" to the future. f (s t+1 |st, a t ) is the state s at time t t and Action A t In the set of t+1 In other words, this policy distribution is t and Action A t State s in the set t+1 In the above formula (7), E[·] indicates the value of state s expected under the policy π when action a is performed in state s, just like in the above formula (1). Note that action a at time t k:∞ is the policy distribution P at time t π (a t |s t-1 ) is determined based on the state s at time t. k+1:∞ is the policy distribution P f (s t+1 |s t, a t Based on this condition, the value function V(s) is defined. Details of this will be explained below.
[0053] State S k In this case, the next action to be taken is a k This is simply called "next action a k " Next action a k The expected value Q when subsequent actions are selected by sampling according to the policy π (s k , a k ) is shown in equation (8).
[0054]
[0055] Comparing the above formula (7) and formula (8), the left sides are different. We will explain in order using this formula (8). Note that the state s in the reward function r(s, a) on the right side of this formula (8) k , action a k is a variable designated by the identification information "k+t".
[0056] By the way, the next action a kThe expected value when is specified regardless of the policy is shown in equation (9). In addition, the second and subsequent terms on the right-hand side are rewritten using the value function V.
[0057]
[0058] When comparing this equation (9) with the equation for the value function V, it can be seen that the first steps shown on the right-hand side are different. k The value function when follows the policy is shown in equation (10), and the policy function is similarly shown in equation (11).
[0059]
[0060]
[0061] Here, for example, instead of using equation (10) for estimating the value function from the data of dataset D, the above relationship can be regarded as equation (12) since the value of the reward is unknown. In other words, the policy distribution can be regarded as a normalized Q function. The softmax term in equation (12) aims to normalize the expected cumulative reward. Q on the left side of equation (12) π (s k , a) is an example of the state-action value function Q(state, action) in reinforcement learning. P on the right side of Equation (12) π (a|s k ) is an example of a policy distribution P(action|state) in reinforcement learning. The policy distribution P(action|state) is the probability that, as the state transitions through multiple stages, the state is in state s at a certain stage k, for example, and action a is selected in that state s.
[0062]
[0063] In this way, it is best to think that "normalization is necessary because the reward function does not take normalization into account." As a result of these considerations, we arrived at the following policy (Policy 1).
[0064] Policy 1: "Consider the policy distribution of offline data as a Q function." In other words, the policy distribution P of reinforcement learning shown in Equation (12) π (a|s k ) is the state s at a certain stage as the state transitions through multiple stages. kThis relationship is shown in the following equation (13):
[0065]
[0066] If the number of data registered in dataset D is small, it is advisable to consider methods such as adding noisy data as data in dataset D, or adding data with noise added to normal value data as data in dataset D. For example, when estimating the distributions of the value function and the policy function, the control unit 13 may add noise to either or both of the state and the action, and apply regularization that reduces the frequency of states or actions to which noise has been added.
[0067] (Method for estimating value from a limited number of state-action data) Next, an example of estimating value from a limited number of state-action data will be described with reference to Fig. 6. Fig. 6 is a diagram for explaining an example of estimating value from a limited number of state-action data according to some embodiments of the present disclosure.
[0068] Policy 2: "Consider the policy distribution of offline data as a Q function." In other words, the state value function V (state) of reinforcement learning is set using the TD (Temporal Difference) error of reinforcement learning. This relationship is shown in the following formula (14). In V(s) of formula (14), for example, when state s is identified by identification information "k", V(s k ) is the state s k This is an example of a state value function V (state) in
[0069] For example, if the target system is applied to a specific business, the usage pattern will be one in which there is little variation in the observed state s and the selected action a.
[0070] In Figure 6(a), the state-action data is allocated to a two-dimensional plane with the components of the action space and the state space as axes. Figure 6(b) shows the distribution during the startup task and the modified driving task at startup. Figure 6(c) shows the distribution during 90% steady driving. Figure 6(d) shows the distribution during 95% steady driving. Figure 6(e) shows the distribution during 100% steady driving.
[0071] The axes of each figure in Figure 6 are the same, with the horizontal axis representing the distribution of the action space and the vertical axis representing the distribution of the state space. The distributions shown in each figure are distributions on a two-dimensional plane of the action space and the state space.
[0072] For "business data" used in this type of application, there will be many records of well-operated states, and few records of poorly operated states. Therefore, the distribution results shown in the figure are often highly correlated. As mentioned above, the distribution trends will differ depending on the operating conditions (type of task).
[0073] As mentioned above, the state probability P(s) may not follow a strict Gaussian distribution. In such cases, for example, by dividing the analysis target range along the time axis, different types of tasks can be separated, and the observed state s k and selected action a k Separating different types of work (tasks) is an example of task division.
[0074] Alternatively, it is advisable to determine whether the data distribution can be considered to be a Gaussian distribution or a mixed Gaussian distribution. If it is possible to do so using the above method, it is advisable to follow this. If it is difficult to do so using the above method, it is advisable to nonlinearly map the original data using an autoencoder or the like and determine whether it is possible to consider the data to be a Gaussian distribution in the mixed space generated by this mapping. If the desired distribution can be obtained by using such a nonlinear mapping, it is advisable to use this.
[0075] However, if the performance of each individual task is unknown, it may be difficult to perform classification associated with each individual task.
[0076] For example, each task that specifies a control target is executed to perform specialized operations such as "steady operation at 95% production" or "changed operation to 100% production." Even for tasks with such conditions, the target values of each individual task are often not recorded as additional information in the data. This makes it difficult to classify the performance of each task.
[0077] In contrast, classification becomes easier by associating and recording information that can identify task switching points or information that allows each task to be classified. It is recommended to use classification rules that can be applied to such classification. If it is not possible to use information that facilitates classification, clustering can be performed based on state distribution.
[0078] In the above method, if the frequency distribution shows a Gaussian distribution or can be regarded as a Gaussian distribution, the frequency distribution can be regarded as a value. This will be explained below.
[0079] More specifically, as shown in equation (15), the distribution of the Q function is considered to follow the policy distribution. Also, as shown in equation (16), the distribution of the value is considered to follow the distribution of the value function. V(s) when the state s in equation (16) is identified by the identification information "k" is k ) is an example of a state-value function V (state) in reinforcement learning.
[0080]
[0081]
[0082] The above equation (15) is derived from the above equations (8) and (9). Note that the right side of equation (15) is k It is preferable to approximate it as a probability distribution determined for each. An example of this is shown in the following equation (17).
[0083]
[0084] This is because if the denominator and numerator on the right side of the above-mentioned equation (15) are calculated separately, the estimation error may increase.
[0085] Furthermore, in the optimal solution of the Q function, the relationship of the following equation (19) can be derived by utilizing the precondition shown in the following equation (18).
[0086]
[0087]
[0088] The right-hand side of the formula (19) in the brackets is the instance (s k , a k , s k+1 ) where the possible values of equation (19) are restricted to the following equation (20).
[0089]
[0090] (Method of Estimating Distributions) Next, estimation of the policy distribution and the distributions of the value function and reward function will be described.
[0091] As input information, a data set D and a discount rate γ (0<γ≦1) are used. As output information, a policy distribution and distributions of a value function and a reward function are output.
[0092] First, the policy distribution and the value function distribution are fitted from the data. The reward function distribution is calculated using the following equation (22) from the data extracted from dataset D (see equation (21)), recorded, and fitted again.
[0093]
[0094]
[0095] Note that a feature of some embodiments of the present disclosure is the use of the following equation (23). In this case, the TD error δ is expressed as equation (24), and by using this, the above equation (23) is converted to equation (25). Ideal strategy A π is expressed as the following equation (26), which can be calculated using the value of the Q function.
[0096]
[0097]
[0098]
[0099]
[0100] (Method for approximating state distribution) A method for approximating state distribution will be described with reference to FIG. 7. FIG. 7 is a diagram for explaining a method for approximating state distribution according to some embodiments of the present disclosure. The method for approximating state distribution described here uses a nonlinear mapping from the state space to make the distribution in the latent space a Gaussian distribution or a mixed Gaussian distribution. The nonlinear mapping nonlinearly maps from the state space to a one-dimensional latent space R. This mapping is shown in Equation (27). Note that the distribution in the latent space is considered to be, for example, a standard normal distribution N(0,1), and is expressed as a probability value between 0 and 1. For example, this distribution is shown in Equation (28).
[0101]
[0102]
[0103] An example of a method for approximating the state distribution is given below. First, the parameter u of the nonlinear mapping g is updated using a gradient ascent method or the like to approximate the distribution. The gradient ascent method is shown in the following equation (29). This instance is the same as the minimization of equation (30). The relationship for this minimization is shown in equation (31).
[0104]
[0105]
[0106]
[0107] An approximation method using a nonlinear mapping g has been explained. Furthermore, an autoencoder constraint may be added to the dimensionality reduction. Adding or removing this constraint can be determined arbitrarily. When an autoencoder is used, the state characteristics are preserved even in a low-dimensional space. The "nearness of the state" is memorized by learning the nonlinear mapping g. An example of an autoencoder is shown in Equation (32).
[0108]
[0109] (When applying the policy gradient method to the policy distribution) When applying the policy gradient method to the policy distribution, it may not be possible to set valid values for the coefficients in the computational formula for the policy gradient method. The computational formula for the policy gradient method is shown in Equation (33). This corresponds to the case where the value of the coefficient A in this formula is, for example, 1.
[0110]
[0111] In such a case, it is advisable to add noise to the data set D. The probability value of this noise can be defined by equation (34).
[0112]
[0113] It is advisable to make the probability value of the noise small. For example, the recorded state s t Noise s in the vicinity of noise To add t Then, noise ε is added to the state s t Noise s based on noise This noise s noise is shown in equation (35).
[0114]
[0115] In this case, state s t Noise s in the vicinity of noise Since there exists a state s t In areas where the density is relatively high, noise s noise The influence of state s is relatively small. t The density of t In the region around the noise s noise In order to adjust for such differences in influence, the S / N ratio (see Equation (36)) and noise width σ 2 (see equation (37)) may be used as a hyperparameter for restricting the magnitude of the eigenvalue within a predetermined range.
[0116]
[0117]
[0118] Kernel density estimation is considered to be neural approximation. The calculation formula for kernel density estimation is shown in Equation (38). An example of the kernel function in this case is shown in Equation (39).
[0119]
[0120]
[0121] An example of approximating a density function by holding the sample instances shown in Equation (40) is shown in Fig. 8. Fig. 8 is a diagram for explaining approximation of a density function according to some embodiments of the present disclosure.
[0122]
[0123] For example, one instance s i , u is updated once by the operation of equation (41). In this case, it is assumed that one kernel is added. However, the added instance s i The effect of may not be uniform. Note that the amount of change in the output of each function per update will differ depending on the model and the current parameters.
[0124]
[0125] (About behavioral cloning using commercial dataset D)
[0126] The evaluation criterion for reward r is defined using the following equation (42). Reward r is converted into a real number set R formed by the product of state set S and action set A. Value function V π (s k ), the value of the Q function Q π (s k , a k ) (referred to as Q-function), and TD error δ (referred to as TD-error). Note that when the TD error δ is optimized to be 0, the value function V π (s k ) can be approximated using the following formula:
[0127]
[0128] The optimal policy at runtime is determined using the Q function. For example, the policy function P π (a t |s t ) is shown in equation (43). If the Q function is normalized, the policy function P π (a t |s t ) to the normalized Q function value Q π (s k , a k ) is defined using the Q function value Q π (s k , a k ) takes values between 0 and 1. The value of the Q function, Q π (s k , a k ) sums to 1. In this situation, we can use an approximate reward and γ.
[0129]
[0130] As described above, in the case of business data set D, it is expected that the actions and states contained therein, which are empirical values, will contain a relatively large number of cases in which the correct action was selected and a transition to a good state occurred as a result. Therefore, the frequency of actions and states can be considered desirable.
[0131] It is desirable that the following probabilities are proportional to each other. For example, the policy function P π (a t |s t ), value function V π (s k ) and the state function P(s t ) should be proportional to each other.
[0132] From the above, the reward r(S t , a t ) is shown in the following equation (44).
[0133] The policy function P in the above equation (44) π (a t |s t ) and the state function P(s t+1 ) can be approximated using the elements of the data set D as described above. This allows us to calculate the reward r(St , a t ) can be predicted.
[0134] (Control for Balancing Processes) An adjustment technique for balancing processes will be described with reference to Figs. 9A to 9C. Figs. 9A to 9C are diagrams for explaining an adjustment technique for balancing processes. Fig. 9A shows a state s 0 , s 1 , s 2 , s 3 In this case, the selected action a 0 , a 1 , a 2 and the reward obtained in each state r 0 , r 1 , r 2 , r 3 Shows.
[0135] The graph shown in FIG. 9B shows an example of a probability distribution. Two distributions are shown within this graph. The first graph shows a case where the samples are relatively dispersed, and the second graph shows a case where the samples are relatively concentrated. Some embodiments of the present disclosure utilize the distribution trend characteristics in this way. When the samples are relatively concentrated, as in the second graph, the distribution has a median or central value that is easy to determine. The graph shown in FIG. 9C shows how much discounting is performed to estimate future reward r depending on the value of the discount rate γ (denoted as gamma in FIG. 9C ). The discount rate γ shown here defines seven levels ranging from 1 to 8. For example, when the discount rate γ is 1, there is no discounting, and the graph appears as a nearly horizontal line. On the other hand, when the discount rate γ is less than 1, the amount of discounting is determined by the value of the discount rate γ.
[0136] In order to balance the calculation of the value V and the reward r, for example, the following estimation processes are performed alternately. This results in suitable parameters. The first estimation process is a process related to estimating the reward value, and the second estimation process is a process using TD (Temporal Difference) error, for example, a process for minimizing a specified TD error. This optimization process is not limited to a minimization process, and may be a maximization process depending on the calculation formula used. Each process will be explained below in order.
[0137] Reward Value Estimation: Reward value estimation involves estimating a probability distribution. To estimate the reward value, it is advisable to minimize the sample energy of the distribution function.
[0138] The calculation formula for deriving the value of the Q function is shown in formula (45). π is defined as in equation (46).
[0139]
[0140]
[0141] In the above formula, D p is the set of positive column samples that are the samples recorded in the data set D, and D n denotes the set of negative column samples, which are samples not recorded in the dataset D. p is the number of positive column samples (number of positive column samples), and K n is the number of negative column samples (number of negative column samples). λ is the negative column ratio, which takes a value between 0 and 1. The values of the above variables may be determined appropriately. As an example, the number of positive column samples K p and the number of negative column samples K n may be set to the same number and the negative column ratio λ may be set to 0.5.
[0142] Minimizing the TD error: The reward value r is calculated from the current value of the Q function and the value of the value function V using the above equation (44). The result of this calculation is treated as a constant. For example, in at least three state transitions shown in FIG. 9A, the state s 0 From state s 3In this case, the result obtained from the evaluation formula (formula (47)) at the final step (t=3) is minimized.
[0143]
[0144] This equation (47) is the constraint that balances the value V and the reward γ. t+1 ) or the correct value of the above equation (47) is treated as a constant, and the result obtained from the following evaluation equation (equation (48)) is minimized.
[0145]
[0146] This equation (48) becomes the constraint that balances the value V and the value of the Q function.
[0147] (Optimization Process) The optimization process according to some embodiments of the present disclosure will be described with reference to Fig. 10. Fig. 10 is a diagram for explaining the optimization process according to this embodiment.
[0148] The input for this optimization process is a dataset D (dataset D={(s 0 , a 1 , s 1 , a 2 , ..., a T , s T ) k} K k=0 This dataset D contains state and behavior data for a total of K episodes. Each episode is identified by an episode index k. t is the state at time t. t+1 is the state s at time t t This is the action chosen when
[0149] The output of this optimization process is the state probability P(s t ) and the action probability P(a t |s t-1 ) and
[0150] An example of the algorithm for this optimization process is shown below: (1) Prepare a function with parameters (equation (49)). (2) Sample a set of data from data set D. This process is shown in equation (50).
[0151]
[0152]
[0153] In addition, the behavior a obtained from the above sample t+1 It is recommended to use this as the control output.
[0154] (3) The current log probability value is calculated using equation (51).
[0155]
[0156] The measure π in the above formula (51) is the measure π shown in formula (52). μ and strategy π σ Includes:
[0157]
[0158] (4) Update the function parameters to increase the probability value of the sample. For example, the function parameters include d v (s t ), d π (a t |s t-1 ), θ v , and θ μ、 θ σ Includes:
[0159]
[0160] (5) Points around the sample are regarded as negative examples, and the function parameters are updated to reduce the probability values of negative example points as shown in equation (54).
[0161]
[0162] Repeat steps (2) to (5) above.
[0163] (Regarding the gradient of the value function in the state space) The gradient of the value function in the state space will be described with reference to FIGS. 11 and 12. FIG. 11 is a diagram for describing the gradient of the value function in the state space of some embodiments of the present disclosure. As shown in FIG. 11, a state s and a policy π are mapped onto a two-dimensional plane including a time axis t (horizontal axis) in the environment and an axis indicating the update history. The data string of state s and policy π shown in the top row of this two-dimensional plane is when the update history index is 0, the data string in the next row is when the update history index is 1, and the data string in the next row is when the update history index is 2. The update history index is 0, 1, 2, and the larger the value, the more repeated updates are indicated. In the above single data string, states s and policies π are arranged alternately. The horizontal axis indicates that time t has passed as it moves to the right. The time axis may be defined from the origin 0 to any time τ. For example, the "state s" shown on the left side of the top row 0、0 " indicates the state when the update history is 0 and the time t is 0, and "State s 1、0 " indicates the state when the update history is 1 and time t is 0. The second one from the left in the top row, "Policy π θ0 (s 0、0 ) is the state s when the update history is 0 and the parameter θ is 0 and the time t is 0. 0、0 The fourth one from the left shows the value of the policy function at θ0 (s 0、1 ) is the state s when the update history is 0 and the parameter θ is 0 and the time t is 1. 0、1 The value of the policy function at is shown below.
[0164] The gradient of the value function V in the state space is determined using the value function and an inverse model.
[0165] The following shows a process for obtaining the gradient of the value function V using this value function V and the inverse model. The input for this process is the inverse model P(a t |s t+1 , s t ) and the value function V(s t ) The output of this process is the distribution of policies P(a t|s t )
[0166] First, we will explain the case where a change occurs in the time axis direction within a given environment. It is assumed that the value function is optimized. In this case, the TD error δ becomes 0, as shown in equation (55).
[0167]
[0168] ε is a variable that takes a real value. When this variable ε is used, equations (55) to (56) hold.
[0169]
[0170] According to the above equation (56), the state S t The state after a small change Δs occurs from t+1 ) in state S t It has been shown that the value can be derived using the same value function as
[0171] Using equation (57) including the environment function f, the state s t+1 The policy function π is defined as follows: if it is optimized, then the state s t+1 The value V in state s t The value is the sum of the value V at the time of the calculation and ΔV. ΔV is a real number greater than or equal to 0.
[0172]
[0173] Equation (58) holds true if the small change Δs is not 0. Equation (59) is obtained by normalizing both sides of this equation (58).
[0174]
[0175]
[0176] Usage: In the learning process, the constraint of the above formula (59) is introduced. That is, this formula is equivalent to the minimum value indicated by formula (60).
[0177]
[0178] In the learning process, the constraint of the above equation (59) is introduced. That is, this equation is equivalent to the minimum value indicated by equation (60). For example, if equation (61) holds, updating of either or both of the parameters θv and θπ may be canceled.
[0179]
[0180] 12 is a diagram for explaining the gradient of the value function in a state space showing the relationship between the parameter θv and the state s according to some embodiments of the present disclosure. v The vertical axis shows the change in state s. t , parameter θ v The value function V(s t , θ v ), the state s t The gradient of the value function V at a given time can be specified for each axis component. t is state s t+1 , the gradient of the value function V may be corrected to reflect the influence of the state change.
[0181] Processing by the arithmetic processing device according to some embodiments of the present disclosure will be described with reference to Fig. 13. Fig. 13 is a flowchart illustrating processing by the arithmetic processing device according to some embodiments of the present disclosure.
[0182] The data acquisition unit 131 receives a user operation performed on the input / output unit 16 (S11). The mode control unit 132 identifies the control mode requested by the user and assigns it to the next process (S12). The learning preparation mode, learning mode, and control execution mode are examples of modes related to the next process.
[0183] If the request is for the learning preparation mode, the data acquisition unit 131 acquires a pair of information on the environment s and the action a (S21), assigns identification information to the information, and adds the information on the environment s and the action a to the data set D. This data acquisition may be performed, for example, by acquiring the information on the environment s and the action a at a predetermined interval, or by acquiring all data that has already been stored. This process ends when a predetermined number of data items have been acquired.
[0184] If the request is for the learning mode, the probability distribution estimation unit 134 performs a process of estimating a probability distribution based on the data of the accumulated dataset D. By estimating this probability distribution, the probability distribution estimation unit 134 derives, for example, a policy distribution and distributions of a value function and a reward function (S31).
[0185] Next, if the probability distribution (first probability distribution) generated by the estimation process by the probability distribution estimation unit 134 is dissociated from a Gaussian distribution, the probability distribution approximation unit 135 generates a probability distribution (second probability distribution) by converting the probability distribution (first probability distribution) generated by the estimation process by the probability distribution estimation unit 134 (S32). The process of step S32 may be performed as needed.
[0186] Next, the balance adjustment unit 136 may execute a balance process in which the items for the estimation process are alternately selected so as to prevent bias in the probability distribution estimated by the probability distribution estimation unit 134 (S33).
[0187] Next, the value function adjustment unit 137 adjusts the gradient of the value function V in the state space (S34). This step S34 may be performed as needed. The control unit 13 terminates this offline learning process once each variable has been determined through the above-described processes. In the learning process, the control unit 13 (arithmetic processing unit) uses an internal evaluation function in reinforcement learning to calculate at least one of an evaluation value for an action in a state in an environment including the controlled object and an evaluation value for the next state of that state. The control unit 13 determines an action for the controlled object based on the calculated evaluation value. For example, the internal evaluation function may be a value function in reinforcement learning. In this case, the control unit 13 may configure the value function for reinforcement learning using trajectory data indicating a frequency distribution based on the frequency of states. Alternatively, in this case, the control unit 13 may configure the policy function for reinforcement learning using trajectory data based on a conditional probability depending on the state of the action determined in reinforcement learning. Examples of trajectory data here include the curve in the distribution diagram shown in FIG. 6A and the curve showing the distribution after mapping shown in FIG. 7. For example, the value function in reinforcement learning may be a function based on trajectory data showing a frequency distribution based on the frequency of states. In this case, the control unit 13 may configure the policy function in reinforcement learning using trajectory data based on conditional probabilities depending on the state of the behavior determined in reinforcement learning.
[0188] If the request is for the control execution mode, the control execution unit 133 uses the results of the previously executed learning process to control the plant P. After operating the control system 1 for the required period, the control system is stopped.
[0189] According to some embodiments of the present disclosure, the control device 10 (learning device) includes a processing unit that uses an internal evaluation function in reinforcement learning to calculate at least one of an evaluation value for an action in a state in an environment including a controlled object and an evaluation value for a next state of the state, and determines an action for the controlled object based on the calculated evaluation value. The internal evaluation function is defined to generate an evaluation value that indicates a higher evaluation for a state occurring more frequently in the desired action of the controlled object. This allows the control system 1 to reduce the amount of calculation required for reinforcement learning. Furthermore, the control system 1 is a processing unit for optimizing a control system including states, actions, and values as control variables, and includes a processing unit that configures a reinforcement learning value function or policy function for controlling the control system using trajectory data based on empirical values of states related to the control system. This allows the control system 1 to be realized with a configuration that can optimize a control system that requires safety control.
[0190] A more simply configured embodiment will be described with reference to Figures 14 and 15. Figure 14 is a diagram illustrating an example of the configuration of a processing device according to some embodiments of the present disclosure. In the configuration shown in Figure 14, the processing device 610 includes a processing unit 611. In this configuration, the processing unit 611 configures a value function or a policy function for reinforcement learning for controlling the control system 1 using trajectory data based on empirical values of states related to the control system 1.
[0191] 15 is a flowchart illustrating an example of a processing procedure of a computational method according to some embodiments of the present disclosure. The computational method illustrated in FIG. 15 includes constructing a value function or a policy function for reinforcement learning using trajectory data based on empirical values of states (step S610). This allows optimization of a control system that requires control safety.
[0192] 16 is a schematic block diagram illustrating the configuration of a computer according to at least one embodiment. In the configuration shown in FIG. 16, a computer 700 includes a CPU 710, a main memory device 720, an auxiliary memory device 730, an interface 740, and a non-volatile recording medium 750.
[0193] One or more of the above-described arithmetic processing device 10 and control system 1, or a part thereof, may be implemented in a computer 700. In this case, the operation of each of the above-described processing units is stored in the auxiliary storage device 730 in the form of a program. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program. The CPU 710 also allocates storage areas in the main storage device 720 corresponding to each of the above-described storage units in accordance with the program. Communication between each device and other devices is executed by an interface 740 having a communication function and performing communication under the control of the CPU 710.
[0194] When the arithmetic processing device 10 is implemented in a computer 700, the operations of the neural network device and each of its components are stored in the form of a program in an auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.
[0195] Furthermore, the CPU 710 allocates a storage area for processing by the arithmetic processing device 10 in the main memory device 720 in accordance with the program. Communication between the arithmetic processing device 10 and other devices is performed by the interface 740, which has a communication function and operates under the control of the CPU 710. Interaction between the arithmetic processing device 10 and a user is performed by the interface 740, which has a display device and an input device, displaying various images under the control of the CPU 710 and accepting user operations.
[0196] It is also possible to record a program for executing all or part of the processing performed by the arithmetic processing unit 10 on a computer-readable recording medium, and have the computer system read and execute the program to perform the processing of each unit. Note that the term "computer system" here includes hardware such as the OS and peripheral devices.
[0197] Furthermore, "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs (Read Only Memory), and CD-ROMs (Compact Disc Read Only Memory), as well as storage devices such as hard disks built into computer systems. The program may be one that realizes part of the functions described above, or may be one that can realize the functions described above in combination with a program already recorded in the computer system.
[0198] Although several embodiments of the present disclosure have been described in detail above with reference to the drawings, the specific configurations are not limited to these embodiments, and designs within the scope of the gist of this disclosure are also included.
[0199] FIG. 3B is a diagram illustrating another application example of some embodiments of the present disclosure. The time chart shown in this figure illustrates the procedure for restoring operational status in railway traffic management. For example, a train is operating according to the timetable when an accident or other reason causes a service interruption. After this interruption, the train undergoes a traffic rescheduling phase and then returns to normal operation. Such an interruption may occur again. As described above, when the operational status changes, control is implemented to stabilize the state during unsteady state control. While states during such unsteady state control tend to have an indeterminate trend, applying some embodiments of the present disclosure enables stable control. Some control systems experience internal state changes during operation, and even if the same input is received, the output is determined by the changed internal state. For example, according to a reinforcement learning technique in a comparative example, when optimizing a controller that controls a controlled object to match the state of the controlled object, a calculation to search for a more optimal solution is performed for each control cycle that controls the controlled object. Applying such a comparative example would require a long time to obtain an optimal solution. In contrast, according to the reinforcement learning techniques of some embodiments of the present disclosure, optimal solutions corresponding to each internal state can be prepared in advance through offline learning, making it possible to perform control without the time required in the comparative example.
[0200] Some or all of the above embodiments may be described as, but are not limited to, the following notes. (Note) [1] One aspect is a learning device including a processing unit that uses an internal evaluation function in reinforcement learning to calculate at least one of an evaluation value for an action in a state in an environment including a controlled object and an evaluation value for a next state of the state, and determines an action for the controlled object based on the calculated evaluation value, wherein the internal evaluation function is defined to generate the evaluation value indicating that the higher the frequency of a state in a desired action of the controlled object, the higher the evaluation. [2] The internal evaluation function in the processing unit of [1] above may be a value function in reinforcement learning, and the processing unit may configure the value function for reinforcement learning using trajectory data indicating a frequency distribution based on the frequency of the state. [3] The internal evaluation function in the processing unit of [1] above may be a value function in reinforcement learning, and the processing unit may configure the policy function for reinforcement learning using trajectory data based on a conditional probability of an action determined in the reinforcement learning depending on the state. [4] In the processing device of [1] above, the value function in reinforcement learning may be a function based on trajectory data indicating a frequency distribution based on the frequency of the state, and the processing unit may configure the policy function for reinforcement learning using the trajectory data based on the conditional probability of the state of an action determined in the reinforcement learning. [5] In the processing device of [1] above, the policy distribution for reinforcement learning may be the probability that action a is selected in state s at a certain stage among multiple stages of state transitions, and the state value function for reinforcement learning may be set using a TD (Temporal Difference) error of the reinforcement learning. [6] In estimating the distribution of the value function and the policy function, the processing unit of [4] above may add noise to either or both of the state and the action, and apply regularization to reduce the frequency of the state to which the noise has been added or the action to which the noise has been added. [7] In the processing unit of [4] above, the control system may be controlled using the value function or policy function for reinforcement learning.[8] The arithmetic processing unit of [4] above may normalize the magnitude of the gradient of the value function of the reinforcement learning and use the normalized value as an initial value. [9] The arithmetic processing unit of [1] above may record the state of the environment when controlling the control object through an action determined based on an operation or command of an expert related to control of the control object, and obtain a frequency of the state based on the record.
[10] A control system of one aspect includes the learning device of any of [1] to [9] above and a control unit that has undergone offline reinforcement learning by the learning device based on the frequency of the state, and the control unit may control the control object using a result of the offline reinforcement learning.
[11] One aspect may be a learning method that includes using an internal evaluation function in reinforcement learning to determine at least one of an evaluation value for an action in a state in an environment including the control object and an evaluation value for a next state of the state, and determining an action for the control object based on the obtained evaluation value, and the internal evaluation function is defined to generate the evaluation value indicating that a higher evaluation is given to a state with a higher frequency in the desired action of the control object.
[12] One aspect may be a program for causing an apparatus that executes a program for a learning process related to reinforcement learning to use an internal evaluation function in reinforcement learning to determine at least one of an evaluation value for an operation in a state in an environment including a controlled object and an evaluation value for a next state of the state, and to determine an operation for the controlled object based on the obtained evaluation value, and to generate, using the internal evaluation function, the evaluation value that indicates that the higher the frequency of a state in a desired operation of the controlled object, the higher the evaluation.
[0201] This application claims priority based on Japanese Patent Application No. 2022-180039, filed November 10, 2022, the disclosure of which is incorporated herein in its entirety by reference.
[0202] The present disclosure may be applied to a learning device, a control system, a learning method, and a recording medium.
[0203] 1 Control system 10, 610 Processing device (learning device) 13, 613 Control unit 15 Storage unit
Claims
1. Using an internal evaluation function in reinforcement learning, at least one of an evaluation value for an action in a state within an environment including a control target and an evaluation value for a next state of the state is obtained, An arithmetic processing unit that determines an action related to the control target based on the obtained evaluation value is provided, The internal evaluation function is defined to generate the evaluation value indicating that the higher the frequency of the state in the desired action of the control target, the higher the evaluation. A learning device.
2. The internal evaluation function is a value function in the reinforcement learning, The arithmetic processing unit configures the value function of the reinforcement learning using trajectory data indicating a frequency distribution based on the frequency of the state. The learning device according to claim 1.
3. The internal evaluation function is a value function in the reinforcement learning, The arithmetic processing unit configures the policy function of the reinforcement learning using trajectory data based on the conditional probability of the action determined in the reinforcement learning by the state. The learning device according to claim 1.
4. The value function in the reinforcement learning is a function based on trajectory data indicating a frequency distribution based on the frequency of the state, The arithmetic processing unit configures the policy function in the reinforcement learning using trajectory data based on the conditional probability of the action determined in the reinforcement learning by the state. The learning device according to claim 1.
5. The policy distribution of the reinforcement learning is the probability that action a is selected in state s at a certain stage during the transition of the state through a plurality of stages, The state value function of the reinforcement learning is set using the TD (Temporal Difference) error of the reinforcement learning. The learning device according to claim 4.
6. The arithmetic processing unit in estimating the distribution of the value function and the policy function, adds noise to either or both of the state and the action, and applies regularization to reduce the frequency of the state with the added noise and the action with the added noise. The learning device according to claim 4.
7. The arithmetic processing unit controls a control system using the value function or the policy function of the reinforcement learning. The learning device according to claim 4.
8. The learning device according to claim 1, a control unit that is offline reinforcement learned based on the frequency of the state by the learning device, is provided, The control unit controls the control target using the result of the offline reinforcement learning. A control system.
9. Using an internal evaluation function in reinforcement learning, at least one of an evaluation value for an action in a state within an environment including a control target and an evaluation value for a next state of the state is obtained. Determining an action related to the control target based on the obtained evaluation value; The internal evaluation function is defined to generate the evaluation value indicating that the higher the frequency of the state in the desired action of the control target, the higher the evaluation; A learning method including the above.
10. In an apparatus that executes a program for a learning process related to reinforcement learning, Using an internal evaluation function in reinforcement learning, at least one of an evaluation value for an action in a state within an environment including a control target and an evaluation value for a next state of the state is obtained. Causing an action related to the control target to be determined based on the obtained evaluation value; Causing the internal evaluation function to generate the evaluation value indicating that the higher the frequency of the state in the desired action of the control target, the higher the evaluation; A program for causing the above to be executed.