Learning device, control device, control system, learning method, and recording medium

By using a Q-function with set upper and lower limit values for cumulative reward values and incorporating regularization, the learning device stabilizes reinforcement learning, addressing issues of unstable learning and overfitting.

WO2025120859A1PCT designated stage expired Publication Date: 2025-06-12NEC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/044096
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-08
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

In reinforcement learning, the stabilization of learning is challenging due to the small proportion of time steps showing significant reward values in the training data, leading to unstable learning and potential overfitting.

Method used

A learning device and method that utilize a Q-function with upper and lower limit values set for the cumulative reward value, combined with regularization techniques, to stabilize learning and prevent overfitting.

Benefits of technology

The proposed solution stabilizes learning by ensuring convergence of the Q-function value and preventing overfitting, thereby improving the accuracy and reliability of reinforcement learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023044096_12062025_PF_FP_ABST
    Figure JP2023044096_12062025_PF_FP_ABST
Patent Text Reader

Abstract

This learning device comprises a learning means that uses a function indicating a prediction value in which either an upper limit value or a lower limit value is set for a cumulative prediction value of a reward value which is an index value indicating the evaluation of an action to be implemented by a control target in an operation environment of the control target, that learns a policy which is a determination rule for the action, and that learns the function.
Need to check novelty before this filing date? Find Prior Art

Description

Learning device, control device, control system, learning method, and recording medium

[0001] The present invention relates to a learning device, a control device, a control system, a learning method, and a recording medium.

[0002] One type of machine learning is reinforcement learning (see, for example, Patent Document 1).

[0003] Japanese Patent Application Laid-Open No. 2020-187742

[0004] In reinforcement learning, it is desirable to be able to stabilize learning.

[0005] An example of an object of the present disclosure is to provide a learning device, a control device, a control system, a learning method, and a recording medium that can solve the above-mentioned problems.

[0006] According to a first aspect of the present disclosure, a learning device includes a learning means for learning a strategy, which is a decision rule for the behavior, and learning the function, using a function indicating a predicted value of the cumulative reward value, which is an index value indicating an evaluation of the behavior performed by the controlled object in the operating environment of the controlled object, with at least either an upper limit or a lower limit set for the predicted value.

[0007] According to a second aspect of the present disclosure, a control device includes a control means for controlling the control object by learning a policy, which is a decision rule for the behavior, using a function indicating a predicted value in which at least either an upper limit or a lower limit is set for a predicted cumulative value of a reward value, which is an index value indicating an evaluation of the behavior performed by the control object in the operating environment of the control object, and using the policy obtained by learning the function.

[0008] According to a third aspect of the present disclosure, a control system includes a control object and a control device, and the control device includes a control means for controlling the control object by learning a policy, which is a decision rule for the behavior, using a function indicating a predicted value in which at least either an upper limit or a lower limit is set for a predicted cumulative value of a reward value, which is an index value indicating an evaluation of the behavior performed by the control object in the operating environment of the control object, and using the policy obtained by learning the function.

[0009] According to a fourth aspect of the present disclosure, the learning method includes a computer learning a strategy, which is a decision rule for the behavior, and learning the function, using a function that indicates a predicted value of the cumulative reward value, which is an index value that indicates an evaluation of the behavior performed by the controlled object in the operating environment of the controlled object, with at least either an upper limit value or a lower limit value set for the predicted value.

[0010] According to a fifth aspect of the present disclosure, the recording medium is a recording medium having recorded thereon a program that causes a computer to learn a policy, which is a decision rule for the behavior, and learn the function, using a function that indicates a predicted value of the cumulative reward value, which is an index value that indicates an evaluation of the behavior performed by the controlled object in the operating environment of the controlled object, with at least either an upper limit or a lower limit set for the predicted value.

[0011] According to one aspect of the present disclosure, learning can be stabilized in reinforcement learning.

[0012] FIG. 1 is a diagram showing an example of the configuration of a learning device according to at least one embodiment; FIG. 2 is a diagram showing an example of a processing procedure in which a learning device according to at least one embodiment learns control of a control object; FIG. 3 is a diagram showing an example of a task performed by a control object according to at least one embodiment; FIG. 4 is a diagram showing an example of a difference in the variation in Q function values ​​due to differences in learning settings; FIG. 5 is a diagram showing an example of a difference in performance due to differences in learning settings; FIG. 6 is a diagram showing an example of the configuration of a control system according to at least one embodiment; FIG. 7 is a diagram showing another example of the configuration of a learning device according to at least one embodiment; FIG. 8 is a diagram showing another example of the configuration of a control device according to at least one embodiment; FIG. 9 is a diagram showing another example of the configuration of a control system according to at least one embodiment; FIG. 10 is a diagram showing an example of a processing procedure in a learning method according to at least one embodiment; FIG. 11 is a schematic block diagram showing the configuration of a computer according to at least one embodiment.

[0013] Hereinafter, embodiments of the present invention will be described, but the following embodiments do not limit the invention according to the claims. Furthermore, not all of the combinations of features described in the embodiments are necessarily essential to the solution of the invention. In the following, characters with an overline (overline) may be indicated by adding a superscript "-" to the character. For example, an overlined φ is represented as φ - It can also be written as:

[0014] First Embodiment Fig. 1 is a diagram illustrating an example of the configuration of a learning device according to at least one embodiment. In the configuration illustrated in Fig. 1, the learning device 100 includes a communication unit 110, a display unit 120, an operation input unit 130, a storage unit 180, and a processing unit 190. The processing unit 190 includes a rewriting unit 191 and a learning unit 192.

[0015] The learning device 100 learns control of a control object through reinforcement learning. The term "study" used here refers to updating parameter values ​​of a machine learning model based on training data. The term "learning" used here can also be referred to as "training."

[0016] The reinforcement learning referred to here is machine learning that learns a policy, which is a behavioral rule of an agent that takes an action in a certain environment, based on a state in the environment and a reward that represents an evaluation of the state or the action. The learning device 100 may be configured using a computer such as a personal computer (PC).

[0017] The controlled object here is not limited to a specific one, and can be a variety of objects that can be learned to control by reinforcement learning. For example, the controlled object here may be a stand-alone device such as an industrial robot or a machine tool. Alternatively, the controlled object here may be equipment such as a plant or a power plant, or a system such as a production line in a factory. Alternatively, the controlled object here may be a moving object such as an automobile, a railroad vehicle, an airplane, a ship, or a self-propelled mobile robot, or a transportation system such as a railroad or air traffic control system.

[0018] In addition, in learning by the learning device 100, the control target is assumed to execute a set task. The task here is not limited to a specific one, and can be any task for which success or failure can be determined. In the following, time is represented by time steps, and is represented as time 0, 1, .... The state s at time t is t When the controlled object performs the action a determined based on the policy π at time t, the state changes to state s at time t+1. t+1 Here, t is an integer t≧0.

[0019] The communication unit 110 communicates with other devices. For example, the communication unit 110 may receive training data from a device that stores the training data. The display unit 120 has a display screen, such as a liquid crystal panel or an LED (Light Emitting Diode) panel, and displays various images. For example, the display unit 120 may display data related to the learning performed by the learning device 100, such as the number of times an episode has been executed in reinforcement learning. An episode here refers to the period from the start to the end of task execution. An episode is also called a trial.

[0020] The operation input unit 130 includes input devices such as a keyboard and a mouse, and receives user operations. For example, the operation input unit 130 may receive user operations for setting the learning rate and other settings related to the learning performed by the learning device 100.

[0021] The storage unit 180 stores various types of data. For example, the storage unit 180 may store training data. The storage unit 180 may also store a machine learning model that is the target of learning performed by the learning device 100, such as a policy function and its parameter values. The storage unit 180 is configured using a storage device included in the learning device 100.

[0022] Processing unit 190 performs various processes by controlling each unit of learning device 100. The functions of processing unit 190 may be performed by a CPU (Central Processing Unit) included in learning device 100 reading and executing a program from storage unit 180.

[0023] The rewriting unit 191 rewrites the training data. In particular, the rewriting unit 191 rewrites at least one step of the training data into training data that represents the achievement of the goal of the task performed by the controlled object. Achieving the goal of the task is also referred to as the task being successful. The goal of the task is also referred to as the goal. The rewriting unit 191 is an example of a rewriting means.

[0024] In reinforcement learning, when a reward is obtained only when a task is successful, the proportion of time steps in the training data that show significant reward values ​​may be small. This small proportion of time steps that show significant reward values ​​may result in slow progress in learning. Specifically, this may mean that the task success rate does not improve much during learning, or that it takes a long time for the task success rate to improve.

[0025] In response to this, the rewriting unit 191 rewrites part of the training data with (part of) the training data when the task is successful, so that the reward value when the task is successful is indicated. This is expected to increase the proportion of time steps where significant reward values ​​are indicated, making it easier for learning to progress.

[0026] The method by which the rewriting unit 191 rewrites the training data is not limited to a specific method. For example, the rewriting unit 191 may rewrite the training data based on a Final Strategy, a Future Strategy, or a Random Strategy, but is not limited to this.

[0027] In the Final Strategy, the final state of the episode is treated as the state when the task is successful. In other words, in the Final Strategy, the task is treated as successful when its execution ends. In this case, the end of the task execution may also be when the task is aborted without being successful.

[0028] In the Future Strategy, a state randomly selected from future states in the same episode is treated as the state when the task is successful. In the Random Strategy, a state randomly selected from the states included in the training data is treated as the state when the task is successful.

[0029] The learning unit 192 controls the control object by reinforcement learning. In particular, the learning unit 192 learns a policy and a function using a function indicating a predicted value in which at least either an upper limit or a lower limit is set for the predicted value of the cumulative reward value. The learning unit 192 corresponds to an example of a learning means.

[0030] In the following, an example will be described in which the learning unit 192 uses a Q function with upper and lower limits set. The Q function here is a function of a certain policy π θ Q is a function that outputs the cumulative value of the likelihood of reward obtained under Q up to an infinite time. The function value of the Q function is also referred to as the Q value. The Q function used by the learning unit 192 is expressed, for example, as in Equation (1).

[0031]

[0032] y indicates the Q value. That is, y indicates the function value of the Q function (the output value of the Q function). r indicates the reward value obtained at the time step when the Q value is calculated. γ indicates a constant coefficient of 0≦γ≦1. γ is also called the learning rate.

[0033] min is a function that outputs the minimum value. min(A, B) outputs the smaller value of either A or B. min i∈M Q φ-i (s', a', g) denotes the minimum value that the Q function can take under the state s', action a', and task goal g.

[0034] Here, as the value of the parameter φ of the Q function, N values ​​φ 1 , ..., φ N where N is an integer greater than or equal to 2. Also, the target parameter φ of the Q function is - As the value of - 1 , ..., φ - N shall be prepared.

[0035] The target parameter values ​​here are copies of the parameter values. The target parameter values ​​are introduced to prevent the learning of the Q function from becoming unstable. In equation (1), the learning device 100 calculates N target parameter values ​​φ - 1 , ..., φ - N The selected target parameter values ​​are then randomly selected from φ - i It is written as follows.

[0036] min i∈M Q φ-i (s′, a′, g) is the set of two or more target parameter values ​​φ selected by the learning device 100. - iThis process is intended to prevent the Q function value (y) from becoming too large and the Q function learned using it from becoming over-learned, and can be considered as regularizing the value of the parameter φ of the Q function.

[0037] However, the regularization method used by the learning device 100 is not limited to a specific one. For example, when the learning device 100 learns the Q function, the value of the parameter φ and the target parameter φ - Alternatively, some of the values ​​of Q may be randomly changed. This regularization method is also called Dropout. The learning device 100 may also limit the value of the parameter φ of the Q function. This regularization method is also called Weight Decay. The learning device 100 may use a combination of multiple regularization methods.

[0038] max is a function that outputs the maximum value, and max(A, B) outputs the larger value of A or B.

[0039] α denotes a constant coefficient in the range of 0≦α≦1. θ (a'|s', g) is the policy π θ is obtained by applying a logarithm to the likelihood that the learning device 100 outputs the action a'. This subexpression is introduced to prevent a high probability value from being assigned to a specific action. This is expected to promote exploration during reinforcement learning and make learning more efficient. However, when the learning device 100 uses the subexpression αlogπ θ A Q function that does not include (a'|s', g) may be used.

[0040] Here, a' is the policy π under the state s' and the task goal g. θ Here, θ denotes the parameter of the policy π. a′ denotes the parameter of the policy π under the state s′ and the task goal g. θ The selection based on the above is shown in equation (2).

[0041]

[0042] Q minindicates a constant value set as the lower limit of the Q function value. max indicates a constant value set as the upper limit of the Q function value. For example, if the reward value when the task is successful is 0 and the reward value when the task is not successful is -1, then Q max The value of may be expressed as in equation (3).

[0043]

[0044] Also, in this case, Q min The value of may be expressed as in equation (4).

[0045]

[0046] Q shown in formula (3) max and Q shown in formula (4) min are examples of the theoretical minimum and maximum values ​​of the Q function value when learning is stable. However, the upper and lower limit values ​​of the Q function value used by the learning unit 192 are not limited to specific values.

[0047] As mentioned above, in reinforcement learning, the proportion of time steps in which significant reward values ​​are shown in the training data is small, which may prevent learning from progressing. In contrast, rewriting the training data is expected to increase the proportion of time steps in which significant reward values ​​are shown, making it easier for learning to progress.

[0048] On the other hand, it was found that rewriting the training data to increase the proportion of time steps showing significant reward values ​​can cause learning to become unstable. In particular, it was found that rewriting the training data can cause the Q function value to not converge, which can prevent learning from converging. Thus, rewriting the training data can progress learning in the sense that the frequency of parameter values ​​being updated using significant reward values ​​increases, but it can also prevent learning from converging.

[0049] Therefore, the learning unit 192 performs learning of control of the control target using a Q function in which upper and lower limit values ​​of the Q function value are set. This is expected to stabilize learning. For example, it is expected that the Q function value will converge more easily, and learning will converge more easily. Since learning will converge more easily, it is expected that a policy with a relatively high probability of succeeding in the task will be obtained in a relatively short learning time.

[0050] Alternatively, the learning unit 192 may use a Q function in which only one of the upper and lower limits of the Q function value is set. In this case, too, it is expected that the Q function value will converge more easily than when a Q function in which neither the upper nor lower limit is set is used, and that learning will converge more easily.

[0051] Furthermore, the learning unit 192 regularizes at least one of the parameter values ​​of the machine learning model of the policy or the parameter values ​​of the machine learning model of the Q function. The machine learning model of the policy may be expressed in the form of a function (as a policy function). In this case, the parameters of the policy function correspond to the parameters of the machine learning model of the policy. Furthermore, the machine learning model of the Q function may be expressed in the form of a function (as a Q function). In this case, the parameters of the Q function correspond to the parameters of the machine learning model of the Q function.

[0052] Here, it has been found that overfitting may occur when training data is rewritten as described above and learning is performed using a Q-function with an upper limit and a lower limit (or either one of them). Therefore, it is expected that overfitting can be avoided or reduced by having the learning unit 192 regularize at least one of the parameter values ​​of the machine learning model of the policy or the parameter values ​​of the machine learning model of the Q-function.

[0053] However, when the learning unit 192 learns control of a control object using a Q function in which upper and lower limit values ​​of the Q function value are set, this is not limited to the case where the rewriting unit 191 rewrites the training data. Even in cases other than when the rewriting unit 191 rewrites the training data, it is expected that learning will be stabilized by the learning unit 192 learning control of a control object using a Q function in which upper and lower limit values ​​of the Q function value are set. In this case, the learning device 100 may be configured without the rewriting unit 191.

[0054] Furthermore, even when the rewriting unit 191 does not rewrite the training data, the learning unit 192 may use a Q function for which upper and lower limits are set, and may also regularize the parameter values ​​of the machine learning model. This is expected to prevent or reduce overlearning.

[0055] 2 is a diagram showing an example of a processing procedure in which the learning device 100 learns control of a control object. In the processing in FIG. 2, the learning unit 192 performs initial setting for learning (step S101). Specifically, the learning unit 192 sets an initial value for a parameter θ of the policy π. The initial value of the parameter θ may be determined in advance.

[0056] The learning unit 192 also learns the parameter value φ of the Q function. 1 , ..., φ N As described above, N indicates the number of values ​​of the parameter φ of the Q function. 1 , ..., φ N The initial value of may be determined in advance. The learning unit 192 also empties a buffer D that stores training data. The training data stored in the buffer D is referred to as a data set D. The learning unit 192 also empties a target parameter value φ - i As the parameter value φ i Set a copy of i = 1,...,N.

[0057] Next, the learning unit 192 samples the goal and the initial state (step S102). Specifically, the learning unit 192 samples the goal distribution p g The goal distribution p is sampled according to g The distribution of goals p g Sampling the goal g according to [mathematical formula - see original document] is shown in equation (5).

[0058]

[0059] The learning unit 192 also calculates the initial distribution p s0 According to the initial state s 0 The initial state distribution p s0 The initial distribution p s0 According to the initial state s 0 Sampling is shown as in equation (6).

[0060]

[0061] Next, the learning device 100 starts a loop that performs processing for times t = 0, ..., T (step S111). Here, T is an integer greater than or equal to 1 that indicates the end time of one episode. In other words, T indicates the number of time steps included in one episode. The time step being processed in loop L11 is represented as time t.

[0062] In the processing of loop L11, the learning unit 192 determines whether the control target is behavior a t Reward when you do t and the next state s t+1 (Step S112). t is the state s t Under the policy π θ Action a is determined based on t But, state s t Under the policy π θ The fact that it is determined based on the above is shown in equation (7).

[0063]

[0064] The controlled object actually behaves at By doing this, you can get reward r t and the next state s t+1 Alternatively, the learning unit 192 may obtain the behavior a t By simulating the reward r t and the next state s t+1 may be calculated.

[0065] Next, the learning unit 192 determines whether the time t has reached the episode end time T (step S113). If it is determined that the time t has reached the episode end time T (step S113: YES), the learning unit 192 updates the dataset D stored in the buffer D (step S114).

[0066] Specifically, the learning unit 192 adds the training data of the episode that has reached the end time T to the dataset D stored in the buffer D. Adding the training data of the episode that has reached the end time T to the dataset D is expressed as in Equation (8).

[0067]

[0068] s t indicates the state at time t. t indicates the action taken by the controlled object at time t. t denotes the reward obtained at time t. t+1 indicates the state at time t+1. That is, state s t+1 is the state s t The rewriting unit 191 writes a new goal g' t Then, the rewriting unit 191 determines the new goal g' t Reward value r' under t Calculate the new goal g'. t Reward value r' under t The calculation is expressed as in equation (9).

[0069]

[0070] r(s t , a t , g't ) is the goal g' t and state s t Under this condition, the control object is the action a t Then, the learning unit 192 adds the training data for one episode, in which the goal and reward have been rewritten by the rewriting unit 191, to the dataset D. The addition of the training data for one episode, in which the goal and reward have been rewritten by the rewriting unit 191, to the dataset D is expressed as in equation (10).

[0071]

[0072] Dataset {(s t , a t , r' t , s t+1 , g' t ) t=0 T corresponds to an example in which the training data is rewritten to training data obtained when the objective of the task performed by the controlled object is achieved.

[0073] Next, the learning device 100 starts a loop L12 that repeats the process G times (step S121). G is an integer greater than or equal to 1. Repeating the process G times is also referred to as G Updates. In the process of loop L12, the learning unit 192 samples a mini-batch from the training data (step S122). Mini-batch B can be expressed as B = {(s, a, r, s', g)}. Here, s represents a state, a represents an action to be taken by the controlled object, r represents a reward, s' represents a next state, and g represents a goal.

[0074] Next, the learning unit 192 samples a set M of one or more different indices from the set {1, ..., N} of indices of the target parameter values ​​of the Q function (step S123).Then, the learning unit 192 calculates a Q function value y using, for example, the Q function shown in equation (1) (step S124).

[0075] Next, the learning device 100 starts a loop L13 in which processing is performed for index i=1, ..., N (step S131). As described above, N indicates the number of parameter values ​​of the Q function. It can also be said that N indicates the number of target parameter values ​​of the Q function.

[0076] In the processing of loop L13, the learning unit 192 calculates the parameter value φ of the Q function. i (step S132). The learning unit 192 updates the parameter value φ using the gradient descent method based on equation (11). i As described above with respect to the learning device 100, the learning unit 192 regularizes the value of the parameter φ during learning.

[0077]

[0078] However, the learning unit 192 determines the parameter value φ i The method for updating the target parameter value φ is not limited to a specific method. - i is updated (step S133).

[0079]

[0080] ρ is the target parameter value φ - i indicates a constant coefficient of 0≦ρ≦1 to adjust the degree of change of ρ.

[0081] After step S133, the learning device 100 performs termination processing of loop L13 (step S134). Specifically, the learning device 100 determines whether or not the processing of loop L13 has been performed for all values ​​of i = 1, ..., N. If it is determined that there are values ​​for which the processing of loop L13 has not been performed, the processing returns to step S132, and the learning device 100 continues to perform the processing of loop L13 for the unprocessed values. On the other hand, if it is determined that the processing of loop L13 has been performed for all values ​​of i = 1, ..., N, the learning device 100 terminates loop L13.

[0082] If loop L13 is ended in step S134, the learning device 100 performs termination processing of loop L12 (step S141). Specifically, the learning device 100 determines whether the processing of loop L12 has been repeated G times. If the learning device 100 determines that the processing of loop L12 has not yet been repeated G times, the processing returns to step S122. In this case, the learning device 100 continues processing loop L12. On the other hand, if it determines that the processing of loop L12 has been repeated G times, the learning device 100 ends loop L12.

[0083] When the learning device 100 finishes the loop L12 in step S141, the learning unit 192 updates the value of the parameter θ of the policy (step S142). The learning unit 192 may update the value of the parameter θ using the gradient descent method based on equation (13).

[0084]

[0085] Action a is a policy π under state s and goal g. θ The action a is determined based on the policy π under the state s and goal g. θ The fact that it is determined based on the above is shown in equation (14).

[0086]

[0087] However, the method by which the learning unit 192 updates the value of the parameter θ is not limited to a specific method.

[0088] After step S142, the learning device 100 performs termination processing of loop L11 (step S151). Specifically, the learning device 100 determines whether or not processing of loop L11 has been performed for all times t = 0, ..., T. If it is determined that there is a time at which processing of loop L11 has not been performed, the process returns to step S112, and the learning device 100 continues processing of loop L11 for the unprocessed time. On the other hand, if it is determined that processing of loop L11 has been performed for all times t = 0, ..., T, the learning device 100 terminates loop L11.

[0089] If loop L11 is ended in step S151, the learning device 100 ends the processing in Fig. 2. On the other hand, if it is determined in step S113 that time t has not reached the end time T of the episode (step S113: NO), the processing proceeds to step S121.

[0090] 3 is a diagram showing an example of a task performed by a control target. In the example of FIG. 3, the control target 910 is configured as a robot having a robot arm. The control target 910 executes a task of moving an object 920 placed on a platform 930 with the robot arm, for example by flicking the object 920 with the robot arm, to a target point P11.

[0091] In this task, when the object 920 is located at point P11, the task is considered successful and a reward value of 0 is given. On the other hand, when the object 920 is not located at point P11, the task is considered unsuccessful and a reward value of −1 is given.

[0092] In this task, for example, in the early stages of learning, it is conceivable that the task will rarely be successful and learning will not progress easily. Therefore, it is expected that learning will progress more easily if the rewriting unit 191 rewrites the training data. For example, the rewriting unit 191 may rewrite the training data so that the reward value is 0 even when the object 920 is at a predetermined position other than point P11.

[0093] On the other hand, rewriting the training data by the rewriting unit 191 may cause the learning to become unstable. Therefore, the learning unit 192 performs learning of control over the control object 910 using a Q function in which upper and lower limit values ​​of the Q function value are set. This is expected to stabilize the learning.

[0094] On the other hand, if the learning unit uses a Q function in which upper and lower limits of the Q function value are set to learn the control of the control target 910, overlearning may occur. Therefore, regularization is performed on at least one of the parameter values ​​of the machine learning model of the policy or the parameter values ​​of the machine learning model of the Q function. This is expected to prevent or reduce overlearning.

[0095] Fig. 4 is a diagram showing an example of differences in the variation of Q function values ​​due to differences in learning settings. The horizontal axis of the graph in Fig. 4 represents the number of time steps, and the vertical axis represents the Q function value. Fig. 4 shows examples of the average, maximum, and minimum values ​​of the Q function value when learning is performed multiple times.

[0096] Line L111 shows the average value of the Q function value when neither the training data is rewritten nor the upper and lower limits of the Q function value are set. Line L112 shows the maximum value of the Q function value when neither the training data is rewritten nor the upper and lower limits of the Q function value are set. Line L113 shows the minimum value of the Q function value when neither the training data is rewritten nor the upper and lower limits of the Q function value are set.

[0097] Line L121 shows the average value of the Q function value when the training data has been rewritten and the upper and lower limits of the Q function value have not been set. Line L122 shows the maximum value of the Q function value when the training data has been rewritten and the upper and lower limits of the Q function value have not been set. Line L123 shows the minimum value of the Q function value when the training data has been rewritten and the upper and lower limits of the Q function value have not been set.

[0098] Line L131 shows the average value of the Q function value when both the training data has been rewritten and the upper and lower limits of the Q function value have been set. Line L132 shows the maximum value of the Q function value when both the training data has been rewritten and the upper and lower limits of the Q function value have been set. Line L133 shows the minimum value of the Q function value when both the training data has been rewritten and the upper and lower limits of the Q function value have been set.

[0099] Line L141 indicates the upper limit of the Q function value. Line L142 indicates the lower limit of the Q function value. Figure 4 shows an example in which the control target executes the task in Figure 3, and the reward value is 0 if the task is successful, and -1 if the task is not successful.

[0100] In the example of Fig. 4, the upper limit of the Q function value is set to the upper limit value shown in formula (3), and the lower limit of the Q function value is set to the lower limit value shown in formula (4). In the example of Fig. 4, the Q function shown in formula (1) is used. In formula (1), the Q function Q used in learning the Q function is φ-i Upper and lower limits are set for the value of (s', a', g), and the calculated Q function value y may be greater than the upper limit or smaller than the lower limit.

[0101] The case where neither the training data is rewritten nor the upper and lower limits of the Q function value are set is also referred to as the case of reinforcement learning only. The Q function values ​​in the case of reinforcement learning only are indicated by lines L111, L112, and L113.

[0102] The case where the training data is rewritten and the upper and lower limits of the Q function value are not set is also referred to as the case of reinforcement learning + rewriting. The Q function values ​​in the case of reinforcement learning + rewriting are shown by lines L121, L122, and L123.

[0103] The case where both the rewriting of training data and the setting of upper and lower limits of the Q function value are performed is also referred to as the case of reinforcement learning + rewriting + upper and lower limit setting. The Q function values ​​in the case of reinforcement learning + rewriting + upper and lower limit setting are shown by lines L131, L132, and L133.

[0104] In the example of Figure 4, the variance in the Q function values ​​is larger in the case of reinforcement learning + rewriting than in the case of reinforcement learning alone. It is thought that the larger variance in the Q function values ​​makes it difficult for learning to converge.

[0105] In contrast, in the case of reinforcement learning + rewriting + upper / lower limit setting, the variance in the Q function value is smaller than in either the case of reinforcement learning alone or the case of reinforcement learning + rewriting. In the case of reinforcement learning + rewriting + upper / lower limit setting, the Q function value falls roughly between the upper limit value indicated by line L141 and the lower limit value indicated by line L142. It is expected that the smaller variance in the Q function value will make it easier for learning to converge.

[0106] FIG. 5 is a diagram showing an example of differences in performance due to differences in learning settings. The horizontal axis of the graph in FIG. 5 represents the number of time steps. The vertical axis represents performance. FIG. 5 shows an example of average performance values ​​when learning is performed multiple times. In FIG. 5, the vertical axis represents a numerical value corresponding to the success rate of the task as an index value indicating performance. More specifically, the vertical axis represents the cumulative value of reward values ​​within one episode as an index value indicating performance.

[0107] Line L211 indicates the average value of performance (average value of index values ​​indicating performance) when neither the training data is rewritten nor the upper and lower limits of the Q function value are set. As described above, the case where neither the training data is rewritten nor the upper and lower limits of the Q function value are set is also referred to as the case of reinforcement learning only.

[0108] Line L221 shows the average performance when upper and lower limit values ​​of the Q function value are set and the training data is not rewritten. The case where upper and lower limit values ​​of the Q function value are set and the training data is not rewritten is also referred to as the case of reinforcement learning + upper and lower limit value setting.

[0109] Line L231 shows the average performance when the training data is rewritten but the upper and lower limits of the Q function value are not set. As described above, the case where the training data is rewritten but the upper and lower limits of the Q function value are not set is also referred to as the case of reinforcement learning + rewriting.

[0110] Line L241 shows the average performance value when both the training data is rewritten and the upper and lower limits of the Q function value are set. As described above, the case where both the training data is rewritten and the upper and lower limits of the Q function value are set is also referred to as the case of reinforcement learning + rewriting + upper and lower limit setting.

[0111] In the example of Figure 5, the number of time steps is 10 x 10 5When the number of times is more than 1, the relationship between the average performance values ​​is as follows: reinforcement learning only < reinforcement learning + upper and lower limit setting < reinforcement learning + rewriting < reinforcement learning + rewriting + upper and lower limit setting.

[0112] In particular, the average performance value was higher in the case of reinforcement learning + rewriting than in the case of reinforcement learning alone, and the average performance value was even higher in the case of reinforcement learning + rewriting + upper and lower limit setting.

[0113] It is expected that the larger the average performance, the more likely the task will be successful. In this respect, the larger the average performance, the higher the learning accuracy can be evaluated. From this, it is expected that reinforcement learning + rewriting will enable more accurate learning than reinforcement learning alone. In the case of reinforcement learning + rewriting + upper and lower limit setting, it is expected that even more accurate learning will be possible.

[0114] Furthermore, the average performance is higher when reinforcement learning is performed with upper and lower limit settings than when reinforcement learning is performed alone. This suggests that even if the rewriting unit 191 does not rewrite the training data, the learning unit 192 can obtain relatively high-performance learning results by using a Q function with upper and lower limit values.

[0115] Fig. 6 is a diagram showing an example of the configuration of a control system according to at least one embodiment. In the configuration shown in Fig. 6, the control system 1 includes a control device 200 and a control target 950. The control device 200 includes a communication unit 210, a display unit 220, an operation input unit 230, a storage unit 280, and a processing unit 290. The processing unit 290 includes a control unit 291.

[0116] The control system 1 is a system that controls a control object 950. As described above with respect to the learning performed by the learning device 100, the control object 950 is not limited to a specific one and can be various objects capable of learning control through reinforcement learning. For example, the control object 950 may be a stand-alone device such as an industrial robot or a machine tool. Alternatively, the control object 950 may be equipment such as a plant or a power plant, or a system such as a production line in a factory. Alternatively, the control object 950 may be a moving object such as an automobile, a railroad vehicle, an airplane, a ship, or a self-propelled mobile robot, or a transportation system such as a railway or air traffic control system. Alternatively, the control object 950 may be the same as the control object 910 in FIG. 3.

[0117] The control device 200 controls the control target 950 using the policy acquired by the learning device 100 through learning. For example, the control device 200 determines an action to be taken by the control target 950 based on the policy. Then, the control device 200 transmits a control command to the control target 950, causing the control target 950 to take the determined action.

[0118] Control device 200 may be configured using a computer such as a personal computer. For example, control device 200 may be implemented in the same computer as learning device 100. Alternatively, control device 200 may be implemented in a computer separate from the computer in which learning device 100 is implemented.

[0119] The communication unit 210 communicates with other devices. For example, the communication unit 210 may be configured to transmit a control command to the control target 950. The display unit 220 has a display screen such as a liquid crystal panel or an LED panel, and displays various images. For example, the display unit 220 may be configured to display information related to the control of the control target 950, such as a control command for the control target 950 or the progress of tasks performed by the control target 950.

[0120] The operation input unit 230 includes input devices such as a keyboard and a mouse, and receives user operations. For example, the operation input unit 230 may receive user operations that instruct the control of the control target 950, such as a user operation to start control of the control target 950, or a user operation to select a task to be executed by the control target 950.

[0121] The storage unit 280 stores various data. For example, the storage unit 280 may store a policy acquired by the learning device 100 through learning. The storage unit 280 is configured using a storage device provided in the control device 200. The processing unit 290 controls each unit of the control device 200 to perform various processes. The functions of the processing unit 290 may be performed by a CPU provided in the control device 200 reading and executing a program from the storage unit 280.

[0122] The control unit 291 controls the control object 950 using a strategy acquired by the learning device 100 through learning. For example, the control unit 291 determines an action to be taken by the control object 950 based on the strategy. The control unit 291 then causes the control object 950 to perform the determined action by transmitting a control command to the control object 950 via the communication unit 210. The control unit 291 corresponds to an example of a control means. By the control unit 291 controlling the control object 950 using the strategy acquired by the learning device 100 through learning, it is expected that the control object 950 will be relatively more likely to achieve the goal of the task, as described above with respect to the learning performed by the learning device 100.

[0123] As described above, the learning unit 192 learns the policy and the Q function using a Q function in which at least either an upper limit or a lower limit is set for the predicted value of the cumulative reward value. The learning device 100 can stabilize learning in reinforcement learning. In particular, the learning device 100 is expected to make it easier for the Q function value to converge, which is expected to make it easier for learning to converge.

[0124] Furthermore, the rewriting unit 191 rewrites at least one step of the training data used for learning into training data when the goal of the task performed by the controlled object is achieved. The learning unit 192 uses the rewritten training data to learn the policy and the Q-function.

[0125] According to the learning device 100, even if the proportion of time steps showing significant reward values ​​in the training data is small, it is expected that the proportion of time steps showing significant reward values ​​will increase, making it easier for learning to progress.

[0126] Furthermore, it has been found that rewriting training data can make learning unstable. However, according to the learning device 100, the learning unit 192 uses a Q function in which at least either an upper limit or a lower limit is set for the predicted value of the cumulative reward value to learn the policy and the Q function, and it is expected that learning will be stabilized.

[0127] Furthermore, the learning unit 192 regularizes at least one of the parameter values ​​of the machine learning model of the policy and the parameter values ​​of the machine learning model of the Q function. It has been found that overfitting may occur when training the policy and the Q function using a Q function with at least one of an upper limit and a lower limit set. However, it is expected that the learning device 100 can avoid or reduce overfitting.

[0128] Furthermore, the control unit 291 controls the control object 950 using the policy obtained by learning the policy and learning the Q function, which are performed using a Q function in which at least either an upper limit or a lower limit is set for the predicted value of the accumulation of reward values. According to the control device 200, even if the proportion of time steps in which significant reward values ​​are shown in the training data is small, it is possible to control the control object 950 using a policy learned with relatively high accuracy. In this respect, according to the control device 200, it is expected that there is a relatively high possibility that the control object 950 will be able to achieve the task goal.

[0129] Second Embodiment Fig. 7 is a diagram showing another example of the configuration of a learning device according to at least one embodiment. In the configuration shown in Fig. 7, the learning device 610 includes a learning unit 611. In this configuration, the learning unit 611 learns a policy, which is a decision rule for behavior, and a function, using a function indicating a predicted value in which at least either an upper limit or a lower limit is set for a predicted value of the cumulative reward value, which is an index value indicating an evaluation of the behavior of a control object in the operating environment of the control object. The learning unit 611 corresponds to an example of a learning means.

[0130] The learning device 610 can stabilize learning in reinforcement learning. In particular, the learning device 100 is expected to facilitate convergence of the values ​​of the above functions, thereby facilitating convergence of learning. The learning unit 611 can be realized, for example, using the functions of the learning unit 192 in FIG. 1 .

[0131] Third Embodiment Fig. 8 is a diagram showing another example of the configuration of a control device according to at least one embodiment. In the configuration shown in Fig. 8, the control device 620 includes a control unit 621. In this configuration, the control unit 621 controls the control object using a policy obtained by learning a policy, which is a decision rule for behavior, and learning a function, which is performed using a function indicating a predicted value in which at least either an upper limit or a lower limit is set for a predicted value of the cumulative reward value, which is an index value indicating an evaluation of the behavior performed by the control object in the operating environment of the control object. The control unit 621 corresponds to an example of control means.

[0132] According to the control device 620, even if the proportion of time steps showing significant reward values ​​in the training data is small, it is possible to control the control object using a policy learned with relatively high accuracy. In this respect, according to the control device 620, it is expected that the control object will be relatively likely to achieve the task goal. The control unit 621 can be realized, for example, using the functions of the control unit 291 in FIG. 6 .

[0133] <Fourth embodiment> Fig. 9 is a diagram showing another example of the configuration of a control system according to at least one embodiment. In the configuration shown in Fig. 9, a control system 630 includes a controlled object 631 and a control device 632. The control device 632 includes a control unit 633.

[0134] In this configuration, the control unit 633 controls the control object 631 using a policy obtained by learning a policy, which is a decision rule for behavior, and learning a function, which is performed using a function indicating a predicted value in which at least either an upper limit or a lower limit is set for a predicted value of the accumulation of reward values, which are index values ​​indicating an evaluation of the behavior performed by the control object 631 in the operating environment of the control object 631. The control unit 633 corresponds to an example of control means.

[0135] According to the control system 630, even if the proportion of time steps showing significant reward values ​​in the training data is small, it is possible to control the control object using a relatively highly accurate learned policy. In this respect, according to the control system 630, it is expected that the control object will be relatively likely to achieve the task goal. The control unit 633 can be realized, for example, using the functions of the control unit 291 in FIG. 6 .

[0136] Fifth Embodiment Fig. 10 is a diagram showing an example of a processing procedure in a learning method according to at least one embodiment. The learning method shown in Fig. 10 includes performing learning (step S611).

[0137] In learning (step S611), the computer learns the policy, which is the decision rule for behavior, and the function, using a function that indicates a predicted value in which at least either an upper limit or a lower limit is set for the predicted cumulative value of the reward value, which is an index value that indicates the evaluation of the behavior performed by the controlled object in the operating environment of the controlled object.

[0138] According to the learning method shown in Fig. 10, learning can be stabilized in reinforcement learning. In particular, according to the learning method shown in Fig. 10, it is expected that the value of the above function will converge more easily, and therefore learning will converge more easily.

[0139] 11 is a diagram illustrating an example of a computer configuration according to at least one embodiment. In the configuration shown in FIG. 11, a computer 700 includes a CPU 710, a main memory device 720, an auxiliary memory device 730, an interface 740, and a non-volatile recording medium 750.

[0140] One or more of the learning device 100, control device 200, learning device 610, control device 620, and control device 632, or a portion thereof, may be implemented in a computer 700. In this case, the operation of each of the above-described processing units is stored in the auxiliary storage device 730 in the form of a program. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program. The CPU 710 also allocates storage areas in the main storage device 720 corresponding to each of the above-described storage units in accordance with the program. Communication between each device and other devices is performed by an interface 740 having a communication function and performing communication under the control of the CPU 710. The interface 740 also has a port for a non-volatile recording medium 750, and reads information from the non-volatile recording medium 750 and writes information to the non-volatile recording medium 750.

[0141] When learning device 100 is implemented in computer 700, the operations of processing unit 190 and each of its units are stored in the form of a program in auxiliary storage device 730. CPU 710 reads the program from auxiliary storage device 730, loads it into main storage device 720, and executes the above-described processing in accordance with the program.

[0142] Furthermore, the CPU 710 allocates a storage area for the storage unit 180 in the main storage device 720 in accordance with the program. Communication with other devices by the communication unit 110 is performed by the interface 740 having a communication function and operating under the control of the CPU 710. Display of images by the display unit 120 is performed by the interface 740 having a display device and displaying various images under the control of the CPU 710. Reception of user operations by the operation input unit 130 is performed by the interface 740 having an input device and receiving user operations under the control of the CPU 710.

[0143] When the control device 200 is implemented in a computer 700, the operations of the processing unit 290 and each of its units are stored in the form of a program in an auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.

[0144] Furthermore, the CPU 710 allocates a storage area for the storage unit 280 in the main storage device 720 in accordance with the program. Communication with other devices by the communication unit 210 is performed by the interface 740 having a communication function and operating under the control of the CPU 710. Display of images by the display unit 220 is performed by the interface 740 having a display device and displaying various images under the control of the CPU 710. Reception of user operations by the operation input unit 230 is performed by the interface 740 having an input device and receiving user operations under the control of the CPU 710.

[0145] When the learning device 610 is implemented in the computer 700, the operation of the learning unit 611 is stored in the form of a program in the auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.

[0146] Furthermore, CPU 710 allocates a storage area in main memory 720 for learning device 610 to perform processing in accordance with the program. Communication between learning device 610 and other devices is achieved by interface 740 having a communication function and performing communication under the control of CPU 710. Interaction between learning device 610 and a user is achieved by interface 740 having a display device and an input device, displaying various images under the control of CPU 710, and accepting user operations.

[0147] When the control device 620 is implemented in the computer 700, the operation of the learning unit 611 is stored in the form of a program in the auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.

[0148] Furthermore, the CPU 710 allocates a storage area in the main memory device 720 for the control device 620 to perform processing in accordance with the program. Communication between the control device 620 and other devices is achieved by the interface 740 having a communication function and performing communication under the control of the CPU 710. Interaction between the control device 620 and a user is achieved by the interface 740 having a display device and an input device, displaying various images under the control of the CPU 710, and accepting user operations.

[0149] When the control device 632 is implemented in the computer 700, the operation of the control unit 633 is stored in the form of a program in the auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.

[0150] Furthermore, the CPU 710 allocates a storage area in the main storage device 720 for the control device 632 to perform processing in accordance with the program. Communication between the control device 632 and other devices is achieved by the interface 740 having a communication function and performing communication under the control of the CPU 710. Interaction between the control device 632 and a user is achieved by the interface 740 having a display device and an input device, displaying various images under the control of the CPU 710, and accepting user operations.

[0151] One or more of the above-described programs may be recorded on nonvolatile recording medium 750. In this case, interface 740 may read the programs from nonvolatile recording medium 750. Then, CPU 710 may directly execute the programs read by interface 740, or may temporarily store the programs in main storage device 720 or auxiliary storage device 730 and then execute them.

[0152] Note that a program for executing all or part of the processing performed by learning device 100, control device 200, learning device 610, control device 620, and control device 632 may be recorded on a computer-readable recording medium, and the program recorded on the recording medium may be loaded into a computer system and executed to perform the processing of each unit. Note that the term "computer system" here includes hardware such as an operating system (OS) and peripheral devices. Furthermore, the term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, read-only memories (ROMs), and compact disc read-only memories (CD-ROMs), as well as storage devices such as hard disks built into computer systems. The program may be designed to implement part of the aforementioned functions, or may be capable of implementing the aforementioned functions in combination with a program already recorded on the computer system.

[0153] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, each embodiment may be combined with other embodiments as appropriate.

[0154] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes.

[0155] (Supplementary Note 1) A learning device comprising: a learning means for learning a policy, which is a decision rule for the behavior of a controlled object, and learning the function, using a function indicating a predicted value in which at least either an upper limit or a lower limit is set for the predicted cumulative value of a reward value, which is an index value indicating an evaluation of the behavior of the controlled object in the operating environment of the controlled object.

[0156] (Supplementary Note 2) The learning device according to Supplementary Note 1, further comprising a rewriting means for rewriting at least one step of the training data used for the learning into training data when the goal of the task performed by the controlled object is achieved, and the learning means uses the rewritten training data to learn the policy and the function.

[0157] (Supplementary Note 3) The learning device according to Supplementary Note 1 or Supplementary Note 2, wherein the learning means regularizes at least one of parameter values ​​of a machine learning model of the policy or parameter values ​​of a machine learning model of the function.

[0158] (Supplementary Note 4) A control device comprising: a control means for controlling the controlled object by learning a policy, which is a decision rule for the action, using a function indicating a predicted value in which at least either an upper limit or a lower limit is set for a predicted cumulative value of a reward value, which is an index value indicating an evaluation of the action taken by the controlled object in the operating environment of the controlled object, and using the policy obtained by learning the function.

[0159] (Supplementary Note 5) The control device according to Supplementary Note 4, wherein the control means controls the control object using the policy obtained by learning the policy and learning the function using training data in which at least one step of the training data used for the learning has been rewritten to training data when the goal of the task performed by the control object has been achieved.

[0160] (Supplementary Note 6) The control device according to Supplementary Note 4 or Supplementary Note 5, wherein the control means controls the control target using the policy obtained by learning the policy and learning the function, in which at least one of parameter values ​​of a machine learning model of the policy or parameter values ​​of a machine learning model of the function is regularized.

[0161] (Supplementary Note 7) A control system comprising a controlled object and a control device, wherein the control device comprises: control means for controlling the controlled object by learning a policy, which is a decision rule for the action, using a function indicating a predicted value in which at least either an upper limit or a lower limit is set for a predicted cumulative value of a reward value, which is an index value indicating an evaluation of the action taken by the controlled object in the operating environment of the controlled object, and using the policy obtained by learning the function.

[0162] (Supplementary Note 8) The control system according to Supplementary Note 7, wherein the control means controls the control object using the policy obtained by learning the policy and learning the function using training data in which at least one step's worth of training data among the training data used for the learning has been rewritten to training data when the goal of the task performed by the control object has been achieved.

[0163] (Supplementary Note 9) The control system according to Supplementary Note 7 or Supplementary Note 8, wherein the control means controls the control target using the policy obtained by learning the policy and learning the function, in which at least one of parameter values ​​of a machine learning model of the policy or parameter values ​​of a machine learning model of the function is regularized.

[0164] (Supplementary Note 10) A learning method including a computer learning a policy, which is a decision rule for the behavior, and learning the function, using a function that indicates a predicted value of the cumulative reward value, which is an index value that indicates an evaluation of the behavior performed by the controlled object in the operating environment of the controlled object, with at least either an upper limit or a lower limit set for the predicted value.

[0165] (Supplementary Note 11) The learning method according to Supplementary Note 10, further comprising the computer rewriting at least one step of the training data used for the learning with training data obtained when the goal of the task performed by the controlled object is achieved, and performing the learning includes the computer using the rewritten training data to learn the policy and the function.

[0166] (Supplementary Note 12) The learning method according to Supplementary Note 10 or Supplementary Note 11, wherein the performing of the learning includes the computer performing regularization on at least one of parameter values ​​of a machine learning model of the policy or parameter values ​​of a machine learning model of the function.

[0167] (Supplementary Note 13) A control method including: a computer learning a policy, which is a decision rule for the behavior, using a function indicating a predicted value in which at least either an upper limit or a lower limit is set for a predicted cumulative value of a reward value, which is an index value indicating an evaluation of the behavior performed by the controlled object in the operating environment of the controlled object, and controlling the controlled object using the policy obtained by learning the function.

[0168] (Supplementary Note 14) The control method according to Supplementary Note 13, wherein the performing of the control includes the computer performing control over the control object using the policy obtained by learning the policy and learning the function using training data in which at least one step of the training data used for the learning has been rewritten to training data when a goal of a task performed by the control object has been achieved.

[0169] (Supplementary Note 15) The control method according to Supplementary Note 13 or Supplementary Note 14, wherein performing the control includes performing control of the control object by the computer using the policy obtained by learning the policy and learning the function, in which at least one of parameter values ​​of a machine learning model of the policy or parameter values ​​of a machine learning model of the function is regularized.

[0170] (Supplementary Note 16) A recording medium having recorded thereon a program for causing a computer to execute the following: learning a policy, which is a decision rule for the behavior, and learning the function, using a function indicating a predicted value in which at least either an upper limit or a lower limit is set for a predicted value of the cumulative reward value, which is an index value indicating an evaluation of the behavior performed by the controlled object in the operating environment of the controlled object.

[0171] (Supplementary Note 17) The recording medium described in Supplementary Note 16, wherein the program further causes the computer to rewrite at least one step of the training data used for the learning into training data when a goal of a task performed by the control object is achieved, and in performing the learning, the program causes the computer to use the rewritten training data to learn the policy and the function.

[0172] (Supplementary Note 18) The recording medium described in Supplementary Note 16 or Supplementary Note 17, wherein, in performing the learning, the program causes the computer to perform regularization of at least one of parameter values ​​of a machine learning model of the policy or parameter values ​​of a machine learning model of the function.

[0173] (Supplementary Note 19) A recording medium having recorded thereon a program for causing a computer to execute the following: learning a policy, which is a decision rule for the behavior, using a function indicating a predicted value in which at least either an upper limit or a lower limit is set for a predicted cumulative value of a reward value, which is an index value indicating an evaluation of the behavior performed by the controlled object in the operating environment of the controlled object; and controlling the controlled object using the policy obtained by learning the function.

[0174] (Supplementary Note 20) The recording medium described in Supplementary Note 19, wherein in performing the control, the program causes the computer to perform control of the control object by using the policy obtained by learning the policy and learning the function using training data in which at least one step's worth of training data among the training data used for the learning has been rewritten to training data when the goal of the task performed by the control object is achieved.

[0175] (Supplementary Note 21) The recording medium described in Supplementary Note 19 or Supplementary Note 20, wherein in performing the control, the program causes the computer to perform control of the control object using the policy obtained by learning the policy and learning the function, in which regularization is performed on at least one of parameter values ​​of a machine learning model of the policy or parameter values ​​of a machine learning model of the function.

[0176] The present invention may be applied to a learning device, a control device, a control system, a learning method, and a recording medium.

[0177] REFERENCE SIGNS LIST 1 Control system 100, 610 Learning device 110, 210 Communication unit 120, 220 Display unit 130, 230 Operation input unit 180, 280 Storage unit 190, 290 Processing unit 191 Rewriting unit 192, 611 Learning unit 200, 620, 632 Control device 291, 621, 633 Control unit 631, 910, 950 Control target

Claims

1. A learning device comprising learning means for performing learning of a policy which is a decision rule of the action and learning of the function by using a function showing a predicted value in which at least one of an upper limit value or a lower limit value is set for a predicted value of accumulation of a reward value which is an index value showing an evaluation of the action performed by a control target in its operating environment.

2. The learning device according to claim 1, further comprising rewriting means for rewriting at least one step of training data among the training data used for the learning into training data when the control target achieves the goal of the task performed by the control target, wherein the learning means performs learning of the policy and learning of the function by using the rewritten training data.

3. The learning device according to claim 1 or 2, wherein the learning means performs regularization on at least one of a parameter value of a machine learning model of the policy or a parameter value of a machine learning model of the function.

4. A control device comprising control means for performing control on the control target by using the policy obtained by learning of the policy which is a decision rule of the action and learning of the function performed by using a function showing a predicted value in which at least one of an upper limit value or a lower limit value is set for a predicted value of accumulation of a reward value which is an index value showing an evaluation of the action performed by the control target in its operating environment.

5. A control system comprising a control target and a control device, wherein the control device comprises control means for performing control on the control target by using the policy obtained by learning of the policy which is a decision rule of the action and learning of the function performed by using a function showing a predicted value in which at least one of an upper limit value or a lower limit value is set for a predicted value of accumulation of a reward value which is an index value showing an evaluation of the action performed by the control target in its operating environment.

6. A learning method including a computer performing learning of a policy which is a decision rule of the action and learning of the function by using a function showing a predicted value in which at least one of an upper limit value or a lower limit value is set for a predicted value of accumulation of a reward value which is an index value showing an evaluation of the action performed by a control target in its operating environment.

7. A recording medium storing a program that causes a computer to perform learning of a policy that is a decision rule for the action and learning of the function by using a function indicating a predicted value in which at least one of an upper limit value or a lower limit value is set for a predicted value of the accumulation of reward values that are index values indicating an evaluation of an action performed by a control target in its operating environment.