Learning device, learning method, and program
The learning device addresses accuracy issues in data-scarce states by constructing models and refining policies using feedback, improving control performance.
Patent Information
- Application Number
- JP2024528209
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-06-23
- Publication Date
- 2025-10-15
- Estimated Expiration
- 2042-06-23
AI Technical Summary
Existing learning control methods face accuracy issues when insufficient data is available for certain states of a controlled object, leading to suboptimal control performance.
A learning device that utilizes a model acquisition means to construct a model from available data, incorporates feedback information to refine the model, and employs a policy management means to improve control strategies, even in data-scarce states.
Enhances the accuracy of learning control in data-scarce states by integrating user feedback to enhance model precision and policy optimization.
Smart Images

Figure 0007754309000017 
Figure 0007754309000018 
Figure 0007754309000019
Abstract
Description
[Technical Field]
[0001] The present invention relates to a learning device. , studies How to learn Law Call program Regarding. [Background technology]
[0002] There are cases where learning of control for a controlled object is performed offline. Here, offline learning means learning using data that has been obtained in advance. For example, Patent Document 1 describes a method of adjusting the parameters of a dynamics model by performing offline repetitive learning control using motion data obtained by making an actual robot arm trace a circular trajectory and motion data obtained by making a simulator to which a dynamics model of the robot arm is applied trace a circular movement. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2020-032481 Summary of the Invention [Problem to be solved by the invention]
[0004] When learning control of a controlled object using previously obtained data, if there is a state for which sufficient data is not available among the possible states of the controlled object, the accuracy of learning control in that state may be low. It is preferable to be able to improve the accuracy of learning control in that state, even for states for which sufficient data is not available.
[0005] An example of the object of the present invention is to provide a learning device that can solve the above-mentioned problems. , studies How to learn Law Call program The purpose is to provide [Means for solving the problem]
[0006] According to a first aspect of the present invention, a learning device includes: a model acquisition means for acquiring a model that takes a state and an action as input and outputs a next state through learning using data that links the state of the environment in which an agent performs an action, the actions that can be performed in that state, and the next state when the action is performed in that state; a feedback information acquisition means for acquiring feedback information based on the acquired model, which is information used for learning the model or for learning a new model that takes the state and the action as input and outputs the next state; and a policy management means for learning a policy that indicates the behavior of the agent according to the state, using the model acquired through learning using the feedback information.
[0010] The present invention two According to this aspect, the learning method includes a computer learning using data that links the state of the environment in which an agent performs an action, the actions that can be performed in that state, and the next state when the action is performed in that state, thereby obtaining a model that takes the state and the action as input and the next state as output, obtaining feedback information that is information used to learn the model based on the obtained model, or to learn a new model that takes the state and the action as input and the next state as output, and using the model obtained by learning using the feedback information to learn a policy that indicates the action of the agent depending on the state.
[0013] The present invention three According to the embodiment, programis a program for causing a computer to execute the following steps: acquiring a model in which the state and the action are input and the next state is output by learning using data linking the state of the environment in which an agent performs an action, the action executable in that state, and the next state when that action is performed in that state; acquiring feedback information based on the acquired model, which is information used for learning the model or for learning a new model in which the state and the action are input and the next state is output; and learning a policy indicating the action of the agent according to the state, using the model acquired by learning using the feedback information. In be. [Effects of the Invention]
[0016] According to the present invention, when learning control of a control target using data obtained in advance, it is expected that the accuracy of learning control in a state where sufficient data is not available can be improved. [Brief explanation of the drawings]
[0017] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of a learning system according to an embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of a configuration when a data collection device according to an embodiment acquires data from a control target. [Figure 3] FIG. 2 is a diagram illustrating an example of a data flow related to the data collection device according to the embodiment. [Figure 4] FIG. 10 is a diagram illustrating an example of the flow of data when the learning device according to the embodiment learns a model of a controlled object and a policy πθ. [Figure 5] FIG. 10 is a diagram illustrating an example of a processing procedure in which the learning device according to the embodiment learns a model of a controlled object and a policy πθ. [Figure 6] FIG. 2 is a diagram illustrating an example of a system configuration during operation of a control target according to an embodiment. [Figure 7] FIG. 10 is a diagram illustrating an example of a hopper. [Figure 8]FIG. 10 is a diagram illustrating an example of a configuration of data indicating a state in the embodiment. [Figure 9] FIG. 2 is a diagram illustrating an example of a configuration of data indicating behavior in the embodiment. [Figure 10] FIG. 10 is a diagram illustrating an example of input and output of data in a model and a strategy when a learning device according to an embodiment generates a pseudo trajectory. [Figure 11] FIG. 10 is a diagram illustrating an example of a display screen of the importance of state elements displayed by a display unit according to the embodiment. [Figure 12] FIG. 10 is a diagram showing an example of a display screen of the importance of a pseudo trajectory displayed by a display unit according to the embodiment. [Figure 13] FIG. 10 is a diagram showing an example of an editing screen for a pseudo trajectory τ1 displayed by a display unit according to the embodiment. [Figure 14] FIG. 10 is a diagram showing an example of an editing screen for a pseudo trajectory after correction by a user, displayed on a display unit according to the embodiment. [Figure 15] FIG. 10 is a diagram showing an example of a setting screen for weights for elements of a pseudo trajectory τ1, displayed by a display unit according to the embodiment. [Figure 16] FIG. 10 is a diagram showing an example of an editing screen for a pseudo trajectory τ3 displayed by a display unit according to the embodiment. [Figure 17] FIG. 10 is a diagram showing an example of a setting screen for weights for elements of a pseudo trajectory τ3, displayed by a display unit according to the embodiment. [Figure 18] FIG. 10 is a diagram illustrating another example of the configuration of the learning device according to the embodiment. [Figure 19] 1 is a diagram illustrating an example of a configuration of a display device according to an embodiment. [Figure 20] FIG. 10 is a diagram illustrating another example of the configuration of the display device according to the embodiment. [Figure 21] FIG. 10 is a diagram illustrating an example of a processing procedure in a learning method according to an embodiment. [Figure 22] 10A to 10C are diagrams illustrating an example of a processing procedure in a display method according to an embodiment. [Figure 23]FIG. 10 is a diagram illustrating an example of another procedure of the processing in the display method according to the embodiment. [Figure 24] FIG. 1 is a schematic block diagram illustrating the configuration of a computer according to at least one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0018] The following describes embodiments of the present invention, but the following embodiments do not limit the scope of the invention as claimed. Furthermore, not all of the combinations of features described in the embodiments are necessarily essential to the solution of the invention. FIG. 1 is a diagram illustrating an example of the configuration of a learning system according to an embodiment. In the configuration illustrated in FIG. 1, learning system 1 includes a data collection device 300 and a learning device 100. Learning device 100 includes a communication unit 110, a display unit 120, an operation input unit 130, a storage unit 180, and a processing unit 190. Storage unit 180 includes a data storage unit 181, a model storage unit 182, and a policy storage unit 183. Processing unit 190 includes a data management unit 210, a learning unit 220, an analysis unit 230, and a feedback information acquisition unit 240. Learning unit 220 includes a model management unit 221 and a policy management unit 222.
[0019] The learning system 1 learns a control method for a control object. Specifically, a data collection device 300 acquires data from the control object in advance. The learning device 100 acquires a model of the control object using the acquired data, and uses the acquired model to learn the control method.
[0020] Here, "before" refers to before the learning device 100 learns the control method for the control object. As will be described later in the description of "environment," the data collection device 300 may acquire data from the operating environment of the control object in addition to the control object. The learning device 100 may also construct a model that includes the operating environment of the control object in addition to the control object.
[0021] The control target for which the learning device 100 learns a control method is not limited to a specific one. Various controllable objects can be the control target. For example, the control target may be equipment such as a plant or a power plant, a system such as a production line in a factory, or a standalone device. Alternatively, the control target may be a moving object such as an automobile, an airplane, a ship, or a self-propelled mobile robot.
[0022] The learning of a control method for a control target performed by the learning device 100 can be considered as a type of reinforcement learning. Reinforcement learning here is machine learning that learns a policy, which is a behavioral rule of an agent that takes action in a certain environment, based on a state observed in the environment and a reward that represents an evaluation of the state or action.
[0023] When a control object itself operates in accordance with a control rule, the control object corresponds to an example of an agent, the operation of the control object corresponds to an example of an action, and the operation rule corresponds to an example of a policy. For example, if the controlled object is a chemical plant and the control mechanism is incorporated into the chemical plant and operates automatically or semi-automatically, the chemical plant is an example of an agent, the operation of the chemical plant is an example of an action, and the operating rules for the chemical plant to operate automatically or semi-automatically are an example of a policy.
[0024] When a control device that controls the control object is provided separately from the control object, the control device corresponds to an example of an agent, the control of the control object performed by the control device corresponds to an example of an action, and the control rules correspond to an example of a policy. For example, if the object to be controlled is a chemical plant and a control device that controls the chemical plant is installed externally to the chemical plant, the control device is an example of an agent, the control of the chemical plant performed by the control device is an example of an action, and the control rules by which the control device controls the chemical plant are an example of a policy.
[0025] In both cases where the controlled object itself operates according to a control rule and where a control device that controls the controlled object is provided separately from the controlled object, the controlled object, or the controlled object and its operating environment, are examples of the environment. That is, the state of the controlled object may be the subject of state observation, or in addition to the state of the controlled object, the state of the operating environment of the controlled object may also be the subject of state observation. Furthermore, the value of the reward may be obtained by observing the state, or may be obtained by calculation, etc.
[0026] For example, if the control target is a chemical plant, the chemical plant or the chemical plant and its operating environment are examples of environments in reinforcement learning. Examples of states in reinforcement learning include the state of the chemical plant, such as the values of pressure sensors and flow rate sensors installed in the chemical plant, or the state of the chemical plant and the operating environment of the chemical plant, such as the temperature of the air surrounding the chemical plant.
[0027] Another example of behavior in reinforcement learning is a control command value for a chemical plant, such as a PID (Proportional-Integral-Differential) control command value for the opening degree of a predetermined valve. A measurable value, such as the production amount of a product such as ethylene or gasoline measured by a sensor, may be used as the reward value. Alternatively, a value obtained by calculation, such as calculating the production amount of a product from the consumption amount of raw materials, may be used as the reward value without providing a sensor for measuring the production amount of the product.
[0028] The user may be a chemical plant designer, operator, or practitioner. The model of the controlled object may be a simulator used in the operation of the chemical plant. Control rules for controlling the chemical plant are an example of a policy in reinforcement learning.
[0029] The learning device 100 uses the interaction between the simulator and the policy to construct a pseudo trajectory, which is data showing a time series of control over the chemical plant and the state of the chemical plant. For example, the learning device 100 can improve the accuracy of the simulator by having the user return feedback information on the generated pseudo trajectory.
[0030] Furthermore, the learning device 100 can automatically operate a chemical plant by using a policy constructed using the improved simulator. Furthermore, by presenting the generated pseudo trajectories to practitioners, it can also assist in creating operation plans for chemical plants.
[0031] Hereinafter, the controlled object, or the controlled object and its operating environment, will also be referred to as the "environment" and represented by p. The state of the controlled object, or the state of the controlled object and the state of the operating environment of the controlled object, will also be referred to as the "state" and represented by s. The behavior of the controlled object, or the control of the controlled object, will also be referred to as the "action" and represented by a. The operating rule that prescribes the behavior of the controlled object, or the control rule that prescribes the control of the controlled object, will also be referred to as the "policy" and represented by π. The policy that is the target of learning by the learning device 100 will be referred to as π. θ It is expressed as:
[0032] "π θ "θ" in " is the policy π θ The parameter values in the model representing the information gathering strategy π are shown below. β To distinguish it from θ " is written as follows. Strategy π θ may be configured as a function that receives the input of state s and outputs the action a. In this case, the policy π θ is expressed as in equation (1).
[0033]
number
[0034] In the following, time will be expressed in time steps of a fixed time Δt, and will be expressed as time step 0, time step 1, time step 2, etc. Note that the time Δt may be different for each time step. The time step is expressed by setting the time step defined as the reference time step, such as the time step at which the controlled object starts to operate or the current time step T, as time step 0. When the current time step T is used as the reference, time step t can also be expressed as "(current time step T) + t". Each step in the time step is also referred to as each time step.
[0035] Also, state observation is performed at each time step, and the state at time step t (t is an integer t≧0) is defined as s t The state s at time step t+1 is t+1 Let us consider the state s at time step t. t It is also called the next state of .
[0036] Furthermore, the learning device 100 calculates a policy π for determining an action for each time step. θ The action at time step t is a t It is expressed as: Furthermore, the learning device 100 θ The reward value used for learning is obtained at each time step. The reward at time step t is defined as r t Reward r t is the action a at time step t t This value can be said to indicate the evaluation of the quality of the product.
[0037] Strategy π θIn the learning of the algorithm, for example, the cumulative reward r0+r1+r2+··· for each time step is set as the objective function, and the policy π is chosen so that the evaluation indicated by this objective function becomes higher. θ Alternatively, the reward may be discounted as time passes, and the cumulative reward r0 + αr1 + α 2 The objective function can be r2+··· (where 0<α<1). In other words, the cumulative reward can be said to be a value that indicates the evaluation of the quality of the action at each time step. In this case, the higher the cumulative reward (e.g., the larger the cumulative reward value), the higher the quality of the action, and the lower the cumulative reward (e.g., the smaller the cumulative reward value), the lower the quality of the action. The learning device 100 may use a reward in which the larger the value, the higher the evaluation. Alternatively, the learning device 100 may use a reward in which the smaller the value, the higher the evaluation (i.e., so-called loss).
[0038] As described above, the data collection device 300 acquires data in advance from the control target. 1 shows an example in which data collection device 300 is configured as a device separate from learning device 100. In this case, data collection device 300 may be configured using a computer such as a personal computer (PC) or a workstation. Alternatively, the data collection device 300 may be a part of the learning device 100 .
[0039] 2 is a diagram showing an example of a configuration when data collection device 300 acquires data from a control target. As shown in Fig. 2, the control target is also referred to as control target 810. In addition, a system with the configuration shown in Fig. 2 is also referred to as data collection system 2. 2 shows an example in which the data collecting device 300 transmits data to the learning device 100 after completing data acquisition from the control target 810. In this case, the data collecting device 300 does not need to be connected to the learning device 100 for communication while acquiring the data.
[0040] Alternatively, the data collecting device 300 may transmit data to the learning device 100 while acquiring data from the control target 810. In this case, the data collecting device 300 may be communicatively connected to both the control target 810 and the learning device 100 at the same time.
[0041] Fig. 3 is a diagram showing an example of the flow of data related to the data collection device 300. In the example shown in Fig. 3, the data collection device 300 performs an action a when the state of the environment p is s. s' is the next state of state s (i.e., the state at the next time step), and the action a causes a transition from state s to the next state s'.
[0042] The data collecting device 300 then acquires the value of the reward r at the time step when the next state s' is reached and the observed value of the next state s'. The data collecting device 300 may be equipped with a sensor to observe the state and acquire the observed value. Alternatively, the control object 810 as the environment p may be equipped with a sensor to observe the state, and the data collecting device 300 may acquire the observed value of the state from the control object 810. Furthermore, the value indicating the state may include a value obtained by calculation. The value indicating the status corresponds to an example of information indicating the status.
[0043] A value indicating a state is also simply referred to as a state. For example, obtaining a value indicating a state is also referred to as obtaining a state. A value of a reward is also simply referred to as a reward. For example, obtaining a value of a reward is also referred to as obtaining a reward. A value indicating an action is also simply referred to as an action. For example, obtaining a value indicating an action is also referred to as obtaining an action.
[0044] As described above, Fig. 3 shows an example in which the data collection device 300 performs the action a. However, an entity other than the data collection device 300 may perform the action a. For example, an operator may operate a chemical plant, which is an example of environment p, and the data collection device 300 may record the operations performed by the operator and the state of the chemical plant for each time step. In this case, the operations performed by the operator for each time step correspond to an example of action a, and the state of the chemical plant at the start of the operation corresponds to an example of state s. Furthermore, the state of the chemical plant after the operation is completed (specifically, the state of the chemical plant in the next time step) corresponds to an example of next state s'.
[0045] Furthermore, as described above, the reward r may be obtained by observing the environment p, or the data collection device 300 or the learning device 100 may calculate the reward r. The rule for determining the action a when the data collection device 300 acquires data from the controlled object is called a data collection strategy, and π β The data collection strategy may be, for example, a function that receives an input of a state s and determines an action a. In this case, the data collection strategy π β As in the case of the above formula (1), β (a|s)".
[0046] The data collection device 300 generates data that combines a state s, an action a, a reward r, and a next state s' for each time step and transmits the generated data to the learning device 100. Data that combines a state, an action, a reward, and a next state is also called quadruple data, and these four pieces of data are expressed in parentheses "()". For example, quadruple data of a state s, an action a, a reward r, and a next state s' is also expressed as "(s, a, r, s')". The quadruple data is an example of data that links the state of the environment in which an agent performs an action, the actions that can be performed in that state, the reward that represents the quality of the action, and the next state when the action is performed in that state.
[0047] The data collection device 300 may be configured to repeatedly acquire data from the control object 810 until a predetermined number of quadruplets of data is obtained. In this way, the learning device 100 may be configured to acquire and store a predetermined number of quadruplets of data from the data collection device 300.
[0048] In the learning device 100, the data management unit 210 stores the quadruple data transmitted by the data collection device 300 in the model storage unit 182. This set of quadruple data is referred to as the data set D env It can also be written as:
[0049] The learning device 100 may calculate the reward r. In this case, the data collection device 300 transmits to the learning device 100 triplet data that combines the state s, the action a, and the next state s′. The data collection device 300 may transmit data that does not include the next state to the learning device 100 in the form of time-series data, and the learning device 100 may insert the next state.
[0050] As described above, the learning device 100 acquires a model of the control object 810 using data acquired from the control object 810 by the data collecting device 300, and uses the acquired model to learn a control method. The model of the control object 810 corresponds to an example of a model of the environment p. The model of the control object 810 acquired by the learning device 100 is also referred to as a model of the environment p.
[0051] Here, it is conceivable that sufficient data regarding certain states and actions cannot be obtained from the data acquired by the data collecting device 300 from the control target 810. For example, when the data collecting device 300 acquires data during actual operation of the control target 810, it is conceivable that data regarding states and actions other than the states that appeared during the actual operation and the actions performed during the actual operation in response to those states cannot be obtained.
[0052] When the learning device 100 acquires a model of the control object 810 by learning using this data, it is conceivable that the accuracy of the model will be low for states and actions other than those shown in the data. The learning device 100 uses this model to acquire a policy π for controlling the control object 810. θ When learning the policy π, the state and action for which the model is not accurate cannot be learned with sufficient accuracy. θ However, it is possible that appropriate actions cannot be suggested.
[0053] For example, even if there is a more preferable control method than the control method currently being used for the control target 810, the policy π θ Furthermore, the state of the control object 810 when the preferred control method is executed and the control from that state may not be learned because the accuracy of the model of the control object 810 is low, and the policy π θ It is possible that learning will not be possible.
[0054] Therefore, the learning device 100 receives input of information based on the user's knowledge and uses the information to learn a model of the control object 810. As the accuracy of the model of the control object 810 improves, the policy π θ It is expected that the accuracy of the control object 810 will be improved. Here, it is assumed that the user is an expert on the control object 810, for example.
[0055] Learning device 100 analyzes the output data of the model so that the user can input information that is effective in improving the accuracy of the model of control target 810. Then, based on the analysis results, learning device 100 presents to the user information indicating which of the items related to the information input by the user have a large impact on improving the accuracy of the model. The information that the user inputs to learning device 100 is also referred to as feedback information, or simply as feedback information, in response to the information presented to the user by learning device 100. Feedback information can be said to be information for reducing the discrepancy between the environment and the environment model.
[0056] The information that study device 100 presents to the user and the information that the user inputs to study device 100 are not limited to a particular type of information. For example, the learning device 100 may present to the user information indicating the accuracy of the model output for each of the multiple items included in the state. This information can be interpreted as information indicating the importance of each item in terms of improving the accuracy of the model. For items with low accuracy in the model output, updating the model to improve the accuracy of those items is expected to result in a more accurate model.
[0057] The user refers to information indicating the accuracy of the model output for each item included in the state, and inputs, for example, constraint expressions that the model's inputs and outputs must satisfy for items with low accuracy. It is expected that the learning device 100 will learn the model using the input constraint expressions, thereby obtaining a model with higher accuracy.
[0058] Alternatively, the learning device 100 may acquire multiple pieces of time-series data of the input and output of the model of the control object 810, and present information indicating the accuracy of each piece of time-series data to the user. Each piece of time-series data of the input and output of the model of the control object 810 is also called a pseudo trajectory. A set of pseudo trajectories is referred to as a dataset D model It can also be written as:
[0059] The user refers to the information indicating the accuracy of the pseudo trajectories and, for example, corrects pseudo trajectories with low accuracy. The corrected pseudo trajectories can be said to be data indicating the correct answer that the model should output. When the learning device 100 uses all of the multiple pseudo trajectories to train a model, it is expected that a more accurate model can be obtained by using the corrected pseudo trajectories rather than the pseudo trajectories before correction.
[0060] The information indicating the accuracy of each pseudo trajectory can be interpreted as information indicating the importance of each pseudo trajectory in terms of improving the accuracy of the model. It is expected that a user can obtain a more accurate model by correcting a pseudo trajectory with low accuracy rather than correcting a pseudo trajectory that is originally highly accurate. The learning device 100 may be configured using a computer such as a personal computer or a workstation.
[0061] The communication unit 110 communicates with other devices. For example, the communication unit 110 receives data obtained from the control target 810 and transmitted by the data collection device 300. Display unit 120 has a display screen such as a liquid crystal panel or an LED (Light Emitting Diode) panel, and displays various images. For example, display unit 120 displays the information presented to the user by study device 100, as described above. The display unit 120 is an example of a display means.
[0062] The operation input unit 130 includes input devices such as a keyboard and a mouse, and receives user operations. For example, the operation input unit 130 receives a user operation for inputting the above-mentioned feedback information. The operation input unit 130 corresponds to an example of an input means.
[0063] The storage unit 180 stores various data. The storage unit 180 is configured using a storage device included in the learning device 100. The data storage unit 181 stores data used for learning the model of the control object 810 and the policy π θ In particular, the data storage unit 181 stores data set D env and Dataset D model Dataset D env and Dataset D model is an example of data used for learning a model of the control object 810. envand Dataset D model In addition to learning the model of the control object 810, the policy π θ It may also be used for learning.
[0064] The model storage unit 182 stores a model of the control target 810. In particular, the model storage unit 182 stores a plurality of models of the control target 810. In the following, an example will be described in which the model storage unit 182 stores two models, and these two models are represented by p^0 and p^1. However, the model storage unit 182 may store three or more models.
[0065] Both models p^0 and p^1 may be configured as a function that receives a state and an action in that state as input and outputs the next state. If the state input to model p^0 is represented by s, the action by a', and the next state output by model p^0 by s''0, then model p^0 can be expressed as in equation (2).
[0066]
number
[0067] If the state input to model p^1 is represented by s, the action by a', and the next state output by model p^1 by s''1, then model p^1 can be expressed as in equation (3).
[0068]
number
[0069] The accuracy of the model output can be evaluated using multiple models stored in the model storage unit 182. When the same state and action are input to multiple models and the outputs of the multiple models are similar, the accuracy of these models for this input can be evaluated as relatively high. On the other hand, when the same state and action are input to multiple models and there is a large variance in the outputs of the multiple models, the accuracy of these models for this input can be evaluated as relatively low.
[0070] As an index showing the magnitude of variation in the outputs of a plurality of models, for example, the variance of the outputs of these plurality of models can be used. For example, the next state s''0 output by model p^0 and the next state s''1 output by model p^1 are collectively referred to as the next state s'', and the variance between the next state s''0 and the next state s''1 is represented as Var(s''). The variance Var(s'') is expressed as in equation (4).
[0071]
number
[0072] s'' avg represents the average of the next state s''0 and the next state s''1. The learning device 100 uses the reward that reflects this variance Var(s'') to θ For example, the learning device 100 may learn the policy π θ The learning may be performed.
[0073]
number
[0074] "r(s,a')" represents the reward before the variance Var(s'') is reflected. The reward r(s,a') indicates the evaluation of taking action a' in state s from the perspective of the agent's behavioral goal. According to the reward r'', the policy π θ When the accuracy of the models p̂0 and p̂1 for the next state s″ resulting from the action a′ output by the learning device 100 is low, the evaluation indicated by the reward r″ is low. θ By learning the above, the probability of outputting an action a' that will cause the model p^0 and p^1 to transition to a next state with low accuracy is relatively small. θFrom this point of view, the accuracy of the models p^0 and p^1 is such that the policy π θ It is expected that this will reduce the degree to which the accuracy of the
[0075] On the other hand, strategy π θ The action a' output by the model p^0 and p^1 is limited to the action a' that transitions to the next state with low accuracy. θ However, it is possible that appropriate actions cannot be suggested. Therefore, as described above, the learning device 100 presents the user with information indicating which of the items related to the information input by the user have the greatest impact on improving the accuracy of the model, and accepts input of feedback information by the user.
[0076] The policy storage unit 183 stores the policy π θ Remember.
[0077] Processing unit 190 performs various processes by controlling each unit of learning device 100. The functions of processing unit 190 are performed, for example, by a CPU (Central Processing Unit) included in learning device 100 reading and executing a program from storage unit 180.
[0078] The data management unit 210 manages the data stored in the data storage unit 181. For example, the data management unit 210 manages the data of the quadruplets received by the communication unit 110 from the data collection device 300 as a data set D stored in the data storage unit 181. env Store in.
[0079] In addition, the data management unit 210 calculates the model of the control target 810 and the policy π θ The data set D stored in the data storage unit 181 is a set of quadruple data generated using model Store in. When the learning device 100 acquires multiple pseudo trajectories as described above, the data management unit 210 stores the pseudo trajectories in a data set D model Store in.
[0080] When the learning device 100 uses the feedback information to obtain a more accurate model of the control object 810 and generates quadruple data using the model, the data management unit 210 stores the generated quadruple data in the storage unit 180. The data management unit 210 stores the newly obtained data in the already obtained data set D model Alternatively, the data management unit 210 may overwrite the already obtained data set D model The data management unit 210 may store the newly obtained data in the data set D. model You can add it to the dataset D model Alternatively, a different data set may be generated.
[0081] The learning unit 220 learns a model of the control object 810 and calculates the policy π θ The learning unit 220 also learns the model of the control object 810 and the policy π θ Generate the quadruple data using
[0082] The model management unit 221 learns the model of the control target 810. Specifically, the model management unit 221 learns the model of the control target 810. env The model p^0 and the model p^1 are trained using the above. Furthermore, when the user inputs feedback information using the operation input unit 130, the model management unit 221 reflects the obtained feedback information and performs learning of the model p̂0 and the model p̂1 again.
[0083] The model management unit 221 may be configured to train the already obtained models p^0 and p^1 so as to update these models. Alternatively, the model management unit 221 may be configured to restart the training of the model p^0 and the training of the model p^1 from the beginning without using the already obtained models p^0 and p^1.
[0084] That is, the model management unit 221 may use the feedback information to further train an already obtained model, or may use the feedback information to train a new model. The model management unit 221 corresponds to an example of a model acquisition means.
[0085] The policy management unit 222 manages the policy π θ In particular, the policy management unit 222 uses data of four pairs indicating input and output data of the model obtained by learning that reflects feedback information to learn the policy π θ As a result, the policy management unit 222 learns the data set D env For states and actions for which sufficient data could not be obtained using only the strategy π θ In this respect, a more accurate policy π θ It is expected that this will be achieved. The policy management unit 222 corresponds to an example of a policy management means.
[0086] The analysis unit 230 generates information indicating which of the above-mentioned items related to the information input by the user has a large effect on improving the accuracy of the model. For example, as described above for the learning device 100, the analysis unit 230 may calculate, for each item included in the information indicating the state, an evaluation index value for the accuracy of that item in the next state output by each of the models p^0 and p^1. For example, for each item included in the information indicating the state, the analysis unit 230 may calculate, as an evaluation index value for the accuracy of that item, the variance between the value of that item in the output of the model p^0 and the value of that item in the output of the model p^1. In this case, the items included in the information indicating the status correspond to examples of items related to information input by the user.
[0087] Furthermore, as described above for the learning device 100, the analysis unit 230 may calculate an evaluation index value for the accuracy of each of the multiple pseudo trajectories. For example, the analysis unit 230 may calculate a next state according to model p^0 and a next state according to model p^1 for each time step in the pseudo trajectory. Then, the analysis unit 230 may calculate, as an evaluation index value for the accuracy of the pseudo trajectory, a value obtained by summing or averaging the variances of the next states for each time step for all time steps in one pseudo trajectory. In this case, the pseudo trajectory corresponds to an example of an item related to information input by the user. The analysis unit 230 corresponds to an example of an analysis means.
[0088] Feedback information acquisition unit 240 acquires feedback information. Specifically, feedback information acquisition unit 240 reads feedback information that the user inputs using study device 100. As described above, the analysis unit 230 analyzes the output of the model of the control target 810 and presents the analysis result to the user, and the feedback information acquisition unit 240 acquires feedback information that the user inputs by referring to the analysis result. In this respect, it can be said that the feedback information acquisition unit 240 acquires feedback information based on the output of the model of the control target 810. When the analysis unit 230 presents the analysis result to the user, the feedback information acquisition unit 240 may externally prompt (for example, prompt the user) to input feedback information. The feedback information acquisition unit 240 corresponds to an example of a feedback acquisition means.
[0089] As described above with respect to the learning device 100, the feedback information acquisition unit 240 may acquire feedback information indicating constraints that must be satisfied by the input data and output data of the model of the control object 810 in order to improve the accuracy of the model. In particular, the feedback information acquisition unit 240 may acquire feedback information indicating constraints on items that are relatively poorly evaluated for accuracy among items included in the next state output by the model of the control object 810.
[0090] In this case, an item with a relatively low accuracy rating may be, for example, an item with a rating lower than the average (or median) of the ratings for multiple items including that item. Alternatively, an item with a relatively low accuracy rating may be an item with the lowest rating among a selected portion of multiple items. Alternatively, an item with a relatively low accuracy rating may be an item that satisfies a predetermined criterion. In this case, the "predetermined criterion" refers to the criterion for determining whether the rating is low. For example, the "predetermined criterion" may be that the rating is below a threshold.
[0091] The user refers to the evaluation index value for each item included in the state displayed by the display unit 120 and inputs feedback information using the operation input unit 130. In this respect, it can be said that the feedback information acquisition unit 240 acquires feedback information input by a user operation accepted by the operation input unit 130 after the display unit 120 starts displaying the evaluation index value for each item included in the state.
[0092] Furthermore, as described above with respect to the learning device 100, the feedback information acquisition unit 240 may acquire feedback information indicating a correction to a pseudo trajectory. In particular, the feedback information acquisition unit 240 may acquire feedback information indicating a correction to a pseudo trajectory whose accuracy is evaluated as being relatively low.
[0093] The user refers to the evaluation index value for each pseudo trajectory displayed by the display unit 120 and inputs feedback information using the operation input unit 130. In this respect, it can be said that the feedback information acquisition unit 240 acquires feedback information input by a user operation accepted by the operation input unit 130 after the display unit 120 starts displaying the evaluation index value for each pseudo trajectory.
[0094] FIG. 4 shows the process in which the learning device 100 learns the model of the control object 810 and the policy π θ 10 is a diagram illustrating an example of the flow of data when learning is performed. In the example shown in FIG. 4, the data management unit 210 manages the data set D env The model manager 221 uses this quadruple data to train the model p^0 and the model p^1. The policy manager 222 uses this quadruple data to train the policy π θ Here, the policy π θ The learning method of dataset D model Strategy π for generating θ The learning method is not limited to a specific one as long as it can acquire the policy π θ The learning may be performed.
[0095] Dataset D env After the learning using the model p̂0 and the policy π is completed, the data management unit 210 θ and obtain the quadruple (s, a', r'', s'') of the dataset D model accumulates in Specifically, the policy management unit 222 assigns a certain state s to a policy π θ to obtain the action a′, and outputs the state s and the action a′ to the model management unit 221.
[0096] The policy management unit 222 θThe state s to be input to the data set D is not limited to a specific one. env The state s may be read from any of the quadruple data included in the model manager 221 and output to the policy manager 222. Alternatively, the policy manager 222 may generate the state s arbitrarily within the range of possible actions of the control object 810. When the learning device 100 uses a pseudo trajectory, the policy manager 222 uses the next state s'' output by the model manager 221 as the state s in the next time step.
[0097] The model management unit 221 inputs the state s and the action a' into the model p^0 to obtain the next state s''0. The model management unit 221 also inputs the state s and the action a' into the model p^1 to obtain the next state s''1. Then, the model management unit 221 outputs the next state s'' and the reward r'' to the policy management unit 222.
[0098] The next state s'' is not limited to a specific one as long as it is a state obtained from the next state s''0 and the next state s''1. For example, the model management unit 221 may select the next state s''0 as the next state s''. Alternatively, the model management unit 221 may select the next state s''1 as the next state s''. Alternatively, the model management unit 221 may calculate the average of the next state s''0 and the next state s''1 and use this as the next state s''.
[0099] As described above, the analysis unit 230 calculates the variance between the next states s''0 and s''1, or the variance for each item of these states. To this end, the learning unit 220 may associate the next states s''0 and s''1 with the quadruple data (s, a', r'', s'').
[0100] Alternatively, as shown in the above formula (5), the model management unit 221 calculates the variance Var(s'') between the next states s''0 and s''1 when calculating the reward r''. The learning unit 220 may associate the variance Var(s'') with the quadruple data (s, a', r'', s'') instead of the next states s''0 and s''1. The reward r'' may be calculated by a unit other than the model management unit 221. For example, the policy management unit 222 may calculate the reward r''.
[0101] Alternatively, the learning unit 220 may include a pair of next states s''0 and s''1 in the data of the quadruple instead of the next state s''. In this case, when a next state is required, each unit of the learning device 100 obtains the next state according to a predetermined method for obtaining the next state from the next states s''0 and s''1, such as selecting the next state s''0.
[0102] D model After a predetermined amount of quadruple data has been accumulated, the analysis unit 230 analyzes the output of the model of the control target 810. As the analysis result, the analysis unit 230 generates information indicating which of the items related to the information input by the user have the greatest impact on improving the accuracy of the model. The analysis unit 230 then presents the obtained analysis result to the user by displaying it on the display unit 120.
[0103] The user generates feedback information by referring to the analysis results by the analysis unit 230. The user performs an input operation of the feedback information on the operation input unit . The feedback information acquisition unit 240 reads feedback information from a signal output by the operation input unit 130 in response to a user operation, and outputs the feedback information to the model management unit 221 .
[0104] The model management unit 221 uses the obtained feedback information to train the models p^0 and p^1. As described above, the model management unit 221 may update the already obtained models p^0 and p^1, or may train from scratch without using the already obtained models p^0 and p^1.
[0105] The policy management unit 222 calculates the policy π using the newly obtained models p^0 and p^1. θ The policy management unit 222 learns the already obtained policy π θ Alternatively, the already obtained policy π θ Alternatively, learning may be performed from the beginning without using the above.
[0106] FIG. 5 shows the process in which the learning device 100 learns the model of the control object 810 and the policy π θ FIG. 10 is a diagram illustrating an example of a procedure for a process of learning the above.
[0107] (Step S1) The model management unit 221 manages the data set D env Specifically, the model management unit 221 constructs models p̂0 and p̂1 based on the data set D env Based on the quadruple data (s, a, r, s') contained in the model, models p^0 and p^1 are searched for that take input of state s and action a and output next state s'. For example, a known supervised learning method can be used to search for models p^0 and p^1. When model p^0 or p^1 receives input of state s and action a and outputs next state s', this is also referred to as estimating next state s' from state s and action a.
[0108] In addition, the policy management unit 222 env Based on the policy π θ As mentioned above, the strategy π θ The learning method of dataset D model Strategy π for generating θ Any method that can acquire the above knowledge is acceptable, and is not limited to a specific learning method. After step S1, the process proceeds to step S2.
[0109] (Step S2) The data management unit 210 manages the data set D env The data management unit 210 randomly extracts a quadruple of data (s, a, r, s′) from the data management unit 210. The data management unit 210 outputs the extracted quadruple of data to the policy management unit 222. After step S2, the process proceeds to step S3.
[0110] (Step S3) The policy management unit 222 assigns the state s included in the quadruple data (s, a, r, s') extracted in step S2 to the policy π θ to obtain behavior a'. Note that the obtained behavior a' may be different from the behavior a included in the quadruple data. After step S3, the process proceeds to step S4.
[0111] (Step S4) The model management unit 221 inputs the state s and action a' in step S3 into the model p^0 to obtain the next state s''0. Also, the model management unit 221 inputs the state s and action a' in step S3 into the model p^1 to obtain the next state s''1.
[0112] (Step S5) The model management unit 221 calculates the reward r'' based on the above formula (5). After step S5, the process proceeds to step S6.
[0113] (Step S6) The data management unit 210 stores a data set D in the form of a quadruple (s, a', r'', s'') of a state s, an action a', a reward r'', and a next state s''. model Store in. After step S6, the process proceeds to step S7. When the learning device 100 handles a pseudo trajectory, the learning unit 220 may set the next state s'' as a new state s and repeat the processes of steps S3 to S6. In this case, the learning unit 220 repeats the processes of steps S3 to S6 the number of times as many times as the number of time steps in the pseudo trajectory, and then the process proceeds to step S7.
[0114] (Step S7) The learning unit 220 uses the data set D model and D env Using this, policy π θ Update the policy π θ For example, a known reinforcement learning method can be used to update the parameter k. After step S7, the process proceeds to step S8.
[0115] (Step S8) The analysis unit 230 performs an analysis to efficiently improve the models p̂0 and p̂1. For example, the analysis unit 230 may calculate, for each item included in a state, the variance between the value of that item in the next state s''0 and the value of that item in the next state s''1, as described above. Alternatively, the analysis unit 230 may calculate the variance between the next states s''0 and s''1 for each time step in the pseudo trajectory, as described above, and sum up the variances calculated for each time step for all time steps in the pseudo trajectory. After step S8, the process proceeds to step S9.
[0116] (Step S9) The display unit 120 displays the analysis results obtained by the analysis unit 230. Then, the feedback information acquisition unit 240 acquires feedback information based on a user operation received by the operation input unit 130. As described above, the user is an expert on the control target 810, for example. After step S9, the process proceeds to step S10.
[0117] (Step S10) The model management unit 221 manages the data set D env and D model and the feedback information, the models p̂0 and p̂1 are updated. For example, when a constraint expression relating to an item included in a state is input as feedback information, the model management unit 221 env The quadruple data in and dataset D model For each of the quadruple data included in the above, the parameter values of the model p^0 are learned so that the model p^0 outputs the next state s'' in response to the input of the state s and action a included in the quadruple data, and so that the constraint equations indicated in the feedback information are satisfied. The model management unit 221 similarly learns the parameter values for the model p^1. For example, a known constrained supervised learning method can be used to learn the parameter values of the models p̂0 and p̂1. After step S10, the process proceeds to step S11.
[0118] (Step S11) The learning unit 220 determines whether a condition set in advance as a termination condition for updating the models p̂0 and p̂1 is satisfied. The condition for ending the update of the models p^0 and p^1 is not limited to a specific method. For example, the condition for ending the update of the models p^0 and p^1 may be whether or not the processes of steps S2 to S11 in FIG. 4 have been repeated a predetermined number of times.
[0119] Alternatively, the condition for ending the update of models p^0 and p^1 may be whether or not the evaluation of the accuracy of models p^0 and p^1 is higher than a predetermined evaluation. Alternatively, the termination condition for updating models p^0 and p^1 is policy π θ The condition may be whether the evaluation of the item is higher than a predetermined evaluation.
[0120] If the learning unit 220 determines that the termination condition is not met (step S11: NO), the process returns to step S2. On the other hand, if the learning unit 220 determines in step S11 that the termination condition is met (step S11: YES), the learning device 100 terminates the processing of FIG.
[0121] 6 is a diagram showing an example of a system configuration during operation of a control target 810. The system shown in FIG. In the configuration shown in FIG. 6, the control device 400 uses the policy π θ The control object 810 is controlled using the above. The learning device 100 may function as the control device 400. Alternatively, the control device 400 may be provided separately from the learning device 100, and the control device 400 may be configured to receive the policy π θ may be stored.
[0122] Next, the processing performed by the learning device 100 will be described using an example in which a policy for controlling a hopper is learned. Fig. 7 is a diagram showing an example of a hopper. In the example shown in Fig. 7, hopper 910 includes a top portion 911, thigh portions 912, leg portions 913, and foot portions 914. The connection portions of these portions are configured as joints whose angles are adjustable, and each joint is provided with a rotor for adjusting the angle.
[0123] Rotator 921 adjusts the angle between top of head 911 and thigh 912. Rotator 921 is also referred to as thigh rotor 921. The rotor 922 adjusts the angle between the thigh 912 and the leg 913. The rotor 922 is also referred to as a leg rotor 922. The rotor 923 adjusts the angle between the leg 913 and the foot 914. The rotor 923 is also referred to as a foot rotor 923.
[0124] Hopper 910 is a mobile robot that obtains thrust by changing its posture due to the movement of each rotor. Hopper 910 must be moved without tipping over. 7, the x-axis is set on the plane on which the hopper 910 moves, and the z-axis is set so as to be perpendicular to this plane. The plane on which the hopper 910 moves is also referred to as the movement plane.
[0125] FIG. 8 is a diagram showing an example of the structure of data indicating the state s. In the example shown in Figure 8, the state s is represented by an 11-dimensional numerical vector. The elements of the state s are represented as element s e0 , s e1 ,···,s e10 Each of these elements corresponds to an example of an item included in a state. However, the number of elements in a state s is not limited to a specific number.
[0126] element s e0 indicates the z coordinate value of the top of the head 911, that is, the height of the top of the head 911. element s e1 indicates the angle of the top of the head 911 relative to the plane of movement. element s e2 indicates the angle of the thigh 912 relative to the plane of movement. element s e3 indicates the angle of the leg 913 relative to the plane of movement. element s e4 indicates the angle of the foot 914 relative to the plane of movement.
[0127] element s e5 indicates the x-component of the velocity of the top of the head 911. element s e6 indicates the z-component of the velocity of the top of the head 911. element s e7 denotes the angular velocity of the top of the head 911 relative to the plane of movement. element s e8 denotes the angular velocity of the thigh 912 relative to the plane of motion. element s e9 denotes the angular velocity of the leg 913 relative to the plane of movement. element s e10denotes the angular velocity of the foot 914 relative to the plane of motion.
[0128] FIG. 9 is a diagram showing an example of the structure of data indicating action a. In the example shown in Figure 9, action a is represented by a three-dimensional numerical vector. e0 , a e1 , a e2 However, the number of elements of action a is not limited to a specific number.
[0129] Element a e0 represents the torque applied to the thigh rotor 921. Element a e1 represents the torque applied to the rotor 922 of the leg. Element a e2 represents the torque applied to the rotor 923 of the foot.
[0130] In addition, the above formula (5) will be used as the formula for calculating the reward r'', and the value of the ``r(s, a')'' part of formula (5) will be calculated based on formula (6).
[0131]
number
[0132] The right side of equation (6) "s e5 -|a e0 |-|a e1 |-|a e2 |" indicates the difference obtained by subtracting the sum of the magnitudes of the torques applied to the thigh rotor, leg rotor, and foot rotor from the speed of the top 911 of the hopper 910 in the x coordinate. Equation (6) indicates that the faster the speed at which the hopper 910 moves in the positive direction of the x coordinate, the higher the evaluation, and also indicates that the smaller the total power consumption to operate the rotors, the higher the evaluation.
[0133] FIG. 10 shows the relationship between the model p̂0 and the policy π when the learning device 100 generates a pseudo trajectory. θ 1 is a diagram illustrating an example of input and output of data in FIG. In the example of Figure 10, dataset D env A plurality of quadruple time series data obtained by operating the hopper 910 is stored in the memory 910. This quadruple time series data is called a trajectory.
[0134] The state of the hopper 910 at time step t (t is an integer t≧0) is denoted by s t The time step 0 is the start of the operation of the hopper 910, and the initial state of the hopper 910 is represented by s0. The behavior of the hopper 910 at the time step t is expressed as a t and the reward r at time step t is expressed as r t It is expressed as:
[0135] The data management unit 210 manages the data set D env The initial state s0 of the hopper 910 in the trajectory is read out from the trajectory table 221 and output to the policy management unit 222 and the model management unit 221. The policy management unit 222 assigns the initial state s0 to the policy π θ to obtain the behavior a0, and outputs the obtained behavior a0 to the model management unit 221.
[0136] The model management unit 221 inputs the state s0 and the action a0 into the model p^0 to obtain the action s1. Fig. 10 shows an example in which the state output by the model p^0 is used as the next state, and the model management unit 221 outputs the obtained state s1 to the policy management unit 222. The policy management unit 222 assigns the state s1 to the policy π θ to obtain the behavior a1, and outputs the obtained behavior a1 to the model management unit 221.
[0137] In this way, for time step t, the policy manager 222 calculates the state s t Policy π θ Enter and act a t Obtain the action a t to the model management unit 221. The model management unit 221 outputs the state s t and Action a t is input to the model p^0, and the state s t+1and obtain the resulting state s t+1 is output to the policy management unit 222. The strategy management unit 222 and the model management unit 221 continue the action a until a predetermined termination condition is met, for example, the hopper 910 reaches the destination or falls over. t Obtaining and status s t+1 Repeat the acquisition.
[0138] Furthermore, the model management unit 221 calculates the state s to be input to the model p̂0 for each time step. t and Action a t Same state as t and Action a t is also input to model p^1, and the behavior s t+1 Also obtain. Furthermore, the learning unit 220 learns the state s t and Action a t and are input into the reward function shown in equation (6) to obtain the reward r t The learning unit 220 obtains the reward r t and the state s according to the models p^0 and p^1, respectively. t+1 and into equation (5) to obtain the reward r'' at time step t. t Get. The learning unit 220 uses the data of the quadruplets (s t ,a t ,r'' t ,s t+1 ) to the data management unit 210. The data management unit 210 accumulates the quadruple data output by the learning unit 220 as time-series data to generate a pseudo trajectory.
[0139] The content of the analysis by the analysis unit 230 and the information input by the user as feedback information are not limited to specific ones, and can be various depending on the control of the control target, for example. (1) The analysis results indicate the accuracy of each element included in the state predicted by the model, and the feedback information indicates constraints on the input and output of the model; and (2) When the analysis result indicates the accuracy of each of the multiple pseudo trajectories, and the feedback information indicates the corrected value of the pseudo trajectory or information on whether the value indicated in the pseudo trajectory is correct, This will be explained using the following example.
[0140] (1) The analysis results indicate the accuracy of each element included in the state predicted by the model, and the feedback information indicates constraints on the input and output of the model.
[0141] 5, the analysis unit 230 calculates, for each element of the state, the accuracy of that element in the model output based on equation 7. The accuracy of an element in the model output is also referred to as the importance of that element.
[0142]
number
[0143] Let i=0, 1, . . . , 10, and element s of state s e0 , s e1 ,···,s e10 , s e_i Also, s t+1,e_i is the element s at time step t+1 e_i Indicates the importance of. I(s e_i ) is the element s e_i The analysis unit 230 calculates the importance of the element s in the next state output by the model p̂0 for each time step based on the formula (7). e_i and the element s in the next state output by model p^1 e_i The variance of the pseudo trajectory is calculated, and the variance is summed over all time steps in the pseudo trajectory to determine the importance I(s e_i )
[0144] Importance I(s e_i ) is the element s in the next state output by the model p^0 e_i This corresponds to an example of an evaluation index value for accuracy. Equation (7) can be interpreted as indicating that the more ambiguous the elements of the state predicted by the model, the more important they are. The more ambiguous the elements of the state predicted by the model, the lower the accuracy of the model's predictions.
[0145] Note that the ∞ in Σ in equation (7) indicates that the number of time steps in the pseudo trajectory is arbitrary. The analysis unit 230 sums up the variances for each time step up to the last time step in the pseudo trajectory. When there are multiple pseudo trajectories, the analysis unit 230 further calculates the importance I(s) shown in equation (7) for all the pseudo trajectories. e_i ) may be summed.
[0146] Alternatively, the analysis unit 230 may calculate the importance of each element of the state in the model output based on equation (8) instead of equation (7).
[0147]
number
[0148] In equation (8), the variance of the element in the output of model p^0 and model p^1 for each time step and each state element, Var(p^0(s t+1,e_i |a t ,s t ),p^1(s t+1,e_i |a t ,s t ))" and the cumulative reward "Σ t’=t ∞ r t’ " Multiply by
[0149] Equation (8) can be interpreted as indicating that among the elements of the state predicted by the model, the more ambiguous the element and the more it contributes to the cumulative reward, the more important it is. Regarding equation (8), when there are multiple pseudo trajectories, the analysis unit 230 further calculates the importance I(s) shown in equation (8) for all the pseudo trajectories. e_i ) may be summed.
[0150] In this way, the analysis unit 230 calculates the importance of each element of the states output by the models p^0 and p^1, and outputs the calculated importance as the analysis result. However, the importance calculated by the analysis unit 230 is not limited to those shown in the above formulas (7) and (8), and various other values are possible.
[0151] FIG. 11 is a diagram showing an example of a display screen of the importance of the state elements displayed by the display unit 120. In FIG. In the example of FIG. 11, the display unit 120 displays the identification information “s e0 "," "s e1 ",..., "s e10 ", along with an explanation of each element and the importance of each element calculated by the analysis unit 230.
[0152] The user can determine that the greater the importance value of an element, the more important it is to improve the accuracy by referring to the display screen shown in FIG. 11. Specifically, the user can determine that the greater the importance value of an element, the more important it is to improve the accuracy. e0 The most important thing is to improve the accuracy of the element s e2 The improvement of the accuracy of element s is important, and then e3 It can be concluded that improving the accuracy of For example, consider the case where a user is considering introducing one of the following two constraints into a model search: The first constraint equation that the user considers is shown as equation (9).
[0153]
number
[0154] Equation (9) indicates a constraint that the z coordinate value of the top of the head 911 at time step t plus the value obtained by multiplying the velocity in the z-axis direction by coefficient c1 represents the z coordinate value of the top of the head 911 at time step t+1. The second constraint equation that the user considers is shown in equation (10).
[0155]
number
[0156] Equation (10) indicates the constraint that the angle calculated by adding the torque applied to the rotor of the thigh 912 multiplied by coefficient c2 to the angle of the thigh 912 at time step t represents the angle of the thigh 912 at time step t+1. Since both equations (9) and (10) require investigation of coefficient values, it is assumed that the user is considering adopting only one of these two constraint equations. In this case, the user can select one of the two constraint equations by referring to the importance calculated by the analysis unit 230.
[0157] Equation (9) is the state element s e0 and s e6 By referring to these two factors, the user can determine the importance I(s e0 )=1 and I(s e6 )=0, and calculate the importance of equation (9) as 1. On the other hand, equation (10) is a state element s e2 From the reference, the user can see the state element s e2 Importance of I(s e2 )=0.5 is the importance of equation (10). Since the importance of formula (9) is greater than the importance of formula (10), the user adopts formula (9). The user obtains the value of coefficient c1, for example, by actual measurement, and sets it in formula (9).
[0158] The user may generate a constraint expression by focusing on an element with a high degree of importance. For example, the user may generate a constraint expression by focusing on an element s with the highest degree of importance. e0 It may be decided to generate a constraint equation for the z coordinate value of the top of the head 911 by paying attention to the following.
[0159] In step S9 of FIG. 5, the user inputs the selected equation (9) as feedback information to learning device 100, and feedback information acquisition section 240 acquires the input feedback information.
[0160] 5, the model management unit 221 searches for a model by adding equation (9), in which the value of coefficient c1 is set, to the constraints for searching for models p^0 and p^1. The model management unit 221 may also search for model p^0 based on equation (11).
[0161]
number
[0162] "←" represents substitution. In equation (11), it represents substituting the model p^ obtained by the calculation on the right side into model p^0. In other words, it represents adopting model ^p as model p^0. argmin is a function that outputs a parameter value that minimizes the value of the objective function. The model management unit 221 calculates "-logp^(s t+1 |s t ,a t )+(s t+1,e0 -s t,e0 -s t,e6 *c1) 2 A search is performed for model p^ so that the value of " is smaller, and the obtained model p^ is adopted as model p^0. "*" represents multiplication.
[0163] "-logp^(s t+1 |s t ,a t )" is dataset D envThe higher the likelihood of the state-action pair included in the equation, the smaller the value of this term. t+1 |s t ,a t )" is the model p^ in state s t and Action a t The next state that is output after receiving the input of is the data set D env The next state s is denoted by t+1 The closer it is to , the smaller the value becomes.
[0164] "(s t+1,e0 -s t,e0 -s t,e6 *c1) 2 ” means that the model p^ is in state s t and Action a t The value of this term decreases as the degree to which the next state output in response to the input satisfies equation (9) adopted as the constraint equation increases. As in the case of model p^0, the model management unit 221 may search for model p^1 based on equation (12).
[0165]
number
[0166] The model management unit 221 env For example, the model management unit 221 may search for a model using a data set D env The state,s,shown in the quadruplet contained in t , action a t , next state s t+1 , state s t element e0 of s t,e0 , state s t s, which is element e6 of t,e6 , and the next state s t+1 element e0 of s t+1,e0 may be applied to equations (11) and (12) to search for models p̂0 and p̂1. Alternatively, the model management unit 221 may envIn addition to or instead of the dataset D model A model may be searched for based on the above.
[0167] A user may input multiple constraint equations to the learning device 100. In this case, the model management unit 221 may calculate the importance of each of the constraint equations described above and weight the constraint equations based on the calculated importance. For example, the model management unit 221 may search for the model p̂0 based on equation (13).
[0168]
number
[0169] ReLU stands for Ramp Function. t+1,e2 -s t,e2 -c2*a t,e0 )" is 0 if equation (10) holds, and is 0 if equation (10) does not hold. t+1,e2 -s t,e2 -c2*a t,e0 ", that is, the value obtained by subtracting the right side of equation (10) from the left side.
[0170] "(ReLU(s t+1,e2 -s t,e2 -c2*a t,e0 )) 2 ” means that the model p^ is in state s t and Action a t If the next state output in response to the input satisfies equation (10), which is adopted as a constraint equation, then this term is 0; if equation (10) is not satisfied, then the smaller the degree of deviation, the smaller the value becomes.
[0171] α1 is "(s t+1,e0 -s t,e0 -s t,e6 *c1) 2 " is the weighting coefficient for "(ReLU(s t+1,e2 -s t,e2 -c2*a t,e0 ))2 " is the weighting coefficient for For example, similarly to the above, the model management unit 221 calculates the importance of formula (9) as 1 and the importance of formula (10) as 0.5. Then, the model management unit 221 normalizes the calculated importance to calculate α1=1 / (1+0.5) and α2=0.5 / (1+0.5).
[0172] By using the weighting coefficients exemplified in Equation (13), the model management unit 221 can ensure that the more important constraints set by the user are more strongly reflected in the model search. This is expected to result in a highly accurate model.
[0173] Learning device 100 may be configured so that input of feedback information by the user is not required during model training and policy training. For example, the storage unit 180 may store in advance constraint equations based on user knowledge, such as those exemplified by the above equations (9) and (10). Then, in step S9 of Fig. 5, the feedback information acquisition unit 240 may calculate the importance of each constraint equation as described above and select the constraint equation with the highest importance. Alternatively, the feedback information acquisition unit 240 may select multiple constraint equations based on the importance of each constraint equation, such as by selecting a predetermined number of constraint equations in descending order of importance.
[0174] (2) When the analysis results indicate the accuracy of each of multiple pseudo trajectories, and the feedback information indicates the correction value of the pseudo trajectory or information on whether the value indicated in the pseudo trajectory is correct.
[0175] In the following, an example will be described in which the learning unit 220 generates six pseudo trajectories. The pseudo trajectories are denoted as τ0, τ1, . . . , τ5. However, the number of pseudo trajectories generated by the learning unit 220 is not limited to a specific number as long as it is two or more.
[0176] The jth (j is an integer between 0 and 5) pseudo-trajectory τj The states and actions at time step t are respectively t,τj , a t,τj It is written as follows. In step S8 of FIG. 5, the analysis unit 230 calculates the importance of each pseudo trajectory based on, for example, equation (14).
[0177]
number
[0178] I(τ j ) is the pseudo-trajectory τ j The analysis unit 230 calculates the importance of the next state s output by the model p̂0 for each pseudo trajectory and for each time step based on the equation (14). t+1 ,τ j and the next state s output by model p^1 t+1 ,τ j The variance of the pseudo trajectory is calculated, and the variance is summed over all time steps in the pseudo trajectory to determine the importance I(τ j ) Importance I(τ j ) is an example of an evaluation index value for the accuracy of time series data.
[0179] Equation (14) can be interpreted as indicating that the more ambiguous the state predicted by the model, the more important the pseudo-trajectory. A pseudo-trajectory in which the state predicted by the model is ambiguous overall can be said to be a pseudo-trajectory with low prediction accuracy by the model when considering all time steps included in the pseudo-trajectory comprehensively.
[0180] As explained for equation (7), the ∞ in Σ in equation (14) also indicates that the number of time steps in the pseudo trajectory is arbitrary. The analysis unit 230 sums up the variances for each time step up to the last time step in the pseudo trajectory. However, the importance calculated by the analysis unit 230 is not limited to that shown in the above formula (14), and various other values are possible.
[0181] FIG. 12 is a diagram showing an example of a display screen of the importance of pseudo trajectories displayed by the display unit 120. As shown in FIG. In the example of Figure 12, the display unit 120 displays the identification information of each pseudo trajectory, "τ0", "τ1", ... "τ5", the importance of each pseudo trajectory, and a button icon that accepts user operations to request the display of an editing screen for each pseudo trajectory.
[0182] The user can determine that the greater the importance value of a pseudo trajectory, the more important it is to improve its accuracy, by referring to the display screen shown in Fig. 12. Specifically, the user can prioritize the pseudo trajectories in such a way that improving the accuracy of the pseudo trajectory τ1 is the most important, followed by improving the accuracy of the pseudo trajectory τ3, and then improving the accuracy of the pseudo trajectory τ5.
[0183] If correcting all the pseudo trajectories is too much for the user, the user can select the pseudo trajectories to be corrected based on their importance. Here, it is assumed that the user decides to correct the pseudo trajectories τ1 and τ3.
[0184] In step S9 of FIG. 5, the user corrects the pseudo trajectory, and the feedback information acquisition unit 240 acquires feedback information indicating the correction of the pseudo trajectory.
[0185] Fig. 13 is a diagram showing an example of an editing screen for the pseudo trajectory τ1 displayed by the display unit 120. When a pressing operation is performed on the button icon shown in the row of the pseudo trajectory τ1 on the display screen of Fig. 12, the display unit 120 displays the editing screen of Fig. 13.
[0186] 13, the display unit 120 displays the value of each element of the state s and the value of each element of the action a for each time step of the pseudo trajectory τ1. The user modifies the pseudo trajectory by modifying the values of the elements of the displayed states. The display unit 120 displays the value of the element of the behavior a as reference information for the user to obtain the correct value of the element of the state. Alternatively, the value of the element of the behavior a may also be subject to correction by the user.
[0187] FIG. 14 is a diagram showing an example of an editing screen displayed on the display unit 120 for a pseudo trajectory after correction by the user. Fig. 14 shows an example in which a user performs a correction operation on the editing screen shown in Fig. 13. The user modifies element s in state s at time step 3. e0 and s e2 Here, it is assumed that the user has knowledge of the correct values, such as being able to calculate the values of these elements. In the example of FIG. 14, the display unit 120 displays the corrected value with an underline.
[0188] Trajectory τ j The state s of the control object 810 at time step t is t,τj The state after the user makes the corrections to s t,τj,f It is expressed as: In addition, weights based on the reliability of the pseudo trajectory elements are introduced. The pseudo trajectory elements here refer to the state elements shown in the pseudo trajectory. If behavior is also subject to user modification, the behavior elements shown in the pseudo trajectory are also referred to as pseudo trajectory elements.
[0189] The user sets the weight value based on the user's own judgment about the correctness of the elements of the pseudo trajectory. For example, the user sets the weight value to 1 for elements whose values are judged to be correct. The user also sets the weight value to 0 for elements whose values are judged to be unknown. The user also sets the weight value to -1 for elements whose values are judged to be incorrect. For a pseudo trajectory that has been corrected, the user sets the weight value for the corrected pseudo trajectory. The weight setting value is also an example of feedback information.
[0190] FIG. 15 is a diagram showing an example of a setting screen for setting weights for elements of the pseudo trajectory τ1, displayed by the display unit 120. FIG. 15 shows an example of weights set by the user for the values of the elements shown in FIG. 14. The user sets the weights for the element s e0 The corrected value of and element s e2 The corrected value is determined to be correct, and the weight value is set to 1.
[0191] On the other hand, the user can select the element s e1 , s e3 , s e4 ,···,s e10 The weight value for each value is set to 0, as it is determined that the correctness or incorrectness of each value is unknown. This weight is used when the model management unit 221 updates the model using feedback information. For example, by using this weight, the model management unit 221 can filter out element values that the user has determined to be correct or incorrect values.
[0192] Fig. 16 is a diagram showing an example of an editing screen for the pseudo trajectory τ3 displayed by the display unit 120. When a pressing operation is performed on the button icon shown in the row of the pseudo trajectory τ3 on the display screen of Fig. 12, the display unit 120 displays the editing screen of Fig. 16. For the pseudo trajectory τ3, the user does not modify the element values but only sets the weights.
[0193] FIG. 17 is a diagram showing an example of a setting screen displayed by the display unit 120 for setting weights for elements of the pseudo trajectory τ3. FIG. 17 shows an example of weights set by the user for the values of the elements shown in FIG. 16. The user sets the weights for the element s e4 , s e9 Each value is judged to be correct and the weight value is set to 1.
[0194] On the other hand, the user can select the element s e0 , s e2 , s e5 , s e7 , s e8 , s e10 The weight value for each value is set to 0, as it is determined that the correctness or incorrectness of each value is unknown. Also, the user can select the element s e1 , s e3 , s e6 Each value is judged to be incorrect and the weight value is set to -1.
[0195] 5, the model management unit 221 searches for a model using the obtained feedback information. The model management unit 221 may search for a model p̂0 based on equation (15).
[0196]
number
[0197] Status t’+1,τj,f Regarding the state s t’+1,τj If no corrections have been made to the state s t’+1,τj The value of t’+1,τj,f Also, the state s t’+1,τj,f If only some of the elements in are modified, the elements that are not modified are in state s t’+1,τjThe value of that element in is used as is.
[0198] "w t’,τj " is the trajectory τ j The state s of the control object 810 at time step t' shown in t’,τj is a vector indicating the weight set for each element of In equation (15), the weight "w t’,τj " is expressed as a horizontal vector. Also, the model p^ is t’,τj and Action a t’,τj For the input, the corrected trajectory τ j The next state s shown in t’+1,τj,f The likelihood of outputting "p'^(s t’+1,τj,f |s t’,τj ,a t’,τj )" is expressed as a vertical vector for each element of the state.
[0199] "w t’,τj ·p'^(s t’+1,τj,f |s t’,τj ,a t’,τj )" uses the weight vector "w t’,τj ” and likelihood vector “p'^(s t’+1,τj,f |s t’,τj ,a t’,τj )" and take the dot product. Therefore, the value of this inner product formula is that for elements with a weight value of 1, the model p^ calculates the corrected pseudo-trajectory τ for the value of that element in the next state. j The state shown in t’+1,τj,f The higher the likelihood of outputting the value of that element in , the larger the value.
[0200] On the other hand, elements with a weight value set to 0 are filtered out in the calculation of the value of this dot product formula. In other words, for elements with a weight value set to 0, the value output by the model p^ does not affect the value of this dot product formula.
[0201] In addition, for elements whose weight value is set to -1, the value of this inner product formula is calculated as the modified pseudo-trajectory τ j The state shown in t’+1,τj,f The higher the likelihood of outputting the value of that element in , the smaller the value.
[0202] The model management unit 221 searches for the model p̂0 using equation (15), thereby reflecting the user's correction of the pseudo trajectory and the weight setting in the search for the model. In this respect, it is expected that the accuracy of the obtained model will be high. However, it is not essential for the user to set weights. If the user does not set weights, the model management unit 221 may set all weight values to 1 and perform a model search.
[0203] As in the case of model p^0, the model management unit 221 may search for model p^1 based on equation (16).
[0204] [Number]
[0205] The model management unit 221 env Alternatively, the model management unit 221 may search for a model using the data set D env In addition to or instead of the dataset D model A model may be searched for based on the above.
[0206] As described above, the model management unit 221 acquires models p^0 and p^1, which take the state and action as input and the next state as output, through learning using data linking the state of the environment in which the agent performs an action, the actions executable in that state, and the next state when that action is performed in that state. Based on the acquired model, the feedback information acquisition unit 240 acquires feedback information, which is information used for learning the model, or for learning a new model that takes the state and action as input and the next state as output. The policy management unit 222 learns policies that indicate the agent's actions according to the state through learning using the model acquired using the feedback information.
[0207] According to the learning device 100, the model management unit 221 uses a previously obtained data set D env In this case, when sufficient data is not available, the feedback information can be used to compensate for the lack of information, and it is expected that a relatively accurate model can be obtained. By having the policy management unit 222 use this model to learn the policy (learn the control for the control object 810), it is expected that the accuracy of the control learning in the state where sufficient data is not available can be improved.
[0208] Furthermore, the feedback information acquisition unit 240 acquires the feedback information indicating constraints that must be satisfied by the input data and output data of the model to be trained. The model management unit 221 searches for a model using the constraints. According to the learning device 100, a model can be obtained in which constraints are reflected in the relationship between input data and output data, and in this respect, it is expected that a relatively accurate model can be obtained.
[0209] Furthermore, for each item included in the information indicating the state, the analysis unit 230 calculates an evaluation index value of the accuracy of that item in the next state output by the obtained model. The feedback information acquisition unit 240 acquires feedback information indicating constraint conditions related to items with relatively low evaluations of accuracy. According to the learning device 100, for items included in the information indicating the state, for which the accuracy of the values output by the model is relatively low, the accuracy is improved based on constraints, and in this respect, it is expected that the accuracy of the model can be efficiently improved.
[0210] Furthermore, the display unit 120 displays an evaluation index value for the accuracy of the item in the next state output by the obtained model. The operation input unit 130 accepts a user operation for inputting feedback information. After the display unit 120 starts displaying the evaluation index value, the feedback information acquisition unit 240 acquires the feedback information input by the user operation accepted by the operation input unit 130.
[0211] According to learning device 100, model manager 221 uses feedback information to search for a model, allowing user knowledge corresponding to the accuracy of the model to be reflected in the model search. In this respect, learning device 100 is expected to be able to obtain relatively accurate models relatively efficiently.
[0212] Furthermore, the feedback information acquisition unit 240 acquires feedback information indicating corrections to the input / output data of the obtained model. The model management unit 221 performs model learning using the input / output data reflecting the corrections. According to the learning device 100, a model with relatively high accuracy can be expected to be obtained because the model is learned using input / output data that reflects the correction.
[0213] Furthermore, the analysis unit 230 calculates an evaluation index value for the accuracy of the time-series data for each of the multiple time-series data of the input and output of the obtained model. The feedback information acquisition unit 240 acquires feedback information indicating corrections to time-series data with a relatively low evaluation of accuracy.
[0214] According to the learning device 100, among multiple time-series data, time-series data with a relatively low evaluation of accuracy is corrected. The model management unit 221 uses the corrected time-series data to learn the model, which is expected to improve the accuracy of the model relatively efficiently.
[0215] Furthermore, the display unit 120 displays an evaluation index value of the accuracy of the time-series data for each of the multiple time-series data of the input and output of the obtained model. The operation input unit 130 accepts a user operation to input feedback information. After the display unit 120 starts displaying the evaluation index values, the feedback information acquisition unit 240 acquires the feedback information input by the user operation accepted by the operation input unit 130.
[0216] According to the learning device 100, the model management unit 221 uses feedback information to learn a model, which allows the model to be learned using data whose values have been modified by the user for time-series data with relatively low accuracy among multiple time-series data. In this respect, the learning device 100 is expected to be able to relatively efficiently obtain a model with relatively high accuracy.
[0217] In addition, the display unit 120 displays, for each item included in the information indicating the next state that is output by the model simulating the environment in which the agent performs an action in response to input of information indicating the state and information indicating the action, an evaluation index value for the accuracy of that item in the information indicating the next state.
[0218] With learning device 100, the user can refer to the index values displayed by display unit 120 and input feedback information to learning device 100 to improve items with low accuracy in the model output. In this respect, learning device 100 is expected to be able to efficiently improve the accuracy of the model.
[0219] The display unit 120 also displays an evaluation index value for the accuracy of the time series data for each of a plurality of time series data of the state in the environment and the agent's behavior, which is time series data of the input and output of the model that simulates the environment in which the agent performs its behavior. Learning device 100 allows a user to select and correct time-series data with low accuracy by referring to the index values displayed on display unit 120. In this respect, learning device 100 is expected to be able to efficiently improve the accuracy of the model.
[0220] 18 is a diagram showing another example of the configuration of a learning device according to an embodiment. In the configuration shown in FIG. 18, a learning device 610 includes a model acquisition unit 611, a feedback information acquisition unit 612, and a policy management unit 613. With this configuration, the model acquisition unit 611 acquires a model in which the state and action are input and the next state is output through learning using data linking the state of the environment in which the agent performs an action, the actions that can be performed in that state, and the next state when that action is performed in that state.The feedback information acquisition unit 612 acquires feedback information based on the acquired model, which is information used for learning the model or for learning a new model in which the state and action are input and the next state is output.The policy management unit 613 learns policies that indicate the agent's actions according to the state through learning using the model acquired using the feedback information. The model acquisition unit 611 corresponds to an example of a model acquisition means, the feedback information acquisition unit 612 corresponds to an example of a feedback information acquisition means, and the policy management unit 613 corresponds to an example of a policy management means.
[0221] According to the learning device 610, when the model acquisition unit 611 does not have enough data from the previously acquired data, it is expected that the model acquisition unit 611 can compensate for the lack of information with feedback information and obtain a relatively accurate model. When the policy management unit 613 uses this model to learn a policy, it is expected that the accuracy of control learning in a state where sufficient data is not available can be improved.
[0222] The model acquisition unit 611 can be realized, for example, by using the functions of the model management unit 221 in Fig. 1. The feedback information acquisition unit 612 can be realized, for example, by using the functions of the feedback information acquisition unit 240 in Fig. 1. The policy management unit 613 can be realized, for example, by using the functions of the policy management unit 222 in Fig. 1.
[0223] 19 is a diagram showing an example of the configuration of a display device according to an embodiment. In the configuration shown in FIG. With this configuration, the display unit 621 displays, for each item included in the information indicating the next state that is output by the model that simulates the environment in which the agent performs an action in response to input of information indicating a state and information indicating an action, an evaluation index value of the accuracy of that item in the information indicating the next state. The display unit 621 corresponds to an example of a display means.
[0224] According to the display device 620, a user can refer to the index values displayed by the display unit 621 and input feedback information to the display device 620 to improve items with low accuracy in the model output. In this respect, the display device 620 is expected to be able to efficiently improve the accuracy of the model. The display unit 621 can be realized using the functions of the display unit 120 in FIG. 1, for example.
[0225] 20 is a diagram showing another example of the configuration of the display device according to the embodiment. In the configuration shown in FIG. With this configuration, the display unit 631 displays an evaluation index value for the accuracy of the time series data for each of multiple pieces of time series data of the state in the environment and the agent's behavior, which is time series data of the input and output of a model that simulates the environment in which the agent performs its behavior. The display unit 631 is an example of a display means.
[0226] With the display device 630, a user can select and correct time-series data with low accuracy by referring to the index values displayed on the display unit 631. In this respect, the display device 630 is expected to be able to efficiently improve the accuracy of the model. The display unit 631 can be realized using the functions of the display unit 120 in FIG. 1, for example.
[0227] 21 is a diagram illustrating an example of a processing procedure in a learning method according to an embodiment. The learning method illustrated in FIG. 21 includes acquiring a model (step S611), acquiring feedback information (step S612), and learning a policy (step S613).
[0228] In acquiring a model (step S611), the computer acquires a model in which the state and the action are input and the next state is output, based on data linking the state of the environment in which the agent performs an action, the action that can be performed in that state, and the next state when that action is performed in that state.
[0229] In obtaining feedback information (step S612), the computer obtains feedback information, which is information for obtaining a more accurate model, based on the output of the obtained model. In learning a policy (step S613), the computer learns a policy that indicates the behavior of the agent depending on the state, using a model obtained using feedback information.
[0230] According to the learning method shown in Fig. 21, when the previously obtained data is insufficient, the lack of information can be compensated for by feedback information, and it is expected that a relatively accurate model can be obtained. By learning a policy using this model, it is expected that the accuracy of control learning in a state where sufficient data is not available can be improved.
[0231] 22 is a diagram showing an example of a processing procedure in a display method according to an embodiment. The display method shown in FIG. 22 includes performing display (step S621). In displaying (step S621), the computer displays, for each item included in the information indicating the next state that is output by the model simulating the environment in which the agent performs the action in response to input of information indicating the state and information indicating the action, an evaluation index value for the accuracy of that item in the information indicating the next state.
[0232] According to the display method shown in Fig. 22, a user can refer to the displayed index values and input feedback information to the computer to improve items with low accuracy in the model output. In this respect, the display method shown in Fig. 22 is expected to efficiently improve the accuracy of the model.
[0233] 23 is a diagram showing another example of the procedure of the process in the display method according to the embodiment. The display method shown in FIG. 23 includes performing display (step S631). In displaying (step S631), the computer displays an evaluation index value for the accuracy of time series data for each of multiple pieces of time series data of the state in the environment and the agent's behavior, which is time series data of the input and output of a model that simulates the environment in which the agent performs its behavior.
[0234] According to the display method shown in Fig. 23, the user can select and correct time series data with low accuracy by referring to the displayed index values. In this respect, the display method shown in Fig. 23 is expected to efficiently improve the accuracy of the model.
[0235] FIG. 24 is a schematic block diagram illustrating the configuration of a computer according to at least one embodiment. In the configuration shown in FIG. 24, a computer 700 includes a CPU 710, a main memory device 720, an auxiliary memory device 730, an interface 740, and a non-volatile recording medium 750.
[0236] One or more of the learning device 100, learning device 610, display device 620, and display device 630, or a portion thereof, may be implemented in a computer 700. In this case, the operation of each of the above-described processing units is stored in the auxiliary storage device 730 in the form of a program. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program. The CPU 710 also allocates storage areas in the main storage device 720 corresponding to each of the above-described storage units in accordance with the program. Communication between each device and other devices is performed by an interface 740 having a communication function and performing communication under the control of the CPU 710. The interface 740 also has a port for a nonvolatile storage medium 750, and reads and writes information from and to the nonvolatile storage medium 750.
[0237] When learning device 100 is implemented in computer 700, the operations of processing unit 190 and each of its components are stored in the form of a program in auxiliary storage device 730. CPU 710 reads the program from auxiliary storage device 730, loads it into main storage device 720, and executes the above-described processing in accordance with the program.
[0238] Furthermore, the CPU 710 allocates storage areas corresponding to the storage unit 180 and each unit thereof in the main storage device 720 in accordance with the program. Communication by the communication unit 110 is performed by the interface 740 having a communication function and performing communication under the control of the CPU 710. Display of images by the display unit 120 is performed by the interface 740 having a display device and displaying images under the control of the CPU 710. Reception of user operations by the operation input unit 130 is performed by the interface 740 having an input device and receiving the user operations.
[0239] When the learning device 610 is implemented in the computer 700, the operations of the model acquisition unit 611, the feedback information acquisition unit 612, and the policy management unit 613 are stored in the form of a program in the auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-mentioned processing in accordance with the program.
[0240] Furthermore, CPU 710 allocates a storage area in main memory 720 for the learning device 610 to perform processing in accordance with the program. Communication between learning device 610 and other devices is performed by interface 740, which has a communication function and operates under the control of CPU 710. Interaction between learning device 610 and a user is performed by interface 740, which has a display device and an input device, displaying various images under the control of CPU 710 and accepting user operations.
[0241] When the display device 620 is implemented in the computer 700, its operation is stored in the form of a program in the auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.
[0242] Furthermore, the CPU 710 allocates a storage area in the main memory device 720 for the display device 620 to perform processing in accordance with the program. Communication between the display device 620 and other devices is performed by the interface 740, which has a communication function and operates under the control of the CPU 710. Display of images by the display unit 621 is performed by the interface 740 having a display device and displaying images under the control of the CPU 710. Reception of user operations on the display device 620 is performed by the interface 740 having an input device and receiving the user operations.
[0243] When the display device 630 is implemented in the computer 700, its operation is stored in the form of a program in the auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.
[0244] Furthermore, the CPU 710 allocates a storage area in the main storage device 720 for the display device 630 to perform processing in accordance with the program. Communication between the display device 630 and other devices is performed by the interface 740, which has a communication function and operates under the control of the CPU 710. Display of images by the display unit 631 is performed by the interface 740 having a display device and displaying images under the control of the CPU 710. Reception of user operations on the display device 630 is performed by the interface 740 having an input device and receiving user operations.
[0245] One or more of the above-described programs may be recorded on nonvolatile recording medium 750. In this case, interface 740 may read the programs from nonvolatile recording medium 750. CPU 710 may then directly execute the programs read by interface 740, or may temporarily store the programs in main storage device 720 or auxiliary storage device 730 and then execute them.
[0246] Note that a program for executing all or part of the processing performed by learning device 100, learning device 610, display device 620, and display device 630 may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be read into a computer system and executed to perform the processing of each part. Note that the term "computer system" here includes hardware such as an OS (Operating System) and peripheral devices. Furthermore, "computer-readable recording media" refers to portable media such as flexible disks, optical magnetic disks, ROMs (Read Only Memory), and CD-ROMs (Compact Disc Read Only Memory), as well as storage devices such as hard disks built into computer systems. The program may be one that realizes part of the aforementioned functions, or may be one that can realize the aforementioned functions in combination with a program already stored in the computer system.
[0247] Although an embodiment of the present invention has been described above in detail with reference to the drawings, the specific configuration is not limited to this embodiment, and includes designs within the scope of the gist of the present invention.
[0248] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes.
[0249] (Appendix 1) a model acquisition means for acquiring a model in which the state and the action are input and the next state is output by learning using data in which the state of the environment in which the agent performs an action, the action that can be performed in that state, and the next state when that action is performed in that state are linked; a feedback information acquisition means for acquiring feedback information, which is information used for learning the obtained model or learning a new model in which a state and an action are input and a next state is output, based on the obtained model; a policy management means for learning a policy indicating an action of the agent depending on a state by using a model obtained by learning using the feedback information; A learning device comprising:
[0250] (Appendix 2) the feedback information acquisition means acquires the feedback information indicating constraints that must be satisfied by input data and output data of a model to be trained; the model acquisition means searches for a model using the constraints; 2. The learning device of claim 1.
[0251] (Appendix 3) The method further comprises an analysis means for calculating, for each item included in the information indicating the state, an evaluation index value of the accuracy of the item in the next state output by the obtained model; the feedback information acquisition means acquires the feedback information indicating constraints related to items with a relatively low evaluation of accuracy. 3. The learning device according to claim 2.
[0252] (Appendix 4) a display means for displaying the evaluation index value; an input means for receiving a user operation for inputting the feedback information; Furthermore, the feedback information acquisition means acquires the feedback information input by a user operation accepted by the input means after the display means starts displaying the evaluation index values. 4. The learning device according to claim 3.
[0253] (Appendix 5) the feedback information acquisition means acquires the feedback information indicating a modification to the input / output data of the model obtained; the model acquisition means performs model learning using the input / output data in which the correction is reflected. 2. The learning device of claim 1.
[0254] (Appendix 6) The method further comprises an analysis means for calculating an evaluation index value of the accuracy of the time-series data for each of the plurality of time-series data of the input and output of the model obtained, the feedback information acquisition means acquires the feedback information indicating a correction to time-series data whose accuracy is evaluated as relatively low. 6. The learning device according to claim 5.
[0255] (Appendix 7) a display means for displaying the evaluation index value; an input means for receiving a user operation for inputting the feedback information; Furthermore, the feedback information acquisition means acquires the feedback information input by a user operation accepted by the input means after the display means starts displaying the evaluation index values. 7. The learning device according to claim 6.
[0256] (Appendix 8) A display means for displaying, for each item included in the information indicating the next state output by the model simulating the environment in which the agent performs the action in response to input of the information indicating the state and the information indicating the action, an evaluation index value of the accuracy of that item in the information indicating the next state. A display device comprising:
[0257] (Appendix 9) A display means for displaying an evaluation index value for the accuracy of time-series data for each of a plurality of time-series data of states in the environment and actions of the agent, the time-series data being time-series data of inputs and outputs of a model that simulates an environment in which the agent performs actions. A display device comprising:
[0258] (Appendix 10) The computer Learning is performed using data that links the state of the environment in which the agent performs an action, the actions that can be performed in that state, and the next state when that action is performed in that state, to obtain a model in which the state and action are input and the next state is output. Based on the obtained model, feedback information is obtained, which is information used to learn the model or to learn a new model that uses a state and an action as input and a next state as output. learning a policy indicating an action of the agent depending on the state using a model obtained using the feedback information; A learning method that includes:
[0259] (Appendix 11) The computer A model that simulates the environment in which an agent performs an action receives input of information indicating the state and information indicating the action, and outputs the information indicating the next state. For each item included in the information indicating the next state, the model displays an evaluation index value for the accuracy of that item in the information indicating the next state. Display method including
[0260] (Appendix 12) The computer Displaying an evaluation index value for the accuracy of time-series data for each of a plurality of time-series data of the state in the environment and the behavior of the agent, which is time-series data of the input and output of a model that simulates the environment in which the agent performs its behavior Display method including
[0261] (Appendix 13) On the computer, Acquire a model that takes the state and action as input and the next state as output by learning using data that links the state of the environment in which the agent performs an action, the actions that can be performed in that state, and the next state when that action is performed in that state; Obtaining feedback information based on the obtained model, which is information used to learn the model or to learn a new model that uses a state and an action as input and a next state as output; learning a policy indicating an action of the agent depending on a state using a model obtained by learning using the feedback information; A recording medium that records a program for executing the program.
[0262] (Appendix 14) On the computer, A model that simulates an environment in which an agent performs an action receives input of information indicating a state and information indicating an action, and outputs the information indicating a next state, and displays an evaluation index value of the accuracy of that item in the information indicating the next state. A recording medium that records a program for executing the program.
[0263] (Appendix 15) On the computer, Displaying an evaluation index value of the accuracy of time-series data for each of a plurality of time-series data of the state in the environment and the behavior of the agent, which is time-series data of the input and output of a model that simulates the environment in which the agent performs its behavior. A recording medium that records a program for executing the program. [Industrial Applicability]
[0264] The present invention may be applied to a learning device, a display device, a learning method, and a recording medium. [Explanation of symbols]
[0265] 1. Learning System 2. Data Collection System 3. Control System 100, 610 Learning Device 110 Communications Department 120, 621, 631 display section 130 Operation input section 180 Storage section 181 Data storage unit 182 Model Memory Unit 183 Strategy Memory Unit 190 Processing section 210 Data Management Department 220 Learning Department 221 Model Management Department 222, 613 Policy Management Department 230 Analysis Department 240, 612 Feedback information acquisition section 300 Data Collection Device 400 control device 620, 630 display device 611 Model Acquisition Department
Claims
1. a model acquisition means for acquiring a model in which the state and the action are input and the next state is output by learning using data in which the state of the environment in which the agent executes an action, the action executable in that state, and the next state when that action is performed in that state are linked; a feedback information acquisition means for acquiring feedback information, which is information used for learning the obtained model or for learning a new model in which a state and an action are input and a next state is output, based on the obtained model; a policy management means for learning a policy indicating an action of the agent according to a state by using a model obtained by learning using the feedback information; A learning device comprising:
2. the feedback information acquisition means acquires the feedback information indicating constraints that must be satisfied by input data and output data of a model to be trained; the model acquisition means searches for a model using the constraints; The learning device according to claim 1 .
3. The method further comprises an analysis means for calculating, for each item included in the information indicating the state, an evaluation index value of the accuracy of the item in the next state output by the obtained model; the feedback information acquisition means acquires the feedback information indicating constraints related to items with a relatively low evaluation of accuracy. The learning device according to claim 2 .
4. a display means for displaying the evaluation index value; an input means for receiving a user operation for inputting the feedback information; Furthermore, the feedback information acquisition means acquires the feedback information input by a user operation accepted by the input means after the display means starts displaying the evaluation index values. The learning device according to claim 3 .
5. the feedback information acquisition means acquires the feedback information indicating a modification to the input / output data of the model obtained; the model acquisition means performs model learning using the input / output data in which the correction is reflected. The learning device according to claim 1 .
6. The method further comprises an analysis means for calculating an evaluation index value of the accuracy of the time-series data for each of the plurality of time-series data of the input and output of the model obtained, the feedback information acquisition means acquires the feedback information indicating a correction to time-series data whose accuracy is evaluated as relatively low. The learning device according to claim 5 .
7. a display means for displaying the evaluation index value; an input means for receiving a user operation for inputting the feedback information; Furthermore, the feedback information acquisition means acquires the feedback information input by a user operation accepted by the input means after the display means starts displaying the evaluation index values. The learning device according to claim 6.
8. The computer Learning is performed using data that links the state of the environment in which the agent performs an action, the actions that can be performed in that state, and the next state when that action is performed in that state, to obtain a model in which the state and action are input and the next state is output. Based on the obtained model, feedback information is obtained, which is information used to learn the model or to learn a new model that uses a state and an action as input and a next state as output. learning a policy indicating an action of the agent depending on the state using the model obtained by learning using the feedback information; A learning method that includes:
9. On the computer, Acquire a model that takes the state and action as input and the next state as output by learning using data that links the state of the environment in which the agent performs an action, the actions that can be performed in that state, and the next state when that action is performed in that state; Obtaining feedback information based on the obtained model, which is information used to learn the model or to learn a new model that uses a state and an action as input and a next state as output; learning a policy indicating an action of the agent depending on a state using a model obtained by learning using the feedback information; A program to execute.
Citation Information
Patent Citations
GP world model using strategy model to assist in training and training method thereof
CN114492215A
Control device for control target with combustion device and control device for plant with boiler
JP2007271187A
Robot control device, robot device, parameter adjustment method for robot control, and program
JP2020032481A
Calculator system and mathematical model generation support method
JP2021064049A