Learning method and recording medium
By employing model-based reinforcement learning and hybrid model variational inference, the learning method for agent action control is improved, thereby enhancing learning efficiency and the appropriateness of action sequences.
Patent Information
- Application Number
- CN202010606260.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-25
- Filing Date
- 2020-06-29
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2040-06-29
AI Technical Summary
Existing methods for learning agent action control have room for improvement, especially regarding the unreliability of dynamic models.
A model-based reinforcement learning approach is adopted. Supervised learning is performed by obtaining time-series data of the agent's state and actions to build a dynamic model. Variational inference of the hybrid model is used to derive multiple action sequence candidates, and finally one is selected as the agent's action sequence.
It improves the learning efficiency of agent action control, shortens the learning time, outputs more appropriate action sequences, and approaches the optimal action sequence.
Smart Images

Figure CN112183766B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to learning methods and recording media. Background Technology
[0002] As a learning method for controlling intelligent agents, there is a method that employs a dynamic model that considers unreliability (see Non-Patent Document 1). Here, an intelligent agent refers to an agent that takes action in response to its environment.
[0003] Non-patent literature 1: K. Chua, R. Calandra, R. McAllister, and S. Levine. "Deepreinforcement learning in a handful of trials using probabilistic dynamics models. In NeurIPS, 2018."
[0004] However, there is room for improvement in the learning methods used to control the actions of intelligent agents. Summary of the Invention
[0005] Therefore, the present invention provides a learning method for improving the actions of an intelligent agent.
[0006] The learning method of one technical solution of the present invention uses a model-based reinforcement learning method for learning the actions of an agent to obtain time-series data representing the state and actions of the agent when it takes action; supervised learning is performed using the obtained time-series data to construct a dynamic model; based on the dynamic model, multiple candidate action sequences of the agent are derived by using variational inference of a mixture model as a variational distribution; and one of the derived multiple candidates is selected as the action sequence of the agent and output.
[0007] In addition, these inclusive or specific technical solutions can also be implemented by systems, devices, integrated circuits, computer programs or computer-readable CD-ROMs and other recording media, or by any combination of systems, devices, integrated circuits, computer programs and recording media.
[0008] Invention Effects
[0009] The learning method of the present invention can improve the learning methods used to control the actions of intelligent agents. Attached Figure Description
[0010] Figure 1 This is a block diagram illustrating the functional structure of the learning device in the implementation method.
[0011] Figure 2 This is an explanatory diagram showing the time sequence of the states and actions of the intelligent agent in the implementation method.
[0012] Figure 3 This is an explanatory diagram showing the variation distribution of the implementation method.
[0013] Figure 4 This is an explanatory diagram illustrating the concept of variational inference performed by the inference unit in the implementation method.
[0014] Figure 5 This is an explanatory diagram illustrating the concept of variational inference performed by the inference unit in correlation techniques.
[0015] Figure 6 This is an explanatory diagram that compares multiple candidate technologies determined by the inference unit of the implementation method with related technologies.
[0016] Figure 7 This is a flowchart illustrating the learning method of the implementation method.
[0017] Figure 8 This is an explanatory diagram that compares the change of cumulative reward over time in the learning method of the implementation method with the correlation technology.
[0018] Figure 9 This is an explanatory diagram that compares the convergence value of the cumulative reward in the learning method of the implementation method with the correlation technique.
[0019] Label Explanation
[0020] 10 Learning Devices
[0021] 11 Acquisition Department
[0022] 12. Study Department
[0023] 13 Storage Department
[0024] 14. Inference Department
[0025] 15 Output Section
[0026] 17 Dynamic Model
[0027] 20 intelligent agents
[0028] Paths 30, 31, 32, 33, 34, 35 Detailed Implementation
[0029] The learning method of one technical solution of the present invention uses a model-based reinforcement learning method for learning the actions of an agent to obtain time-series data representing the state and actions of the agent when it takes action; supervised learning is performed using the obtained time-series data to construct a dynamic model; based on the dynamic model, multiple candidate action sequences of the agent are derived by using variational inference of a mixture model as a variational distribution; and one of the derived multiple candidates is selected as the action sequence of the agent and output.
[0030] According to the above technical solution, since variational inference using a mixture model as a variational distribution is performed, multiple candidate action sequences for the agent can be derived. Each of these candidate action sequences, compared to the case of variational inference using a single model as a variational distribution, shows a smaller difference from the optimal action sequence, and the convergence speed is faster when deriving these multiple candidates. Furthermore, it is assumed that one of the candidate action sequences is used as the agent's action sequence. In this way, a more appropriate action sequence can be output in a shorter time. Thus, according to the above technical solution, the learning method used for agent control can be improved.
[0031] For example, new time-series data representing the state and actions of the agent when it acts according to the output action sequence can also be obtained as the time-series data.
[0032] According to the above technical solution, since the state and actions of the agent controlled by the action sequence output as the result of learning are used in new learning, the action sequence constructed as the result of learning can be made closer to a more appropriate action sequence. This further improves the learning method used for controlling the agent.
[0033] For example, when deriving the above multiple candidates, the mixing ratio of the probability distributions corresponding to each of the above multiple candidates in the overall mixture model can also be derived together with the above multiple candidates; when selecting one of the above candidates, the candidate corresponding to the probability distribution with the largest mixing ratio is selected from the above multiple candidates as the above candidate.
[0034] According to the above technical solution, since the candidate action sequence corresponding to the probability distribution with the largest mixing ratio in the hybrid model is used as the agent's action sequence, the difference from the optimal action sequence can be further reduced. Therefore, according to the above technical solution, the learning method used for agent control can be improved to output an action sequence that is closer to the optimal action sequence.
[0035] For example, the above mixture model can also be a mixture Gaussian distribution.
[0036] According to the above technical solution, since a mixture Gaussian distribution is specifically used as a mixture model, it is easier to derive multiple candidates based on variational inference. Therefore, it is easier to improve the learning methods used for agent control.
[0037] For example, the dynamic model described above can also be an integration of multiple neural networks.
[0038] According to the above technical solution, because the integration of multiple neural networks is used as a dynamic model, it is easier to construct a dynamic model with high estimation accuracy. Therefore, it is easier to improve the learning methods used for agent control.
[0039] Furthermore, the recording medium of one technical solution of the present invention is a program recording medium used to enable a computer to execute the above-described learning method.
[0040] The above technical solutions achieve the same effect as the learning methods described above.
[0041] In addition, these inclusive or specific technical solutions can also be implemented by systems, devices, integrated circuits, computer programs or computer-readable CD-ROMs and other recording media, or by any combination of systems, devices, integrated circuits, computer programs and recording media.
[0042] Hereinafter, the embodiments will be described in detail with reference to the accompanying drawings.
[0043] Furthermore, the embodiments described below are all inclusive or specific examples. The numerical values, shapes, materials, constituent elements, arrangement and connection methods of constituent elements, steps, and order of steps shown in the following embodiments are examples and are not intended to limit the invention. In addition, any constituent elements in the following embodiments that are not described in the independent claims representing the highest-level concept are described as arbitrary constituent elements.
[0044] (Implementation Method)
[0045] In this embodiment, a learning device and a learning method for improving the learning method used to control the actions of an intelligent agent will be described.
[0046] Figure 1 This is a block diagram illustrating the functional structure of the learning device 10 in this embodiment.
[0047] Figure 1The learning device 10 shown is a device for learning the actions of agent 20 using model-based reinforcement learning. The learning device 10 acquires information about the actions of agent 20, uses the acquired information to perform model-based reinforcement learning, and is able to control the actions of agent 20.
[0048] The intelligent agent 20 is a device that selectively and sequentially adopts one of a plurality of states and selectively and sequentially performs one of a plurality of actions. The intelligent agent 20 is, for example, an industrial machine equipped with a robotic arm that processes an object. In this case, the coordinates of the robotic arm, information related to the object, and information indicating the processing status correspond to the aforementioned "states." Furthermore, the information used to move the robotic arm (the coordinates of the target position of the robotic arm) and the information used to change the state of the industrial machine correspond to the aforementioned "actions."
[0049] The agent 20 provides the learning device 10 with time-series data representing the state of the agent 20 as it performs a series of actions. Furthermore, the agent 20 is able to act according to the action sequence output by the learning device 10.
[0050] like Figure 1 As shown, the learning device 10 includes an acquisition unit 11, a learning unit 12, a storage unit 13, an inference unit 14, and an output unit 15. Each functional unit of the learning device 10 can be implemented by the CPU (Central Processing Unit) (not shown) of the learning device 10 using the memory (not shown) to execute a predetermined program.
[0051] The acquisition unit 11 is a functional unit that acquires time-series data representing information about the agent 20 during its actions. The information acquired by the acquisition unit 11 about the agent 20 includes at least the state and actions of the agent 20. The sequence of actions of the agent 20 acquired by the acquisition unit 11 can be any sequence, either a randomly set sequence of actions or an action sequence output by the output unit 15 (described later).
[0052] Furthermore, the acquisition unit 11 can also control the agent 20 using the action sequence output by the output unit 15 (described later), thereby acquiring new time-series data representing the state and actions of the agent 20 as the aforementioned time-series data. The acquisition unit 11 acquires one or more of the aforementioned time-series data.
[0053] The learning unit 12 is a functional unit that constructs a dynamic model 17 by performing supervised learning using the time series data acquired by the acquisition unit 11. The learning unit 12 saves the constructed dynamic model 17 to the storage unit 13.
[0054] When conducting supervised learning, the learning unit 12 uses reward values set according to the action sequence of the agent 20. For example, if the agent 20 is an industrial machine equipped with a robotic arm, the more appropriate the series of actions of the robotic arm, the higher the reward value is set. In addition, the higher the quality of the object processed by the industrial machine, the higher the reward value is set.
[0055] An example of dynamic model 17 is the ensemble of multiple neural networks. That is, the learning unit 12 determines the coefficients (weights) of the filters in each layer of the multiple neural networks based on either a sequence of actions identical to the sequence of actions of agent 20 acquired by the acquisition unit 11, or a sequence of actions that is a subset of the action sequence of agent 20 acquired by the acquisition unit 11 but different from each other. In this case, the output of the ensemble of multiple neural networks is the result of aggregating the output data of each neural network to the input data using an aggregation function. The aggregation function can contain various operations, such as operations that average the output data of each neural network or operations that take the majority decision of each output data.
[0056] Storage unit 13 is a storage device that stores the dynamic model 17 constructed by learning unit 12. The dynamic model 17 is saved by learning unit 12 and read by inference unit 14. Storage unit 13 is implemented by memory or storage device.
[0057] The inference unit 14 is a functional unit that derives multiple candidate action sequences for agent 20 based on the dynamic model 17. When deriving multiple candidate action sequences for agent 20, the inference unit 14 utilizes variational inference, which employs a mixture model as a variational distribution. Thus, the inference unit 14 derives multiple action sequences as action sequences for agent 20.
[0058] The mixture model used in variational inference in the inference unit 14 is, for example, a mixture Gaussian distribution, but it is not limited to this case. For example, as an example of a mixture model when the action sequence is a vector of discrete values, there is a mixture categorical distribution.
[0059] Output unit 15 is a functional unit that outputs the action sequence of agent 20. Output unit 15 outputs one of the multiple candidates derived from inference unit 14 as the action sequence of agent 20. For example, suppose the action sequence output by output unit 15 is input to agent 20, and agent 20 acts according to the action sequence.
[0060] In addition, the action sequence output by the output unit 15 can also be managed as numerical data by other devices (not shown).
[0061] Furthermore, when deriving multiple candidates, the inference unit 14 can also derive the mixing ratio of the probability distributions corresponding to each candidate in the overall mixture model. Moreover, when selecting a candidate, the output unit 15 can select the candidate corresponding to the probability distribution with the largest mixing ratio in the overall mixture model from among the derived multiple candidates.
[0062] Figure 2 This is an explanatory diagram illustrating an example of a time series representing the state of the agent 20 in this embodiment.
[0063] exist Figure 2 In the example, state S is represented as a sequence of states of agent 20. t and S t+1 Here, let the state of agent 20 at time t be state S. t Let the state of agent 20 at time t+1 be S. t+1 .
[0064] also, Figure 2 Action a shown t This indicates that agent 20 is in state S. t The actions taken below. Figure 2 The reward r shown t This indicates that the agent 20 starts from state S. t Transfer to S t+1 And the reward received.
[0065] In this way, the state and actions of agent 20, as well as the reward value obtained by agent 20, can be graphically represented. The inference unit 14 derives a series of actions that enable agent 20 to take. t The reward r at that time t The action sequence that maximizes the expected value of the cumulative value (also known as cumulative reward).
[0066] The following describes the method for deriving the action sequence of the inference unit 14.
[0067] Figure 3 This is an explanatory diagram showing the cumulative reward for each action sequence in this embodiment. Here, the explanation is based on the case where the action sequence is a vector with continuous values as elements, but the same explanation applies even for discrete values.
[0068] Figure 3 This is a graph with the action sequence as the horizontal axis and the cumulative reward within each action sequence as the vertical axis. For example... Figure 3 As shown, cumulative rewards typically take multiple maxima for the action sequence.
[0069] Figure 4 This is an explanatory diagram illustrating the concept of variational inference performed by the inference unit 14 in this embodiment.
[0070] Figure 4 It is a graph with the action sequence on the horizontal axis and the probability distribution corresponding to each action sequence on the vertical axis. Additionally, Figure 4 The probability distribution along the vertical axis can be based on Figure 3 The cumulative reward on the vertical axis is transformed through a prescribed calculation.
[0071] The inference unit 14 derives candidates for the optimal action sequence of agent 20 by performing variational inference using a Gaussian Mixture Model (GMM) as a variational distribution.
[0072] Figure 4 The solid line shown represents the line for Figure 3 The cumulative reward of the action sequence shown is obtained by transforming it into a probability distribution. As a means of transforming the cumulative reward into a probability distribution, methods such as mapping the cumulative reward using an exponential function can be cited.
[0073] Figure 4 The dashed line shown represents a mixture Gaussian distribution, which is a probability distribution resulting from the mixing (combination) of multiple Gaussian distributions. A mixture Gaussian distribution is determined by a set of parameters including the mean, variance, and peak value of each of the individual Gaussian distributions that constitute the mixture Gaussian distribution.
[0074] The inference unit 14 determines the aforementioned parameters of the mixture Gaussian distribution through variational inference.
[0075] Specifically, the inference unit 14 repeatedly determines multiple parameters of the mixture Gaussian distribution to make it so that... Figure 4 The calculation involves reducing the difference between the probability distribution represented by the dashed line (Gaussian mixture distribution) and the probability distribution represented by the solid line. The number of times the calculation is executed can be set arbitrarily, or the iteration can be terminated when the cumulative improvement rate of each iteration falls within a certain value (e.g., ±1%).
[0076] The inference unit 14 determines the parameters of each of the multiple Gaussian distributions contained in the mixture Gaussian distribution by performing variational inference as described above. Since the multiple Gaussian distributions determined correspond to action sequences respectively, the determination of the above parameters is equivalent to the inference unit 14 deriving multiple candidates for action sequences of the agent 20.
[0077] exist Figure 4 In the example, the inference unit 14 obtains the mean values μ1 and μ2, the variances σ1 and σ2, and the mixing ratios π1 and π2 as parameters of two Gaussian distributions. The parameters of the action sequence determined in this way become multiple candidates for the action sequence of the agent 20.
[0078] The following explanation compares the variational inference performed by the inference unit 14 with that in the correlation technique. The variational inference in the correlation technique uses a single model (e.g., a single Gaussian distribution) instead of a mixture model (mixture Gaussian distribution).
[0079] Figure 5 This is an explanatory diagram illustrating the concept of variational inference in correlation techniques. Figure 5 The display of the vertical and horizontal axes and Figure 4 It's the same.
[0080] The inference unit of the correlation technology derives candidates for the optimal action sequence of agent 20 by performing variational inference using a Gaussian distribution (i.e., a single Gaussian distribution that is not a mixture of Gaussian distributions) as a variational distribution.
[0081] Figure 5 The dashed line shown represents a Gaussian distribution. A Gaussian distribution is determined by the combination of the mean, variance, and peak values of the probability distribution.
[0082] The inference part determines the above parameters of the Gaussian distribution through variational inference.
[0083] Specifically, the inference part repeatedly determines multiple parameters of the mixture Gaussian distribution to make it possible to... Figure 5 The operation reduces the difference between the probability distribution represented by the dashed line (Gaussian distribution) and the probability distribution represented by the solid line. The number of times the operation is performed is the same as in the case of inference unit 14.
[0084] Therefore, the inference department only obtained one action sequence μ0.
[0085] like Figure 4 and Figure 5 As shown, the variational distribution represented by the dashed line (equivalent to...) Figure 4 The mixture Gaussian distribution and Figure 5 The difference between the Gaussian distribution in the diagram and the probability distribution represented by the solid line (the probability distribution corresponding to the action sequence) tends to decrease when using a mixture Gaussian distribution (see [reference]). Figure 4 This is because, when using a mixture Gaussian distribution, by adjusting the parameters of each Gaussian distribution, the difference from the probability distribution represented by the solid line can be reduced.
[0086] Figure 6 This is an explanatory diagram that conceptually represents the multiple candidates determined by the inference unit 14 in this embodiment.
[0087] Here, we will use the pathfinding problem as an example to illustrate this. This problem is... Figure 7The problem involves navigating within a rectangular area while avoiding obstacles marked by black circles (●), and tracing a path from the starting position (“S” in the diagram) to the target position (“G” in the diagram). The direction of travel at each position on the path corresponds to an “action”.
[0088] exist Figure 6 In this process, for the multiple paths derived by the inference unit 14, the explanation is carried out while comparing them with a path derived by the inference unit of the association technique.
[0089] Figure 6 (a) represents a path 30 derived from the inference part of the related technology. Figure 6 The path 30 shown in (a) corresponds to Figure 5 The action sequence μ0 is shown.
[0090] Figure 6 (b) represents the multiple paths 31, 32, 33, 34, and 35 derived by the inference unit 14 in this embodiment. Here, an example is given where the inference unit 14 uses a mixture of five Gaussian distributions as a variational distribution. In this case, the number of derived paths is 5.
[0091] In addition, Figure 6 In (b), the width of the line representing path 31, etc., represents the proportion of the Gaussian distribution that is mixed with the Gaussian distribution corresponding to that path. In this example, the line of path 31 is drawn to be thicker than the other lines, indicating that the proportion of the Gaussian distribution that is mixed with the Gaussian distribution corresponding to path 31 is higher than that of the other paths 32-35.
[0092] Output unit 15 can output from Figure 6 Choose one of the paths shown in (b), for example, path 31.
[0093] Figure 7 This is a flowchart illustrating the learning method performed by the learning device 10 of this embodiment. This learning method is a learning method for the actions of the agent 20 using model-based reinforcement learning.
[0094] like Figure 7 As shown, in step S101, time-series data representing the state and actions of agent 20 during its actions are obtained.
[0095] In step S102, a dynamic model 17 is constructed by using supervised learning through the time series data obtained in step S101.
[0096] In step S103, based on dynamic model 17, multiple candidates for action sequences of agent 20 are derived by using variational inference of a hybrid model as a variational distribution.
[0097] In step S104, one candidate selected from the multiple exported candidates is output as the action sequence of agent 20.
[0098] Through the above series of processes, the learning device 10 improves the learning method used to control the actions of the agent 20.
[0099] The effects of the learning method described in this embodiment will be explained below.
[0100] Figure 8 This is an explanatory diagram that compares the change in cumulative reward of the learning method of this embodiment over time with that of related technologies.
[0101] The cumulative reward for learning about the learning device 10 of this embodiment is represented by an example of the simulation result using an existing simulator. The existing simulator is the physics calculation engine MuJoCo. Figure 8 (a) through (d) are the results of four different tasks (i.e., HalfCheetah, Ant, Hopper, and Walker2d) provided by OpenAI Gym, an evaluation platform for reinforcement learning, which is used to evaluate actions in the simulators described above.
[0102] Figure 8 The thick solid line shown represents the change in cumulative reward of the learning device 10 in this embodiment. Furthermore, Figure 8 The thin solid lines shown represent the changes in the cumulative reward of the learning device related to the technology. Furthermore, dashed lines represent the convergence values for the respective changes in the cumulative reward.
[0103] like Figure 8 As shown, in Figure 8 In all the tasks shown in (a) to (d), the learning device 10 of this embodiment starts increasing the cumulative reward corresponding to the elapsed time earlier than the learning device of the related technology, and converges to the convergence value in a shorter time (i.e., the convergence speed is faster). In addition, the learning device 10 of this embodiment has a larger convergence value of cumulative reward compared with the learning device of the related technology.
[0104] This means that the learning device 10 of this embodiment has improved learning performance compared to learning devices of related technologies.
[0105] Figure 9 This is an explanatory diagram that compares the convergence value of the cumulative reward of the learning method in this embodiment with that of the correlation technique. Figure 9 In the Half-Cheetah task, the convergence value of the cumulative reward is represented by the number M of Gaussian distributions contained in the Gaussian distribution and the mixture of Gaussian distributions.
[0106] Furthermore, the case where M=1 corresponds to a single Gaussian distribution, which corresponds to the learning method of correlation techniques. Additionally, the cases where M=3, 5, or 7 correspond to a mixture of Gaussian distributions, which corresponds to the learning method of this embodiment.
[0107] like Figure 9 As shown, compared to variational inference using a single Gaussian distribution (M=1), the convergence value of the cumulative reward is larger when using a mixture of Gaussian distributions (M=3, 5, or 7). Furthermore, the convergence value of the cumulative reward varies depending on the value of M. Additionally, the convergence value of the cumulative reward exhibits properties that differ depending on the method of defining the reward, the type of task, etc.
[0108] As described above, according to the learning method of this embodiment, since variational inference using a hybrid model as a variational distribution is performed, multiple candidate action sequences for the agent can be derived. Each of these candidate action sequences, compared to the case of variational inference using a single model as a variational distribution, shows a smaller difference from the optimal action sequence, and the convergence speed when deriving these multiple candidates is faster. Furthermore, it is assumed that one of the multiple candidate action sequences is used as the agent's action sequence. In this way, a more appropriate action sequence can be output in a shorter time. Thus, according to the above technical solution, the learning method used for controlling an agent can be improved.
[0109] Furthermore, since the state and actions of the agent controlled using the action sequences output as a result of learning are applied to new learning, the action sequences constructed as a result of learning can be made closer to appropriate action sequences. This allows for further improvement of the learning methods used to control agents.
[0110] Furthermore, by using the candidate action sequence corresponding to the probability distribution with the largest mixing ratio in the hybrid model as the agent's action sequence, the difference from the optimal action sequence can be reduced. Therefore, according to the above technical solution, the learning method used for agent control can be improved to output an action sequence that is closer to the optimal action sequence.
[0111] Furthermore, since a mixture Gaussian distribution is specifically used as a mixture model, it is easier to derive multiple candidates based on variational inference. Therefore, it is easier to improve the learning methods used for agent control.
[0112] Furthermore, because the integration of multiple neural networks is used as a dynamic model, it is easier to construct dynamic models with high estimation accuracy. Therefore, it is easier to improve learning methods used for agent control.
[0113] Furthermore, in the above embodiments, each component may be constructed using dedicated hardware, or implemented by executing software programs suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory. Here, the software for the content management system, etc., implementing the above embodiments is a program like the following.
[0114] That is, the program is a program that enables a computer to execute a learning method, which is a model-based reinforcement learning method for learning the actions of an agent, to obtain time-series data representing the state and actions of the agent when it takes action; to construct a dynamic model by performing supervised learning using the obtained time-series data; to derive multiple candidate action sequences of the agent by using variational inference of a mixture model as a variational distribution based on the dynamic model; and to output one of the multiple candidate sequences as the action sequence of the agent.
[0115] The learning methods for one or more technical solutions have been described above based on the embodiments, but the present invention is not limited to these embodiments. As long as they do not depart from the spirit of the present invention, various modifications that can be conceived by those skilled in the art, or forms constructed by combining the constituent elements of different embodiments, can also be included within the scope of one or more technical solutions.
[0116] The present invention provides a learning device for controlling the actions of intelligent agents.
Claims
1. A learning method, which is an agent-based action learning method using model-based reinforcement learning, characterized in that, Obtain time-series data representing the state and actions of the aforementioned intelligent agent at the time of its actions; A dynamic model is constructed by using the obtained time series data for supervised learning. Based on the above dynamic model, multiple candidate action sequences of the agent are derived by using variational inference of the hybrid model as a variational distribution. One of the exported candidates will be selected as the action sequence for the agent and output. When constructing the dynamic model, the dynamic model is constructed by using the obtained time series data and supervising learning with cumulative rewards. The cumulative rewards are determined based on the action sequence of the agent and are the cumulative value of the rewards when the agent performs a series of actions included in the action sequence.
2. The learning method as described in claim 1, characterized in that, Furthermore, new time-series data representing the state and actions of the aforementioned agent when it acts according to the output action sequence are obtained, and these are used as the aforementioned time-series data.
3. The learning method as described in claim 1, characterized in that, When deriving the above multiple candidates, the mixing ratio of the probability distributions corresponding to each of the above multiple candidates in the overall mixture model is also derived together with the above multiple candidates. When selecting one of the above candidates, among the derived multiple candidates, the candidate corresponding to the probability distribution with the largest mixing ratio is selected as the above candidate.
4. The learning method as described in claim 1, characterized in that, The above mixture model is a mixture Gaussian distribution.
5. The learning method as described in claim 1, characterized in that, The dynamic model described above is an integration of multiple neural networks.
6. A program recording medium, on which a program is recorded, characterized in that, The above procedure is used to enable a computer to execute the learning method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Track predication method based on Gauss mixture time series model
CN107610464A