Agent action generation strategy training method based on optimism principle and deep model
Patent Information
- Application Number
- CN202311725468.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-15
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-12-15
AI Technical Summary
[0006]发明目的:针对机器人行走控制任务中基于模型的方法的探索效率不足导致行走策略陷入次优表现不佳和训练成本高的问题,本方法引入乐观性原则来提升探索效率
[0024]在机器人行走控制决策任务中,已有方法的探索效率不足导致机器人控制策略经常陷入次优且样本利用率不高,限制了行走机器人在现实世界中的应用,为了解决该问题,本发明通过上述的三个模块引入乐观性原则来提升机器人的探索效率,进而提升机器人行走的性能并降低了训练成本。具体地,本发明提供了一种基于乐观性原则和深度模型的机器人控制策略训练方法,可用于机器人行走控制任务,该方法相较于先前方法,首次将乐观性原则以计算可行的方式应用在深度强化学习中,避免了原有乐观性框架中的偏差,降低了优化难度,并提升了机器人行走控制策略的探索效率,能有效地避免机器人行走控制策略陷入次优解,进而提升策略的性能和样本利用率,降低训练成本,该方法具备广泛的应用前景,可用于机器人行动控制的各项任务中。
Smart Images

Figure CN117689039B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a training method for an agent's action generation strategy based on the optimistic principle and a deep model, specifically a training method for a robot walking strategy based on the optimistic principle and a deep model, belonging to the field of robot action control technology. Background Technology
[0002] Deep reinforcement learning is a deep learning paradigm that combines deep neural networks with traditional reinforcement learning algorithms. It has achieved significant results in multiple fields and is considered a highly promising technology. In the field of robot motion control, deep reinforcement learning methods are frequently used to train control strategies for robots, helping them complete various tasks, such as teaching bipedal robots how to walk on the ground. However, deep reinforcement learning suffers from low sample utilization in this area, hindering its further application in real-world robot control scenarios where environmental sampling is costly.
[0003] Model-based methods are an important strategy in deep reinforcement learning for improving sample efficiency. The core idea is to learn a model from the trajectories obtained from the robot's interaction with the real environment, making the model as similar to the real environment as possible. During policy training, this model can be used to simulate the robot's interaction with the real environment, thereby generating multiple planned trajectory samples to improve the policy. Due to its high sample utilization rate, model-based methods have been widely used in the field of robot control.
[0004] While model-based methods improve sample utilization by leveraging models, they often suffer from insufficient exploration efficiency. Exploration efficiency refers to how effectively a robot explores unknown areas during the learning process to discover better strategies. On one hand, if the robot's exploration is insufficient, the learned model may have significant uncertainty, failing to reliably simulate the robot's interaction with the real environment, thus leading to suboptimal control strategies. On the other hand, if the robot tends to explore unknown areas in the real environment, the model's uncertainty is significantly reduced, but the time cost and sample size cost of converging to the optimal strategy increase. Therefore, in robot walking control scenarios, how to efficiently explore based on the learned model is a key challenge.
[0005] The most promising approach currently is to improve exploration efficiency based on the optimism principle. The optimism principle states that when faced with uncertain actions or strategies, robots perceive them as having high potential value, thus encouraging them to try unknown actions or strategies. This principle helps prevent robot control strategies from prematurely falling into local optima or becoming overly conservative. However, applying the optimism principle to deep models presents two challenges: difficulty in estimating model confidence and difficulty in optimization. This has prevented any work from successfully incorporating the optimism principle into deep models to improve robot exploration efficiency, limiting the performance and sample utilization of robot walking control strategies. Summary of the Invention
[0006] Objective: To address the problem of insufficient exploration efficiency in model-based methods for robot walking control, leading to suboptimal performance and high training costs, this invention introduces the optimism principle to improve exploration efficiency. Specifically, this invention provides a training method for agent action generation strategies based on the optimism principle and deep models. This method utilizes the model posterior as an estimate of uncertainty and uses the gradient of the value model to ensure the optimization of the optimistic model in unconstrained objectives. This solves two problems in applying the optimism principle to robot control: difficulty in model confidence estimation and optimization, thus providing a new approach to alleviate the problem of insufficient exploration efficiency in robot walking control tasks.
[0007] This invention overcomes the problem that the original optimistic framework cannot be applied to robot control decision-making scenarios through innovative technical improvements. It introduces the optimistic principle to improve the exploration efficiency in robot walking control tasks, helps the walking strategy escape suboptimal solutions to improve the performance of the robot walking strategy, and also has a high sample utilization rate, reducing the training cost caused by collecting samples in real environment.
[0008] Technical Solution: This paper presents a training method for intelligent agent action generation strategies based on the optimism principle and deep models. Taking robot walking tasks as the specific implementation object of intelligent agent action generation, this method applies the optimism principle to model-based deep reinforcement learning algorithms, encouraging the robot to explore walking strategies by making the model more optimistic. Currently, although model-based deep reinforcement learning is widely used in robot walking control, the exploration efficiency of existing algorithms is low, limiting the sample utilization and performance of robot walking control strategies. Therefore, this method introduces the optimism principle to improve exploration efficiency. To overcome the model uncertainty prediction problem and optimization difficulty when applying the optimism principle to robot walking control, this method uses the model's Bayesian posterior on real samples as an estimate of model confidence, aiming to avoid the performance bias and optimization difficulties caused by manually constructing confidence sets in existing methods. During robot walking strategy optimization, branch programming is used to generate planned trajectory samples based on the model, and the training is performed using a sample set obtained by mixing planned trajectory samples and real trajectory samples, thereby reducing optimization costs and difficulty. The improvements in model and policy optimization reduce the optimization difficulty of the method. The optimistic principle is combined with the model-based deep reinforcement learning algorithm in a computationally feasible way. This novel approach alleviates the problem of insufficient exploration efficiency of existing methods in robot walking control tasks, thereby improving the performance and sample utilization of robot walking control strategies and reducing training costs.
[0009] The training method for agent action generation strategy based on the optimistic principle and deep models takes the robot walking task as the specific implementation object of the agent action generation task. First, it is necessary to model the robot walking task as a Markov decision process.<S,A,T,R,γ> Here, S represents the state space, which needs to contain all the information used in the robot's walking process, including but not limited to map obstacle location information and robot state information; A represents the action space, which contains all actions that the robot can control, such as applying torque to the leg joints; T represents the state transition function, which gives the probability distribution T(·|s,a) of the new state to which the robot transitions after taking any action a∈A in any state s∈S; R represents the reward function, which gives the reward R(s,a) received by the robot after taking any action a∈A in any state s∈S; γ represents the discount factor, which is used to balance long-term and short-term rewards; the robot is used for training. The interactive environment for practicing the walking strategy is the robot walking simulation environment E. It simulates the interaction between a real robot walking and its environment. This robot walking simulation environment E simulates all the information of a real robot walking, providing all the information for the Markov decision process. In this environment, the robot is a two-dimensional, two-legged figure composed of four main body parts: a single torso at the top (with the legs separated behind the torso), two thighs in the middle below the torso, two lower legs at the bottom below the thighs, and two feet connected to the lower legs. The thighs and torso are connected by joints, the thighs and lower legs are connected by joints, and the lower legs and feet are connected by joints. The goal is to coordinate the forward movement of the two sets of feet, lower legs, and thighs by applying torque to the six joints connecting the feet, lower legs, and thighs. In this environment, S contains 17 dimensions of information: the height, angle, and angular velocity of the robot's torso, as well as its velocities along the X and Z axes; the angles and angular velocities of the left and right thigh joints; the angles and angular velocities of the left and right lower leg joints; and the angles and angular velocities of the left and right foot joints. A includes six types of actions, representing the torques applied to the left and right thighs, left and right calves, and left and right feet. The transfer function T is calculated by the simulation engine and provides the robot's state information after completing the current action. The reward function R is designed as follows: First, when the robot does not fall, it receives a small reward at each time step; second, when the robot successfully moves forward, it receives a reward based on the distance, angle, and speed of movement, with higher rewards for faster and straighter movement; finally, a small penalty is applied based on the magnitude of the robot's action, with higher penalties for larger actions. Note that the robot walking simulation environment E is not the only feasible design approach; in actual tasks, different states, actions, transfer functions, and reward functions can be designed according to specific circumstances.
[0010] When this invention is executed in the robot walking simulation environment E, it involves three key modules: model construction, planning using the model, and training the robot walking strategy.
[0011] In the model construction module, to improve the exploration efficiency of the robot's walking task and thus help the walking strategy escape suboptimal conditions, this invention introduces the principle of optimism. Specifically, an optimistic deep model M needs to be constructed, which includes a transition function and a reward function. It accepts state s and action a as inputs and predicts the distribution of reward r and the next state s′, i.e., (s′,r)~M(·|s,a). In this step, the training of model M requires real trajectory samples based on the interaction between the policy and the simulation environment E. The set of real trajectory samples is represented as... The i-th trajectory is denoted as HisTraj i ={(s0,a0,s1,r0),(s1,a1,s2,r1),…} i , where (s k ,a k ,s k+1 ,r k Let M represent the state at time step k, the action taken, the state at the next time step, and the reward collected at the current time step, respectively. The optimistic model M, represented by a neural network, needs to fit the actual trajectory samples of the robot's movement and exhibit sufficient optimism regarding uncertain transitions and rewards, i.e., it should assume these areas will bring significant rewards. In tabular problems with discrete actions and states, a confidence set is usually manually constructed to limit the optimism of the model. However, in robot walking tasks, due to the continuity of the state space and action space, it is impossible to manually construct a confidence set. Therefore, this invention proposes an optimized and feasible method for constructing an optimistic model. Specifically, the optimism of the model can be expressed using the value function of the initial state. This is reflected in the following, where π represents the robot's walking strategy, which specifies the action distribution π(·|s) that the robot should take in any state s, M represents an optimistic model, and s0 represents the robot's initial state. A larger value for M indicates a more optimistic model. Therefore, in order for the model to maintain optimism regarding uncertainty while accurately fitting real trajectory samples, it can be... The model's loss function is formed by combining the model's posterior probability with its prior probability. If the model is assumed to be uniformly distributed, then the model's loss function can be written as:
[0012]
[0013] in It is a set of real trajectory samples, where π represents the robot policy, M represents the model, and s0 represents the initial state. The initial value function reflects the optimism of the model. λ represents the likelihood probability of the true sample set on model M, reflecting the model's confidence level. λ represents the weights, controlling the model's optimism and confidence, thus ensuring that the model's optimism is kept within a set confidence interval. A larger λ indicates higher model confidence, and a smaller λ indicates higher model optimism. Note the value term in the loss function. The optimism level of the model was controlled, while the likelihood probability term... The confidence level of the model on real samples is controlled by using weights λ to manage the model's optimism and confidence, thus ensuring that the model's optimism is kept within a set confidence interval. λ can be set to different values as needed during actual training. To optimize the model using gradient descent based on the aforementioned loss function, this method proposes a value model gradient to calculate the gradient of the value term with respect to the model:
[0014]
[0015] Where (s,a,s′,r) represent the specific state, action, next-moment state, and reward, respectively. Let π(·|s) represent the distribution of state s access under the current policy π and model M, where π(·|s) represents the distribution of action a of the robot in the current state s, and M(·|s,a) represents the distribution of reward and state at the next moment predicted by model M based on the current state s and action a. The symbol represents a partial derivative, indicating that the partial derivative of this term with respect to the model parameters has been calculated. It is a state-action value network. It is a state-value network, where M(s′,r|s,a) represents the probability of predicting (s′,r) based on (s,a) under the current model. In the implementation of this invention, the deep model is represented using an integrated deep neural network, where each neural network outputs the mean and variance of a Gaussian distribution over the predicted state and reward.
[0016] After learning the optimistic model M, this invention requires planning based on this model to improve the robot's walking performance. In the module that uses model M for planning, planned trajectory samples can be generated based on model M, and their set is represented as follows: Where (s) ij ,a ij ) is from the set of real trajectory samples The i-th pair of states and actions in the j-th randomly sampled trajectory, and This refers to the next state and reward predicted by model M. To improve the stability and performance of the model, this module adopts an ensemble model approach, that is, learning N models, given (s) ij ,a ijAfter that, the N models will each output their predictions for the next state and the reward. Randomly select one prediction as the prediction result of model M.
[0017] In this invention, planning using the model is required in two stages: model training and policy training. During model training, on the one hand, it is necessary to use a set of real trajectory samples... The planning process is performed to obtain the planned trajectory sample. This maximizes the likelihood probability of the model. This means aligning the learned model as closely as possible to the actual transfer and reward functions, thereby improving sample utilization and achieving model optimization. Furthermore, based on the aforementioned optimization objectives, this invention requires optimization of the model based on planned trajectory samples. To calculate the value of a state This is to improve the model's optimism in uncertain regions, thereby helping the robot's walking strategy improve exploration efficiency. During the strategy training phase, in order to use the model to guide the strategy training, it is usually necessary to allow the strategy and model to interact to obtain the planned trajectory. And based on The strategy is updated; however, the longer the planning steps, the less reliable the resulting trajectory becomes. In this invention, the optimism of the model further exacerbates this problem. Therefore, to ensure the reliability of the planned trajectory... To assess the confidence level, this invention employs a branching programming approach, where the initial state s for each planning iteration starts from the real sample set. In the middle sampling, the action a is selected based on the initial state s and the current policy π(·|s). The reward r and the state s′ at the next time step are given by the model M(·|s,a). Then the process is repeated based on s′. In order to ensure the quality of the planning sample, the step size of the general planning should not be too long. Short step planning refers to planning within 5 steps.
[0018] In the policy training module, the state-action-value network needs to be updated based on the training samples. This network demonstrates the value obtainable by taking action a in the current state s during robot walking. This invention is based on... To train the robot's walking strategy. In this invention, we train an optimistic model, aiming to use the model's planning samples... Introducing optimism into the strategy. Specifically, after first planning the model using the short-step branching programming method described above, a relatively optimistic planning sample with a certain level of confidence can be obtained. Finally, the planning sample will be... and real samples The sample is mixed into the training sample set B in a certain proportion, and the state-action value network is updated based on the training sample set B. And strategy π:
[0019]
[0020]
[0021] Where L Q and L π These represent the action value network. The loss function of strategy π, where B is the real sample. and planning samples The training sample set is mixed in a certain proportion, and α is the weight parameter of the control policy entropy. In the specific implementation of this invention, two state-action value networks are maintained. This helps to obtain a more stable value function estimate, thereby improving the performance and stability of the algorithm.
[0022] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the agent action generation strategy training method based on the optimistic principle and deep model as described above.
[0023] A computer-readable storage medium storing a computer program that performs the agent action generation strategy training method based on the optimistic principle and deep model as described above.
[0024] In robot walking control decision-making tasks, the insufficient exploration efficiency of existing methods often leads to suboptimal robot control strategies and low sample utilization, limiting the application of walking robots in the real world. To address this issue, this invention introduces the optimism principle through the three modules mentioned above to improve the robot's exploration efficiency, thereby enhancing robot walking performance and reducing training costs. Specifically, this invention provides a robot control strategy training method based on the optimism principle and a deep model, applicable to robot walking control tasks. Compared to previous methods, this method, for the first time, applies the optimism principle in a computationally feasible manner to deep reinforcement learning, avoiding biases inherent in the original optimism framework, reducing optimization difficulty, and improving the exploration efficiency of robot walking control strategies. It effectively prevents robot walking control strategies from falling into suboptimal solutions, thereby improving strategy performance and sample utilization, and reducing training costs. This method has broad application prospects and can be used in various robot motion control tasks. Attached Figure Description
[0025] Figure 1 This is a flowchart of a method according to an embodiment of the present invention;
[0026] Figure 2This is a schematic diagram of the gradient update of the model described in the embodiments of the present invention;
[0027] Figure 3 This is a schematic diagram of the process for obtaining the planned trajectory as described in the embodiments of the invention;
[0028] Figure 4 This is a schematic diagram of the strategy update process described in an embodiment of the present invention. Detailed Implementation
[0029] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. After reading the present invention, any modifications of the present invention in various equivalent forms by those skilled in the art will fall within the scope defined by the appended claims.
[0030] As mentioned earlier, deep reinforcement learning methods have many applications in real-world scenarios, including autonomous driving, intelligent delivery, and robot control. In these fields, effectively utilizing limited samples to train policies is particularly important, especially in robot control where acquiring samples is relatively costly. Model-based deep reinforcement learning improves sample utilization and policy generalization performance by building models. However, model-based methods often face challenges in exploration efficiency in robot walking control tasks. While the optimism principle can be applied to tabular problems by manually designing confidence sets for the model to improve exploration efficiency, in robot walking control tasks, due to the continuity of the state-action space, it is impossible to construct an optimistic model by manually designing confidence sets. Therefore, finding a computationally feasible way to construct optimistic models is key to applying the optimism principle to deep models.
[0031] To address the aforementioned issues, this invention proposes a training method for agent action generation strategies based on the optimism principle and deep models. Taking robot walking tasks as the specific implementation object of agent action generation tasks, the method combines the optimism principle with the model in a computationally feasible manner by introducing the posterior probability of the model on real samples as a confidence estimate. Furthermore, during the strategy training phase, the planned samples obtained from the model's short-step branch planning are trained together with real samples, thereby encouraging the strategy to explore uncertain areas.
[0032] Specifically, this invention proposes an efficient robot walking strategy training method based on the optimistic principle and model. This method can be applied to robot walking control tasks. By combining the optimistic principle with the model, it alleviates the problem of insufficient exploration efficiency. However, it should be noted that this invention is not limited to robot walking control scenarios, but can be applied to any other robot control tasks that can be modeled using model-based deep reinforcement learning.
[0033] This specific embodiment uses a concrete robot walking simulation environment E as an example for detailed explanation. This robot walking simulation environment E simulates all the information of a real robot walking, providing all the information of a Markov decision process. In this environment, the robot is a two-dimensional, two-legged figure composed of four main body parts: a single torso at the top (with the legs separated behind the torso), two thighs in the middle below the torso, two lower legs at the bottom below the thighs, and two feet connected to the lower legs. The goal is to coordinate the forward movement of the two sets of feet, lower legs, and thighs by applying torque to the six joints connecting the four body parts. In this environment, S contains 17 dimensions of information: the height, angle, and angular velocity of the robot torso, as well as the velocities along the X and Z axes; the angles and angular velocities of the left and right thigh joints; the angles and angular velocities of the left and right lower leg joints; and the angles and angular velocities of the left and right foot joints. A contains six types of actions, namely the torques applied to the left and right thighs, left and right lower legs, and left and right feet. The transfer function T is calculated by the simulation engine and provides the robot's state information after executing the current action. The reward function R is designed as follows: First, when the robot does not fall, it receives a small reward at each time step; second, when the robot successfully moves forward, it receives a reward based on the distance, angle, and speed of the movement, with higher rewards for faster and straighter movement; finally, a small penalty is applied based on the magnitude of the robot's actions, with higher penalties for larger actions. Note that the robot walking simulation environment E is not the only feasible design approach. In actual tasks, different states, actions, transition functions, and reward functions can be designed according to specific circumstances. This invention specifically includes the following steps:
[0034] like Figure 1 As shown in the figure, the flowchart is a process of deployment of the intelligent agent action generation strategy training method based on the optimistic principle and deep model according to the embodiment of this application. As shown in the figure, the robot walking strategy training method based on the optimistic principle and model includes steps S101 to S104. The flowchart shows all the processes in one update round. The training of the strategy needs to go through multiple update rounds.
[0035] Step S101: The robot interacts with the robot walking simulation environment E according to the current strategy to obtain the real trajectory D. The simulation environment defines a Markov decision process.<S,A,T,R,γ> The walking robot performs reinforcement learning in this simulation environment to complete the task of normal walking. The walking simulation environment E used in this method generally needs to include the following steps:
[0036] Step 11: Initialize the simulator and initialize the robot's state information, etc.
[0037] Step 12: Based on the robot state information provided by the simulator, make a decision according to the current strategy and submit the action to the simulator. The simulator will execute the corresponding action and, based on the current robot state, provide the state and reward information for the next moment according to the transition function T and the reward function R.
[0038] Step 13: The simulator determines whether the round end condition is met. If it is, the current control task ends; otherwise, proceed to Step 12. In the robot control simulation simulator, the control task ends when the robot has accumulated a certain number of actions or when the robot triggers the end condition, such as when the robot falls.
[0039] Step S102: Train a model M with a certain degree of optimism based on the collected real trajectory samples. In this invention, the loss function for model training is as follows:
[0040]
[0041] Where π represents the robot policy, M represents the model, and s0 represents the initial state. It is a set of real trajectory samples, and the value function of the initial state. The likelihood probability of the model reflects its optimism, while the likelihood probability of the real sample on the model reflects its confidence level. The weight λ controls the balance between the optimistic and confident terms of the model, ensuring that the optimism is kept within a certain confidence interval; it can be set to 0.0003. To optimize the model using gradient descent based on the above loss function, this method uses the value model gradient to calculate the derivative of the value term with respect to the model:
[0042]
[0043] Where (s,a,s′,r) represent the specific state, action, next-moment state, and reward, respectively. Let π(·|s) represent the distribution of state s access under the current policy π and model M, where π(·|s) represents the distribution of action a of the robot in the current state s, and M(·|s,a) represents the distribution of reward and state at the next moment predicted by model M based on the current state s and action a. The symbol represents a partial derivative, indicating that the partial derivative of this term with respect to the model parameters has been calculated. It is a state-action value network. This is a state-value network, where M(s′,r|s,a) represents the probability of predicting (s′,r) based on (s,a) under the current model. In this specific implementation, the model is represented using an ensemble of neural networks, each outputting the mean and variance of a Gaussian distribution over the predicted state and reward. These neural networks share all network layers except the output layer. Figure 2As shown in the figure, the flowchart illustrates the process of calculating the gradient of the model's loss function on a real sample and updating the model's gradient. The specific update process is as follows:
[0044] Step 21: Assuming the true trajectory sample can be denoted as (s, a, s′, r), first start from the true sample set... A real sample is sampled and input into model M (s, a) to obtain the outputs of multiple ensemble models. Based on the distribution parameters of the outputs of each ensemble model, the likelihood probability of the real sample on the distribution of the outputs of each ensemble network is calculated as the confidence term loss.
[0045] Step 22: Based on the state s of the sampled real sample, sample the action a according to the current policy π(·|s). π ~π(·|s), input (s,a) to model M π The outputs of multiple ensemble networks are obtained. One ensemble network's output is randomly selected to model the Gaussian distribution of the next state and reward. The specific (s′) is then sampled from this distribution. M ,r M According to the current strategy π(·|s′), M Sampling yields action a′ M ~π(·|s′ M According to the action value network The dominance value can be calculated. Then, the gradient of the value model on the sample is calculated.
[0046] Step 23: Update the model by combining the confidence term gradient from Step 21 and the value model gradient from Step 22.
[0047] Step S103: In the optimistic model M, plan the trajectory using real trajectory samples to obtain the planned trajectory. In this invention, in order to use the trained optimistic model to guide the learning of the strategy, it is necessary to plan on the model according to the strategy. In this specific embodiment, in order to make the planned trajectory more reliable, on the one hand, branch planning is used to plan from the state distribution in the real sample, and on the other hand, the number of planning steps is limited to run only the planning with a short step length, such as planning within 5 steps. Figure 3 The planning process in this specific embodiment is given, and the specific steps are detailed below:
[0048] Step 31: First, obtain real trajectory samples. The state distribution in the diagram is used to sample states as the starting point for branch planning.
[0049] Step 32: Based on the given state s, sample action a from the current policy π(·|s), and input the given state s and the sampled action a into the model M.
[0050] Step 33: The model obtains the outputs of multiple ensemble models based on the input state s and action a. The output of each ensemble model determines the Gaussian distribution of the next time step state and reward. The output of one model is randomly sampled from the outputs of the multiple ensemble models, and the specific next time step state and reward are obtained by sampling from the Gaussian distribution of the next time step state and reward determined by its output. In this specific embodiment, using ensemble models to represent the model is to improve the stability and performance of model training. This approach is not necessary; a non-ensemble neural network can be used to represent the model.
[0051] Step 34: Determine if the planning has met the termination condition. If yes, save the planning trajectory and end the planning process; otherwise, return to step 32. In this specific embodiment, the termination condition includes two aspects: firstly, the planning will terminate if the planning step size reaches the upper limit; secondly, the planning process will also terminate if the model issues a termination signal, such as when the robot has fallen.
[0052] Step S104: The real trajectory samples obtained in step S101 and the planned samples obtained in step S103 are mixed in a certain proportion to form training samples B, and the gradient of the robot's walking strategy is updated using the training samples. In this invention, in order to guide the strategy exploration according to the model, the planned samples of the robot's walking strategy on the model are mixed with real samples to form training samples. The larger the proportion of planned samples, the more the robot's walking strategy tends to explore; the larger the proportion of real samples, the more the robot's walking strategy tends to utilize. In this way, the exploration tendency of the robot's walking strategy can be adjusted. The robot walking strategy training step can use any deep reinforcement learning algorithm to train the strategy. In this specific embodiment, a state-action value network and the maximum entropy approach are used to learn the strategy. Figure 4 The process of each round of strategy updates is shown, with the specific steps detailed below:
[0053] Step 41: Mix the real trajectory samples obtained in step S101 and the planned samples obtained in step S103 in a certain ratio to form a batch training sample B. This ratio is a hyperparameter that can be adjusted according to the specific task structure. The higher the proportion of the planned samples, the more the strategy tends to explore; the lower the proportion, the more the strategy tends to exploit.
[0054] Step 42: Update the action value network based on the obtained batch training samples and the current policy. In this embodiment, two action value networks are maintained, which helps to obtain a more stable value function estimate, alleviates the overestimation problem of the Q-value network, and thus improves the performance and stability of the algorithm. The loss function of the action value network is as follows:
[0055]
[0056] Where B is the batch training sample after mixing real samples and planned samples in a certain proportion, and α is the weight parameter of the control policy entropy.
[0057] The two state-action value networks are represented by independent parameters, and their respective network parameters are updated according to the loss function described above.
[0058] Step 43: Update the policy based on the obtained batch training samples and the two action value networks. To mitigate the overestimation of Q-values, the smaller Q-value from the outputs of the two action value networks is selected when calculating the policy loss function.
[0059]
[0060] Where B is the batch training sample, α is the weight parameter of the control policy entropy, and Θ1 and Θ2 represent the parameters of the two state-action value networks, respectively.
[0061] Step 43: Repeat the training process until the number of policy updates reaches the pre-defined limit, at which point the policy update for that round ends.
[0062] Step 5: Repeat steps S101 to S104 until the policy training converges, thus completing the training process and obtaining the robot walking policy. This policy has better performance and higher sample utilization than the robot walking policy obtained by existing training methods.
[0063] Obviously, those skilled in the art should understand that the steps of the intelligent agent action generation strategy training method based on the optimistic principle and deep model described in the above embodiments of the present invention can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, which can then be stored in a storage device for execution by the computing device. Furthermore, in some cases, the steps shown or described can be performed in a different order than presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the embodiments of the present invention are not limited to any particular hardware and software combination.
Claims
1. A training method for an agent action generation strategy based on the optimism principle and deep models, characterized in that, Taking robot walking as the specific implementation object of intelligent agent action generation task, the robot walking task first needs to be modeled as a Markov decision process.<S,A,T,R,γ> Here, S represents the state space, which refers to the state information that the robot can perceive during the walking process, including the location information of obstacles on the map and the robot's state information; A represents the action space, which contains all actions that the robot can control; T represents the state transition function, which gives the probability distribution T(·|s,a) of the new state that the robot will transition to after taking any action a∈A in any state s∈S; R represents the reward function, which gives the reward R(s,a) received by the robot after taking any action a∈A in any state s∈S; and γ represents the discount factor, which is used to balance long-term rewards and short-term rewards. The interactive environment used by the robot to train its walking strategy is the robot walking simulation environment E. This robot walking simulation environment E simulates the interaction process between the real robot walking and the environment, providing information for the Markov decision process. When deploying in the robot walking simulation environment E, the robot walking strategy training method includes model construction, planning using the model, and training the robot walking strategy. Construct an optimistic deep model M, which includes a transition function and a reward function. It takes state s and action a as input and predicts the distribution of reward r and the state s′ at the next time step, i.e., (s′,r)~M(·|s,a). Model M is represented by an ensemble of deep neural networks, each of which outputs the mean and variance of the Gaussian distribution on the predicted state and reward. The model's loss function is: Where v represents the robot policy, M represents the model, and s0 represents the initial state. The initial value function reflects the optimism of the model. It is the set of real trajectory samples obtained by the robot interacting with the robot walking simulation environment E using strategy π, and it is represented as The i-th trajectory is denoted as HisTraj i ={(s0,a0,s1,r0),(s1,a1,s2,r1),…,(s k ,a k ,s k+1 ,r k )} i , where (s k ,a k ,s k+1 ,r k ) represent the state at the k-th time step, the action taken, the state at the next time step, and the reward collected at the current time step, respectively; λ represents the likelihood probability of the real sample set on model M, which reflects the confidence level of the model; λ represents the weight, which controls the optimism and confidence of the model, thereby ensuring that the optimism of the model is controlled within the set confidence interval. The larger λ is, the higher the confidence level of the model, and the smaller λ is, the higher the optimism of the model.
2. The method for training an agent action generation strategy based on the optimism principle and deep models according to claim 1, characterized in that, In the robot walking simulation environment, the robot is a two-dimensional, two-legged figure composed of four body parts: a single torso at the top, two thighs in the middle below the torso, two lower legs at the bottom below the thighs, and two feet connected to the lower legs. The torso and thighs, the thighs and lower legs, and the lower legs and feet are all connected by joints. By applying torque to the six joints connecting the feet, lower legs, and thighs, the two sets of feet, lower legs, and thighs move forward in a coordinated manner.
3. The method for training an agent action generation strategy based on the optimism principle and deep models according to claim 1, characterized in that, In the robot walking simulation environment, S contains 17 dimensions of information: the height, angle, and angular velocity of the robot's torso, as well as its velocities along the X and Z axes; the angles and angular velocities of the left and right thigh joints; the angles and angular velocities of the left and right lower leg joints; and the angles and angular velocities of the left and right foot joints. A contains 6 types of actions, namely the torques applied to the left and right thighs, left and right lower legs, and left and right feet. The transfer function T is calculated by the simulation engine, providing the robot's state information after executing the current action. The reward function R is designed as follows: First, when the robot does not fall, it receives a small reward at each time step. Second, when the robot successfully moves forward, it receives a reward based on the distance, angle, and speed of movement; the faster and straighter the robot moves, the higher the reward. Finally, a small penalty is applied based on the magnitude of the robot's action; the larger the magnitude of the action, the higher the penalty.
4. The method for training an agent action generation strategy based on the optimism principle and deep models according to claim 1, characterized in that, Model M is optimized using gradient descent, and a value model gradient is proposed to calculate the gradient of the value term with respect to the model: Where (s,a,s) ′ ,r) represent the specific state, action, next state, and reward, respectively. Let π(·|s) represent the distribution of state s access under the current policy π and model M, where π(·|s) represents the distribution of action a of the robot in the current state s, and M(·|s,a) represents the distribution of reward and state at the next moment predicted by model M based on the current state s and action a. The symbol represents a partial derivative, indicating that the partial derivative of this term with respect to the model parameters has been calculated. It is a state-action value network. It is a state-value network, M(s) ′ ,r|s,a) represents the result of predicting (s) based on (s,a) under the current model. ′ The probability of r).
5. The method for training an agent action generation strategy based on the optimism principle and deep models according to claim 1, characterized in that, After learning the optimistic model M, planning is performed on the model; planning trajectory samples are generated based on model M, and the set of planning trajectory samples is represented as follows. Where (s) ij ,a ij ) is from the set of real trajectory samples The i-th pair of states and actions in the j-th randomly sampled trajectory, and This refers to the next state and reward predicted by model M; an ensemble model approach is used, i.e., learning N models, given (s) ij ,a ij After that, N models output predictions for the next state and reward, respectively. Randomly select one prediction as the prediction result of model M; The model is used for planning in two stages: model training and policy training. During model training, on the one hand, it is necessary to use a set of real trajectory samples... The planning process is performed to obtain the planned trajectory sample. This maximizes the likelihood probability of the model. Optimize the model; on the other hand, based on the planned trajectory samples Calculate the value of the state During the training phase of the strategy, the strategy interacts with the model to obtain the planned trajectory. And based on The strategy is updated using branching programming, with the initial state s at each planning iteration starting from the real sample set. Medium-step planning involves selecting action a based on the initial state s and the current policy π(·|s), using model M(·|s,a) to provide the reward r and the state s′ at the next moment, and then repeating the process based on s′; short-step planning refers to planning within 5 steps. Finally, the planned trajectory will be... and real samples Mixed into the training sample set B at a set ratio, and the state-action value network is updated based on set B. And strategy π: Where L Q and L π These represent the action value network. The loss function of strategy π, where B is the real sample. and planning samples The training sample set is mixed according to a set ratio, and α is the weight parameter of the control policy entropy.
6. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the training method for agent action generation strategy based on the optimistic principle and deep model as described in any one of claims 1-5.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program that performs a training method for an agent action generation strategy based on the optimistic principle and a deep model as described in any one of claims 1-5.
Citation Information
Patent Citations
Unmanned vehicle adaptive path planning method based on dynamic window method and near-end strategy
CN116679719A
Robot skill training-oriented confidence inverse reinforcement learning method
CN116992977A