A training method for the motion strategy of a blind hexapod robot
By simulating the environment and ontology perception signals, combined with proximal strategy optimization algorithms, the motion strategy of blind hexapod robots is trained, which solves the problem that blind hexapod robots are difficult to train in the real world, and realizes adaptive motion and stable walking in complex environments.
Patent Information
- Application Number
- CN202311273174.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-27
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-09-27
AI Technical Summary
Blind hexapod robots are difficult to train in the real world, and external sensors are susceptible to natural factors, limiting their perception capabilities.
By establishing a simulated environment for blind hexapod robots and wild mountainous areas, using ontology-aware signals and proximal strategy optimization algorithms, a policy network model is built and trained to generate adaptive motion strategies.
The blind hexapod robot is realized to adaptively move in complex wild environments, overcome the problem of poor effectiveness of traditional control methods, and reduce the dependence on external sensors.
Smart Images

Figure CN117340876B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hexapod robot motion control, and more specifically, to a training method for the motion strategy of a blind hexapod robot. Background Art
[0002] Hexapod robots often perform tasks in the wild due to their superior traversability. With the development of machine learning and CPG, complex-structured legged robots perform well in unstructured environments such as mountains, sand, and stairs, and are sufficient to replace wheeled robots to perform various field tasks. Hexapod robots have multiple redundant degrees of freedom, can achieve static balance with three or more legs, and also have a large number of gaits, and can select the best gait for different environments to achieve the maximum efficiency of walking.
[0003] Currently, the motion control of hexapod robots mainly includes the following: biologically inspired control, engineering-based control, machine learning-based control, and the combination of two or more of the above. The main method of biologically inspired control is based on CPG control, and various gaits are generated through the rhythm signals generated by CPG. By constructing the kinematic model of the robot, inverse kinematics is used to plan the trajectory of the feet. Machine learning mainly includes methods such as RL and DL, which are used to train various gaits or generate adaptive gaits for specific environments.
[0004] Hexapod robots have 18 joints and a high degree of freedom, resulting in a high-dimensional continuous state space. And the terrain for performing tasks is complex, not only with uneven ground, but also various obstacles. The above two points result in poor effects of traditional control methods. And if the proximal policy optimization algorithm is used, a large amount of exploration data and training rounds are required, and the time consumed is incalculable, so it is difficult to train in the real world. Moreover, currently, hexapod robots obtain environmental information using external sensors, and external sensors are easily affected by natural factors such as light, and in severe cases, the sensors will fail, which imposes certain limitations on the perception ability of hexapod robots. Summary of the Invention
[0005] The present invention provides a training method for the motion strategy of a blind hexapod robot to overcome the problem that it is difficult to train a blind hexapod robot in the real world.
[0006] To solve the above technical problems, the technical solution of the present invention is as follows:
[0007] A training method for the motion strategy of a blind hexapod robot includes the following steps:
[0008] S1. Establish a simulation environment for the blind hexapod robot and the wild mountain;
[0009] S2. Obtain the environmental data of the simulation environment, and the environmental data is the body perception signal of the blind hexapod robot;
[0010] S3. Construct a policy network model with the environmental data as the input;
[0011] S4. Construct a Proximal Policy Optimization algorithm and train the policy network through the optimization algorithm;
[0012] S5. Construct a reward function to evaluate the blind hexapod robot, and provide a feedback signal for the optimization algorithm to guide the training and parameter optimization of the policy network.
[0013] Further, step S1 is specifically as follows:
[0014] Establish a physical model of the blind hexapod robot in Mujoco, and define the body structure, joint parameters and motion control mode of the blind hexapod robot to describe the toe trajectory of the blind hexapod robot;
[0015] Establish a physical model of the wild mountain in Mujoco, and define the roughness of the terrain and the physical positions of the obstacles, so that the established physical model of the wild mountain is similar to the real terrain.
[0016] Further, step S1 also includes setting the parameters of the blind hexapod robot, and the parameters include the central coordinates of the robot's torso plane, the root coordinates of the six hip joints, the length of the hip joint, the length of the knee joint, and the length of the ankle joint.
[0017] Further, with the forward direction of the robot as the positive x-axis direction, the parameters of the blind hexapod robot are specifically as follows:
[0018] The central coordinates of the blind hexapod robot's torso plane are O(0, 0, 0);
[0019] The root coordinates of the six hip joints are respectively: LF(93.6, 50.805, 0), LM(0, 73.575, 0), LR(-93.6, 50.805, 0), RF(93.6, -50.805, 0), RM(0, -73.575, 0), RR(-93.6, -50.805, 0);
[0020] The length of the hip joint L1 = 45.01;
[0021] The length of the knee joint L2 = 77.06;
[0022] The length of the ankle joint L3 = 150.3.
[0023] Further, in step S2, the state information of the blind hexapod robot in the environment is output by an array, and a total of 46-dimensional data is used as the body perception signal of the blind hexapod robot;
[0024] Adopt n iEach element representing the body signal, then the state information s t is represented as s t = [n 1 , n 2 ,......, n 46 .
[0025] Furthermore, among the 46 - dimensional data, 3 dimensions are used to represent the pose information of the blind hexapod robot, 18 dimensions are used to represent the joint angles of the free joints of the blind hexapod robot, 3 dimensions are used to represent the linear velocity when the blind hexapod robot moves, 3 dimensions are used to represent the angular velocity when the blind hexapod robot moves, and 19 dimensions are used to represent the joint velocities of the free joints of the blind hexapod robot.
[0026] Furthermore, the policy network constructed in step S3 is a four - layer fully - connected neural network.
[0027] The first layer of the policy network is the input layer, and the input is the state s t = [n 1 , n 2 ,......, n 46 array;
[0028] The second and third layers of the policy network are hidden layers, and each layer has 96 neurons;
[0029] The fourth layer of the policy network is the output layer, and the output is the joint angle a t = [m 1 ,......, m 18 array, where m i represents the action of each joint.
[0030] Furthermore, the construction process of the proximal policy optimization algorithm in step S4 is as follows:
[0031] S41. Define the Markov decision process:
[0032] In reinforcement learning, the interaction process between the agent and the environment is modeled as a Markov decision process; under the policy π of the agent, the probability that the agent transfers from state s t to state s t+1 is completely determined by state s t and action a t , then this conditional probability is p π (r t , s t+1 |a t , s t ); where the reward r t is the environmental feedback obtained by the agent when it executes an action a t at time t;
[0033] S42. Define the reward function:
[0034] Using the discount factor γ, define the reward R at each time step t t , representing the cumulative future rewards starting from the current time step; then the reward R at time t t is:
[0035]
[0036] In the formula, the discount factor γ determines the impact of future rewards on the current state and plays a role in attenuating future rewards;
[0037] Among them, the learning process of reinforcement learning is a process of maximizing the expected reward J, and the expression of the expected reward J is:
[0038]
[0039] In the formula, denotes the expected operation on the action a under the policy π, and R(a) represents the immediate reward obtained by the action a;
[0040] S43. Define the policy function and the value function, and use the Bellman equation for estimation:
[0041] Introduce the policy function π(θ) to fit the policy probability distribution, and the value function V π and the action-value function Q π , used to evaluate the value of the policy under a given state and action;
[0042] Among them, the expression of the value function V π is:
[0043]
[0044] The expression of the action-value function Q π is:
[0045]
[0046] Use the Bellman equation to iteratively estimate the value function V π and the action-value function Q π ; Denote the value function at the next time step as Denote the action-value function at the next time step as Then:
[0047]
[0048]
[0049] Use the policy network π(θ) to fit the specific policy probability distribution, and use the policy gradient formula to optimize the policy parameters. The optimization goal is to maximize the reward with the agent's policy probability distribution. The expression of the policy gradient formula is:
[0050]
[0051] S44. Define the advantage function:
[0052] Use the advantage function A(a, s) to measure the value of a specific action for a given state. The expression of the advantage function A(a, s) is:
[0053] A π (s t , a t ) = Q π (s t , a t ) - V π (s t ) (8);
[0054] S45. Introduce the proximal optimization rule:
[0055] By setting a confidence interval, limit the KL divergence between the new policy and the old policy to mitigate policy gradient fluctuations and increase the probability that actions with higher advantage values estimated according to the current policy are selected;
[0056] Let the new policy after updating the policy be π′(a t |s t ; θ′), then the expression of the loss function obtained using importance sampling is:
[0057]
[0058] Use the KL divergence to measure the similarity of the probability distributions of the new and old policies. Then the constraint condition for the optimization process is:
[0059]
[0060] S46. Construct the loss function:
[0061] Use the Lagrange multiplier method to integrate Formula 9 and Formula 10 into one term to construct the improved loss function L′. The expression of the improved loss function L′ is:
[0062]
[0063] Among them, β is a Lagrange multiplier; when βπ(a t |s t ; θ) and π′(at |s t ; When θ') approaches, use r t to represent the ratio of the probability of the same action under the new policy to the probability under the old policy, and then use a parameter ∈ to control the ratio of probabilities r t within the range of (1 - ∈, 1 + ∈), then the final expression of the loss function is:
[0064]
[0065] S47. Update the policy using the gradient descent method:
[0066] Update the policy π by performing gradient descent optimization on the loss function θ .
[0067] Furthermore, step S5 is specifically as follows:
[0068] The total reward r of the blind hexapod robot t is composed of the reward r v and the penalty r θ 、r e . The expression of the total reward r t is:
[0069] r t = λ 1 r v - λ 2 r θ - λ 3 r e (13);
[0070] where λ1, λ2, λ3 are weight parameters; the velocity reward r v is used to encourage the blind hexapod robot to move forward, and the closer the actual velocity is to the set velocity, the more reward is obtained; the penalty r θ is used to correct the heading to ensure that the heading is consistent with the movement direction; the penalty term r e plays a penalty role in the reward function and is used to encourage the blind hexapod robot to avoid obstacles in the mountain environment and maintain a stable walking posture;
[0071] Let the distance traveled by the blind hexapod robot in a single step be x d , then the expression of the velocity reward r v is:
[0072]
[0073] Let the scalar component of the quaternion of the current state be q w , then the expression of the penalty r θ is:
[0074]
[0075] Penalty term r for the blind hexapod robot when encountering an obstacle e The expression is as follows:
[0076] r e = (Pitch * 4) 2 + (Roll * 3) 2 (16);
[0077] Update the policy by minimizing the loss function to obtain the policy network π(θ).
[0078] Compared with the prior art, the beneficial effects of the technical solution of the present invention are as follows:
[0079] The present invention provides a training method for the motion strategy of a blind hexapod robot, which uses the body perception signal and the proximal policy optimization algorithm to solve the problem of motion control in the wild terrain; trains the motion strategy of the blind hexapod robot through Mujoco; this strategy evaluates the current environment by sensing the attitude information of the blind hexapod robot itself and outputs the next action, enabling the blind hexapod robot to make adaptive motions according to the current environment.
[0080] Specifically, the present invention discloses a solution for the motion control problem of a blind hexapod robot. Taking the control system of the blind hexapod robot as the controlled object, for the problems of walking and obstacle avoidance of the controlled object in the wild mountain environment, a proximal policy optimization-based reinforcement learning control algorithm is adopted, and a fully connected neural network is used to fit the policy network and the value network, ensuring the correctness of the fitting function;
[0081] Use Mujoco to model the real robot and terrain and then iteratively train the strategy of the blind hexapod robot, ensuring that the blind hexapod robot can walk normally on flat ground and unstructured ground and has a certain obstacle avoidance ability;
[0082] And based on the body perception signal as the state information, the blind hexapod robot is intensively trained. These signals are not affected by weather and light in the wild environment, enabling the blind hexapod robot that often performs tasks in the wild to overcome the influence of weather. Brief Description of the Drawings
[0083] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0084] Figure 1 Is the flowchart of the present invention;
[0085] Figure 2 This is a training process diagram of the motion strategy of the blind hexapod robot disclosed in the present invention;
[0086] Figure 3 This is a structural diagram of the policy network in the present invention;
[0087] Figure 4 This is a diagram of the instantaneous speed change of the motion of the blind hexapod robot in the present invention;
[0088] Figure 5 This is a diagram of the vertical acceleration change of the motion of the blind hexapod robot in the present invention. Specific implementation manners
[0089] In order to better understand the purpose, structure and function of the present invention, the technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and specific preferred embodiments.
[0090] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as "left side", "right side", "upper part", "lower part", etc. is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. "First", "second", etc. do not represent the importance of components, so it cannot be understood as a limitation to the present invention. The specific dimensions adopted in the embodiments are only for illustrating the technical solution by way of example, and do not limit the protection scope of the present invention. For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0091] Unless otherwise clearly specified and defined, terms such as "installation", "setting", "connection", "fixation", etc. should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection, or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium. It can be the communication inside two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.
[0092] Embodiment 1:
[0093] As Figure 1-2 shown, the present invention provides a technical solution: a training method for the motion strategy of a blind hexapod robot, including the following steps:
[0094] S1. Establish a simulation environment for the blind hexapod robot and the wild mountain;
[0095] S2. Obtain the environmental data of the simulation environment, and the environmental data is the body perception signal of the blind hexapod robot;
[0096] S3. Use the environmental data as input to construct a policy network model;
[0097] S4. Build a Proximal Policy Optimization (PPO) algorithm and use the optimization algorithm to train the policy network;
[0098] S5. Build a reward function to evaluate the blind hexapod robot and provide a feedback signal for the optimization algorithm to guide the training and parameter optimization of the policy network.
[0099] Please refer to Figure 2 , which shows the training process diagram of the hexapod robot motion policy disclosed in this application. The motion policy input is the body perception signal s of the hexapod robot obtained from the environment t , and through the forward propagation of the network, the actions a of 18 joints are output t . After the hexapod robot executes the action a t , the environment changes, and a new input s can be provided to the policy network t+1 to generate a new action a t+1 . At the same time, evaluate the quality of the current environment through the reward function and quantify it as r t , and the reward r t is used for the iteration of the motion policy. It can be found that this is a Markov decision process, and finally an optimal motion policy is generated through repeated iteration.
[0100] Specifically, establish a physical model of the hexapod robot in Mujoco, considering the coordination between the single legs of the hexapod robot and the lengths of each joint to describe the toe trajectory of the hexapod robot; establish a physical model of the wild mountain in Mujoco, considering the roughness of the terrain and the physical positions of obstacles, and the established model is similar to the real terrain; obtain environmental data from Mujoco, considering that the information obtained is the body perception signal of the hexapod robot, and this information can also be obtained in real sensors; construct a policy network and a value network to fit the policy function and value function of the hexapod robot, and use the Bellman equation to estimate them through an iterative method, construct a training reward function, and use the obtained body perception signal as the input of deep reinforcement learning training for iteration.
[0101] In summary, the present invention uses the proximal policy optimization algorithm to generate a motion policy, thereby controlling the coordinated motion of each single leg. Even in a high-dimensional continuous state space and a harsh environment, its motion policy can perform well. For the training problem, using a physics engine for training can well solve this problem. Mujoco is a general physics engine that provides a good simulation environment for the field of robots that need to simulate the interaction between joint structures and the environment. Mujoco can be used to implement model-based calculations or as a traditional simulator. Controlling the robot motion using the information sensed by internal sensors can overcome the influence of weather, while reducing the model complexity and eliminating the need for expensive external sensors. The information obtained by internal sensors can meet the minimum requirements for the normal walking of a hexapod robot, just like a person walking with eyes closed.
[0102] Embodiment 2:
[0103] Based on Embodiment 1, step S1 is specifically as follows:
[0104] Establish a physical model of a blind hexapod robot in Mujoco. Considering the coordination between the single legs of the hexapod robot and the lengths of each joint, define the body structure, joint parameters, and motion control mode of the blind hexapod robot to describe the toe trajectory of the blind hexapod robot;
[0105] Establish a physical model of a wild mountain in Mujoco, define the roughness of the terrain and the physical positions of obstacles, so that the established physical model of the wild mountain is similar to the real terrain.
[0106] Furthermore, obtain environmental data from Mujoco, considering that the obtained information is the body perception signal of the hexapod robot, and this information can also be obtained in real sensors.
[0107] Furthermore, step S1 also includes setting parameters of the blind hexapod robot, and the parameters include the central coordinates of the robot's torso plane, the root coordinates of six hip joints, the lengths of the hip joints, the lengths of the knee joints, and the lengths of the ankle joints.
[0108] Further, modeling is carried out using Mujoco, and the model is defined through the local MJCF scene description language. Define the structure of the blind hexapod robot. Taking the forward direction of the blind hexapod robot as the positive x-axis direction, and setting the center coordinates of the trunk plane of the blind hexapod robot as O(0, 0, 0), then the root coordinates of the six hip joints are LF(93.6, 50.805, 0), LM(0, 73.575, 0), LR(-93.6, 50.805, 0), RF(93.6, -50.805, 0), RM(0, -73.575, 0), RR(-93.6, -50.805, 0). The length of the hip joint L1 = 45.01, the length of the knee joint L2 = 77.06, and the length of the ankle joint L3 = 150.3. Terrain modeling considers the roughness of the real terrain and the physical positions of obstacles, including the dimensions of the robot and the terrain, the contact mechanics of the foot-ground interaction, etc., and is defined through the local MJCF scene description language.
[0109] Embodiment 3:
[0110] Based on Embodiment 1, in step S2, the state information of the blind hexapod robot in the environment is output as an array, and a total of 46-dimensional data serves as the body perception signal of the blind hexapod robot;
[0111] Using n i to represent each element of the body signal, then the state information s t is represented as s t = [n 1 , n 2 ,......, n 46 .
[0112] Further, among the 46-dimensional data, 3 dimensions are used to represent the pose information of the blind hexapod robot, 18 dimensions are used to represent the joint angles of the free joints of the blind hexapod robot, 3 dimensions are used to represent the linear velocity when the blind hexapod robot moves, 3 dimensions are used to represent the angular velocity when the blind hexapod robot moves, and 19 dimensions are used to represent the joint velocities of the free joints of the blind hexapod robot;
[0113] Specifically, the pose information of the blind hexapod robot in the environment is output as an array. The 3-6 dimensions of this array represent the rotation (quaternion) of the blind hexapod robot in the world coordinate system, and the 7-24 dimensions represent the joint angles of the free joints. The velocity information of the blind hexapod robot in the environment is also output as an array. The 0-2 dimensions of this array represent the linear velocity when the blind hexapod robot moves, the 3-5 dimensions represent the angular velocity when the blind hexapod robot moves, and the 6-23 dimensions represent the joint velocities of the free joints. In total, there are 46-dimensional data that can serve as the body perception signal. If n i is used to represent each element of the body signal, then the state information s tCan be represented as s t = [n 1 , n 2 ,......, n 46 .
[0114] Example 4:
[0115] Based on Example 1, please refer to Figure 3 , the policy network constructed in step S3 is a four-layer fully connected neural network,
[0116] The first layer of the policy network is the input layer, and the input is the state s t = [n 1 , n 2 ,......, n 46 array;
[0117] The second and third layers of the policy network are hidden layers, and each layer has 96 neurons;
[0118] The fourth layer of the policy network is the output layer, and the output is the joint angle a t = [m 1 ,......, m 18 , where m i represents the action of each joint.
[0119] Specifically, Figure 3 As shown in the policy network structure diagram, the input layer is the state s t = [n 1 , n 2 ,......, n 46 , two hidden layers, each layer has 96 neurons, and the output layer is the joint angle a t = [a 1 ,......, a 18 . It is built through the neural network framework pytorch. First, normalize the input network, and activate the network output with the tanh function to construct a fully connected neural network. Combine the environment and the algorithm. Each training will obtain the current state of the blind hexapod robot from the environment, so as to evaluate the value of the current action and state with the value network, and iterate the policy network. The iterated policy network guides the next action and generates a new state. Until the training iterates the policy network to the expectation of obtaining the maximum reward, the policy network training is completed.
[0120] Example 5:
[0121] Based on Example 1, the construction process of the proximal policy optimization algorithm in step S4 is as follows:
[0122] S41. Define the Markov decision process:
[0123] In reinforcement learning, the interaction process between the agent and the environment is a Markov decision process. Model the interaction process between the agent and the environment as a Markov decision process. Under the policy π of the agent, the probability that the agent transfers from state s t to state s t+1 is completely determined by state s t and action a t . Then this conditional probability is p π (r t , s t+1 |a t , s t ). Among them, the reward r t is the environmental feedback obtained by the agent after executing an action a t at time t.
[0124] S42. Define the return function:
[0125] Use the discount factor γ to define the return R t at each time step t, which represents the accumulation of future rewards starting from the current time step. Then the return R t at time t is:
[0126]
[0127] In the formula, the discount factor γ determines the influence of future rewards on the current state and plays a role in attenuating future rewards;
[0128] Among them, the learning process of reinforcement learning is a process of maximizing the return expectation J. The expression of the return expectation J is:
[0129]
[0130] In the formula, represents the expectation operation on the action a under the policy π, and R(a) represents the immediate reward obtained by the action a;
[0131] S43. Define the policy function and the value function, and use the Bellman equation for estimation:
[0132] Introduce the policy function π(θ) to fit the policy probability distribution, and the value function V π and the action value function Q π to evaluate the value of the policy under the given state and action;
[0133] Among them, the expression of the value function V π is:
[0134]
[0135] Action value function Q π The expression of is as follows:
[0136]
[0137] Use the Bellman equation to iteratively estimate the value function V π and the action value function Q π ; that is, in order to estimate V π and Q π estimation is carried out, the Bellman equation is used, and they are estimated by an iterative method;
[0138] Denote the value function of the next step as Denote the action value function of the next step as Then:
[0139]
[0140]
[0141] Use the policy network π(θ) to fit the specific policy probability distribution, which enables the agent to obtain as much reward as possible; use the policy gradient formula to optimize the policy parameters, and the optimization goal is to maximize the reward for the agent's policy probability distribution. The expression of the policy gradient formula is:
[0142]
[0143] S44. Define the advantage function:
[0144] Use the advantage function A(a, s) to measure the value of a specific action for a given state. The expression of the advantage function A(a, s) is:
[0145] A π (s t , a t ) = Q π (s t , a t ) - V π (s t ) (8);
[0146] S45. Introduce the proximal optimization rule:
[0147] OPP sets a confidence interval near the initial policy and performs optimization within the interval. It can not only alleviate the fluctuation of the policy gradient but also increase the probability of actions with a relatively large advantage function. Let the new policy after updating the policy be π′(a t |s t; θ'), the loss function obtained using importance sampling is:
[0148]
[0149] The KL divergence is used to measure the similarity of the probability distributions of the old and new policies, and the constraint condition for the optimization process is:
[0150]
[0151] S46. Construct the loss function:
[0152] Using the Lagrange multiplier method, formulas 9 and 10 are integrated into one term, effectively reducing the computational complexity of training and making the implementation simpler. The improved loss function is:
[0153]
[0154] where β is a Lagrange multiplier; when βπ(a t |s t ; θ) and π'(a t |s t ; θ') are close, use r t to represent the ratio of the probability of the same action under the new policy to the probability under the old policy, and then use a parameter ∈ to control the ratio of probabilities r t within the range of (1 - ∈, 1 + ∈), then the final expression of the loss function is:
[0155]
[0156] S47. Update the policy using the gradient descent method:
[0157] By performing gradient descent optimization on the loss function, update the policy π θ .
[0158] Example 6:
[0159] Based on Example 1, step S5 is specifically:
[0160] The total reward r t of the blind hexapod robot consists of the reward r v and the penalties r θ , r e . The expression for the total reward r t is:
[0161] r t = λ 1 r v - λ 2 r θ - λ 3 re (13);
[0162] where λ1, λ2, and λ3 are weight parameters; the velocity reward r v is used to encourage the blind hexapod robot to move forward, and the closer the actual velocity is to the set velocity, the more reward is obtained; it is important to set the velocity as the training objective; if the training objective is set as distance, the robot will obtain an unstable and jumping gait, which is not what this invention expects.
[0163] The penalty r θ is used to correct the heading to ensure that the heading is consistent with the movement direction; the penalty term r e plays a penalty role in the reward function and is used to encourage the blind hexapod robot to avoid obstacles in the mountain environment and maintain a stable walking posture;
[0164] Let the distance of a single step of the blind hexapod robot be x d , then the expression of the velocity reward r v is:
[0165]
[0166] Let the scalar component of the quaternion of the current state be q w , then the expression of the penalty r θ is:
[0167]
[0168] It is inevitable for the blind hexapod robot to encounter obstacles in the mountain environment, and most of these obstacles in the mountain environment are manifested as high slopes or low valleys. When the robot is going up a steep slope or down a steep slope, the changes in its pitch angle Pitch and roll angle Roll are relatively large. These changes can be sensed by the body sensors, and the robot can use this to avoid specific obstacles. The penalty term r e when the blind hexapod robot encounters an obstacle is:
[0169] r e = (Pitch * 4) 2 + (Roll * 3) 2 (16);
[0170] Finally, the policy is updated by minimizing the loss function to obtain the policy network π(θ).
[0171] Please refer to Figure 4 , Figure 4 , which shows the changes in the movement speeds of the blind hexapod robot controlled by the trained policy network during straight movement and obstacle avoidance. It can be seen from the figure that the speeds of the blind hexapod robot during straight movement and obstacle avoidance both quickly accelerate from rest to around 0.4 and then fluctuate around 0.4 subsequently.
[0172] Please refer to Figure 5 , the vertical acceleration of the blind hexapod robot during straight movement and obstacle avoidance is as Figure 5 shown. Record the vertical acceleration of the blind hexapod robot as an index of motion stability. It can be seen from the figure that the vertical acceleration of the blind hexapod robot fluctuates between -0.2 and 0.4 during straight movement and obstacle avoidance, and the stability is good.
[0173] Obviously, the above-mentioned embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all implementation manners here. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A training method for the motion strategy of a blind hexapod robot, characterized in that, it includes the following steps: S1. Establish a simulation environment for the blind hexapod robot and the wild mountain; S2. Obtain the environmental data of the simulation environment, and the environmental data is the body perception signal of the blind hexapod robot; S3. Take the environmental data as the input and construct a policy network model; S4. Construct a proximal policy optimization algorithm to train the policy network through the optimization algorithm; The construction process of the proximal policy optimization algorithm is as follows: S41. Define the Markov decision process: In reinforcement learning, the interaction process between the agent and the environment is modeled as a Markov decision process; under the agent's policy π, the probability that the agent transitions from state s t to state s t+1 is completely determined by state s t and action a t , and this conditional probability is p π (r t , s t+1 | a t , s t ); where the reward r t is the environmental feedback obtained by the agent after executing an action a t at time t; S42. Define the reward function: Using the discount factor γ, define the return R at each time step t t , representing the accumulation of future rewards starting from the current time step; then the return R at time t t is given by: In the formula, the discount factor γ determines the influence of future rewards on the current state and plays a role in attenuating future rewards; Among them, the learning process of reinforcement learning is a process of maximizing the reward expectation J, and the expression of the reward expectation J is: wherein represents the expected operation on action a under policy π, and R(a) represents the immediate reward obtained for action a; S43. Define the policy function and the value function, and use the Bellman equation for estimation: Introduce the policy function π(θ) for fitting the policy probability distribution, as well as the value function V π and the action-value function Q π , which is used to evaluate the value of the policy under a given state and action; Among them, the value function V π has the following expression: Action value function Q π The expression of which is as follows: Use the Bellman equation to iteratively estimate the value function V π and the action-value function Q π ; Denote the value function at the next moment as Denote the action-value function at the next moment as Then: Use the policy network π(θ) to fit the specific policy probability distribution, and use the policy gradient formula to optimize the policy parameters. The optimization goal is to make the policy probability distribution of the agent maximize the reward. The expression of the policy gradient formula is: S44. Define the advantage function: Use the advantage function A(a, s) to measure the value of a specific action for a given state. The expression of the advantage function A(a, s) is: A π (s t ,a t ) = Q π (s t ,a t ) - V π (s t ) (8); S45. Introduce the proximal optimization rule: By setting a confidence interval, limit the KL divergence between the new policy and the old policy to alleviate the policy gradient fluctuation and increase the probability of selecting actions with higher advantage values estimated according to the current policy; Let the new policy after the update policy be π′(a t |s t ; θ′), then the expression of the loss function obtained by importance sampling is as follows: Use the KL divergence to measure the similarity of the probability distributions of the new and old policies, then the constraint condition of the optimization process is: S46. Construct a loss function: Use the Lagrange multiplier method to integrate formula 9 and formula 10 into one item to construct an improved loss function L′. The expression of the improved loss function L′ is: where β is a Lagrange multiplier; when βπ(a t |s t ; θ) and π ′ (a t |s t ; θ ′ ) are close, use r t to represent the ratio of the probability of the same action under the new policy and the old policy, and then use a parameter ∈ to control the ratio r t of the probabilities within the range of (1 - ∈, 1 + ∈), then the final expression of the loss function is: S47. Update the policy using the gradient descent method: Update the policy π by performing gradient descent optimization on the loss function θ ; S5. Construct a reward function to evaluate the blind hexapod robot and provide a feedback signal for the optimization algorithm to guide the training and parameter optimization of the policy network.
2. The training method for the motion strategy of a blind hexapod robot according to claim 1, characterized in that, step S1 is specifically: Establish a physical model of the blind hexapod robot in Mujoco, and define the body structure, joint parameters and action control mode of the blind hexapod robot to describe the toe trajectory of the blind hexapod robot; Establish a physical model of the wild mountain in Mujoco, and define the roughness of the terrain and the physical positions of the obstacles to make the established physical model of the wild mountain similar to the real terrain.
3. The training method for the motion strategy of a blind hexapod robot according to claim 1, characterized in that, step S1 further includes setting parameters of the blind hexapod robot, and the parameters include the central coordinates of the robot torso plane, the root coordinates of the six hip joints, the length of the hip joint, the length of the knee joint, and the length of the ankle joint.
4. The training method for the motion strategy of a blind hexapod robot according to claim 3, characterized in that, Taking the forward direction of the robot as the positive x-axis direction, the parameters of the blind hexapod robot are specifically: The central coordinates of the torso plane of the blind hexapod robot are O(0, 0, 0); The root coordinates of the six hip joints are respectively: LF(93.6, 50.805, 0), LM(0, 73.575, 0), LR(-93.6, 50.805, 0), RF(93.6, -50.805, 0), RM(0, -73.575, 0), RR(-93.6, -50.805, 0); The length of the hip joint L1 = 45.01; The length of the knee joint L2 = 77.06; The length of the ankle joint L3 = 150.
3.
5. The training method of the motion strategy of the blind hexapod robot according to claim 1, characterized in that, in step S2, the state information of the blind hexapod robot in the environment is output by an array, and a total of 46 - dimensional data is used as the body perception signal of the blind hexapod robot; Using n i to represent each element of the body signal, the status information s t is represented as s t = [n 1 , n 2 , ……, n 46 .
6. The training method of the motion strategy of the blind hexapod robot according to claim 5, characterized in that, among the 46 - dimensional data, 4 dimensions are used to represent the posture information of the blind hexapod robot, 18 dimensions are used to represent the joint angles of the free joints of the blind hexapod robot, 3 dimensions are used to represent the linear velocity when the blind hexapod robot moves, 3 dimensions are used to represent the angular velocity when the blind hexapod robot moves, and 18 dimensions are used to represent the joint velocities of the free joints of the blind hexapod robot.
7. The training method of the motion strategy of the blind hexapod robot according to claim 1, characterized in that, the policy network constructed in step S3 is a four - layer fully - connected neural network, The first layer of the policy network is the input layer, and the input is the state s t = [n 1 , n 2 , ……, n 46 array; the second and third layers of the policy network are hidden layers, and each layer has 96 neurons; The fourth layer of the policy network is the output layer, and the output is the joint angle a T = [m 1 , ……, m 18 , where m i represents the action of each joint.
8. The training method of the motion strategy of the blind hexapod robot according to claim 1, characterized in that, step S5 is specifically: Total reward r of the blind hexapod robot t Consisting of reward r v And penalty r θ 、r e The total reward r t The expression of is: r t = λ 1 r v - λ 2 r θ - λ 3 r e (13); Among them, λ1, λ2, and λ3 are weight parameters; the speed reward r v is used to encourage the blind hexapod robot to move forward, and the closer the actual speed is to the set speed, the more rewards are obtained; the penalty r θ is used to correct the heading to ensure that the heading is consistent with the movement direction; the penalty term r e plays a punishing role in the reward function and is used to encourage the blind hexapod robot to avoid obstacles in the mountain environment and maintain a stable walking posture; Let the distance traveled by the blind hexapod robot in a single step be \(x\). d , then the speed reward \(r\) v has the following expression: Let the scalar component of the quaternion in the current state be q w , then the penalty r θ has the following expression: Penalty term r for the blind hexapod robot when encountering an obstacle e The expression is as follows: r e = (Pitch * 4) 2 + (Roll * 3) 2 (16); Update the policy by minimizing the loss function to obtain the policy network π(θ).
Citation Information
Patent Citations
Self-adaptive gait planning method, system and device for hexapod robot and medium
CN114326722A
Quadruped robot motion control method and system, storage medium and equipment
CN114609918A