Quadruped robot reinforcement learning motion planning method and system based on depth vision
Through the reinforced learning motion planning method of four-legged robot based on depth vision, the terrain data obtained by the depth camera and the robot's ontology perception information are used to train a motion strategy suitable for irregular stair terrain, solving the problems of poor adaptability and low learning efficiency in the existing technology, and achieving more efficient motion planning and autonomous decision-making capabilities.
Patent Information
- Application Number
- CN202510467417.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The motion planning of existing four-legged robots in irregular stair terrain has problems such as poor adaptability to terrain training, low learning efficiency, and excessive reliance on artificially designated directions of travel, resulting in insufficient independent decision-making and adaptability.
The reinforcement learning motion planning method of four-legged robot based on depth vision is adopted to obtain three-dimensional point cloud data through a depth camera, build a local height map, and combine ontology perception information and privileged information to train the teacher's strategy network, output joint torque values, and combine the moving line speed to evaluate the strategy, dynamically adjust the terrain difficulty, and improve the adaptability and efficiency of training.
It improves the movement ability and decision-making efficiency of four-legged robots in complex terrain, enhances autonomous navigation capabilities, reduces dependence on external perception sensors, and can better cope with unknown changes in dynamic and complex environments.
Smart Images

Figure CN119974024A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of quadruped robot motion planning, and in particular to a quadruped robot reinforcement learning motion planning method and system based on depth vision. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] Compared with wheeled robots, quadruped robots have greater advantages in complex and rugged terrain. This is because quadruped robots can dynamically adjust the position of their footholds according to the terrain environment, thereby better coping with rugged, irregular or dynamically changing terrain. Therefore, quadruped robots have broad application prospects in a variety of complex scenarios such as mountainous areas, urban ruins, rescue missions and outdoor adventures. However, motion planning of quadruped robots based on external perception is still a very challenging task. The motion planner needs to process a large amount of perception data and perform complex calculations and decisions to ensure that the robot can move smoothly in a dynamically changing environment.
[0004] In recent years, with the development of deep learning and reinforcement learning technologies, the research on motion planning of quadruped robots has made significant progress. Through learning-based methods for motion planning and control, robots can explore and learn autonomously without human intervention, thereby achieving more flexible movement capabilities. Such methods not only help robots complete tasks in known environments, but also improve their ability to adapt to changes in unknown environments, especially in complex terrain and unforeseen environmental conditions. However, although reinforcement learning provides quadruped robots with powerful autonomous learning capabilities, existing quadruped robot motion planning methods still face a series of problems in practical applications.
[0005] In the reinforcement learning motion planning of quadruped robots facing complex terrain, especially irregular staircase terrain, existing methods usually have problems such as poor terrain training adaptability, low learning efficiency, and over-reliance on manually specified travel directions. They cannot effectively cope with the dynamic challenges in the staircase environment, resulting in insufficient autonomous decision-making and adaptability of the robot. Therefore, how to effectively learn irregular staircase terrain so that the quadruped robot can adaptively operate stably in complex environments has become an urgent problem to be solved in existing technologies. Summary of the invention
[0006] In view of the shortcomings of the prior art, the purpose of the present invention is to provide a quadruped robot reinforcement learning motion planning method and system based on deep vision, which adds a visual perception module to the motion planning task in an irregular stair environment. The robot can obtain terrain information from visual information, autonomously predict heading and plan actions.
[0007] In order to achieve the above object, the present invention is implemented through the following technical solutions: A first aspect of the present invention provides a quadruped robot reinforcement learning motion planning method based on depth vision, comprising the following steps: A simulation environment is built, stair terrains with different slopes are modeled to obtain a stair terrain model, and robot dynamic parameters are modeled to obtain a robot model, wherein the stair terrain model includes stair terrains with different difficulty differences; The robot model is loaded to move in the stair terrain model, and the robot's motion status is tracked to obtain the speed tracking effect. The speed tracking effect is obtained by evaluating the joint torque value and the corresponding motion linear velocity output by the teacher strategy network during the robot's motion. The speed tracking effect is evaluated and the environmental parameters in the stair terrain model are adaptively adjusted according to the evaluation results.
[0008] Furthermore, the specific steps of loading the robot model to move in the stair terrain model and tracking the robot's motion status are as follows: Obtain robot depth vision information and robot proprioception information; Reconstructing a local height map based on the robot's deep vision information; The teacher strategy network is used to process the local height map, robot proprioception information and privileged information to obtain the torque value of each joint of the robot. The torque value of each joint is combined with the linear velocity of motion to perform strategy evaluation and obtain the speed tracking effect.
[0009] Furthermore, the specific steps of reconstructing the local height map based on the robot's deep vision information are as follows: Initial 3D point cloud data is obtained, and after filtering and coordinate conversion processing are performed on the initial 3D point cloud data, a local 2D height map is generated as a local height map through projection and bilinear interpolation algorithms.
[0010] Furthermore, the specific steps of using the teacher strategy network to process the local height map, robot body perception information and privilege information are as follows: Building a teacher strategy network; The teacher strategy network is trained based on reinforcement learning to obtain a trained teacher strategy network; The teacher strategy network is used to process the local height map, robot proprioception information and privileged information to obtain the torque value of each joint of the robot.
[0011] Furthermore, the specific steps for training the teacher strategy network based on reinforcement learning are: Training teacher strategies using multi-layer perceptrons; Build a supervised reinforcement learning training framework to distill the student strategy network from the trained teacher strategy network.
[0012] Furthermore, the specific steps of evaluating the strategy by combining the torque value of each joint with the linear velocity of motion are as follows: The torque values of each joint of the robot's legs and the corresponding linear velocity of motion output by the teacher strategy network are input into the strategy evaluation network. The reward signal of the strategy evaluation network is used to generate the robot's motion decision in complex terrain as a speed tracking effect.
[0013] Furthermore, the speed tracking effect is evaluated, and the specific steps of adaptively adjusting the environmental parameters in the stair terrain model according to the evaluation results are as follows: At the end of each round of training, the root mean square error between the actual speed of the robot and the target speed is calculated. If the root mean square error is lower than the preset threshold, the terrain difficulty signal is increased; if it is higher than the threshold, the terrain difficulty signal is reduced; otherwise, the current terrain difficulty signal is maintained.
[0014] A second aspect of the present invention provides a quadruped robot reinforcement learning motion planning system based on deep vision, comprising: The simulation building module is configured to build a simulation environment, model stair terrains with different slopes to obtain a stair terrain model, and model robot dynamic parameters to obtain a robot model, wherein the stair terrain model includes stair terrains with different difficulty differences; The motion tracking module is configured to load the robot model to move in the stair terrain model, track the robot's motion status, and obtain a speed tracking effect, wherein the speed tracking effect is obtained by evaluating the joint torque value and the corresponding motion line speed output by the teacher strategy network during the robot's motion; The effect evaluation module is configured to evaluate the speed tracking effect and adaptively adjust the environmental parameters in the stair terrain model according to the evaluation results.
[0015] The third aspect of the present invention provides a medium having a program stored thereon, which, when executed by a processor, implements the steps in the deep vision-based quadruped robot reinforcement learning motion planning method as described in the first aspect of the present invention.
[0016] The fourth aspect of the present invention provides a device, including a memory, a processor, and a program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the deep vision-based quadruped robot reinforcement learning motion planning method as described in the first aspect of the present invention are implemented.
[0017] One or more of the above technical solutions have the following beneficial effects: The present invention discloses a method and system for motion planning of a quadruped robot based on reinforcement learning based on deep vision. Firstly, a local height map centered on the robot is constructed based on the three-dimensional point cloud data obtained by the depth camera carried by the quadruped robot, and the teacher strategy is trained by combining proprioceptive information (such as joint position, speed, angular velocity, etc.) and other privileged information (such as terrain friction coefficient, target path points, etc.) as input; the teacher strategy adopts a model-free reinforcement learning method, and outputs appropriate leg joint torque with the support of local height map and proprioceptive information; in the training process, a double distillation training framework is adopted, and the knowledge transfer of the teacher-student model is used to effectively improve the training efficiency and reduce the system deployment cost; the student strategy gradually optimizes its motion decision-making process by learning the behavior of the teacher strategy, so as to realize efficient motion control of the robot in complex terrain. ; In order to improve the efficiency and adaptability of the training process, a terrain course improvement module is added to dynamically adjust the terrain difficulty of the training to ensure that the robot learns in more difficult terrain and avoid premature training saturation. The terrain difficulty is adjusted according to the robot's performance to ensure the gradual progression of training; the local height map and proprioceptive information are updated in real time during the training process to ensure that the robot obtains the latest environmental information in order to accurately predict and execute the next movement; in addition, the present invention reduces the dependence on external perception sensors and utilizes the image data of the depth camera to reduce the amount of calculation and improve the robot's autonomous navigation capability, so that the robot can better cope with complex and dynamically changing environments; through the comprehensive application of the above-mentioned technologies, the present invention provides an efficient and reliable motion planning method for a quadruped robot, which improves the robot's motion ability and decision-making efficiency in complex terrain.
[0018] Compared with the traditional quadruped robot motion planning method, the present invention adopts a reinforcement learning framework based on deep visual information, which significantly improves the robot's motion ability in complex environments. First, in the teacher strategy training stage, the point cloud data collected by the depth camera is used to reconstruct the local height map, and combined with proprioceptive information and other privileged information, it is used as input to train the robot strategy; at the same time, the target path point is preset in the terrain environment as a guide for the robot's forward direction, which improves the accuracy and adaptability of terrain perception and avoids the problem of over-reliance on artificially specified travel directions; through the three-dimensional point cloud data of the depth camera, a more accurate height map can be reconstructed, which makes up for the possible lack of details or errors when calling the API directly. Secondly, in the student strategy training stage, the CNN-GRU network architecture is used instead of the preset path point input, the depth image is used as input, and the DAgger algorithm is combined for training to autonomously predict the optimal heading, which improves the flexibility and autonomy of path planning; through the distillation method, the student strategy learns the behavior of the teacher strategy to ensure that when relying only on the depth image and proprioceptive information, it generates motion decisions consistent with the teacher strategy. Finally, in response to the problems of static terrain design and non-targeted training in traditional reinforcement learning processes, the present invention designs a terrain course improvement module, which dynamically adjusts the terrain difficulty according to the robot performance, evaluates the strategy performance in real time, optimizes the terrain parameters, enhances the robot's ability to cope with more complex environments, and improves the adaptability and efficiency of training.
[0019] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0021] Figure 1 This is an overall flow chart of a method for motion planning of a quadruped robot based on deep vision according to the first embodiment of the present invention; Figure 2 This is a strategy update control flow chart of Embodiment 1 of the present invention; Figure 3 This is a teacher-student strategy framework diagram of Embodiment 1 of the present invention; Figure 4 This is a schematic diagram of the movement of a quadruped robot on a staircase terrain according to Embodiment 1 of the present invention. DETAILED DESCRIPTION
[0022] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0023] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "include" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or their combinations; Embodiment 1: With the widespread application of quadruped robots in industrial manufacturing, medical services, logistics distribution, military applications and other fields, their intelligence level and environmental adaptability are constantly improving. Compared with wheeled robots, quadruped robots have higher dynamic stability and stronger environmental adaptability, and can cope with complex terrain and perform more challenging tasks. In the face of dynamic and irregular terrain, especially in complex environments such as rugged stairs, quadruped robots need to have a high degree of autonomy and flexibility to complete tasks.
[0024] At present, the motion planning problem of quadruped robots under irregular stair terrain is studied by the reinforcement learning motion planning method of quadruped robots based on deep visual information. To realize the reinforcement learning motion planning problem of quadruped robots, multiple issues need to be considered: The first is the accuracy of terrain data collection. The method of directly collecting terrain data through the terrain height reading function in Isaac Gym and using it as privileged information for motion planning has limitations when dealing with complex and dynamic terrain. In particular, when the environment changes rapidly or details change, the terrain data directly generated in the simulation is difficult to reflect the slight changes in the local terrain or new obstacles, resulting in the robot being unable to perceive and make corresponding motion decisions in a timely manner. In addition, the strategies trained by the terrain data read by the function are usually only applicable to the simulation environment and cannot be effectively applied to the real world. There are differences between the simulation environment and the actual environment, which may cause the robot to have insufficient adaptability and autonomous decision-making capabilities when it is actually deployed.
[0025] The second is the autonomy and flexibility of path planning. Robots usually rely on external input of direction and speed instructions. This reliance on manual designation limits the robot's autonomous navigation capabilities, especially in complex environments, where the robot cannot independently decide on the direction of travel and autonomously avoid obstacles. The robot can only travel along a predetermined path and lacks the ability to respond to environmental changes in real time, resulting in insufficient flexibility and adaptability in dynamic or unknown environments. In particular, when faced with dynamic obstacles or sudden environmental changes, the robot cannot adjust its path in time, affecting its overall performance.
[0026] The third is the dynamic adaptability of the motion strategy. During the reinforcement learning training process, the terrain and environment design are usually static, lacking dynamic adjustment and response mechanisms. The robot completes the training under fixed environmental conditions, which makes it impossible for the robot to effectively adapt to the changing terrain or unknown challenges in the environment. Especially in a dynamic environment where the terrain and obstacles change, the robot may encounter unseen terrain, which makes the generalization ability of the training model insufficient and difficult to cope with complex changes in practical applications.
[0027] In order to solve the problems of poor terrain training adaptability, low efficiency and over-reliance on manually specified travel directions, the first embodiment of the present invention provides a quadruped robot reinforcement learning motion planning method based on deep vision, such as Figure 1 As shown. First, create irregular staircase terrains of different specifications, import the robot model and randomize the parameters. After that, obtain the robot's proprioceptive perception information and deep visual information, build a terrain reconstructor to construct a local height map. In addition, the multi-layer perceptron (MLP) is used to train the teacher strategy in combination with privileged information to construct a terrain improvement module. Finally, a supervised training framework is built to distill the student strategy. The present invention adopts a double distillation training framework, which improves the training efficiency and reduces the system deployment cost by means of knowledge transfer between the teacher and student models; through the deep visual network architecture, the optimal heading can be autonomously inferred from the low-frequency depth image, while reducing the external perception calculation amount, the robot's autonomous navigation ability is improved, so that it can cope with unknown changes in dynamic and complex environments; at the same time, a terrain course improvement module is introduced, which encourages the robot to learn in more difficult terrains by dynamically adjusting the terrain difficulty during training, thereby improving the targeted training and avoiding premature training saturation.
[0028] The method specifically comprises the following steps: Step 1: Build a simulation environment, model the stairs with different slopes to obtain the stair terrain model, model the robot dynamic parameters to obtain the robot model.
[0029] Step 1.1: Model the staircase terrains with different slopes to obtain a staircase terrain model, wherein the staircase terrain model includes staircase terrains with different difficulty differences.
[0030] In a specific implementation, the Isaac Gym physics engine is used to build a simulation environment, including modeling of stairs with different slopes and robot dynamics modeling. In the simulation environment, a staircase terrain is built, and the staircase terrain is composed of rectangular blocks of different sizes, friction coefficients, and numbers defined in the code. This embodiment designs three types of staircase terrains with obvious differences in difficulty as components of complex terrains, namely low stairs, ordinary stairs, and high stairs. Each type of terrain is divided into two types, mild inclination and severe inclination, according to the inclination of the stairs. There are six types of terrains in total, and a target path point is set at the center of each step.
[0031] Step 1.2: Identify the robot dynamic parameters and build the robot model.
[0032] The robot model is loaded by importing the URDF file, and the domain randomization method is used to dynamically adjust the robot's dynamic parameters, including the mass of each joint of the robot, the mass of the body, the body inertia and the terrain friction coefficient, and additional torque and force are added as disturbance terms to improve the robot's anti-interference ability.
[0033] According to the simulated environment in which the quadruped robot is located, the depth camera information equipped on the robot head and the robot proprioception information are obtained to train the autonomous movement strategy for irregular stair terrain.
[0034] Step 2: Load the robot model and move it in the stair terrain model, track the robot's motion status, and obtain the speed tracking effect. The speed tracking effect is obtained by evaluating the joint torque value and the corresponding linear velocity output by the teacher strategy network during the robot's motion.
[0035] Step 2.1: Obtain robot depth vision information and robot proprioception information.
[0036] The robot model is loaded at the starting position of each type of terrain, and the robot's current proprioception information is obtained, and the three-dimensional point cloud information in the depth camera worn by the robot is used as the depth vision information.
[0037] In a specific implementation, proprioceptive information is obtained by sensors inside the robot. The IMU, joint encoders, and torque sensors inside the robot body are used to provide proprioceptive information such as the position and speed of each joint of the robot's legs, the angular velocity of the body, the posture angle, and the sole contact judgment, thereby providing the strategy network with an accurate numerical value of the robot's current state.
[0038] The robot proprioceptive information can be represented as a 53-dimensional feature vector, denoted as: .
[0039] in, is a set of feature vectors, including 3D angular velocity , 2D pitch and roll angles read by imu , yaw angle , Yaw angle scaling , the yaw angle at the next moment , 3D control commands , 2D environment type judgment , 12-dimensional joint position , 12-dimensional joint velocity , 12-dimensional history commands and 4-dimensional foot contact judgment .
[0040] In the teacher strategy stage, the deep visual information is presented as 3D point cloud data. After filtering and coordinate transformation of the initial point cloud data, a local 2D height map is generated through projection and bilinear interpolation algorithm to reflect the change of terrain height. In the student strategy stage, 8 consecutive images of size The depth camera on the robot's head can provide terrain environment information and help the robot plan its actions.
[0041] Step 2.2: Reconstruct the local height map based on the robot's deep vision information.
[0042] In a specific implementation, the three-dimensional point cloud information in the equipped depth camera is read and a local height map centered on the robot is reconstructed. Specifically, the initial three-dimensional point cloud data is obtained from the depth camera equipped on the robot head, and after filtering and coordinate conversion processing, a local two-dimensional height map is generated as a local map through projection and bilinear interpolation algorithm to reflect the change in terrain height.
[0043] The generated local height map is combined with proprioceptive information and other privileged information and input into the policy network composed of MLP to train the teacher motion strategy. During each training or operation, the module continuously updates the local map to ensure that the robot's motion decisions are based on the latest terrain data.
[0044] Step 2.3: Use the teacher strategy network to process the local height map, robot body perception information and privileged information to obtain the torque value of each joint of the robot and combine it with the linear velocity of the robot body to evaluate the strategy and obtain the speed tracking effect.
[0045] like Figure 3 As shown, it includes the teacher strategy network in the first stage Training process and student strategy network in the second stage Training process. In the first stage, the local height map, preset path points, robot proprioceptive information (53 dimensions), and other privileged information such as body inertia encoded by the privileged information encoder are used as input, and the teacher strategy is trained using MLP to output actions. Among them, the depth camera obtains the local height map through the terrain reconstructor, and the robot proprioceptive information is obtained through the IMU and joint encoder. In the second stage, the CNN-GRU network is used to extract the direction and speed from the depth image obtained by the depth camera, combined with the robot proprioceptive information as input, and the initialization parameters of the student strategy network are set to the deep copy from the first stage. The output of the first stage is used for supervised learning to obtain the output action.
[0046] Among them, the privileged information includes terrain friction coefficient, target path point, body inertia, body mass, motor strength, motor torque output ratio and communication delay time.
[0047] In a specific implementation, during the teacher strategy network training process, the environment privilege information, the local height map centered on the robot, and the robot's current proprioceptive information are input into the teacher strategy, and then the torque of each joint of the robot's leg is output and transmitted together with the corresponding linear velocity of motion to the strategy evaluation network, and the strategy evaluation network determines the next state and action of the robot based on the reward value feedback. In addition, the network weight parameters in the teacher strategy are initialized to the student strategy to speed up the training speed, and supervised training is performed. The student strategy needs to take the encapsulated proprioceptive information and the depth image collected by the depth camera as input, and use the CNN-GRU visual network architecture to predict the target heading from the depth image. The final output of the strategy is the joint torque value.
[0048] Step 2.3.1: Construct the teacher-student network. The teacher-student network includes the teacher-strategy network and Student Policy Network .
[0049] Step 2.3.2: Train the teacher strategy network based on reinforcement learning to obtain the trained teacher-student network. The training process of the teacher strategy network is the first stage, and the training process of the student strategy network is the second stage.
[0050] Step 2.3.2.1: Train the teacher strategy using MLP.
[0051] The teacher policy network consists of a gated recurrent unit (GRU) and an MLP. The GRU is used to process time series data and extract motion features, and the MLP is used as a policy network to calculate the torque value of the robot joint. The parameters of the policy network are optimized using the proximal policy optimization network architecture (PPO).
[0052] Since the robot motion decision involves time series data (historical states affect current decisions), using only MLP cannot capture time dependencies. Therefore, GRU is used to process time series, and weights are optimized through back propagation through time (BPTT), backpropagating from the final time step to the past 24 time steps, and calculating the gradient of the loss function to the GRU and MLP weights to update the time series modeling parameters and the output parameters of the policy network to optimize the robot motion strategy.
[0053] In the first stage, the local height map is used, combined with the robot's proprioceptive information and other privileged information, and input into the teacher policy network. The teacher policy network is trained using model-free reinforcement learning. The policy input is processed by GRU and MLP to extract terrain features and motion states, and output the robot's leg joint torque. Specifically, the teacher policy network is trained using model-free reinforcement learning. , the teacher policy network has access to the local height map , proprioceptive information and environmental parameters Privileged information such as GRU and MLP are used to extract feature vectors and , the final output action , that is, the torque values of the twelve joints of the robot's legs at each moment.
[0054] First, the local height map The point cloud in is compressed into , and then passed to the GRU for predicting joint angles along with the rest of the observations, using the GRU to process the environment point cloud , respectively and , as input to the basic feed-forward strategy.
[0055] The basic feedforward network is the strategy structure used to generate joint torques in the teacher strategy. It combines GRU and MLP for processing and is used to extract features from three-dimensional point cloud data and robot proprioception information, and finally outputs the torques of the 12 joints of the robot's legs.
[0056] The specific formula is: , , .
[0057] This embodiment uses PPO for training. During the strategy training process, BPTT is used to calculate the loss function and backpropagate from the final time step to the past 24 time steps. After that, the gradient of the loss function with respect to the GRU and MLP weights is calculated to update the time series modeling parameters of the GRU and the output parameters of the MLP to optimize the robot motion strategy.
[0058] The output joint torque and linear velocity of motion will serve as the input of the strategy evaluation network, which will evaluate and adjust the robot behavior based on the reward value feedback.
[0059] During the training of the teacher strategy network, the double distillation method is used to transfer the knowledge of the teacher strategy network to the student strategy network. The student strategy network uses the weights of the teacher strategy network when it is initialized, and is trained through supervised learning to optimize the parameters of the student strategy network to improve learning efficiency. In this process, the action output of the teacher strategy network generates the robot's motion decision in complex terrain through the reward signal of the strategy evaluation network.
[0060] Step 2.3.2.2: Build a supervised reinforcement learning training framework to distill the student strategy network from the trained teacher strategy network.
[0061] In the second phase, in the student strategy network During the training, the student policy network is trained using DAgger, and the time-truncated back-propagation algorithm is used to minimize the mean square error between the travel speed predicted by the student policy network and the travel speed output by the teacher policy network. The student policy network is trained through supervised learning, so that the quadruped robot can accurately track the speed command. The output of the student policy network includes the required linear velocity and the required yaw angle, thereby ensuring the precise movement of the robot under different stair terrain conditions. Among them, the required linear velocity includes the basic velocity of forward and lateral movement. In the teacher policy training stage, in order to obtain more accurate environmental information, the deep visual information uses three-dimensional point cloud data to train the optimal strategy. At the same time, a state estimation network based on the CNN-GRU network architecture is constructed for feature extraction in the student policy training stage; in the student policy training stage, in order to reduce the external perception calculation cost, the deep visual information uses two-dimensional depth images, and the state estimation network generated in the first stage is embedded in the student policy network as a trained module to extract terrain features and time series features from continuous depth images, and predict the target direction and speed.
[0062] The student policy network receives the encapsulated proprioceptive information and the depth image collected by the depth camera as input, and uses the CNN-GRU visual network architecture to extract spatial features from the depth image and predict the target heading. The output joint torque value is optimized by the regression loss function so that the action of the student policy network is consistent with that of the teacher policy network. The weights of the student policy network are initialized by the teacher policy network, and supervised learning is performed during the training process, gradually adjusting the strategy to adapt to different terrain conditions. During training, the student policy network uses the depth image from the depth camera, the speed command output by the state estimation network, and the robot proprioceptive information as input, and uses DAgger to optimize the output of the student policy network to be close to the output of the teacher policy network. Through continuous iterative updates, the student policy network learns more precise motion control, thereby achieving efficient autonomous decision-making in a variety of terrain environments.
[0063] The input of the student policy network includes external perception information (depth image ) and proprioceptive information And use the state estimation network generated in the first stage as the convolution layer in the second stage to extract the deep image information, GRU to extract the timing features, and finally predict the output direction and speed instructions.
[0064] The main goal of the second stage is to use supervised learning to distill the teacher policy network generated in the first stage into an architecture that only relies on onboard sensor information, without retraining the entire controller, only the feature vector and The estimator of is used, and the base policy trained in the first stage is used.
[0065] In the second stage, DAgger is used for training to minimize the mean square error between the student strategy prediction action and the teacher strategy network output action through BPTT. ,in, Output actions for the teacher strategy network, i.e., real actions, To predict actions. Specifically, the depth image is first preprocessed and visual features are extracted through a convolutional network. Afterwards, the GRU network is used to process the historical proprioceptive information and historical deep visual information to estimate the latent features of the terrain geometry. Since the camera is located in front of the robot, the proprioceptive information combined with the depth enables the GRU to implicitly track and estimate the terrain below the robot, obtaining the estimated latent features. Afterwards, the external features are estimated using the historical proprioceptive information and the direction and speed instructions predicted by the state estimation network. Finally, the MLP network combines the estimated latent features and external features to generate the predicted action .
[0066] In order to minimize the behavioral deviation between the teacher strategy and the student strategy, this embodiment sets the initialization parameters of the student strategy network to the copy from the first stage. The teacher-student network architecture is as follows: Figure 3 Doing so ensures that the student strategy starts out behaving as similarly as possible to the teacher strategy, thus accelerating the learning process.
[0067] Through the trained student strategy, environmental features are extracted from the depth image collected by the robot, and the current motion state estimation is generated in combination with proprioceptive information. Afterwards, the student strategy outputs the torque value of each joint to control the robot's motion according to the current proprioceptive information and external environmental information. Through reinforcement learning optimization, the student strategy is continuously adjusted to maximize the long-term cumulative reward. In the second stage, this embodiment can combine real-time environmental information and motion planning goals to ensure that the robot can make autonomous decisions and move efficiently in dynamic and irregular stair terrain.
[0068] Since the CNN-GRU network has not been trained in the initial stage of training, if the target path point predicted by the CNN-GRU network is directly used as the robot's heading, the robot will walk aimlessly. Therefore, a mixed method of teacher and student strategies is selected to predict the direction of movement. Specifically, when the angle between the direction of movement predicted by the student strategy network and the direction of movement predicted by the teacher strategy network is less than a preset threshold, the direction of movement predicted by the student strategy network is used; otherwise, the direction of movement predicted by the teacher strategy network is used.
[0069] Step 2.3.3: According to the training process, output the torque value of each joint of the robot.
[0070] Step 2.3.4: Input the torque values of each joint of the robot's legs and the corresponding linear velocity of motion output by the teacher strategy network into the strategy evaluation network, and generate the robot's motion decision in complex terrain as a speed tracking effect through the reward signal of the strategy evaluation network.
[0071] In a specific implementation, the torque values of each joint of the robot's leg and the linear velocity of motion output by the teacher-student strategy network are input into the strategy evaluation network, which consists of four modules: input layer, feature extraction module, reward calculation module and output layer. First, the input layer receives the torque values of the robot's leg joints and the linear velocity of motion output by the teacher strategy network. After that, MLP is used to extract motion features, calculate the reward value of the robot's current behavior, and the output layer provides the reward value to the PPO network. PPO aims to maximize the action reward and optimizes the parameters of the teacher strategy network through iterative training.
[0072] The basic principle of designing the reward function in this embodiment is to guide the robot to move stably in a complex environment through the robot's ability to track heading, control body balance, and perform precise movements, and to encourage the robot to make necessary dynamic adjustments during the task, to ensure that it can autonomously optimize its motion strategy and demonstrate high autonomy and intelligence.
[0073] In order to meet the design requirements, the reward function consists of multiple different reward items, and the total reward item for each control time step is defined as the weighted sum of each reward item. The total reward items in this embodiment are mainly divided into two categories: tracking rewards and regularization rewards. The tracking rewards include target speed tracking rewards and target yaw rewards, and the regularization rewards include vertical linear velocity penalties, posture error penalties, collision penalties, hip position penalties, and foot edge penalties. The sum of each reward item multiplied by its corresponding weight coefficient is the total reward function value.
[0074] In this embodiment, the end step length for each round is designed to be 50,000, and the training process of this round is considered to be completed when the robot reaches the end position of the current terrain, and the quadruped robot is reset to the starting coordinates of the training. If the quadruped robot fails to move to the desired position, it will be reset to the starting coordinate position of the movement after the end step length is reached in this round. The overall control process of the strategy update is as follows: Figure 2 As shown in the figure, the quadruped robot uses the deep visual information and proprioceptive information collected by the sensor as the input state s of reinforcement learning. The PPO network calculates and outputs the control signal a of the robot in the action space based on the input state s. Subsequently, the control signal a is mapped to the 12 joint angles of the robot, and the position control is used to enable the quadruped robot to complete the target action. At the same time, the system will decide whether to reset the robot to the starting position based on whether the robot reaches the end point or falls. All states, actions, and reward information are stored in the experience replay pool. When the data accumulates to the set threshold, the PPO strategy update is triggered to optimize the robot's movement ability in complex terrain until the reward value stabilizes and the training is completed.
[0075] Step 3: Evaluate the speed tracking effect and adaptively adjust the environmental parameters in the stair terrain model according to the evaluation results.
[0076] In a specific embodiment, Figure 4As shown, by evaluating the robot's performance in the current terrain partition, the terrain difficulty is dynamically adjusted. At the end of each round of training, the root mean square error (RMSE) between the actual speed of the robot's movement and the target speed is calculated. If the RMSE is lower than the preset threshold, the terrain difficulty signal is increased to increase the terrain complexity; if it is higher than the threshold, the terrain difficulty signal is reduced to reduce the terrain complexity; otherwise, the current terrain difficulty signal is maintained. According to the performance of the robot in different terrain partitions, the terrain difficulty is gradually increased or decreased to ensure that the training process is challenging and adaptable. In this embodiment, in each new stair terrain partition, the corresponding terrain features, such as slope, number of steps, ground friction coefficient, etc., are adjusted to achieve the adjustment of terrain complexity. The specific terrain corresponding to the difficulty signal can be customized according to actual conditions, which can help the robot adapt to more complex environments and update the target difficulty of training in real time according to actual performance.
[0077] Specifically, the speed tracking effect is evaluated. For each terrain, if the root mean square error of speed tracking is less than 0.2, the difficulty of the terrain is increased. Set to 1. If it is higher than 0.5, the terrain difficulty is increased. Set to -1. Otherwise, the terrain difficulty will be increased. Set to 0.
[0078] Embodiment 2: Embodiment 2 of the present invention provides a quadruped robot reinforcement learning motion planning system based on deep vision, including: The simulation building module is configured to build a simulation environment, model stair terrains with different slopes to obtain a stair terrain model, and model robot dynamic parameters to obtain a robot model, wherein the stair terrain model includes stair terrains with different difficulty differences; The motion tracking module is configured to load the robot model to move in the stair terrain model, track the robot's motion status, and obtain a speed tracking effect, wherein the speed tracking effect is obtained by evaluating the joint torque value and the corresponding motion line speed output by the teacher strategy network during the robot's motion; The effect evaluation module is configured to evaluate the speed tracking effect and adaptively adjust the environmental parameters in the stair terrain model according to the evaluation results.
[0079] Embodiment three: Embodiment 3 of the present invention provides a medium on which a program is stored. When the program is executed by a processor, the steps in the deep vision-based quadruped robot reinforcement learning motion planning method as described in Embodiment 1 of the present invention are implemented.
[0080] Embodiment 4: Embodiment 4 of the present invention provides a device, including a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, the steps in the deep vision-based quadruped robot reinforcement learning motion planning method as described in Embodiment 1 of the present invention are implemented.
[0081] The steps involved in the above embodiments 2, 3 and 4 correspond to the method embodiment 1. For the specific implementation methods, please refer to the relevant description part of embodiment 1.
[0082] Those skilled in the art should understand that the modules or steps of the present invention described above can be implemented by a general-purpose computer device, or alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0083] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.
Claims
1. A deep vision-based quadruped robot reinforcement learning motion planning method, characterized in that: The following steps are involved: A simulation environment is built, stair terrains with different slopes are modeled to obtain a stair terrain model, and robot dynamic parameters are modeled to obtain a robot model, wherein the stair terrain model includes stair terrains with different difficulty differences; The robot model is loaded to move in the stair terrain model, and the robot's motion status is tracked to obtain the speed tracking effect. The speed tracking effect is obtained by evaluating the joint torque value and the corresponding motion linear velocity output by the teacher strategy network during the robot's motion. The speed tracking effect is evaluated and the environmental parameters in the stair terrain model are adaptively adjusted according to the evaluation results.
2. The method for motion planning of a quadruped robot based on deep vision according to claim 1, characterized in that: The specific steps to load the robot model and track the robot's motion in the stair terrain model are as follows: Obtain robot depth vision information and robot proprioception information; Reconstructing a local height map based on the robot's deep vision information; The teacher strategy network is used to process the local height map, robot proprioception information and privileged information to obtain the torque value of each joint of the robot. The torque value of each joint is combined with the linear velocity of motion to perform strategy evaluation and obtain the speed tracking effect.
3. The method for motion planning of a quadruped robot based on deep vision according to claim 2, characterized in that: The specific steps for reconstructing the local height map based on the robot's deep vision information are: Initial 3D point cloud data is obtained, and after filtering and coordinate conversion processing are performed on the initial 3D point cloud data, a local 2D height map is generated as a local height map through projection and bilinear interpolation algorithms.
4. The method for motion planning of a quadruped robot based on deep vision according to claim 2, characterized in that: The specific steps of using the teacher strategy network to process the local height map, robot proprioception information and privilege information are as follows: Building a teacher strategy network; The teacher strategy network is trained based on reinforcement learning to obtain a trained teacher strategy network; The teacher strategy network is used to process the local height map, robot proprioception information and privileged information to obtain the torque value of each joint of the robot.
5. The method for motion planning of a quadruped robot based on deep vision according to claim 4, characterized in that: The specific steps for training the teacher strategy network based on reinforcement learning are: Training teacher strategies using multi-layer perceptrons; Build a supervised reinforcement learning training framework to distill the student strategy network from the trained teacher strategy network.
6. The method for motion planning of a quadruped robot based on deep vision according to claim 4, characterized in that: The specific steps of evaluating the strategy by combining the torque value of each joint with the linear velocity of motion are as follows: The torque values of each joint of the robot's legs and the corresponding linear velocity of motion output by the teacher strategy network are input into the strategy evaluation network. The reward signal of the strategy evaluation network is used to generate the robot's motion decision in complex terrain as a speed tracking effect.
7. The method for motion planning of a quadruped robot based on deep vision according to claim 1, characterized in that: The specific steps for evaluating the speed tracking effect and adaptively adjusting the environmental parameters in the stair terrain model according to the evaluation results are as follows: At the end of each round of training, the root mean square error between the actual speed of the robot and the target speed is calculated. If the root mean square error is lower than the preset threshold, the terrain difficulty signal is increased; if it is higher than the threshold, the terrain difficulty signal is reduced; otherwise, the current terrain difficulty signal is maintained.
8. A quadruped robot reinforcement learning motion planning system based on deep vision, characterized in that: include: The simulation building module is configured to build a simulation environment, model stair terrains with different slopes to obtain a stair terrain model, and model robot dynamic parameters to obtain a robot model, wherein the stair terrain model includes stair terrains with different difficulty differences; The motion tracking module is configured to load the robot model to move in the stair terrain model, track the robot's motion status, and obtain a speed tracking effect, wherein the speed tracking effect is obtained by evaluating the joint torque value and the corresponding motion line speed output by the teacher strategy network during the robot's motion; The effect evaluation module is configured to evaluate the speed tracking effect and adaptively adjust the environmental parameters in the stair terrain model according to the evaluation results.
9. A computer-readable storage medium, characterized in that: A plurality of instructions are stored therein, and the instructions are suitable for being loaded by a processor of a terminal device and executing the deep vision-based quadruped robot reinforcement learning motion planning method described in any one of claims 1-7.
10. A terminal device, characterized in that: It includes a processor and a computer-readable storage medium, the processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are suitable for being loaded by the processor and executing the deep vision-based quadruped robot reinforcement learning motion planning method described in any one of claims 1-7.
Citation Information
Patent Citations
Motion control method and system for quadruped robot under terrain subareas
CN118192254A
Human motion gradient recognition method based on GRU-CNN
CN118364374A
Motion control method for quadruped robot with damaged legs
CN118915802A
Ship collision avoidance optimization method under condition of uncertain obstacle ship motion information
CN119207166A
Deep reinforcement learning quadruped robot motion control method and system based on constraint reward
CN119512184A
Cited By
Quadruped robot robust adaptive multi-skill learning method based on key frame guidance
CN121523059A
A robust self-adaptive multi-skill learning method for quadruped robots based on key frame guidance
CN121523059B
Quadruped robot eye-leg cooperative obstacle crossing method and system facing complex terrain
CN121722121A
A complex terrain-oriented four-legged robot eye-leg cooperative obstacle crossing method and system
CN121722121B