A method for omnidirectional motion control of a quadruped robot in a three-dimensional environment based on hybrid representation learning
Through the hybrid representation learning method, the quadruped robot achieves omnidirectional motion control in a three-dimensional environment, solves the problems of field of view occlusion and inconsistent field of view, improves its ability to pass in complex environments, and is suitable for a variety of quadruped robot platforms.
Patent Information
- Application Number
- CN202411747494.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2044-12-02
AI Technical Summary
Existing quadruped robot motion control methods have difficulty achieving omnidirectional motion in a three-dimensional environment, especially when the field of view is blocked or the motion direction is inconsistent with the field of view, and cannot effectively pass through complex obstacles.
Using a hybrid representation learning method, an asymmetric action network-strategy network structure reinforcement learning framework is designed. Combining supervised learning, contrastive learning and reinforcement learning, an end-to-end training framework is constructed through the proprioception and deep visual information of the quadruped robot to achieve a full understanding of the local environment and omnidirectional motion control.
The quadruped robot can achieve omnidirectional navigation on a variety of challenging terrains and has excellent obstacle-crossing capabilities. It is suitable for various quadruped robot platforms, including Jueying Lite3, Jueying X30, ANYmal C, Spot, Cheetah3, etc., and can complete tasks such as jumping off platforms, climbing stairs, and passing through low holes.
Smart Images

Figure CN119620757B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of quadruped robot motion control based on reinforcement learning, and in particular to a method for omnidirectional motion control of a quadruped robot in a three-dimensional environment based on hybrid representation learning. Background Art
[0002] In recent years, with the continuous development of motion control algorithms for quadruped robots, quadruped robots have garnered increasing attention due to their superior locomotion capabilities. Quadruped robots, with their animal-like body structure, can flexibly adjust their foothold and posture across various terrains to adapt to the demands of navigating complex environments. Furthermore, compared to humanoid robots and multi-legged robots like hexapods, quadruped robots, thanks to their stable and efficient body design, can achieve more robust and rapid locomotion, making them more suitable for deployment in challenging environments such as mountains, forests, tunnels, and caves for exploration and search and rescue missions. However, unlike structured urban environments, quadruped robots operating in harsh environments such as the wild and confined underground spaces require the assistance of visual information to navigate uneven terrain and three-dimensional obstacles above and below ground. They also demonstrate mobility even when their field of vision is obstructed or when their movement direction exceeds their field of vision.
[0003] Traditional motion control methods for quadruped robots, typified by optimization-based model predictive control, offer acceptable stability for walking tasks on flat terrain. However, when faced with highly uneven terrain and complex obstacles, traditional control methods rely on offline planning using pre-generated elevation maps. Consequently, when faced with the inaccurate mapping often encountered in unstructured outdoor environments, traditional control methods are severely degraded in robustness and struggle to adaptively select dynamic, agile behaviors to navigate unstructured terrain.
[0004] Data-driven approaches based on reinforcement learning are currently a popular approach in the field of quadruped robot motion control. Inspired by how animals learn robust motor skills, reinforcement learning methods bypass the complex modeling process and train a neural network in a simulated environment built for a specific task. The goal is to maximize the return of a set reward function, learning a motion strategy that meets the task requirements with high disturbance tolerance and high dynamic performance, and ultimately deploying it on a physical robot. Nahrendra et al. proposed a motion control learning framework that relies solely on the proprioception of a quadruped robot. By building an action network to implicitly infer terrain properties such as terrain height and friction, the robot can adaptively adjust its gait on rough terrain to achieve stable terrain navigation. Fan Yong et al. (Publication No.: CN118192254A) proposed a teacher-student network trained using robot state observations and privileged information. By using a state estimation network module that predicts linear velocity based on proprioception observations and an incremental terrain partitioning training method, they achieved motion control for a physical quadruped robot in various complex terrains. Zhao Xiaoqing et al. (Publication No. CN118012077A) proposed a two-stage reinforcement learning method that incorporates motion imitation learning. This method improves the efficiency of learning a quadruped robot's motion strategy through joint tracking rewards. However, most of these control methods focus solely on the quadruped robot's body state and fail to utilize visual feedback, such as from depth cameras. This makes it difficult for the quadruped robot to navigate obstacles of comparable size or even higher.
[0005] In addition, Agarwal et al. proposed a teacher-student network learning method that uses robot-centered depth camera visual information. The network is guided by privileged information such as elevation maps to learn the terrain perception required for the movement of quadruped robots from depth maps with unclear features, thus achieving movement skills in complex terrains such as going up and down stairs and passing through cracks. Zhang Wei et al. (publication number: CN113867333A) proposed a method for quadruped robots to climb stairs based on visual perception. The geometric properties of the stairs are obtained through three-dimensional point clouds, and the robot's reference speed, landing position, target posture and climbing trajectory are planned accordingly. However, these methods only achieve forward motion with the assistance of visual information, ignoring the characteristics of the quadruped robot's omnidirectional and flexible movement, and are unable to cope with harsh field environment conditions where visual perception is disturbed or even the field of view is completely blocked.
[0006] In summary, current quadruped robot motion control methods have achieved good traversability and stability on both flat and rough terrain. However, these methods lack the ability to construct an efficient end-to-end training framework, and thus are unable to achieve omnidirectional motion skills in a three-dimensional environment using only limited visual perception. Summary of the Invention
[0007] In view of the fact that existing quadruped robot motion control methods are limited by the robot's narrow forward field of view and external sensory noise, making it difficult to obtain omnidirectional motion capabilities in a three-dimensional environment, the present invention proposes a three-dimensional omnidirectional motion control method for a quadruped robot based on hybrid representation learning, especially for generating motion skills of a quadruped robot through an environment containing three-dimensional obstacles when the robot's field of view is blocked or the motion direction is inconsistent with the field of view.
[0008] The present invention provides a method for controlling the omnidirectional motion of a quadruped robot in a three-dimensional environment based on hybrid representation learning, which specifically includes the following steps:
[0009] S1. Design an overall reinforcement learning framework with an asymmetric action network-policy network structure, and design the input and output content of each part of the neural network in the reinforcement learning overall framework; the neural network includes a standard input feature extraction encoder, a privileged information input feature extraction encoder, an action network, and a value network;
[0010] S2. Based on the overall framework of reinforcement learning in step S1, a hybrid representation learning training process is designed; the hybrid representation learning includes supervised learning, contrastive learning and reinforcement learning methods;
[0011] S3. Based on the training process in step S2, a neural network model is built for the standard input feature extraction encoder, the privileged information input feature extraction encoder, the action network, and the value network to implement the reasoning and gradient update functions required for training;
[0012] S4. Build a simulation training environment and design a training terrain and reward function suitable for a quadruped robot to learn omnidirectional locomotion in a three-dimensional environment under different conditions.
[0013] S5. Based on the overall reinforcement learning framework of step S1, the training process of step S2, the neural network model of step S3, and the simulation training environment of step S4, reinforcement learning training is performed to obtain an omnidirectional motion control strategy for a three-dimensional environment suitable for limited visual perception, and the strategy is migrated and deployed on a quadruped robot.
[0014] Furthermore, the step S1 includes the following sub-steps:
[0015] S11. Design an overall reinforcement learning framework, in which the standard input feature extraction encoder is used to process the standard input directly obtained from the physical quadruped robot and output a hidden layer vector as the extracted feature for the action network to generate the next action; the privileged information input feature extraction encoder uses the input obtained in the simulation to provide an encoding of the surrounding environment; the value network receives the privileged information input, evaluates the results output by the action network, and guides and optimizes the action network;
[0016] S12. Based on the internal and external sensors of the quadruped robot, design the standard input content of the overall reinforcement learning framework serving the motion control task, including the quadruped robot's proprioception and depth vision image;
[0017] S13. Based on the standard input content of step S12, design privileged information input content that can be obtained in the simulation environment, which is used to represent the local information of the three-dimensional environment and the state of the robot body, including body privileged observation and privileged external perception;
[0018] S14. Design the output content of the overall reinforcement learning framework, including the action network output for quadruped robot motion control and the value network output for guiding policy gradient updates.
[0019] Furthermore, step S2 includes the following sub-steps:
[0020] S21. Design state estimation learning based on supervised learning methods: Construct an encoder-decoder structure before the action network, where the encoder accepts standard input and extracts features, reducing the input dimension to a hidden layer vector; the decoder then decodes the hidden layer vector into various state estimates that meet the requirements of the motion control task, calculates the mean squared error between the state estimates obtained from the simulation, and then implements gradient propagation to guide the update of the standard input encoder network;
[0021] S22. Design a terrain understanding learning algorithm based on contrastive learning: Build an encoder before the action network to accept privileged input and extract features, reducing the input dimension to a hidden layer vector. Then, based on contrastive learning, calculate the contrast error between the hidden layer vector and the contrast vector in the output of the standard input feature extraction encoder, thereby implementing gradient propagation to guide the update of the standard input feature extraction encoder network and the privileged information input feature extraction encoder network.
[0022] S23. Design motion control learning based on reinforcement learning method: Optimize the reinforcement learning algorithm through proximal strategy, calculate the value function loss, alternative strategy loss and entropy loss, and their sum constitutes PPO loss, thereby realizing gradient propagation and guiding the update of action network and value network.
[0023] Furthermore, step S3 includes the following sub-steps:
[0024] S31. Build a neural network model for a standard input feature extraction encoder: The standard input feature extraction encoder uses a self-attention mechanism to fuse feature extraction from the quadruped robot's proprioception and depth vision images, and then feeds its output into a gated recurrent unit to generate a hidden layer vector.
[0025] S32. Build a neural network model for a privileged information input feature extraction encoder: The privileged information input feature extraction encoder uses a cross-attention mechanism to extract features from the privileged ontology observation and privileged external perception obtained in the simulation, and obtains privileged external perception features extracted with the assistance of privileged ontology perception.
[0026] S33. Build an action network model and a value network model with asymmetric input: The action network extracts the hidden layer vector output by the encoder based on the standard input and the standard input feature, and outputs the expected position of each joint of the quadruped robot; the value network outputs the value of the current state based on the privileged information input.
[0027] Furthermore, in step S4, the training terrain includes: learning to effectively use vision to pass through high platforms, gullies and stair environments with challenging obstacles in the front, learning rugged terrain where vision is disturbed and blocked, learning to crawl through floating obstacle terrain with obstacles placed above the robot, and learning to pass through comprehensive three-dimensional environment terrain in the form of combined obstacles in all directions.
[0028] Furthermore, in step S4, the reward function mainly includes a task reward function item and a stability reward function item; wherein, the task reward function item is used to satisfy the motion control task instructions; the stability reward function item is used to keep the quadruped robot's body posture stable during the movement process and constrain the amplitude of the action output by the strategy.
[0029] Furthermore, in step S5, the motion control strategy obtained through training is exported into a format executable by the quadruped robot and deployed on the quadruped robot for actual machine testing.
[0030] Compared with the prior art, the method of the present invention has the following beneficial effects: (1) The one-stage end-to-end training framework including hybrid representation learning proposed by the present invention enables the quadruped robot to fully understand the local environment information with limited visual perception due to the efficient and comprehensive input encoding method, thereby being able to achieve omnidirectional mobility on a variety of challenging terrains; (2) The multimodal feature fusion extraction method with the introduction of an asymmetric attention mechanism proposed by the present invention has good environmental adaptability. When the movement direction is consistent with the field of view, the quadruped robot can accept visual guidance and demonstrate excellent obstacle passing ability. When the movement direction is inconsistent with the field of view or the field of view is blocked, the quadruped robot can overcome the interference caused by vision and complete the movement task; (3) The present invention is theoretically applicable to various quadruped robot platforms, such as Jueying Lite3, Jueying X30, ANYmal C, Spot, Cheetah3, Go1, etc., and is adaptable to various common motion control tasks, such as jumping onto a high platform, jumping over a ravine, climbing stairs, passing through a small hole, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is a flow chart of the present invention for describing a method for omnidirectional control of a quadruped robot in a three-dimensional environment based on hybrid representation learning;
[0032] Figure 2 This is a diagram of the overall framework structure of reinforcement learning in step S1 according to an embodiment of the present invention;
[0033] Figure 3 is a schematic diagram for describing visual perception in standard input of a quadruped robot according to an embodiment of the present invention;
[0034] Figure 4 is a schematic diagram for describing privileged external perception in privileged information input of a quadruped robot according to an embodiment of the present invention;
[0035] Figure 5 Schematic diagram of the neural network framework of each module in step S3 according to an embodiment of the present invention;
[0036] Figure 6 is a schematic diagram of an embodiment of the present invention used to describe a quadruped robot passing through various obstacles with good vision;
[0037] Figure 7 This is a schematic diagram of an embodiment of the present invention used to describe a quadruped robot passing through an obstacle when its field of vision is blocked and its movement direction is inconsistent with its field of vision. DETAILED DESCRIPTION
[0038] The present invention will be described in detail below based on the accompanying drawings and preferred embodiments. The purpose and effects of the present invention will become more apparent. The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining the present invention and are not intended to limit the present invention.
[0039] This paper proposes a method for controlling the omnidirectional motion of a quadruped robot in a three-dimensional environment based on hybrid representation learning. Figure 1 As shown, the specific steps include:
[0040] S1. Design the overall framework of reinforcement learning with an asymmetric action network-strategy network structure, and design the input and output content of each part of the neural network in the framework.
[0041] S11. Model the reinforcement learning training environment as an infinite-horizon partially observable Markov decision process (POMDP) and design an overall reinforcement learning framework. During training, the neural network framework is divided into two asymmetric input components: an action network pathway that primarily accepts standard inputs and a policy network pathway that accepts privileged information inputs. This stimulates the implicit imagination of local environmental information when learning movement strategies.
[0042] In the embodiment of the present invention, Figure 2 As shown in Figure 1, the overall reinforcement learning framework consists of four main components: a standard input feature extraction encoder, a privileged information input feature extraction encoder, an action network, and a value network. The standard input feature extraction encoder processes standard input directly from the physical quadruped robot and outputs a hidden layer vector as the extracted features for the action network to generate actions. The privileged information input feature extraction encoder utilizes inputs that can only be obtained in simulation to provide a more comprehensive encoding of the surrounding environment. Through the interaction between this encoder and the standard input feature extraction encoder, the hidden layer vector output by the standard input feature extraction encoder contains sufficient understanding of the local environment around the quadruped robot to guide the action decision-making process. The value network also receives privileged information input and uses more comprehensive and realistic input to evaluate the output of the action network, guiding the action network to optimize in a more optimal direction.
[0043] S12. Based on the internal and external sensors carried by the quadruped robot, design standard input content in the overall reinforcement learning framework serving the motion control task, including the quadruped robot's proprioception and depth vision images.
[0044] In the embodiment of the present invention, the standard input consists of two parts: the quadruped robot body observation o t and deep visual image d t . Where the ontological observation is defined as:
[0045]
[0046] This includes the robot body angular velocity ω t , the direction of gravity g in the robot coordinate system t , Linear velocity instruction on the xy plane in the robot coordinate system Z-axis angular velocity command in the world coordinate system The position of each joint of the robot θ t , the speed of each joint of the robot and the joint position instruction a issued in the previous frame t-1In order to retain transient memory, the proprioception of the quadruped robot in the past H frames is stacked. In the embodiment of the present invention, the stacked proprioception As the internal perception part in the standard input. Deep visual image d t The single frame image captured by the front depth camera on the quadruped robot is acquired by the simulated depth camera placed in the same position as the real object in the simulation environment, such as Figure 3 shown.
[0047] S13. Based on the standard input content described in step S12, design privileged information input content that can be obtained in the simulation environment to further characterize the local information of the three-dimensional environment and the state of the robot body, including body privileged observation and privileged external perception.
[0048] In the embodiment of the present invention, the privileged information also consists of two parts: the privileged observation s of the quadruped robot body and the privileged observation s of the quadruped robot body. t and privileged external perception m t . Where privileged ontology observation is defined as:
[0049] s t =[o t ,v t ] T #(2)
[0050] This includes standard ontological observations o t and the body velocity v in the robot coordinate system t , the latter is privileged information that can only be obtained in simulation. In omnidirectional motion tasks in a three-dimensional environment, it is difficult to fully describe the terrain characteristics by relying solely on the elevation map. Since the 2.5D elevation map information is incomplete and the computation and storage overhead of the 3D voxel map is large, the embodiment of the present invention proposes a new privileged external perception m t , which provides a lightweight and efficient three-dimensional terrain representation, defined as:
[0051]
[0052] Among them, c t Represents a cubic depth map, which consists of five depth images centered on the quadruped robot facing different directions, facing forward above below Left and right This cubic depth map sampling method captures depth information more uniformly across space than the equirectangular projection obtained by lidar-like sampling, which tends to be oversampled at the poles and undersampled at the equator. Contains the sparse depth information around the robot’s four legs. Similarly, mt It can only be obtained in a simulation environment, such as Figure 4 The above setting of privileged information input is intended to guide the standard input feature extraction encoder to better extract environmental information through rich and efficient 3D terrain representation.
[0053] S14. Design the output content of the overall reinforcement learning framework, including the action network output for quadruped robot motion control and the value network output for guiding policy gradient updates.
[0054] In the embodiment of the present invention, the output of the overall reinforcement learning framework can be mainly regarded as two parts: the action network output and the value network output. The action network outputs a vector of size 12×1, denoted as a t , represents the desired position of each joint of the quadruped robot. In order to improve learning efficiency, the embodiment of the present invention will a t Let the joint positions θ be the same as those of the quadruped robot when it is standing still. stand The relative value of , that is, the final expected position of each joint of the quadruped robot is defined as:
[0055] θ desired =θ stand +a t #(5)
[0056] The desired position of each joint of the quadruped robot is tracked by a proportional-derivative (PD) controller, which is converted into the desired torque of each joint of the quadruped robot. The value network outputs an estimate of the value of the current state V(s t ).
[0057] S2. Design a hybrid representation learning training process to obtain an overall training method that includes supervised learning, contrastive learning, and reinforcement learning methods.
[0058] S21. Design a state estimation learning model based on supervised learning. Construct an encoder-decoder structure before the action network. The encoder accepts standard input and extracts features, reducing the input dimension to a hidden layer vector. The decoder then decodes this hidden layer vector into individual state estimates that meet the requirements of the motion control task. The mean squared error (MSE) is calculated between the estimated state and the true state obtained in the simulation. Gradient propagation can then be performed to guide the update of the standard input encoder network.
[0059] In the embodiment of the present invention, the state estimation learning based on the supervised learning method focuses on the features that are relatively easy to extract. In the standard input feature extraction encoder, the standard input is encoded into a hidden layer vector, which is composed of the quadruped robot body velocity estimation value. Hidden state estimation vector of quadruped robot Latent vector estimation of forward-looking depth map for quadruped robots and contrastive learning of latent vectors The features in the hidden layer vector can contain the true value of the estimated value e t Specifically, the body velocity estimate in the hidden layer vector Directly return to the true value v t , the hidden state estimation vector The variational self-encoder structure is used, and the decoder is reconstructed into And return to the true value o t+1 In order to enhance the robot's intuitive understanding of depth vision, another forward-looking depth map estimates the latent vector Decoding into forward depth map from cube depth map Then with the truth value Calculate the regression error. The above regression loss L supervise Calculated by Mean-Square Error (MSE), the formula is:
[0060]
[0061] Among them, D KL represents the Kullback-Leibler (KL) divergence used in the variational autoencoder to calculate the similarity between two distributions, q represents the posterior distribution of the hidden state estimation vector, and is determined by the input with d t It is calculated that p represents the prior distribution of the hidden state vector. The embodiment of the present invention adopts the standard normal distribution as the prior distribution of the hidden state vector.
[0062] S22. Design a terrain understanding learning algorithm based on contrastive learning. Independent of the standard input feature extraction encoder, construct an encoder before the action network to accept the privileged input and extract features. Similarly, the input is reduced to a hidden layer vector. Then, based on contrastive learning, the contrast error between this hidden layer vector and the contrast vector in the output of the standard input feature extraction encoder is calculated. Gradient propagation can then be performed to guide the updates of the standard input feature extraction encoder network and the privileged input feature extraction encoder network.
[0063] Although the supervised learning method can make the encoding result of the standard input contain actual clear meaning, the supervised learning method alone will lose the function of feature extraction when there is no direct connection between the input and the estimated value. In the embodiment of the present invention, the terrain understanding learning based on the contrastive learning method focuses on establishing an abstract connection between the standard input and the privileged information input of the quadruped robot, aiming to establish a certain understanding of the state of the quadruped robot and the local environment through limited and noisy proprioception and visual observation. Specifically, the contrastive learning hidden vector in the hidden layer vector of the standard input encoding result is Encoding result with privileged information input The shared characteristics between standard input and privileged information input should be captured. To achieve this, and Transformed by a predictor, the two output vectors are recorded as and The embodiment of the present invention achieves the contrastive learning optimization goal by minimizing their negative cosine similarity, and the formula is:
[0064]
[0065] Among them, L contrast It represents the loss function used in the contrastive learning method, and stopgrad represents the gradient truncation operator, which aims to prevent the occurrence of collapse in the contrastive learning process.
[0066] S23. Design motion control learning based on reinforcement learning. Implement the Proximal Policy Optimization (PPO) reinforcement learning algorithm, calculate the value function loss, alternative policy loss and entropy loss, and the sum of them constitutes the PPO loss L ppo , and then gradient propagation can be performed to guide the update of the action network and the value network.
[0067] In summary, the comprehensive loss function L based on hybrid representation learning used in the embodiment of the present invention is total Defined as:
[0068] L total =L supervise +L contrast +L ppo #(8)
[0069] S3. Build a neural network model for each module to implement the inference and gradient update functions required for training.
[0070] S31. Build a neural network model for a standard input feature extraction encoder. This encoder uses a self-attention mechanism to fuse features from the quadruped robot's proprioception and depth vision imagery. It also uses a recurrent neural network (RNN) to maintain long-term memory for limited observations, preserving understanding of the local environment.
[0071] In an embodiment of the present invention, the standard input feature extraction encoder uses a multilayer perceptron (MLP) to process the observation history To obtain proprioceptive features, and use Convolutional Neural Network (CNN) from the deep visual image d t These proprioceptive and exteroceptive features are then concatenated and processed through a Transformer encoder with a self-attention mechanism to improve the efficiency of extracting fused features from cross-modal inputs. To preserve the quadruped's long-term memory of the input, the output of the Transformer encoder is fed into a Gated Recurrent Unit (GRU) to generate a hidden layer vector.
[0072] S32. Build a neural network model for a privileged information input feature extraction encoder. Based on the privileged ontology observations and privileged external perceptions obtained in the simulation, the privileged information input feature extraction encoder uses a cross-attention mechanism to extract features from the two modalities, obtaining privileged external perception features extracted with the assistance of privileged ontology perception.
[0073] In the embodiment of the present invention, the privileged information input feature extraction encoder is designed to have a network structure similar to the standard input feature extraction encoder to introduce an inductive bias to output similar feature encodings in contrastive learning. However, a key difference between the two is that a cross-attention mechanism is introduced in the privileged information input feature extraction encoder to focus on the environmental information features when processing multimodal inputs. Specifically, the query input q is obtained by transforming the privileged ontology perception observation s into t The key input k and value input v are calculated by projecting them into a multilayer perceptron, and the key input k and value input v are obtained by transforming the privileged visual observation m t The projected visual embedding (extracted by a convolutional neural network (CNN)) is obtained. In the subsequent skip connection and normalization (Add&Norm) module, only the visual embedding participates in the skip connection. This mechanism is designed to ensure that the privileged information input feature extraction encoder focuses on its visual input, while the privileged ontological perception observations provide support as auxiliary information through the attention mechanism.
[0074] S33. Build an action network model and a value network model with asymmetric input. The action network extracts the hidden layer vectors output by the encoder based on the standard input and standard input features, and outputs the desired position of each joint of the quadruped robot. The value network outputs the value of the current state based on the privileged information input.
[0075] In the embodiment of the present invention, the action network is based on ontology perception. t The hidden layer vector output by the standard input feature extraction encoder is used as input, and a 12-dimensional action vector a is output through a multi-layer perceptron. tThe value network uses multi-layer perceptron and convolutional neural network to respectively t and privileged external perception m t Perform projection processing and then use multilayer perceptron to calculate the state value V(s t ).
[0076] The neural network framework and training method used in the embodiments of the present invention are as follows Figure 5 shown.
[0077] S4. Build a simulation training environment and design a training terrain and reward function suitable for a quadruped robot to learn omnidirectional movement skills in a three-dimensional environment under different conditions.
[0078] S41. In an embodiment of the present invention, the Isaac Gym simulator is used to train the entire neural network framework. In order to develop a comprehensive strategy, an embodiment of the present invention designs a series of simulation environments, covering the learning of skills such as long jump, crawling, climbing stairs, and omnidirectional walking on discrete terrain. Based on the omnidirectional motion control task of a quadruped robot in a three-dimensional environment, four simulation training terrains are constructed, including: learning to effectively use vision to pass through high platforms, gullies and stair environments with challenging obstacles in the front; learning rugged terrain where vision is disturbed and blocked; learning to crawl through floating obstacle terrain placed above the robot; learning to pass through a comprehensive three-dimensional environment terrain in the form of combined obstacles in an omnidirectional manner. In addition, the embodiment of the present invention introduces regular noise into the depth image during the simulation process to simulate the noise that may be generated in the visual output due to possible field of view occlusion of the quadruped robot in actual scenarios. In each environment, the embodiment of the present invention introduces terrain courses with gradually increasing difficulty to achieve a smooth motor skill learning process.
[0079] S42. Based on the simulation training environment described in step S41, a reasonable reward function is designed to motivate the quadruped robot to learn to follow the speed instruction through the combination of three-dimensional obstacles.
[0080] In this embodiment of the present invention, a simple yet efficient reward function is designed, primarily comprising a task reward function term and a stability reward function term. The former is designed to satisfy the motion control task instructions, specifically tracking the linear velocity in the xy plane and the z-axis angular velocity in the quadruped robot's coordinate system. The latter is designed to maintain the quadruped robot's body posture stability during movement and constrain the amplitude of the policy's output movements, thereby achieving better physical robot deployment results.
[0081] S5. Complete reinforcement learning training to obtain an omnidirectional motion control strategy for traversing a three-dimensional environment that can adapt to limited visual perception situations, such as when the field of view is blocked or the movement direction is inconsistent with the field of view, and migrate and deploy it on a quadruped robot.
[0082] S51. In an embodiment of the present invention, based on the training framework and simulation environment described above, reinforcement learning training is carried out to obtain a motion control strategy that meets the task requirements.
[0083] S52. In this embodiment of the present invention, the trained strategy is exported into a format that can be run by a quadruped robot and deployed on the quadruped robot for real-machine testing to verify the method's ability to migrate from zero-sample simulation to reality. The robot used in this embodiment of the present invention is the quadruped robot "Jueying Lite3", which has a body height of 0.35m when standing, weighs 11kg, and has 12 degrees of freedom. Figure 6 As shown in the figure, in forward motion under good visual conditions, the strategy implemented by the embodiment of the present invention can enable the quadruped robot to jump onto a 0.6m high platform, cross a 0.9m long ravine, climb a 0.25m high continuous staircase, and drill through a 0.2m high hole, demonstrating the extremely superior obstacle-navigation motion control capability using visual information. Figure 7 As shown, in omnidirectional movement, the strategy implemented by the embodiment of the present invention enables the quadruped robot to climb continuous stairs 0.15m high and pass through a low hole 0.25m high, and can quickly respond to visual interference through proprioception feedback such as collision to overcome the problem of difficulty in effectively obtaining visual information.
[0084] Those skilled in the art will understand that the foregoing descriptions are merely preferred embodiments of the invention and are not intended to limit the invention. Although the invention has been described in detail with reference to the foregoing examples, those skilled in the art will still be able to modify the technical solutions described in the foregoing examples or substitute equivalents for some of the technical features therein. Any modifications, equivalent substitutions, etc. made within the spirit and principles of the invention shall be included within the scope of protection of the invention.
Claims
1. A method for omnidirectional motion control of a quadruped robot in a three-dimensional environment based on hybrid representation learning, characterized in that: The following steps are involved: S1. Design an overall reinforcement learning framework with an asymmetric action network-policy network structure, and design the input and output content of each part of the neural network in the reinforcement learning overall framework; the neural network includes a standard input feature extraction encoder, a privileged information input feature extraction encoder, an action network, and a value network; S2. Based on the overall reinforcement learning framework in step S1, design a hybrid representation learning training process; The hybrid representation learning includes supervised learning, contrastive learning and reinforcement learning methods; S21. Design state estimation learning based on supervised learning methods: Construct an encoder-decoder structure before the action network, where the encoder accepts standard input and extracts features, reducing the input dimension to a hidden layer vector; the decoder then decodes the hidden layer vector into various state estimates that meet the requirements of the motion control task, calculates the mean squared error between the state estimates obtained from the simulation, and then implements gradient propagation to guide the update of the standard input encoder network; S22. Design a terrain understanding learning algorithm based on contrastive learning: Build an encoder before the action network to accept privileged input and extract features, reducing the input dimension to a hidden layer vector. Then, based on contrastive learning, calculate the contrast error between the hidden layer vector and the contrast vector in the output of the standard input feature extraction encoder, thereby implementing gradient propagation to guide the update of the standard input feature extraction encoder network and the privileged information input feature extraction encoder network. S23. Design motion control learning based on reinforcement learning: Optimize the reinforcement learning algorithm through proximal policy, calculate the value function loss, alternative policy loss and entropy loss, and their sum constitutes the PPO loss, thereby realizing gradient propagation and guiding the update of the action network and value network; S3. Based on the training process in step S2, a neural network model is built for the standard input feature extraction encoder, the privileged information input feature extraction encoder, the action network, and the value network to implement the reasoning and gradient update functions required for training; S4. Build a simulation training environment and design a training terrain and reward function suitable for a quadruped robot to learn omnidirectional locomotion in a three-dimensional environment under different conditions. S5. Based on the overall reinforcement learning framework of step S1, the training process of step S2, the neural network model of step S3, and the simulation training environment of step S4, reinforcement learning training is performed to obtain an omnidirectional motion control strategy for a three-dimensional environment suitable for limited visual perception, and the strategy is migrated and deployed on a quadruped robot.
2. The method for controlling omnidirectional motion of a quadruped robot in a three-dimensional environment based on hybrid representation learning according to claim 1, characterized in that: The step S1 includes the following sub-steps: S11. Design an overall reinforcement learning framework, in which the standard input feature extraction encoder is used to process the standard input directly obtained from the physical quadruped robot and output a hidden layer vector as the extracted feature for the action network to generate the next action; the privileged information input feature extraction encoder uses the input obtained in the simulation to provide an encoding of the surrounding environment; the value network receives the privileged information input, evaluates the results output by the action network, and guides and optimizes the action network; S12. Based on the internal and external sensors of the quadruped robot, design the standard input content of the overall reinforcement learning framework serving the motion control task, including the quadruped robot's proprioception and depth vision image; S13. Based on the standard input content of step S12, design privileged information input content that can be obtained in the simulation environment, which is used to represent the local information of the three-dimensional environment and the state of the robot body, including body privileged observation and privileged external perception; S14. Design the output content of the overall reinforcement learning framework, including the action network output for quadruped robot motion control and the value network output for guiding policy gradient updates.
3. The method for controlling omnidirectional motion of a quadruped robot in a three-dimensional environment based on hybrid representation learning according to claim 1, characterized in that: The step S3 includes the following sub-steps: S31. Build a neural network model for a standard input feature extraction encoder: The standard input feature extraction encoder uses a self-attention mechanism to fuse feature extraction from the quadruped robot's proprioception and depth vision images, and then feeds its output into a gated recurrent unit to generate a hidden layer vector. S32. Build a neural network model for a privileged information input feature extraction encoder: The privileged information input feature extraction encoder uses a cross-attention mechanism to extract features from the privileged ontology observation and privileged external perception obtained in the simulation, and obtains privileged external perception features extracted with the assistance of privileged ontology perception. S33. Build an action network model and a value network model with asymmetric input: The action network extracts the hidden layer vector output by the encoder based on the standard input and the standard input feature, and outputs the expected position of each joint of the quadruped robot; the value network outputs the value of the current state based on the privileged information input.
4. The method for controlling omnidirectional motion of a quadruped robot in a three-dimensional environment based on hybrid representation learning according to claim 1, characterized in that: In step S4, the training terrain includes: learning to effectively use vision to pass through high platforms, gullies and stair environments with challenging obstacles in the front, learning rugged terrain where vision is disturbed and blocked, learning to crawl through floating obstacle terrain with obstacles placed above the robot, and learning to pass through comprehensive three-dimensional environment terrain in the form of combined obstacles in all directions.
5. The method for controlling omnidirectional motion of a quadruped robot in a three-dimensional environment based on hybrid representation learning according to claim 1, characterized in that: In step S4, the reward function mainly includes a task reward function item and a stability reward function item; wherein, the task reward function item is used to satisfy the motion control task instructions; the stability reward function item is used to keep the quadruped robot's body posture stable during the movement process and constrain the amplitude of the action output by the strategy.
6. The method for controlling omnidirectional motion of a quadruped robot in a three-dimensional environment based on hybrid representation learning according to claim 1, characterized in that: In step S5, the motion control strategy obtained through training is exported into a format executable by the quadruped robot and deployed on the quadruped robot for actual machine testing.
Citation Information
Patent Citations
Quadruped robot stair climbing planning method based on visual perception and application thereof
CN113867333A
Quadruped robot motion control method and system based on reinforcement learning action simulation
CN118012077A
Motion control method and system for quadruped robot under terrain subareas
CN118192254A
Method and system for adjusting dynamic bandwidth of communication between core particles
CN117155792A
Quadruped robot visual motion control method based on forward kinematics
CN118838388A