A humanoid robot self-supervised state estimation method based on variational autoencoder
By constructing a neural network state estimation model based on variational autoencoders and utilizing self-supervised learning methods, the problem of insufficient state estimation accuracy of quadruped robots in complex environments was solved, achieving high-precision state perception and real-time control.
Patent Information
- Application Number
- CN202511309213.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-09-15
AI Technical Summary
Existing technologies for state estimation in quadrupedal or humanoid robots suffer from decreased accuracy, especially during rapid movement or in complex terrain, and are highly dependent on contact states, making it difficult to achieve high-precision state perception.
A self-supervised state estimation method for humanoid robots based on variational autoencoders is adopted. By constructing a state estimation model of a neural network, the robot's own sensor information and historical observation data are used to perform self-supervised learning to predict the robot's speed, pose and foot contact state.
It improves the accuracy and robustness of state estimation, enables high-precision state perception in complex environments, requires no external sensor support, has closed-loop deployment capability, and supports real-time control of robots in dynamic environments.
Smart Images

Figure CN120821200B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of robot perception and control, and particularly relates to a humanoid robot self-supervised state estimation method based on a variational autoencoder. BACKGROUND
[0002] In the research of legged robots such as quadruped robots or humanoid robots, state estimation technology is a key link to realize stable motion control, path planning and perception fusion. This technology aims to estimate the position, velocity and foot contact state of the robot in the world coordinate system in real time through limited body sensor information, so as to provide reliable feedback information for the motion control module. In the traditional implementation, it is usually necessary to construct the dynamics and kinematics model of the robot and combine IMU, joint speed, position and other multi-source information, and use recursive estimation algorithms such as Extended Kalman Filter (EKF) for fusion to obtain the optimal estimation of the system state variable.
[0003] Although such methods can achieve robust full-state estimation without relying on external positioning systems (such as GPS or VICON), such methods still have problems such as decreased accuracy under fast motion or complex terrain and strong dependence on contact state. Therefore, it is urgent to further develop more robust and accurate state estimation methods to improve the state perception ability in unstructured environments. SUMMARY
[0004] In view of the above defects or improvement needs of the prior art, the present application provides a humanoid robot self-supervised state estimation method based on a variational autoencoder, thereby solving the basic state estimation problem in the model-based control method of legged robots.
[0005] To achieve the above purpose, the embodiment provides the following scheme:
[0006] A humanoid robot self-supervised state estimation method based on a variational autoencoder, comprising the following steps:
[0007] Defining the local coordinate system and the global coordinate system of the robot; wherein the point between the two feet of the robot is set as the coordinate origin, the forward direction of the robot is set as the positive direction of the x-axis, the left direction is set as the positive direction of the y-axis, and the vertical upward direction is set as the positive direction of the z-axis. The xy-axis coordinates of the initial standing position of the robot are set to 0, and the rotation direction follows the right-hand rule;
[0008] Based on the defined local coordinate system and global coordinate system, a state estimation model based on a neural network is constructed; the state estimation model comprises a state encoding module, a state estimation module and a state decoding module;
[0009] state estimation of the robot is completed by using the state estimation model.
[0010] Preferably, after the state estimation model is constructed, motion data of the robot under several environmental conditions is collected to form a robot motion data set; the robot motion data set includes robot observation information and robot privileged information; wherein the robot observation information includes speed instruction information for controlling the motion of the robot, absolute position and speed of each joint of the robot, angular velocity and attitude of the floating base; the robot privileged information includes yaw angle of the floating base attitude, position of the floating base, speed of the floating base, and contact state of the robot foot.
[0011] Preferably, the step of completing the state estimation of the robot by using the state estimation model comprises:
[0012] inputting a historical observation sequence of the robot at the previous h time into the state estimation model;
[0013] extracting features in the historical observation sequence by using the state encoding module, and mapping the input data to a latent vector;
[0014] dividing the latent vector into explicit variables and implicit variables;
[0015] performing state estimation and observation data reconstruction based on the explicit variables and the implicit variables respectively, to realize prediction of the floating base speed, position, attitude and foot contact state of the robot.
[0016] Preferably, the step of collecting motion data of the robot under several environmental conditions to form a robot motion data set comprises: collecting the motion process of a high-precision physical simulator or a real robot; when the data is collected from a real robot, the privileged information of the robot is obtained by means of external measuring equipment.
[0017] Preferably, during the training process, the explicit variable part is supervised to learn the real privileged information; for continuous variables, mean square error is used to construct the state estimation loss, and for discrete variables, cross-entropy loss is used to construct the state estimation loss.
[0018] Preferably, after the state estimation of the robot is completed, the predicted robot state information is input into a motion control model; the motion control model generates control instruction signals according to the current state estimation and control targets, to control the precise motion control of each actuator of the robot; the robot body executes specific actions according to the control signals, and continuously generates new observation information by sensors during the execution process for re-inputting the state estimation model, to complete the closed-loop control process.
[0019] Preferably, the step of constructing the state estimation loss comprises:
[0020] For the real continuous privilege state , the estimated value The loss of the continuous variable is:
[0021] ,
[0022] Where, , , Real speed, position and yaw angle of the floating base are represented by , , Network-predicted corresponding states are represented by
[0023] Preferably, for the real state of the i-th foot end of the robot The estimated state The cross-entropy loss is:
[0024] ,
[0025] Where n represents the number of foot ends of the robot.
[0026] Compared with the prior art, the present application has the following beneficial effects:
[0027] 1. Robust control scenario against delay and observation noise: Since the state estimation makes full use of historical information, it can alleviate the state disturbance caused by sensor delay or noise.
[0028] 2. Improve state estimation accuracy: By encoding the historical observation data through a neural network, it is converted into latent space features, which significantly improves the state perception ability of the robot in complex working conditions. The explicit variable and the privilege information construct a supervised loss term, so that the model can accurately learn the state mapping relationship during training.
[0029] 3. Decouple prediction and reconstruction tasks: The explicit variables in the latent space vector output by the encoding module are used for state estimation, and the implicit variables are used for observation reconstruction, so that the state estimation and reconstruction tasks do not interfere with each other, improving the stability and generalization ability of the model.
[0030] 4. High-precision state estimation without external sensors: In the training phase, the privilege information in the real or simulation data is combined, and in the prediction phase, only the observation information provided by the robot's own sensors is needed to achieve comprehensive estimation of the floating base speed, position, yaw angle and foot end contact.
[0031] 5. Closed-loop deployment capability: The model can be directly integrated into the robot's motion control framework as a state estimation module to provide real-time feedback, support high-precision controllers based on state feedback, and achieve autonomous and robust humanoid robot motion planning. Attached Figure Description
[0032] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0033] Figure 1 This is a schematic diagram illustrating the relationship between several main objects in the robot control system of this invention.
[0034] Figure 2 This is a schematic diagram illustrating the definition of the robot coordinate system in an embodiment of the present invention;
[0035] Figure 3 This is a schematic diagram of the model composition of a self-supervised state estimation method for a humanoid robot based on a variational autoencoder in an embodiment of the present invention;
[0036] Figure 4 This is a schematic diagram of the robot motion data acquisition process in an embodiment of the present invention;
[0037] Figure 5 This is a schematic diagram of the robot motion dataset data structure in an embodiment of the present invention;
[0038] Figure 6 This is a model training block diagram of a self-supervised state estimation method for humanoid robots based on variational autoencoders in an embodiment of the present invention;
[0039] Figure 7 This is a flowchart of the reasoning application in an embodiment of the present invention.
[0040] Explanation of reference numerals in the attached figures:
[0041] 1. State estimation model; 2. Robot motion dataset; 3. Motion control model; 4. Robot body; 11. State encoding module; 12. State estimation module; 13. State decoding module; 21. Robot observation information; 22. Robot privileged information. Detailed Implementation
[0042] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0043] In order to make the above objectives, characteristics and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0044] Embodiment:
[0045] The core of this embodiment is to construct an encoder-decoder architecture based on neural network, and train it through self-supervised learning method, so that the final state estimation model can predict accurate robot state information from historical observation data, realize comprehensive and accurate estimation and prediction of robot speed, pose and foot end contact state, and provide basic data for subsequent robot motion control. In order to better describe the method provided by the present application, first, the robot control system is briefly introduced and the local and global coordinate systems of the robot are defined.
[0046] As shown in Figure 1 , the control system of the robot in the embodiment of the present application is simplified and described as a state estimation model 1, a motion control model 3 and a robot body 4, and there is data interaction between the several main objects. The state estimation model 1 is trained based on the robot motion data set, and the trained model provides state feedback for the motion control model 3 based on the robot's own sensor information, and the motion control model controls the robot body 4 to move according to the feedback information, and the data generated during the movement of the robot is also used as the input of the state estimation model 1, forming a closed-loop system.
[0047] S1. Define the local and global coordinate systems of the robot.
[0048] In this embodiment, the definition of the direction of the robot world coordinate system is as shown in Figure 2 , when the robot stands on the ground with both feet, the point between the two feet is set as the coordinate origin, the direction of the robot moving forward is set as the positive direction of the x-axis, the left direction is set as the positive direction of the y-axis, and the vertical upward direction is set as the positive direction of the z-axis. The xy-axis coordinates of the initial standing position of the robot are set to 0, and the zero position of each joint is defined according to the direction of this coordinate system, and the rotation direction follows the right-hand rule. Therefore, the position and pose of the robot in the world coordinate system and the joint position can be uniquely and accurately represented, which provides a reference for accurate description and calculation of the state of the robot.
[0049] S2. Based on the defined local and global coordinate systems, construct a state estimation model based on neural network.
[0050] As shown in Figure 3As shown, the constructed state estimation model mainly consists of a state encoding module 11, a state estimation module 12, and a state decoding module 13. The state encoding module 11 and the state decoding module 13 are each a neural network, such as a multilayer perceptron (MLP), a long short-term memory (LSTM) network, or a gated recurrent unit (GRU), depending on the type of state estimation. The core function of the state encoding module 11 is to utilize the powerful feature extraction capability of neural networks to transform the robot's historical observation data sequence into a latent feature representation, which serves as the information input to the state decoding module 13. The state estimation module 12 is further subdivided into several sub-modules: a floating base velocity estimation module predicts the linear velocity of the robot's floating base in the world coordinate system; a floating base position estimation module predicts the position and height changes (global position) of the robot's floating base relative to its initial state on the plane; and a foot contact estimation module predicts whether the robot's feet are in contact with the ground. The state decoding module 13 takes the feature information extracted into the latent space by the state encoding module 11 as input and aims to predict the same data as input to the state encoding module 11.
[0051] like Figure 4 As shown, training the state estimation model requires first constructing an accurate and rich dataset, and the figure illustrates the data collection process. Data can be collected from high-precision physical simulators such as Mujoco and Webots, or directly from the motion of real robots. If data is collected from a real robot, privileged information that cannot be obtained from the robot's own sensors requires external measurement equipment, such as motion capture devices, to capture the robot's global pose and velocity. During data collection, to obtain rich, diverse, and representative data, it is necessary to collect the robot's motion data under various environmental conditions, such as different terrains (flat ground, uneven surfaces, slopes, stairs, etc.) and surfaces with different friction levels, simulating complex scenarios the robot might encounter in real-world applications. Simultaneously, the robot's speed is randomly changed during its motion, and random push / pull operations and loads are applied to increase the diversity and complexity of the data, collecting relevant data about the robot system during this process. This comprehensive data collection method yields a large amount of data that fully reflects the robot's motion characteristics under different working conditions, providing sufficient and high-quality data support for subsequent model training.
[0052] like Figure 5As shown, the robot motion dataset 2 required for training the state estimation model in this embodiment includes robot observation information 21 and robot privileged information 22. The robot observation information 21 refers to data that can be directly obtained from the robot's own sensors, including speed command information (linear speed v x , linear speed v y , angular speed w yaw ) for controlling the robot's motion, absolute position and speed of each joint of the robot, angular speed and pose of the floating base. Among them, the joint position and joint speed are obtained by the robot joint encoder, and the floating base angular speed and pose are obtained by the IMU (inertial sensor), but due to the poor numerical accuracy of the yaw direction of the IMU, there will be a large cumulative error as the running time increases, so it cannot be used in actual motion control, and needs to be estimated and calculated by using the state estimation model.
[0053] The robot privileged information 22 refers to information that cannot be directly obtained from the robot's own sensors and needs to be estimated by using the state estimation model, including the yaw angle of the floating base pose, the position of the floating base (where the position in the xy direction is zero at the initial position of the robot, and the zero position in the z direction is the ground), the speed of the floating base, and the contact state of the robot foot, which is 1 for contact and 0 for non-contact. Unlike other estimated information, the contact state information is discrete data.
[0054] As shown in Figure 6 , the model training overall architecture of this embodiment is composed of a state encoding module 11 and a state decoding module 13, and is trained in an end-to-end manner through a self-supervised manner without relying on additional manually labeled data.
[0055] The input data is composed of the historical observation sequence of the robot at the previous h time points , which is input into the state encoding module 11, which is composed of a neural network (such as LSTM, GRU, MLP, etc.) and is used to extract features in the observation sequence. The state encoding module 11 maps the input data to a latent vector , where the latent vector is divided into two parts: explicit variable and implicit variable , that is, where the explicit variable corresponds to the privileged information of the robot at the current time, such as the position, speed, pose of the floating base, and the foot contact state, etc., as an explicit supervised estimation target; while the implicit variable is used to model the implicit time series features that cannot be directly observed but are helpful for reconstructing the observation data. In the training process, the explicit variable part will be compared with the true privileged information The supervised learning is performed, and the state estimation loss is constructed in the form of mean squared error (MSE) for continuous variables and binary cross-entropy (BCE) for discrete variables.
[0056] Let the real continuous privilege state be , and the estimated value be , then the loss of the continuous variable is:
[0057] ,
[0058] wherein, , , respectively represent the real speed, position and yaw angle of the floating base; , , respectively represent the corresponding states predicted by the network.
[0059] Let the real state of the i-th foot end of the robot be , and the estimated state be , then the cross-entropy loss is:
[0060] ,
[0061] wherein, n=2 represents the number of foot ends of the robot. At the same time, the entire latent variable vector (including the explicit and implicit parts) is input into the state decoding module 13 to reconstruct the next observation of the robot .
[0062] S3. The state estimation of the robot is completed by using the state estimation module.
[0063] As shown in Figure 7 , the model inference process of the humanoid robot self-supervised state estimation method based on the variational autoencoder provided by the embodiment of the application is constructed on the information transmission mechanism in the actual operation control loop of the robot. The inference architecture includes three core modules: a state estimation model 1, a motion control model 3 and a robot body 4, and a closed-loop feedback control mechanism is formed between the modules.
[0064] During the inference process, the robot collects observation information in the time period through its own sensor system The state estimation model 1 is fed into the state estimation model 1, which predicts the three-dimensional linear velocity, spatial position change (especially the xy position and z height), the attitude angle of the yaw direction of the floating base, and the contact state (contact or non-contact) of the robot foot end and the ground. The estimation result is then input into the motion control model 3, which generates control instruction signals according to the current state estimation and control targets (such as speed instructions, gait patterns, etc.), and controls the robot actuators to achieve precise motion control. This module can be based on classical control (such as MPC, QP) or learning-based policy controller (such as reinforcement learning policy network). Finally, the robot body 4 executes specific actions according to the control signal, and continuously generates new observation information for re-input into the state estimation model during the execution process, completing the closed-loop control process.
[0065] Through this reasoning structure design, the state estimation model plays a core role in the entire control loop, enabling the robot to accurately infer its complete state using only partial observable sensor information, providing stable and reliable data support for subsequent high-performance and robust motion control. The reasoning process can also be deployed online, supporting real-time state perception and control decision-making of the robot in complex and dynamic environments.
[0066] The above-described embodiments are only descriptions of the preferred modes of the present application and do not limit the scope of the present application. Various modifications and improvements to the technical solutions of the present application made by those of ordinary skill in the art without departing from the design spirit of the present application shall fall within the protection scope of the present application as defined by the claims.
Claims
1. A method for self-supervised state estimation of humanoid robots based on variational autoencoder, characterized by the steps of The method comprises the following steps: Defining a local coordinate system and a global coordinate system of the robot; wherein a point between the two feet of the robot is set as the coordinate origin, the direction in which the robot moves forward is set as the positive direction of the x-axis, the left direction is set as the positive direction of the y-axis, and the vertical upward direction is set as the positive direction of the z-axis, the xy-axis coordinates of the initial standing position of the robot are set as 0, and the rotation direction follows the right-hand rule; Based on the defined local coordinate system and global coordinate system, a state estimation model based on a neural network is constructed; the state estimation model comprises a state encoding module, a state estimation module and a state decoding module; The state estimation of the robot is completed by using the state estimation model, and the steps comprise: Inputting the historical observation sequence of the robot at the previous h time into the state estimation model; Using the state encoding module to extract the features in the historical observation sequence and mapping the input data to a latent vector; Dividing the latent vector into explicit variables and implicit variables; Based on the explicit variables and the implicit variables, state estimation and observation data reconstruction are respectively performed to realize the prediction of the floating base velocity, position, attitude and foot end contact state of the robot.
2. The variational autoencoder-based humanoid robot self-supervised state estimation method according to claim 1, wherein, After the state estimation model is constructed, the motion data of the robot under several environmental conditions is collected to form a robot motion data set; the robot motion data set comprises robot observation information and robot privileged information; wherein the robot observation information comprises velocity command information for controlling the motion of the robot, the absolute position and velocity of each joint of the robot, the angular velocity and attitude of the floating base; the robot privileged information comprises the yaw angle of the floating base attitude, the position of the floating base, the velocity of the floating base and the contact state of the robot foot.
3. The variational autoencoder-based humanoid robot self-supervised state estimation method according to claim 2, wherein, The step of collecting the motion data of the robot under several environmental conditions to form a robot motion data set comprises: collecting the motion process of a high-precision physical simulator or a real robot; when the data is collected from a real robot, the robot privileged information is obtained with the aid of external measuring equipment.
4. The variational autoencoder-based humanoid robot self-supervised state estimation method according to claim 1, wherein, During the training process, the explicit variable part is supervised to learn the real privileged information; For continuous variables, mean square error is used to construct the state estimation loss, and for discrete variables, cross-entropy loss is used to construct the state estimation loss.
5. The variational autoencoder-based humanoid robot self-supervised state estimation method according to claim 1, wherein, After the state estimation of the robot is completed, the predicted robot state information is input into a motion control model; the motion control model generates control instruction signals according to the current state estimation and control targets to control the precise motion control of each actuator of the robot; the robot body executes specific actions according to the control signals, and continuously generates new observation information through sensors during the execution process for re-inputting the state estimation model to complete the closed-loop control process.
6. The variational autoencoder-based humanoid robot self-supervised state estimation method according to claim 4, wherein, The step of constructing the state estimation loss comprises: For real continuous privilege states , estimate , loss for continuous variables is: , wherein, , , respectively represent the true velocity, position and yaw angle of the floating base; , , respectively represent the corresponding states predicted by the network.
7. The variational autoencoder-based humanoid robot self-supervised state estimation method according to claim 6, wherein, For the real state of the robot's ith foot end contact , the estimated state , the cross-entropy loss is: , Wherein, n represents the number of foot ends of the robot.
Citation Information
Patent Citations
Motion control method for quadruped robot based on topographic map reinforcement learning
CN120503210A
Semi-Supervised Variational Autoencoder for Indoor Localization
US20210097387A1