Humanoid robot motion control method based on sequence feature processing

By combining local timing feature extraction and global motion trend analysis methods, the problems of high calculation overhead, poor real-time performance and insufficient stability in complex environments are solved, and higher control accuracy and stability are achieved.

CN120215366APending Publication Date: 2025-06-27BEIJING UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510351992.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-27

Smart Images

  • Figure CN120215366A_ABST
    Figure CN120215366A_ABST
Patent Text Reader

Abstract

The invention discloses a humanoid robot motion control method based on sequence feature processing, and the method comprises the steps: firstly, collecting motion data of a robot in different task scenes, and carrying out the preprocessing, so as to improve the data quality; then, a dynamic updating mechanism is adopted to extract local motion features, and the capturing capacity of short-time motion changes is enhanced; the motion trend of the robot is analyzed through an autoregression method, and the continuity and stability of a control instruction are improved. Based on the processing result, a self-adaptive control strategy is generated, so that the robot can optimize motion parameters in real time according to environment changes. The method shows high gait stability, energy consumption efficiency and motion adaptability in a simulation environment and an actual robot platform, and the motion control level of the humanoid robot in a complex task scene can be effectively improved. The method is suitable for various humanoid robot motion control tasks, and is especially suitable for an intelligent robot system with high stability and real-time performance requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robot motion control, and particularly to a humanoid robot motion control method based on sequential data processing. By optimizing the modeling of robot motion data, the control accuracy and motion stability are improved. Background Art

[0002] The motion control of humanoid robots in complex environments is a key technology. Traditional control methods mainly include optimization control based on mathematical models, feedback control based on experience, and reinforcement learning methods. Among them, the optimization control method can achieve good results under precise modeling conditions, but it is limited by computational complexity and difficult to cope with dynamic environmental changes. The feedback control method has good real-time performance, but its adaptability to unknown environments is limited. The reinforcement learning method has made certain progress in recent years, but it has a high dependence on a large amount of training data and is difficult to ensure stability.

[0003] In the motion control of humanoid robots, motion sequence data contains rich time-related information, such as joint angles, joint velocities, contact torques, etc. How to efficiently utilize this information to improve the continuity and stability of motion decisions is the focus and difficulty of current research. The present invention proposes a method that combines local sequential feature extraction and global motion trend analysis to enhance the reliability and accuracy of humanoid robot motion control. Summary of the Invention

[0004] The present invention aims to provide a humanoid robot motion control model based on sequence feature processing. By optimizing the modeling of robot motion data, its motion stability, control accuracy, and adaptability in different environments are improved. This method combines local sequential feature extraction and global motion trend analysis to solve the problems of large computational overhead, poor real-time performance, and insufficient stability existing in traditional motion control methods when dealing with complex environments, enabling the robot to execute complex motion tasks, such as walking, running, jumping, obstacle avoidance, etc., more precisely.

[0005] The technical solution adopted by the present invention is a humanoid robot motion control method based on sequence feature processing, including the following steps:

[0006] S1, perform feature embedding processing on the motion data of the humanoid robot, map the state, action, reward, and expected return to a hidden space of the same dimension respectively to form feature vectors;

[0007] S2, perform local sequential feature extraction on the feature vectors through a gating mechanism, and use the update gate and reset gate to dynamically adjust the degree of dependence on historical information to extract local sequential information;

[0008] S3. Input the features after local temporal feature extraction into the global dependency modeling module based on Transformer, and use the autoregressive feature to model the global dependency relationship;

[0009] S4. Refine the predicted features of the output features of the Transformer module, optimize the predicted features through the parameter sharing strategy, and finally generate the action control instructions;

[0010] S5. Drive the humanoid robot to execute corresponding actions according to the generated action control instructions.

[0011] Further, the gating mechanism includes:

[0012] An update gate, which is used to control the retention degree of the current hidden state to historical information. By dynamically adjusting the retention ratio, the ability to capture the features of time continuity is enhanced;

[0013] A reset gate, which is used to selectively forget historical information, enabling the model to ignore irrelevant past information and focus on the current input features;

[0014] The calculation of the candidate hidden state combines the current input and the adjusted historical information, captures the local temporal information through a non-linear activation function, and provides a balanced feature of the history and the current input.

[0015] Further, the Transformer module adopts an autoregressive encoder structure, and ensures that the current time step only depends on the previous time series through a masking mechanism to avoid future information leakage. Specifically, it includes:

[0016] a. Use the self-attention mechanism to calculate the relationship weights between all time steps in the sequence, and capture the long-distance sequence information relationship;

[0017] b. Mask the unobserved future time steps to ensure that the output of the current time step is based on historical information, and achieve the logical coherence of dynamic generation.

[0018] Further, in the predicted feature refinement step, the parameter sharing strategy is adopted, and the weight matrix in the feature extraction stage is inherited to the feature refinement stage. Specifically, it includes:

[0019] a. Share the weight matrix in the gating mechanism in the local feature extraction and decoding stages, including the input weight matrix and the hidden state weight matrix;

[0020] b. Reduce the number of model parameters through parameter sharing, reduce the computational complexity, enhance the regularization effect of the model, and improve the prediction accuracy and generalization ability.

[0021] Furthermore, in the feature embedding step, it also includes enhancing the processing of time features, and integrating time dimension information into the feature vector through time embedding, specifically including:

[0022] a. Append time embedding vectors to the input data at each time step to enhance the model's perception ability of time series;

[0023] b. The time embedding vectors are generated through linear transformation and added element-wise to the feature vectors to achieve the fusion of time information and motion features.

[0024] Furthermore, the method is applicable to various humanoid robot motion tasks, including jumping, walking, and running tasks, and can dynamically adjust the action control strategy according to different task requirements.

[0025] Furthermore, the method is optimized through an end-to-end training method, using the action error loss function to minimize the error between the model-predicted action and the actual executed action, and improving the accuracy of the control task.

[0026] A humanoid robot motion control system includes:

[0027] a. A data acquisition module, which is used to collect the state, action, and reward data of the humanoid robot, including joint angle, angular velocity, position, and speed information;

[0028] b. A feature processing module, which is used to execute the motion control model described in any one of claims 1 to 7;

[0029] c. A control execution module, which is used to drive the humanoid robot to execute corresponding actions according to the generated action control instructions, including converting the action instructions into joint torques or motor control signals.

[0030] Furthermore, the data acquisition module includes a sensor unit, which is used to obtain the motion state information of the humanoid robot in real time and transmit the data to the feature processing module for processing.

[0031] The control execution module includes a driving unit, which is used to convert the action control instructions into specific joint torques or motor control signals to drive the humanoid robot to complete the motion task.

[0032] The system also includes a feedback module, which is used to monitor the motion effect of the humanoid robot in real time and use the feedback information for dynamic adjustment and optimization of the model.

[0033] Furthermore, the implementation of the humanoid robot motion control system includes the following steps:

[0034] a. Collect the state, action, and reward data of the humanoid robot;

[0035] b. Generate action control instructions using the described motion control method;

[0036] c. Drive the humanoid robot to perform corresponding actions according to the action control instructions;

[0037] d. Monitor the motion effect in real time, and dynamically adjust the model parameters according to the feedback information to optimize the motion control strategy.

[0038] In the step of collecting the state, actions, and reward data of the humanoid robot, it also includes preprocessing the data to remove noise and outliers to improve the input quality.

[0039] In the step of dynamically adjusting the model parameters, use the gradient descent method to update the model weights according to the feedback error to improve adaptability and stability.

[0040] The present invention proposes a robot motion control model composed of four key parts: data collection, local temporal modeling, global motion prediction, and action decision optimization. The core process is as follows:

[0041] Collect the robot's motion parameters, including joint angles, joint angular velocities, joint torques, sole forces, accelerations, and angular velocities, etc.;

[0042] Filter and denoise the collected robot motion parameters to remove environmental interference and improve the signal quality;

[0043] Standardize the numerical values of different types of robot motion parameters to ensure the consistency of data input and avoid calculation errors;

[0044] Organize the data at fixed time intervals (such as 5 ms or 10 ms) so that the control system can use past motion information for decision-making;

[0045] Record the joint motion data for a period of time;

[0046] Calculate the change rate of each joint at the current moment relative to the previous moment;

[0047] Combine historical motion data to determine whether the current motion state needs to be adjusted;

[0048] Optimize the joint angle and torque distribution according to the calculation results to ensure smooth motion.

[0049] Compared with the prior art, the present invention significantly improves the accuracy, stability, and adaptability of humanoid robot motion control by combining local temporal feature extraction and global sequence modeling. Compared with the prior art, the present invention has the following remarkable technical effects, and these effects are achieved through the technical principle of the present invention.

[0050] The present invention dynamically adjusts the degree of dependence on historical information through a gating mechanism (update gate and reset gate). Combining with the global dependence modeling module of Transformer, it can capture both local temporal features and global motion trends simultaneously. This combination enables the model to more accurately predict and control the motion state of the robot.

[0051] When dealing with complex environments, traditional methods often struggle to balance global and local features, resulting in insufficient control accuracy and stability. Through the combination of local temporal feature extraction and global sequence modeling, the present invention significantly improves the accuracy and stability of motion control, especially showing stronger robustness in dynamic environments.

[0052] The present invention adopts a gating mechanism and a parameter sharing strategy, reducing the number of model parameters and the computational complexity. At the same time, the self-attention mechanism of Transformer can process sequence data in parallel, significantly improving the computational efficiency.

[0053] Due to the sequential processing characteristics, traditional methods have low computational efficiency and are difficult to handle high-dynamic tasks. The present invention significantly reduces the computational complexity and improves the real-time performance through parallel processing and parameter sharing, enabling quick response to environmental changes.

[0054] The present invention reduces the number of model parameters and the computational complexity through a parameter sharing strategy, while enhancing the regularization effect of the model.

[0055] Deep reinforcement learning often requires a large number of parameters to capture complex motion laws, resulting in high computational complexity and easy overfitting. The parameter sharing strategy of the present invention significantly reduces the number of model parameters, enhances the regularization effect, and improves the stability and generalization ability of the model. Brief Description of the Drawings

[0056] Figure 1 It is a flowchart of the method for the walking task of a humanoid robot according to an embodiment of the present application; Detailed Embodiments

[0057] The present invention will be described in detail below with reference to the drawings and embodiments.

[0058] Walking task of a humanoid robot:

[0059] A humanoid robot with multi-degree-of-freedom joints is used, and the experimental environment includes flat ground, slopes, irregular terrains, etc. to simulate the complexity in real scenarios.

[0060] The humanoid robot is run on different terrains, and joint angles, angular velocities, positions, velocity state information, executed actions, and corresponding reward signals are collected through sensors. The collected data is stored as a serialized state-action-reward dataset.

[0061] Normalize the collected data, scale the value ranges of states and actions to the range [-1, 1] to improve the convergence speed of model training.

[0062] Construct sequences of quadruples of expected returns, states, actions, and rewards, and divide them into training sets and test sets at a ratio of 8:2.

[0063] Input the preprocessed sequence data, and extract local temporal features through a gated recurrent unit (GRU). The hidden state dimension of the GRU is set to 128, and the weights of the update gate and the reset gate are automatically learned through training.

[0064] The output of this module serves as the input to the global sequence modeling module.

[0065] Use a Transformer-based decoder architecture, set 3 layers of self-attention mechanisms, with each layer containing 1 attention head. The hidden layer dimension is 128, and the dimension of the feed-forward network is 512.

[0066] Implement autoregressive modeling through the masked self-attention mechanism to ensure that the prediction at the current time step depends only on previous time steps.

[0067] Refine the global features output by the Transformer, and adopt a parameter sharing strategy to share the weight matrices W and U.

[0068] Map the refined features to the action space through a linear layer to generate predicted actions.

[0069] Use the mean squared error MSE as the loss function to optimize the difference between the predicted actions and the true actions of the model.

[0070] Adopt the AdamW optimizer, set the learning rate to 1e-4, and the training batch size to 64.

[0071] During the training process, the model performs 100 iterations on the training set, with each iteration containing 1000 training steps.

[0072] Build a test scenario similar to the training environment in the MuJoCo simulation platform, including flat ground, slopes, and irregular terrains.

[0073] In each test scenario, run the trained model to control the humanoid robot to complete the walking task.

[0074] Collect metrics such as the walking speed, balance stability, and number of falls of the robot to evaluate the model performance.

[0075] On flat terrain, the average walking speed of the humanoid robot reaches 1.2 m / s, and the number of falls is 0.

[0076] On sloping terrain, the walking speed is 1.0 m / s and the number of falls is 1 time per 100 steps.

[0077] On irregular terrain, the walking speed is 0.8 m / s and the number of falls is 3 times per 100 steps.

[0078] Compared with traditional methods, the model of the present invention shows significant advantages in both walking speed and number of falls.

[0079] Humanoid robot running task:

[0080] Select a humanoid robot with high dynamic performance and build a test environment for the running task in the simulation platform, including a straight runway, curves, and obstacles.

[0081] Run the humanoid robot in the simulation environment and collect the state information of joint angles, angular velocities, and body postures, the executed actions, and the corresponding reward signals during the running process.

[0082] Normalize the collected data and scale the value ranges of the states and actions to the range [-1, 1].

[0083] Construct a quadruple sequence of expected return, state, action, and reward, and divide it into a training set and a test set with a ratio of 8:2.

[0084] Input the preprocessed sequence data and extract local temporal features through GRU. The hidden state dimension is set to 128, and the weights of the update gate and the reset gate are automatically learned through training.

[0085] Use a Transformer-based decoder architecture, set 3 layers of self-attention mechanisms, and each layer contains 1 attention head. The hidden layer dimension is 128, and the dimension of the feed-forward network is 512.

[0086] Implement autoregressive modeling through the masked self-attention mechanism.

[0087] Refine the global features output by the Transformer, adopt a parameter sharing strategy, and share the weight matrices W and U.

[0088] Map the refined features to the action space through a linear layer to generate predicted actions.

[0089] Use the mean squared error MSE as the loss function to optimize the difference between the model's predicted actions and the true actions.

[0090] Adopt the AdamW optimizer, set the learning rate to 1e-4, and the training batch size to 64.

[0091] During the training process, the model performs 100 iterations on the training set, and each iteration contains 1000 training steps.

[0092] Build a test scenario similar to the training environment in the MuJoCo simulation platform, including a straight runway, a curve, and obstacles.

[0093] In each test scenario, run the trained model to control the humanoid robot to complete the running task.

[0094] Collect indicators such as the running speed, energy consumption, and number of falls of the robot to evaluate the model performance.

[0095] On the straight runway, the average running speed of the humanoid robot reaches 3.5 m / s, and the energy consumption is reduced by 20%.

[0096] On the curve, the running speed is 3.2 m / s, and the number of falls is 0.

[0097] In the obstacle scenario, the running speed is 2.8 m / s, and the number of falls is 2 times per 100 steps.

[0098] Compared with the traditional method, the model of the present invention shows significant advantages in both running speed and energy consumption.

[0099] Humanoid robot jumping task:

[0100] Select a humanoid robot with high dynamic performance and build a test environment for the jumping task in the simulation platform, including platforms and obstacles of different heights.

[0101] Run the humanoid robot in the simulation environment and collect the state information of joint angles, angular velocities, and body postures, the executed actions, and the corresponding reward signals during the jumping process.

[0102] Normalize the collected data and scale the value ranges of the states and actions to the range of [-1, 1].

[0103] Construct a quadruple sequence of expected return, state, action, and reward, and divide it into a training set and a test set with a ratio of 8:2.

[0104] Input the preprocessed sequence data and extract local temporal features through GRU. The hidden state dimension is set to 128, and the weights of the update gate and the reset gate are automatically learned through training.

[0105] Use a decoder architecture based on Transformer, set 3 layers of self-attention mechanisms, and each layer contains 1 attention head. The hidden layer dimension is 128, and the dimension of the feed-forward network is 512.

[0106] Implement autoregressive modeling through the masked self-attention mechanism.

[0107] Refine the global features output by the Transformer, adopt a parameter sharing strategy, and share the weight matrices W and U.

[0108] Map the refined features to the action space through a linear layer to generate predicted actions.

[0109] Use the mean squared error (MSE) as the loss function to optimize the difference between the model's predicted actions and the true actions.

[0110] Adopt the AdamW optimizer, set the learning rate to 1e-4, and the training batch size to 64.

[0111] During the training process, the model performs 100 iterations on the training set, and each iteration contains 1000 training steps.

[0112] Build a test scenario similar to the training environment in the MuJoCo simulation platform, including platforms and obstacles at different heights.

[0113] In each test scenario, run the trained model to control the humanoid robot to complete the jumping task.

[0114] Collect indicators such as the jumping height, landing stability, and energy consumption of the robot to evaluate the model performance.

[0115] On the low-height platform (0.5 meters), the average jumping height of the humanoid robot reaches 0.6 meters, and the landing stability is 95%.

[0116] On the high-height platform (1.0 meter), the jumping height reaches 1.1 meters, and the landing stability is 90%.

[0117] In the obstacle scenario, the jumping height is 0.8 meters, and the landing stability is 85%.

[0118] Compared with traditional methods, the model of the present invention shows significant advantages in both jumping height and landing stability.

[0119] The present invention provides a technical solution for the motion control of humanoid robots by combining local temporal feature extraction and global sequence modeling.

[0120] It should be noted that in the specification provided here, a large number of specific details are described. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known structures and technologies are not shown in detail so as not to obscure the understanding of this specification.

[0121] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising said element.

[0122] The above are only examples of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A humanoid robot motion control method based on sequence feature processing, characterized in that: The following steps are involved: S1, feature embedding processing is performed on the motion data of the humanoid robot, and the state, action, reward and expected return are mapped to the latent space of the same dimension to form a feature vector; S2, extract local time series features from feature vectors through a gating mechanism, dynamically adjust the degree of dependence on historical information using update gates and reset gates, and extract local time series information; S3, input the features extracted from local time series features into the Transformer-based global dependency modeling module, and use the autoregressive characteristics to model the global dependency relationship; S4, refine the prediction features of the output features of the Transformer module, optimize the prediction features through parameter sharing strategy, and finally generate action control instructions; S5, driving the humanoid robot to perform corresponding actions according to the generated action control instructions.

2. The humanoid robot motion control method based on sequence feature processing according to claim 1 is characterized in that: The gating mechanism includes: The update gate is used to control the degree of retention of historical information in the current hidden state. By dynamically adjusting the retention ratio, the ability to capture time continuity features is enhanced. The reset gate is used to selectively forget historical information, ignoring irrelevant past information and focusing on current input features; The calculation of candidate hidden states combines the current input and adjusted historical information, captures local temporal information through a nonlinear activation function, and provides balanced features for historical and current inputs.

3. The humanoid robot motion control method based on sequence feature processing according to claim 1 is characterized in that: The Transformer module adopts an autoregressive encoder structure and uses a masking mechanism to ensure that the current time step only depends on the previous time series to avoid future information leakage, including: a. Use the self-attention mechanism to calculate the relationship weights between all time steps in the sequence and capture long-distance sequence information relationships; b. Masking is performed on unobserved future time steps to ensure that the output of the current time step is based on historical information and to achieve logical coherence of dynamic generation.

4. The humanoid robot motion control method based on sequence feature processing according to claim 1 is characterized in that: In the prediction feature refinement step, a parameter sharing strategy is adopted to inherit the weight matrix of the feature extraction stage to the feature refinement stage, which specifically includes: a. The weight matrix in the shared gating mechanism during the local feature extraction and decoding stages includes the input weight matrix and the hidden state weight matrix; b. Reduce the number of parameters through parameter sharing, reduce computational complexity, enhance regularization effect, and improve prediction accuracy and generalization ability.

5. The humanoid robot motion control method based on sequence feature processing according to claim 1 is characterized in that: The feature embedding step also includes an enhancement process of the time feature, in which the time dimension information is integrated into the feature vector through time embedding, specifically including: a. Add a time embedding vector to the input data of each time step to enhance the perception of time series; b. The temporal embedding vector is generated through linear transformation and added to the feature vector element by element to achieve the fusion of temporal information and motion features.

6. The humanoid robot motion control method based on sequence feature processing according to claim 1 is characterized in that: The method is applicable to a variety of humanoid robot motion tasks, including jumping, walking and running tasks, and can dynamically adjust the motion control strategy according to different task requirements.

7. The humanoid robot motion control method based on sequence feature processing according to claim 1 is characterized in that: The method is optimized through an end-to-end training method, and uses the action error loss function to minimize the error between the predicted action and the actual executed action, thereby improving the accuracy of the control task.

8. A humanoid robot motion control system implementing the method as claimed in any one of claims 1 to 7, characterized in that: include: a. Data acquisition module, used to collect the status, action and reward data of the humanoid robot, including joint angle, angular velocity, position and speed information; b. a feature processing module for executing the motion control method according to any one of claims 1 to 7; c. A control execution module, used to drive the humanoid robot to perform corresponding actions according to the generated action control instructions, including converting the action instructions into joint torques or motor control signals.

9. The humanoid robot motion control system according to claim 8, characterized in that: The data acquisition module includes a sensor unit, which is used to obtain the motion state information of the humanoid robot in real time and transmit the data to the feature processing module for processing. The control execution module includes a driving unit, which is used to convert the motion control instructions into specific joint torques or motor control signals to drive the humanoid robot to complete the motion task.

10. The humanoid robot motion control system according to claim 8, characterized in that: The system also includes a feedback module for real-time monitoring of the motion effect of the humanoid robot and using the feedback information for dynamic adjustment and optimization.

Citation Information

Cited By

  • Industrial production method and production system based on humanoid robot and humanoid robot

    CN121411344A