A foot-type robot control method and device, electronic equipment and storage medium

By employing a closed-loop mechanism combining generative motion prior models, dynamic models, and residual reinforcement learning, the problem of deviation between planning and execution actions in existing legged robot control methods is solved, achieving high-precision, stable, and robust motion control.

CN122185241BActive Publication Date: 2026-07-14HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
Filing Date
2026-05-12
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing control methods for legged robots rely on pre-set dynamic models, which leads to discrepancies between planned actions and actual actions, affecting control accuracy.

Method used

By acquiring state vectors and morphological parameter encodings, candidate action sequences are generated using a generative motion prior model. A forward simulation and risk assessment are then performed using a dynamic model to select the optimal action sequence. The sequence is then corrected using a residual reinforcement learning model, and finally, an executable action sequence is generated using a projection constraint module.

Benefits of technology

It significantly improves the motion control accuracy, stability and robustness of legged robots in complex environments, reduces the impact of environmental disturbances, and ensures the physical feasibility of motion execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122185241B_ABST
    Figure CN122185241B_ABST
Patent Text Reader

Abstract

The application provides a kind of foot type robot control method, device, electronic equipment and storage medium, it is related to robot control technical field, the foot type robot control method includes: obtaining the state vector of foot type robot, form parameter coding and control condition, input training good generative motion prior model, generate multiple candidate action sequences;Again candidate action sequence is substituted into the dynamics model that is constructed in advance, corresponding prediction state is obtained;Then risk assessment is carried out in combination with candidate action sequence and prediction state, obtains each sequence evaluation value, selects the minimum as optimal action sequence;Then the action residual correction quantity of trained residual reinforcement learning model is output, and the optimal action sequence is corrected to obtain the modified action sequence;Finally, the robot executable action sequence is obtained by the preset projection constraint module projection processing.The application can effectively improve the efficiency of foot type robot control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot control technology, and more specifically, to a control method, device, electronic device, and storage medium for a legged robot. Background Technology

[0002] Legged robots, with their excellent adaptability to complex terrain and dynamic obstacle-crossing performance, have been widely used in industrial inspection, disaster relief, and service interaction scenarios, serving as core equipment for achieving autonomous movement in unstructured environments. Motion control technology, as the core key to legged robots, directly determines their gait stability, posture accuracy, and environmental robustness, playing a crucial role in ensuring efficient and safe robot operation and expanding the boundaries of practical applications.

[0003] In related technologies, existing legged robot control methods mostly rely on preset dynamic models. However, the preset dynamic models differ from the actual operating characteristics of the robot, resulting in deviations between planned actions and actual actions, which affects the control accuracy of the legged robot. Summary of the Invention

[0004] The problem addressed by this invention is how to improve the control accuracy of legged robots.

[0005] To address the above problems, the present invention provides a legged robot control method, device, electronic device, and storage medium.

[0006] In a first aspect, the present invention provides a control method for a legged robot, comprising:

[0007] Obtain the state vector, morphological parameter encoding, and control conditions of the legged robot;

[0008] The state vector, the morphological parameter encoding, and the control conditions are input into a trained generative motion prior model to obtain multiple candidate action sequences.

[0009] The candidate action sequence is input into a pre-built dynamic model to obtain the predicted state corresponding to the candidate action sequence;

[0010] Based on the candidate action sequence and the corresponding predicted state, an evaluation value corresponding to the candidate action sequence is obtained through risk assessment;

[0011] The candidate action sequence corresponding to the smallest evaluation value is determined as the optimal action sequence;

[0012] Based on the optimal action sequence, the corresponding action residual correction amount is obtained through the trained residual reinforcement learning model;

[0013] The optimal action sequence is corrected based on the action residual correction amount to obtain the corrected action sequence;

[0014] The executable action sequence of the legged robot is obtained by projecting the modified action sequence through a preset projection constraint module.

[0015] Optionally, the candidate motion sequence includes joint torque and joint angular velocity corresponding to each joint; the predicted state includes predicted motion velocity, zero torque point, capture point, foot contact force, foot position, and predicted pose; obtaining the evaluation value corresponding to the candidate motion sequence through risk assessment based on the candidate motion sequence and the corresponding predicted state includes:

[0016] Obtain the location of obstacles;

[0017] The speed error is obtained by the difference between the predicted speed and the preset expected speed.

[0018] The attitude stability margin is obtained through stability analysis based on the predicted pose, the zero torque point, the capture point, the foot contact force, and the foot position.

[0019] Energy consumption is obtained by integrating the product of the joint torque and the joint angular velocity corresponding to all joints.

[0020] The minimum geometric distance between the preset key point corresponding to the predicted pose and the position of the obstacle is determined as the collision distance;

[0021] The sliding risk index is obtained by determining the friction constraint based on all the foot contact forces and the corresponding foot positions.

[0022] The evaluation value is obtained based on the speed error, the attitude stability margin, the energy consumption, the collision distance, and the sliding risk index through a preset risk assessment relationship.

[0023] Optionally, the step of obtaining the attitude stability margin through stability analysis based on the predicted pose, the zero-moment point, the capture point, the foot contact force, and the foot position includes:

[0024] The foot end with a contact force greater than a preset contact threshold is defined as the contact foot end;

[0025] A supporting polygon is generated by using a convex hull based on the foot position corresponding to all the contact feet.

[0026] The attitude stability margin is obtained by determining the stability margin based on the predicted pose, the zero moment point, the capture point, and the support polygon.

[0027] Optionally, obtaining the attitude stability margin based on the predicted pose, the zero-moment point, the capture point, and the support polygon through stability margin determination includes:

[0028] The zero-moment point is projected onto the support plane corresponding to the support polygon to obtain the corresponding zero-moment point projection position;

[0029] The distance between the captured point and the projected position of the zero torque point is defined as the captured point position deviation;

[0030] The attitude tilt angle is obtained by attitude decomposition based on the predicted pose.

[0031] The minimum distance between the projection position of the zero moment point and the boundary of the supporting polygon is determined as the initial stable distance;

[0032] The attitude stability margin is obtained by weighted summation of the initial stable distance, the capture point position deviation, and the attitude tilt angle.

[0033] Optionally, the energy consumption satisfies:

[0034] ;

[0035] Where P is the energy consumption, [t0, t1] is the integration interval, and L i (t) represents the joint torque corresponding to the i-th joint at time t, W i (t) represents the joint angular velocity of the i-th joint at time t, and η i The preset motor efficiency parameter is the parameter corresponding to the i-th joint.

[0036] Optionally, the step of determining the sliding risk index based on all the foot contact forces and the corresponding foot positions through friction constraints includes:

[0037] Based on the foot contact force, the corresponding ground normal reaction force and ground tangential force at the foot are obtained through vector analysis.

[0038] The foot friction force corresponding to the foot is obtained by multiplying the ground normal reaction force and the preset friction coefficient.

[0039] The difference between the foot friction force and the corresponding ground tangential force is determined as the foot slip risk index corresponding to the foot.

[0040] The largest foot slip risk index is determined as the slip risk index.

[0041] Optionally, the risk assessment relationship satisfies:

[0042] ;

[0043] Where J is the evaluated value, E is the velocity error, M is the attitude stability margin, P is the energy consumption, D is the collision distance, S is the sliding risk index, U is the preset perception uncertainty, and ω E ω is the preset error weighting coefficient. M ω is the preset stable weight coefficient. P ω is the preset energy consumption weighting coefficient. D ω is the preset collision weight coefficient. S ω is the preset sliding weight coefficient. U These are the preset perception weight coefficients.

[0044] In a second aspect, the present invention provides a control device for a legged robot, comprising:

[0045] The acquisition module is used to acquire the state vector, morphological parameter encoding, and control conditions of the legged robot.

[0046] The generation module is used to input the state vector, the morphological parameter encoding, and the control conditions into a trained generative motion prior model to obtain multiple candidate action sequences;

[0047] The prediction module is used to input the candidate action sequence into a pre-built dynamic model to obtain the predicted state corresponding to the candidate action sequence;

[0048] The evaluation module is used to obtain the evaluation value corresponding to the candidate action sequence through risk assessment based on the candidate action sequence and the corresponding predicted state;

[0049] The filtering module is used to determine the candidate action sequence corresponding to the smallest evaluation value as the optimal action sequence;

[0050] The optimization module is used to obtain the corresponding action residual correction amount based on the optimal action sequence through a trained residual reinforcement learning model;

[0051] The correction module is used to correct the optimal action sequence according to the action residual correction amount to obtain a corrected action sequence;

[0052] The projection module is used to project the modified action sequence through a preset projection constraint module to obtain the executable action sequence of the legged robot.

[0053] Thirdly, the present invention provides an electronic device, including a memory and a processor;

[0054] The memory is used to store computer programs;

[0055] The processor is configured to implement the legged robot control method as described in the first aspect when executing the computer program.

[0056] Fourthly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the legged robot control method as described in the first aspect.

[0057] The beneficial effects of the legged robot control method, device, electronic device, and storage medium of the present invention are as follows: They acquire the state vector, morphological parameter encoding, and control conditions of the legged robot, providing comprehensive and accurate initial information support for subsequent model input; inputting the state vector, morphological parameter encoding, and control conditions into a trained generative motion prior model yields multiple candidate action sequences, providing rich and motion-compliant reference options for control decisions through multimodal action supply; inputting the candidate action sequences into a pre-constructed dynamic model yields the predicted states corresponding to the candidate action sequences, and obtaining the future execution effects of each action sequence through risk-free forward simulation, providing a reliable basis for quantitative evaluation; based on the candidate action sequences and their corresponding predicted states, risk assessment yields the evaluation values ​​corresponding to the candidate action sequences, realizing the evaluation of each action sequence in... Objective quantitative comparisons were conducted across dimensions such as stability, energy efficiency, and safety. The candidate action sequence corresponding to the lowest evaluation value was determined as the optimal action sequence, selecting a benchmark action that balances multi-objective performance from the source, providing a high-quality foundation for subsequent corrections. Based on the optimal action sequence, a corresponding action residual correction amount was obtained through a trained residual reinforcement learning model, offsetting errors caused by disturbances such as unknown friction, load changes, and perception biases with small-scale compensation. The optimal action sequence was then corrected based on the action residual correction amount to obtain a corrected action sequence, achieving dynamic adaptation to real-world working conditions while retaining the advantages of the benchmark action. The corrected action sequence was then projected through a preset projection constraint module to obtain the executable action sequence of the legged robot. Physical safety constraints strictly limited the actions within the feasible domain, avoiding problems such as overtravel and instability. Through a closed-loop mechanism of multimodal candidate supply, accurate simulation evaluation, optimal benchmark selection, adaptive residual correction, and safety constraint projection, control errors were progressively eliminated, the impact of environmental disturbances was reduced, and the physical feasibility of action execution was ensured, significantly improving the motion control accuracy, stability, and robustness of the legged robot in complex and unknown environments. Attached Figure Description

[0058] Figure 1 This is a flowchart illustrating a legged robot control method according to an embodiment of the present invention;

[0059] Figure 2 This is a schematic diagram of the structure of a legged robot control device according to an embodiment of the present invention;

[0060] Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0061] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the accompanying drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.

[0062] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.

[0063] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to"; the term "based on" means "at least partially based on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments"; and the term "optionally" means "optional embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first," "second," etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.

[0064] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0065] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0066] In related technologies, existing legged robot control methods mostly rely on preset dynamic models, but these models differ significantly from the robot's actual operating characteristics. In real-world scenarios, factors such as ground contact characteristics, load variations, foot slippage, and external disturbances can cause the robot's actual dynamic behavior to deviate from the theoretical model, making it impossible for the preset model to accurately reflect the real changes in body posture, foot force, and motion state. Under these circumstances, the action sequence planned based on the ideal model is difficult to adapt to real-time working conditions, directly leading to a significant deviation between the planned and actual actions, and a continuous accumulation of motion tracking errors. Simultaneously, model mismatch can also cause problems such as inaccurate posture prediction and distorted stability margin assessment, making the robot prone to posture jitter, gait imbalance, and foot positioning deviation during movement. This not only reduces motion stability and response speed but also directly affects the overall control accuracy of the legged robot, making it difficult to meet the requirements for high-precision, high-robust autonomous motion control in complex environments.

[0067] To address the problems existing in the aforementioned related technologies, embodiments of the present invention provide a legged robot control method, device, electronic device, and storage medium.

[0068] like Figure 1 As shown, an embodiment of the present invention provides a legged robot control method, comprising:

[0069] S110: Obtain the state vector, morphological parameter encoding, and control conditions of the legged robot.

[0070] Specifically, the state vector, morphological parameter encoding, and control conditions of the legged robot are obtained. The unified state vector includes the robot's body posture, joint state, contact event terms, externally perceived terrain and risk features, and equipment health diagnosis information. The morphological parameter encoding is a learnable vector obtained by normalizing the physical dimensions and topologically masking the morphological parameters such as the number of robot legs, link length, mass distribution, joint limits, rated torque, and foot size. The control conditions include task instruction information such as target speed, target orientation, gait style terms, risk preference, and constraint set.

[0071] S120, the state vector, the morphological parameter encoding, and the control conditions are input into the trained generative motion prior model to obtain multiple candidate action sequences.

[0072] Specifically, the unified state vector, morphological parameter encoding, and control conditions are input into a generative motion prior model that has been trained on offline data from multiple platforms and terrains. The model generates a specified number of future temporal action sequences based on the conditional distribution using a few-step sampling method, thereby obtaining multiple candidate action sequences that satisfy different temporal styles, stability margins, and energy consumption preferences, providing a basis for subsequent multi-target screening and benchmark action selection.

[0073] It should be noted that the generative motion prior model is a conditional generative model pre-trained on offline data from multiple platforms and terrains. Its core function is to output high-quality, multimodal baseline motion sequences for legged robots, serving as the fundamental motion supply module for the entire cross-morphological adaptive control system. Using a diffusion model, flow model, or autoregressive model as its network architecture, it has fully learned the motion patterns of different legged robots (quadrupedal, humanoid, and multi-legged) under various terrains during the training phase, mastering the conditional distribution patterns from state to morphology to command to motion sequences. During runtime, it receives three types of inputs: a unified state vector, morphological parameter encoding, and control conditions. Through low-step sampling and conditional guidance, it rapidly generates multiple candidate motion sequences. After multi-objective cost and uncertainty filtering, it outputs stable, low-risk baseline motions adapted to the current scenario, providing a safe and reliable initial reference for subsequent residual compensation. Simultaneously, it meets the robot's real-time control cycle requirements through rolling caching and frequency division inference.

[0074] S130, input the candidate action sequence into the pre-built dynamic model to obtain the predicted state corresponding to the candidate action sequence.

[0075] Specifically, multiple candidate action sequences output by the generative motion prior model are input one by one into a dynamic model pre-constructed based on the link parameters, mass distribution, contact constraints and environmental characteristics of the legged robot. Forward simulation calculations are used to obtain the predicted states of each candidate action sequence in the future control time domain, such as the base posture, velocity, joint state, foot contact condition and stability margin. This provides a quantitative basis for subsequent candidate action evaluation, multi-objective cost calculation and selection of the optimal benchmark action.

[0076] It should be noted that this pre-built dynamic model is a high-precision forward computation model established in advance based on the physical structure, mechanical characteristics, and environmental constraints of the legged robot. It is used to simulate and predict the robot's state after the execution of actions without actually driving the robot. Using the robot's link length, mass, moment of inertia, joint limits, and foot contact characteristics as core parameters, it integrates environmental and safety conditions such as ground friction, force cone constraints, zero-moment point (ZMP) and capture point (CP) stability criteria. It can perform forward dynamic calculations on any input action sequence, quickly outputting complete predicted states such as base posture, linear velocity, angular velocity, joint angles / torques, foot contact forces, and stability margins. This provides a reliable, quantitative, and risk-free simulation basis for multi-objective evaluation, cost calculation, and optimal selection of candidate action sequences, and is a key simulation foundation module for ensuring action safety and improving decision-making efficiency.

[0077] S140, based on the candidate action sequence and the corresponding predicted state, obtain the evaluation value corresponding to the candidate action sequence through risk assessment.

[0078] Specifically, based on each candidate action sequence and its predicted state obtained through forward calculation using the dynamic model, a comprehensive risk assessment is conducted across multiple dimensions, including motion stability, terrain adaptability, energy consumption level, collision risk, slip probability, and perception uncertainty. Quantitative calculations are performed according to a pre-defined weighted cost function or Pareto ranking rule, ultimately yielding a comprehensive evaluation value for each candidate action sequence. This provides an objective and comparable numerical basis for the subsequent selection of benchmark action sequences. The comprehensive multi-dimensional risk assessment includes motion stability assessment based on the ZMP / CP stability domain and the coverage area of ​​the supporting polygon; terrain adaptability assessment based on terrain height maps and drivability characteristics; energy consumption level assessment based on joint torque and power consumption; collision risk assessment combining the minimum safe distance between the fuselage and the foot; slip probability assessment based on the ground friction coefficient and foot contact cone constraints; and perception uncertainty assessment incorporating sensor noise and environmental feature ambiguity. Through comprehensive quantification and weighted calculations of these dimensions, a comprehensive risk evaluation result objectively reflects the safety level and execution effect of each candidate action sequence.

[0079] S150, the candidate action sequence corresponding to the smallest evaluation value is determined as the optimal action sequence.

[0080] Specifically, the evaluation values ​​of each candidate action sequence obtained through multi-dimensional risk assessment are compared horizontally. The candidate action sequence with the smallest value, lowest overall cost, and best safety and execution effect is selected as the optimal action sequence under the current control cycle. This optimal action sequence serves as the benchmark action sequence for subsequent minor compensation adjustments by the residual reinforcement learning model. This method can quickly screen out the optimal reference action that balances stability, energy efficiency, safety, and terrain adaptability from multiple candidate schemes, reducing online adaptive pressure and control risks from the source, improving the reliability and real-time performance of robot motion decisions, and reducing energy waste and instability caused by ineffective actions. It provides a stable, efficient, and safe benchmark input for the overall control system.

[0081] S160, based on the optimal action sequence, obtain the corresponding action residual correction amount through the trained residual reinforcement learning model.

[0082] Using the optimal action sequence determined within the current control cycle as a benchmark, the robot's real-time unified state vector, online error estimation information, and environmental and task context features are input into a pre-trained residual reinforcement learning model. This model outputs action residual correction amounts used only for minor compensation adjustments to offset the tracking errors and disturbances caused by uncertainties such as unknown ground friction, unknown load changes, actuator performance degradation, and perception bias. This enables the robot to achieve fast, stable, and safe online adaptation based on the benchmark actions.

[0083] It should be noted that this residual reinforcement learning model is a dedicated adaptive policy network that uses the optimal action sequence as a benchmark and outputs only small-amplitude compensation signals. It is the core module for realizing online disturbance rejection and rapid adaptation of legged robots. It takes the robot's unified state vector, online error estimation (friction, load, actuator degradation, perception bias, etc.), and task context features as input. It does not directly generate complete actions, but only outputs small-amplitude, highly safe action residual corrections to offset disturbances caused by environmental and platform uncertainties. During training, anti-drift control is achieved through KL divergence (Kullback-Leibler Divergence), residual norm regularization, and constraint penalties to ensure that the correction amount is always within a safe range and does not compromise the stability of the benchmark action. It can rapidly and robustly improve motion accuracy and robustness under real-world conditions such as unknown terrain, variable loads, and actuator performance degradation, while meeting the real-time and safety requirements of engineering deployment.

[0084] S170, the optimal action sequence is corrected according to the action residual correction amount to obtain the corrected action sequence.

[0085] Specifically, based on the calculated motion residual correction, the determined optimal motion sequence is precisely corrected step-by-step with small amplitudes using a superimposed compensation method. The baseline motion and the residual correction are then fused to form the final corrected motion sequence, effectively offsetting the tracking errors and disturbances caused by uncertainties such as unknown ground friction, load variations, actuator drift, and perception bias. This correction method, while preserving the original stability, energy efficiency, and terrain adaptability of the optimal motion sequence, achieves dynamic adaptation through only small compensations. This avoids the instability risks caused by large motion adjustments and significantly improves the robot's motion accuracy, robustness, and anti-interference capabilities in complex working conditions and unknown environments, making the overall control more closely match the actual operating state.

[0086] S180, the executable action sequence of the legged robot is obtained by projecting the modified action sequence through a preset projection constraint module.

[0087] Specifically, based on the corrected action sequence with completed residual compensation, it is input into a pre-constructed projection constraint module that includes joint limits, velocity / acceleration constraints, torque / power limits, foot cone constraints, ZMP / CP stability domain constraints, and collision safety distance constraints. Through convex optimization calculations, the corrected action sequence is rigorously projected into a physically feasible and safe constraint space, ultimately yielding an executable action sequence that meets robot hardware limitations and motion safety requirements. This processing method can fundamentally prevent dangerous action outputs such as overtravel, overload, instability, and collisions, ensuring that all commands are executed within safe limits. While retaining the adaptive correction effect, it significantly improves the system's engineering safety, reliability, and feasibility, providing hard constraint guarantees for stable robot motion.

[0088] It's important to note that the pre-defined projection constraint module is the core hard constraint layer in the legged robot control system, ensuring physical safety and motion stability. Employing a slightly convex optimization architecture, it pre-constructs all physical constraints and stability rules for the robot, "tailoring" the corrected motion sequence to an absolutely safe and executable range. It uniformly handles multiple key constraints, including joint limits, velocity and acceleration constraints, torque and power limits, foot cone constraints, ZMP / CP stability domain constraints, and collision distance constraints between the body and feet. Through quadratic programming or cone programming, it projects any input motion into a feasible action that satisfies all constraints, while preserving the original motion intent as much as possible. The module also supports contact consistency verification, automatically triggering conservative backoff when the prediction and measurement deviation are too large, avoiding risks such as sudden motion changes, exceeding the range, instability, slippage, or collisions. This module serves as the safety gate connecting the algorithm strategy and the actual actuator, ensuring that adaptive corrections never exceed limits, significantly improving the system's reliability and engineering feasibility in real-world scenarios.

[0089] For example, an executable sequence of actions satisfies:

[0090] ;

[0091] Constraints:

[0092] ;

[0093] Among them, u t For an executable sequence of actions, a t To correct the action sequence, argmin u To optimize the solution operator, its function is to find the optimal variable u that minimizes the objective function from all feasible solutions that satisfy the constraints, i.e., the executable action vector u that minimizes the deviation from the original action. t W is the weight matrix, T is the transpose, and g i (u) represents the i-th constraint, and m represents the number of constraints.

[0094] The objective function is argmin u The function on the right-hand side of the operator is optimized as follows:

[0095] .

[0096] The constraints include joint limit and velocity constraints, acceleration constraints, torque constraints, power constraints, foot cone constraints, stability domain constraints, and collision distance constraints.

[0097] In this embodiment, the state vector, morphological parameter encoding, and control conditions of the legged robot are acquired to provide comprehensive and accurate initial information support for subsequent model input. The state vector, morphological parameter encoding, and control conditions are input into a pre-trained generative motion prior model to obtain multiple candidate action sequences. Multimodal action supply provides rich and motion-compliant reference options for control decisions. The candidate action sequences are input into a pre-constructed dynamics model to obtain the predicted states corresponding to the candidate action sequences. Risk-free forward simulation is used to obtain the future execution effects of each action sequence, providing a reliable basis for quantitative evaluation. Based on the candidate action sequences and their corresponding predicted states, risk assessment is conducted to obtain the evaluation values ​​corresponding to the candidate action sequences, achieving evaluation of each action sequence in terms of stability, energy efficiency, and safety. Objective quantitative comparison is used to determine the optimal action sequence by identifying the candidate action sequence corresponding to the lowest evaluation value. This process selects benchmark actions that balance multi-objective performance from the source, providing a high-quality foundation for subsequent corrections. Based on the optimal action sequence, a trained residual reinforcement learning model is used to obtain the corresponding action residual correction amount, which offsets errors caused by disturbances such as unknown friction, load changes, and perception biases through small-scale compensation. The optimal action sequence is then corrected based on the action residual correction amount to obtain a corrected action sequence, achieving dynamic adaptation to real-world conditions while retaining the advantages of the benchmark action. Finally, the corrected action sequence is projected using a pre-defined projection constraint module to obtain the executable action sequence of the legged robot. Physical safety constraints strictly limit the actions to the feasible domain, avoiding problems such as overtravel and instability. Through a closed-loop mechanism of multimodal candidate supply, accurate simulation evaluation, optimal benchmark selection, adaptive residual correction, and safety constraint projection, control errors are progressively eliminated, the impact of environmental disturbances is reduced, and the physical feasibility of action execution is ensured. This significantly improves the motion control accuracy, stability, and robustness of the legged robot in complex and unknown environments.

[0098] Optionally, the candidate motion sequence includes joint torque and joint angular velocity corresponding to each joint; the predicted state includes predicted motion velocity, zero torque point, capture point, foot contact force, foot position, and predicted pose; obtaining the evaluation value corresponding to the candidate motion sequence through risk assessment based on the candidate motion sequence and the corresponding predicted state includes:

[0099] Obtain the location of obstacles;

[0100] The speed error is obtained by the difference between the predicted speed and the preset expected speed.

[0101] The attitude stability margin is obtained through stability analysis based on the predicted pose, the zero torque point, the capture point, the foot contact force, and the foot position.

[0102] Energy consumption is obtained by integrating the product of the joint torque and the joint angular velocity corresponding to all joints.

[0103] The minimum geometric distance between the preset key point corresponding to the predicted pose and the position of the obstacle is determined as the collision distance;

[0104] The sliding risk index is obtained by determining the friction constraint based on all the foot contact forces and the corresponding foot positions.

[0105] The evaluation value is obtained based on the speed error, the attitude stability margin, the energy consumption, the collision distance, and the sliding risk index through a preset risk assessment relationship.

[0106] Optionally, the energy consumption satisfies:

[0107] ;

[0108] Where P is the energy consumption, [t0, t1] is the integration interval, and L i (t) represents the joint torque corresponding to the i-th joint at time t, W i (t) represents the joint angular velocity of the i-th joint at time t, and η i The preset motor efficiency parameter is the parameter corresponding to the i-th joint.

[0109] Optionally, the risk assessment relationship satisfies:

[0110] ;

[0111] Where J is the evaluated value, E is the velocity error, M is the attitude stability margin, P is the energy consumption, D is the collision distance, S is the sliding risk index, U is the preset perception uncertainty, and ω E ω is the preset error weighting coefficient. M ω is the preset stable weight coefficient. P ω is the preset energy consumption weighting coefficient. D ω is the preset collision weight coefficient. S ω is the preset sliding weight coefficient. U These are the preset perception weight coefficients.

[0112] Specifically, firstly, obstacle locations are acquired to define environmental safety boundaries, providing spatial reference for subsequent collision risk assessment; secondly, the speed error is calculated based on the difference between the predicted motion speed and the preset expected speed, used to quantify the tracking accuracy of the action sequence for the target motion command; thirdly, stability analysis is conducted based on the predicted pose, zero-torque point, capture point, foot contact force, and foot position to obtain the posture stability margin, thereby measuring the robot's tipping risk and dynamic balance capability when performing corresponding actions; fourthly, the joint torque and joint angular velocity of each joint are multiplied to obtain the corresponding power, and the power of all joints is added together and integrated to obtain the energy consumption, achieving an objective quantification of the energy consumption level of the action sequence; fifthly, the minimum geometric distance between the preset key points on the predicted pose and the obstacle positions is determined as the collision distance, directly reflecting the safe distance between the robot body and obstacles during action execution. Specifically, firstly, on the body of the legged robot in the predicted pose... Several pre-defined key points representing collisions are selected from the legs and feet. Combined with the three-dimensional position information of obstacles obtained from environmental perception, the Euclidean distance from each pre-defined key point to the obstacle surface is calculated. The minimum value among all calculated distances is then selected as the collision distance. This directly reflects the closest safe distance between the robot's body structure and surrounding obstacles when executing the current candidate action sequence, accurately quantifying the potential collision risk level and providing a direct safety judgment basis for subsequent risk assessment and action selection. Friction constraints are judged based on the contact force of each foot and the corresponding foot position, and a sliding risk index is obtained to assess the possibility of slippage between the foot and the support surface. Finally, the evaluation value is calculated uniformly through a pre-defined risk assessment relationship, integrating speed error, posture stability margin, energy consumption, collision distance, and sliding risk index, to achieve a normalized and quantitative evaluation of the multi-dimensional performance of the candidate action sequence.

[0113] In this optional embodiment, a comprehensive and detailed evaluation is completed from key dimensions such as motion tracking accuracy, dynamic stability, energy consumption, collision safety and slip risk. This ensures the integrity and objectivity of the risk assessment and provides accurate quantitative basis for the selection of the optimal action sequence, thereby effectively improving the safety, stability and energy efficiency of legged robot motion control, while enhancing the adaptability and control reliability in complex environments.

[0114] Optionally, the step of obtaining the attitude stability margin through stability analysis based on the predicted pose, the zero-moment point, the capture point, the foot contact force, and the foot position includes:

[0115] The foot end with a contact force greater than a preset contact threshold is defined as the contact foot end;

[0116] A supporting polygon is generated by using a convex hull based on the foot position corresponding to all the contact feet.

[0117] The attitude stability margin is obtained by determining the stability margin based on the predicted pose, the zero moment point, the capture point, and the support polygon.

[0118] Specifically, feet with contact force greater than a preset contact threshold are identified as contacting feet, effectively distinguishing between feet actually participating in ground support and suspended feet, thus eliminating interference from invalid contact data on stability assessment. Next, convex hull generation calculations are performed based on the foot positions corresponding to all contacting feet to construct the robot's current support polygon, forming the core stability reference region used to determine the tipping boundary. Then, stability margin is determined by combining the predicted pose, zero-moment point, capture point, and the relative positional relationship with the support polygon. By calculating indicators such as the minimum distance between the zero-moment point and the capture point and the support polygon boundary, a posture stability margin that can quantify the magnitude of the robot's tipping risk is obtained. The larger the value, the more stable the posture and the stronger the anti-disturbance capability.

[0119] It should be noted that the convex hull generation operation is a geometric construction operation for the coordinates of the contact foot positions. It is based on the planar projection coordinates of all effective contact feet, and uses the convex hull algorithm to filter and connect the vertices that form the outer contour boundary, eliminating internal redundant points to form the smallest convex polygon that surrounds all supporting feet, i.e., the support polygon. This operation can quickly and accurately determine the effective support area of ​​the robot on the ground, providing a unified and rigorous geometric boundary for the subsequent determination of the stability margin of zero torque points and capture points, ensuring the objectivity and accuracy of attitude stability judgment.

[0120] In this optional embodiment, by effectively identifying the actual supporting foot end, accurately constructing the supporting area, and making quantitative judgments based on multiple stability criteria, an objective and accurate assessment of the robot's dynamic stability is achieved. This provides a reliable basis for stability performance for subsequent risk assessment and optimal action selection, and significantly improves the safety and reliability of legged robot motion control in complex terrain.

[0121] Optionally, obtaining the attitude stability margin based on the predicted pose, the zero-moment point, the capture point, and the support polygon through stability margin determination includes:

[0122] The zero-moment point is projected onto the support plane corresponding to the support polygon to obtain the corresponding zero-moment point projection position;

[0123] The distance between the captured point and the projected position of the zero torque point is defined as the captured point position deviation;

[0124] The attitude tilt angle is obtained by attitude decomposition based on the predicted pose.

[0125] The minimum distance between the projection position of the zero moment point and the boundary of the supporting polygon is determined as the initial stable distance;

[0126] The attitude stability margin is obtained by weighted summation of the initial stable distance, the capture point position deviation, and the attitude tilt angle.

[0127] Optionally, the attitude stability margin satisfies:

[0128] ;

[0129] Where M is the attitude stability margin, S n Let θ be the initial stable distance, θ be the attitude tilt angle, and d be the initial stable distance. i To capture the position deviation, k d k is the preset distance weighting coefficient. θ k is the preset tilt angle weighting coefficient. i The preset deviation weighting coefficient is k, which is pre-set according to the actual situation. θ and k i Depending on the actual situation, it can be a negative value.

[0130] Specifically, the zero-moment point is projected onto the support plane containing the support polygon to obtain the corresponding zero-moment point projection position, thereby normalizing the spatial stability reference point to the planar support domain and providing a unified benchmark for subsequent stability boundary determination. Then, the distance between the capture point and the projection position of this zero-moment point is determined as the capture point position deviation, which quantifies the degree of deviation between the robot's inertial tendency and the static equilibrium benchmark during dynamic motion, reflecting the risk of dynamic instability. Subsequently, the predicted pose is decomposed to obtain the pitch and roll angles of the robot, directly characterizing the degree of body tilt and reflecting static attitude stability. Specifically, the predicted pose of the legged robot is decomposed in a three-dimensional spatial attitude description form (such as rotation matrix, quaternion, or Euler angles), extracting the pitch and roll angles of the robot relative to the support plane. The tilt angle represents the degree of tilt in the forward and backward direction of the fuselage, while the roll tilt angle represents the degree of tilt in the left and right direction. Together, they constitute the attitude tilt angle, which is used to directly quantify the magnitude of the fuselage's deviation from the horizontal attitude. The larger the value, the more obvious the tilt and the worse the attitude stability. This decomposition step can transform the abstract spatial attitude into an intuitive and calculable angle index, providing a direct and accurate basis for attitude deviation for the subsequent comprehensive calculation of attitude stability margin. Then, the minimum distance from the zero moment point projection position to the boundary of the support polygon is determined as the initial stability distance, which is used to measure the robot's basic stability margin in the current support area. Finally, the initial stability distance, the capture point position deviation, and the attitude tilt angle are weighted and summed using preset weights to calculate the attitude stability margin, realizing the multi-factor fusion quantification of static stability margin, dynamic offset trend, and fuselage tilt state.

[0131] In this optional embodiment, by decomposing the stability influencing factors in layers, constructing quantitative indicators separately, and performing weighted fusion, the overall stability level of the robot under complex motion states can be comprehensively and accurately reflected. This provides an objective and robust basis for stability judgment for risk assessment, effectively improves the accuracy of stability judgment and environmental adaptability of legged robots, and ensures the safety and reliability of motion control.

[0132] Optionally, the step of determining the sliding risk index based on all the foot contact forces and the corresponding foot positions through friction constraints includes:

[0133] Based on the foot contact force, the corresponding ground normal reaction force and ground tangential force at the foot are obtained through vector analysis.

[0134] The foot friction force corresponding to the foot is obtained by multiplying the ground normal reaction force and the preset friction coefficient.

[0135] The difference between the foot friction force and the corresponding ground tangential force is determined as the foot slip risk index corresponding to the foot.

[0136] The largest foot slip risk index is determined as the slip risk index.

[0137] Specifically, the contact force at the foot is analyzed by vector analysis, decomposing it into a ground normal reaction force perpendicular to the support surface and a ground tangential force parallel to the support surface. This allows for precise differentiation of different components of the contact force, providing basic physical parameters for slip determination. Next, the maximum static friction force at the corresponding foot is calculated based on the product of the ground normal reaction force and the preset friction coefficient, thereby determining the anti-slip limit threshold that the foot can provide under the current contact state. Then, the difference between the maximum static friction force at the foot and the ground tangential force is calculated to obtain the foot slip risk index for a single foot. The smaller the difference, the closer the tangential driving force is to the anti-slip limit and the higher the probability of slippage. Finally, the largest slip risk index among all feet is selected as the overall slip risk index, ensuring that the state of the most dangerous foot characterizes the overall slip safety of the robot.

[0138] In this optional embodiment, by decomposing the contact force vector, calculating the friction limit, quantifying the risk of single-leg movement, and screening the global extreme value, a refined and rigorous judgment of the risk of foot slippage is achieved. This can accurately reflect the slippage hazards during the robot's movement, provide reliable slippage safety indicators for subsequent risk assessment and motion optimization, and effectively improve the stability and control safety of legged robot movement in complex terrain.

[0139] like Figure 2 As shown, an embodiment of the present invention provides a legged robot control device 200, comprising:

[0140] The acquisition module 210 is used to acquire the state vector, morphological parameter encoding, and control conditions of the legged robot;

[0141] The generation module 220 is used to input the state vector, the morphological parameter encoding and control conditions into the trained generative motion prior model to obtain multiple candidate action sequences;

[0142] Prediction module 230 is used to input the candidate action sequence into a pre-built dynamic model to obtain the predicted state corresponding to the candidate action sequence;

[0143] Evaluation module 240 is used to obtain the evaluation value corresponding to the candidate action sequence through risk assessment based on the candidate action sequence and the corresponding predicted state;

[0144] The filtering module 250 is used to determine the candidate action sequence corresponding to the smallest evaluation value as the optimal action sequence;

[0145] Optimization module 260 is used to obtain the corresponding action residual correction amount based on the optimal action sequence through a trained residual reinforcement learning model;

[0146] Correction module 270 is used to correct the optimal action sequence according to the action residual correction amount to obtain a corrected action sequence;

[0147] The projection module 280 is used to project the modified action sequence through a preset projection constraint module to obtain the executable action sequence of the legged robot.

[0148] The legged robot control device of this embodiment is used to implement the legged robot control method described above. Its advantages over the prior art are the same as the advantages of the legged robot control method compared to the prior art, and will not be repeated here.

[0149] like Figure 3 As shown, an electronic device 300 provided in this embodiment of the invention includes a memory 310 and a processor 320; the memory 310 is used to store a computer program; the processor 320 is used to implement the legged robot control method described above when the computer program is executed.

[0150] Alternatively, an electronic device 300 includes a memory 310 and a processor 320 coupled to the memory 310; the memory 310 is configured to store a computer program; and the processor 320 is configured to perform the following operations when the computer program is executed:

[0151] Obtain the state vector, morphological parameter encoding, and control conditions of the legged robot;

[0152] The state vector, the morphological parameter encoding, and the control conditions are input into a trained generative motion prior model to obtain multiple candidate action sequences.

[0153] The candidate action sequence is input into a pre-built dynamic model to obtain the predicted state corresponding to the candidate action sequence;

[0154] Based on the candidate action sequence and the corresponding predicted state, an evaluation value corresponding to the candidate action sequence is obtained through risk assessment;

[0155] The candidate action sequence corresponding to the smallest evaluation value is determined as the optimal action sequence;

[0156] Based on the optimal action sequence, the corresponding action residual correction amount is obtained through the trained residual reinforcement learning model;

[0157] The optimal action sequence is corrected based on the action residual correction amount to obtain the corrected action sequence;

[0158] The executable action sequence of the legged robot is obtained by projecting the modified action sequence through a preset projection constraint module.

[0159] This invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the legged robot control method described above.

[0160] Alternatively, a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the following operations:

[0161] Obtain the state vector, morphological parameter encoding, and control conditions of the legged robot;

[0162] The state vector, the morphological parameter encoding, and the control conditions are input into a trained generative motion prior model to obtain multiple candidate action sequences.

[0163] The candidate action sequence is input into a pre-built dynamic model to obtain the predicted state corresponding to the candidate action sequence;

[0164] Based on the candidate action sequence and the corresponding predicted state, an evaluation value corresponding to the candidate action sequence is obtained through risk assessment;

[0165] The candidate action sequence corresponding to the smallest evaluation value is determined as the optimal action sequence;

[0166] Based on the optimal action sequence, the corresponding action residual correction amount is obtained through the trained residual reinforcement learning model;

[0167] The optimal action sequence is corrected based on the action residual correction amount to obtain the corrected action sequence;

[0168] The executable action sequence of the legged robot is obtained by projecting the modified action sequence through a preset projection constraint module.

[0169] The present invention will now be described an electronic device 300 that can serve as a server or client of the present invention, which is an example of a hardware device that can be applied to various aspects of the present invention. Electronic device 300 is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic device 300 can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0170] Electronic device 300 includes a computing unit that can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) or a computer program loaded from a storage unit into random access memory (RAM). The RAM may also store various programs and data required for device operation. The computing unit, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0171] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc. In this application, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of the present invention according to actual needs. Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units can be implemented in hardware or as software functional units.

[0172] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and all such changes and modifications will fall within the scope of protection of the present invention.

Claims

1. A control method for a legged robot, characterized in that, include: Obtain the state vector, morphological parameter encoding, and control conditions of the legged robot; The state vector, the morphological parameter encoding, and the control conditions are input into a trained generative motion prior model to obtain multiple candidate action sequences. The candidate action sequence is input into a pre-built dynamic model to obtain the predicted state corresponding to the candidate action sequence; Based on the candidate action sequence and the corresponding predicted state, an evaluation value corresponding to the candidate action sequence is obtained through risk assessment; The candidate action sequence corresponding to the smallest evaluation value is determined as the optimal action sequence; Based on the optimal action sequence, the corresponding action residual correction amount is obtained through the trained residual reinforcement learning model; The optimal action sequence is corrected based on the action residual correction amount to obtain the corrected action sequence; The executable action sequence of the legged robot is obtained by projecting the modified action sequence through a preset projection constraint module. The candidate motion sequence includes the joint torque and joint angular velocity corresponding to each joint; the predicted state includes the predicted motion velocity, zero torque point, capture point, foot contact force, foot position, and predicted pose; The step of obtaining the evaluation value corresponding to the candidate action sequence through risk assessment based on the candidate action sequence and the corresponding predicted state includes: Obtain the location of obstacles; The speed error is obtained by the difference between the predicted speed and the preset expected speed. The attitude stability margin is obtained through stability analysis based on the predicted pose, the zero torque point, the capture point, the foot contact force, and the foot position. Energy consumption is obtained by integrating the product of the joint torque and the joint angular velocity corresponding to all joints. The minimum geometric distance between the preset key point corresponding to the predicted pose and the position of the obstacle is determined as the collision distance; The sliding risk index is obtained by determining the friction constraint based on all the foot contact forces and the corresponding foot positions. The evaluation value is obtained based on the speed error, the attitude stability margin, the energy consumption, the collision distance, and the sliding risk index through a preset risk assessment relationship.

2. The legged robot control method according to claim 1, characterized in that, The process of obtaining the attitude stability margin through stability analysis based on the predicted pose, the zero-moment point, the capture point, the foot contact force, and the foot position includes: The foot end with a contact force greater than a preset contact threshold is defined as the contact foot end; A supporting polygon is generated by using a convex hull based on the foot position corresponding to all the contact feet. The attitude stability margin is obtained by determining the stability margin based on the predicted pose, the zero moment point, the capture point, and the support polygon.

3. The legged robot control method according to claim 2, characterized in that, The step of obtaining the attitude stability margin based on the predicted pose, the zero-moment point, the capture point, and the support polygon through stability margin determination includes: The zero-moment point is projected onto the support plane corresponding to the support polygon to obtain the corresponding zero-moment point projection position; The distance between the captured point and the projected position of the zero torque point is defined as the captured point position deviation; The attitude tilt angle is obtained by attitude decomposition based on the predicted pose. The minimum distance between the projection position of the zero moment point and the boundary of the supporting polygon is determined as the initial stable distance; The attitude stability margin is obtained by weighted summation of the initial stable distance, the capture point position deviation, and the attitude tilt angle.

4. The legged robot control method according to claim 1, characterized in that, The energy consumption satisfies: ; Where P is the energy consumption, [t0, t1] is the integration interval, and L i (t) represents the joint torque corresponding to the i-th joint at time t, W i (t) represents the joint angular velocity of the i-th joint at time t, and η i The preset motor efficiency parameter is the parameter corresponding to the i-th joint.

5. The legged robot control method according to claim 1, characterized in that, The step of determining the sliding risk index based on all foot contact forces and corresponding foot positions through friction constraints includes: Based on the foot contact force, the corresponding ground normal reaction force and ground tangential force at the foot are obtained through vector analysis. The foot friction force corresponding to the foot is obtained by multiplying the ground normal reaction force and the preset friction coefficient. The difference between the foot friction force and the corresponding ground tangential force is determined as the foot slip risk index corresponding to the foot. The largest foot slip risk index is determined as the slip risk index.

6. The legged robot control method according to claim 1, characterized in that, The risk assessment relationship satisfies: ; Where J is the evaluated value, E is the velocity error, M is the attitude stability margin, P is the energy consumption, D is the collision distance, S is the sliding risk index, U is the preset perception uncertainty, and ω E ω is the preset error weighting coefficient. M ω is the preset stable weight coefficient. P ω is the preset energy consumption weighting coefficient. D ω is the preset collision weight coefficient. S ω is the preset sliding weight coefficient. U These are the preset perception weight coefficients.

7. A control device for a legged robot, characterized in that, include: The acquisition module is used to acquire the state vector, morphological parameter encoding, and control conditions of the legged robot. The generation module is used to input the state vector, the morphological parameter encoding, and the control conditions into a trained generative motion prior model to obtain multiple candidate action sequences; The prediction module is used to input the candidate action sequence into a pre-built dynamic model to obtain the predicted state corresponding to the candidate action sequence; The evaluation module is used to obtain the evaluation value corresponding to the candidate action sequence through risk assessment based on the candidate action sequence and the corresponding predicted state; The candidate motion sequence includes the joint torque and joint angular velocity corresponding to each joint; the predicted state includes predicted motion velocity, zero torque point, capture point, foot contact force, foot position, and predicted pose; obtaining the evaluation value corresponding to the candidate motion sequence through risk assessment based on the candidate motion sequence and the corresponding predicted state includes: Obtain the location of obstacles; The speed error is obtained by the difference between the predicted speed and the preset expected speed. The attitude stability margin is obtained through stability analysis based on the predicted pose, the zero torque point, the capture point, the foot contact force, and the foot position. Energy consumption is obtained by integrating the product of the joint torque and the joint angular velocity corresponding to all joints. The minimum geometric distance between the preset key point corresponding to the predicted pose and the position of the obstacle is determined as the collision distance; The sliding risk index is obtained by determining the friction constraint based on all the foot contact forces and the corresponding foot positions. The evaluation value is obtained based on the speed error, attitude stability margin, energy consumption, collision distance, and sliding risk index through a preset risk assessment relationship. The filtering module is used to determine the candidate action sequence corresponding to the smallest evaluation value as the optimal action sequence; The optimization module is used to obtain the corresponding action residual correction amount based on the optimal action sequence through a trained residual reinforcement learning model; The correction module is used to correct the optimal action sequence according to the action residual correction amount to obtain a corrected action sequence; The projection module is used to project the modified action sequence through a preset projection constraint module to obtain the executable action sequence of the legged robot.

8. An electronic device, characterized in that, Including memory and processor; The memory is used to store computer programs; The processor is configured to implement the legged robot control method as described in any one of claims 1 to 6 when executing the computer program.

9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the legged robot control method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Spine quadruped robot energy consumption control method based on complex terrain perception

    CN118884984A

  • Humanoid robot multi-mode environment sensing and self-adaptive chassis control method

    CN120533719A