Complex terrain self-adaptive motion control method and system for double-wheel-foot robot

By adopting asymmetric training strategies and segmented training mechanisms in the two-wheeled foot robot and combining reinforcement learning models for multi-objective collaborative optimization training, the efficiency and stability of the two-wheeled foot robot in complex terrain environments is solved, and efficient, stable and robust adaptive motion control is achieved.

CN120233686AActive Publication Date: 2025-07-01SHANDONG UNIV

Patent Information

Application Number
CN202510724540.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-01
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

The existing two-wheeled foot robots are difficult to achieve efficient, stable and robust motion control in complex terrain environments, especially in environments with multiple terrain characteristics, and the problems of control failure or excessive energy consumption are more prominent.

Method used

By designing asymmetric training strategies and segmented training mechanisms during reinforcement learning, complex terrain models and motion decision controllers are built, and multi-objective collaborative optimization training is used to achieve adaptive motion control of dual-wheeled foot robots.

Benefits of technology

It realizes efficient and stable motion control of the two-wheeled foot robot on complex terrain, enhances dynamic response ability and motion flexibility, and improves the efficiency and stability of motion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120233686A_ABST
    Figure CN120233686A_ABST
Patent Text Reader

Abstract

The invention discloses a complex terrain self-adaptive motion control method and system for a double-wheel-foot robot, and relates to the technical field of robot motion control. The method comprises the following steps: constructing a complex terrain model comprising a robot model, obstacles and environmental constraints; a motion decision controller is constructed, a reinforcement learning model is utilized to perform multi-target collaborative optimization training on the motion decision controller by adopting an asymmetric training strategy for different complex terrain models, and risk constraints are introduced during updating of the asymmetric training strategy to constrain behaviors of the strategy; strategy gradient back propagation is carried out according to a training result, and an asymmetric training strategy is optimized by using a self-adaptive learning rate adjustment method based on performance feedback. According to the method, an asymmetric training strategy and a segmented training mechanism are designed in the reinforcement learning process, and the efficient, robust and self-adaptive motion control problem of the double-wheel-foot robot in various complex terrain environments is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robot motion control, and in particular to a complex terrain adaptive motion control method and system for a two-wheeled foot robot. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] As the application scenarios of mobile robots continue to expand, how to achieve efficient, stable, and robust motion control on unknown or complex terrains has become a research hotspot in academia and industry. Existing wheeled robots, legged robots, and wheeled-legged robots mostly use classic traditional model predictive control (MPC), sliding mode control, or experience-based trajectory planning methods. These methods can achieve good performance in structured flat scenes, but when faced with unstructured, dynamically changing, or environments with multiple terrain features (such as loose gravel, gullies, steps, slopes, etc.), it is often difficult to balance real-time and stability, and it is easy to cause control failure or excessive energy consumption.

[0004] In recent years, reinforcement learning has attracted widespread attention due to its adaptive decision-making capabilities in high-dimensional, nonlinear systems. Motion control methods based on reinforcement learning can learn end-to-end control strategies in a simulated environment through the collaborative training of policy networks and evaluation networks. However, in the face of various complex terrains, existing reinforcement learning methods do not fully consider the deep fusion and risk constraints of multi-source sensor data, resulting in insufficient robustness of the learned strategies when migrating to real environments, and their training mechanisms cannot be adaptively changed based on unstructured complex scenarios. Therefore, how to use reinforcement learning to achieve efficient adaptive control of two-wheeled and legged robots in complex terrains has become a technical problem that needs to be solved urgently in existing technologies. Summary of the invention

[0005] In view of the shortcomings of the prior art, the purpose of the present invention is to provide a complex terrain adaptive motion control method and system for a two-wheeled legged robot. By designing an asymmetric training strategy and a segmented training mechanism in the reinforcement learning process, the problem of efficient, robust and adaptive motion control of the two-wheeled legged robot in a variety of complex terrain environments is solved.

[0006] In order to achieve the above object, the present invention is implemented through the following technical solutions: A first aspect of the present invention provides a complex terrain adaptive motion control method for a two-wheeled legged robot, comprising the following steps: Construct a robot model according to the posture parameters and motion parameters of the two-wheeled foot robot; Build complex terrain models that include robot models, obstacles, and environmental constraints; Construct a motion decision controller, and use a reinforcement learning model to adopt an asymmetric training strategy for different complex terrain models to perform multi-objective collaborative optimization training on the motion decision controller. Among them, the asymmetric training strategy introduces risk constraints during update to constrain the behavior of the strategy. Perform policy gradient backpropagation according to the training results, and use an adaptive learning rate adjustment method based on performance feedback to optimize the asymmetric training strategy, so as to achieve the adaptive motion control of the two-wheeled foot robot.

[0007] Furthermore, the motion decision controller includes a multi-source sensor data fusion module and an action decision network module. Among them, the multi-source sensor data fusion module fuses the data collected on the two-wheeled foot robot with the environmental terrain information, generates a feature vector and inputs it into the action decision network module to perform real-time prediction on the terrain environment and generate corresponding action decisions.

[0008] Furthermore, the reinforcement learning model includes a policy network and an evaluation network. The policy network is used to output the actions of the robot according to the input of the current state observation of the two-wheeled foot robot, and the evaluation network is responsible for evaluating the value of the robot's actions and states and outputting the state-action value function.

[0009] Furthermore, adopt an asymmetric training strategy, and use the method of parallel learning and collaborative decision-making of multiple two-wheeled foot robots for training.

[0010] Furthermore, the complete reward function of reinforcement learning is: , Among them, is the main reward function, is the speed tracking reward function, is the height stability reward function, is the touchdown penalty function, is the motion stability reward function, is the energy efficiency reward function, is the reward function for special terrain, is the weight parameter specialized for the structured step terrain.

[0011] Furthermore, during the update process of the asymmetric training strategy, use sample estimation of the policy gradient, update the policy using the clipped objective function, and introduce KL divergence to control the amplitude of policy update.

[0012] Furthermore, multi-objective collaborative optimization is achieved by setting multiple optimization objectives and setting different weights for weighted combination to obtain the globally optimal decision-making information. The expression of the global loss function designed based on multi-objective collaborative optimization is as follows: ; Among them, is the global loss function, is the main policy loss, is the risk cost penalty term, is the policy value function loss, is the risk value loss, is the entropy regularization term.

[0013] The second aspect of the present invention provides a complex terrain adaptive motion control system for a two-wheeled and legged robot, including: A parameter acquisition module configured to construct a robot model based on the attitude parameters and motion parameters of the two-wheeled and legged robot; A model construction module configured to construct a complex terrain model including the robot model, obstacles, and environmental constraints; A reinforcement learning training module configured to construct a motion decision controller, and use a reinforcement learning model to perform multi-objective collaborative optimization training on the motion decision controller for different complex terrain models by using an asymmetric training strategy, wherein the asymmetric training strategy introduces risk constraints during update to constrain the behavior of the strategy; A backpropagation optimization module configured to perform policy gradient backpropagation according to the training results, and optimize the asymmetric training strategy by using an adaptive learning rate adjustment method based on performance feedback, so as to realize the adaptive motion control of the two-wheeled and legged robot.

[0014] The third aspect of the present invention provides a computer-readable storage medium storing a computer program, which is suitable for being loaded and executed by a processor to perform the steps in the complex terrain adaptive motion control method for a two-wheeled and legged robot as described in the first aspect of the present invention.

[0015] The fourth aspect of the present invention provides a computer device, which includes: A processor suitable for executing a computer program; A computer-readable storage medium storing a computer program, and when the computer program is executed by the processor, it realizes the complex terrain adaptive motion control method for a two-wheeled and legged robot as described in the first aspect of the present invention.

[0016] The above one or more technical solutions have the following beneficial effects: The present invention discloses a complex terrain adaptive motion control method and system for a two-wheeled and legged robot. By constructing a complex terrain model and a motion decision controller, and using a reinforcement learning model for multi-objective collaborative optimization training, the present invention realizes efficient and stable motion control of the two-wheeled and legged robot on complex terrains, and has remarkable technical effects. It solves the problems of efficient, robust, and adaptive motion control of two-wheeled and legged robots in various complex terrain environments.

[0017] The setting of the motion decision controller of the present invention further improves the motion control ability of the robot. The multi-source sensor data fusion module can effectively fuse various types of data collected on the bipedal robot with the environmental terrain information, generate feature vectors and transmit them to the action decision network module. This process realizes the real-time prediction of the terrain environment and generates corresponding action decisions, enabling the robot to quickly make reasonable motion adjustments according to the real-time environmental changes, enhancing the robot's dynamic response ability and motion flexibility in complex terrains.

[0018] The control method based on reinforcement learning of the present invention enables the robot to find the optimal motion strategy through continuous learning and optimization in complex terrains, thereby improving the efficiency and stability of motion. Its asymmetric training strategy restricts the behavior of the strategy by introducing risk constraints and adopts the method of parallel learning and collaborative decision-making of multiple bipedal robots for training, further improving the training effect and efficiency. During the training process, sample estimation policy gradient is used, the clipped objective function is used to update the policy, and the KL divergence is introduced to control the amplitude of policy update. These technical means effectively ensure the stability and convergence of the training process, enabling the robot to quickly adapt and make optimal decisions in complex and changeable terrain environments.

[0019] The multi-objective collaborative optimization training of the present invention realizes the generation of globally optimal decision-making information, and through weighted combination by setting different weights, enables the robot to comprehensively consider various factors in complex terrains, find the optimal motion path and action strategy, thereby realizing efficient and stable motion control of the bipedal robot on complex terrains, providing reliable technical support for the application of bipedal robots in complex environments, and having broad application prospects and important practical significance.

[0020] Advantages of additional aspects of the present invention will be partly given in the following description, partly will become obvious from the following description, or will be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings.

[0022] Figure 1 It is a flowchart of the complex terrain adaptive motion control method for the bipedal robot in the first embodiment of the present invention; Figure 2Schematic diagram of the two-wheeled and two-legged robot continuously crossing a structured step terrain in Embodiment 1 of the present invention; Figure 3 Schematic diagram of the two-wheeled and two-legged robot crossing a gully terrain in Embodiment 1 of the present invention; Figure 4 Process diagram of the two-wheeled and two-legged robot continuously crossing a structured step terrain in Embodiment 1 of the present invention; Figure 5 Process diagram of the two-wheeled and two-legged robot continuously descending steps in Embodiment 1 of the present invention; Figure 6 Process diagram of the two-wheeled and two-legged robot crossing a gully ground in Embodiment 1 of the present invention; Figure 7 Process diagram of the two-wheeled and two-legged robot moving on a rough ground in Embodiment 1 of the present invention. Detailed implementation manners

[0023] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0024] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary implementation manners according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof; The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0025] Embodiment 1: Embodiment 1 of the present invention provides a complex terrain adaptive motion control method for a two-wheeled and two-legged robot, as Figure 1 shown, including the following steps: Step 1: Construct a robot model according to the attitude parameters and motion parameters of the two-wheeled and two-legged robot.

[0026] In this embodiment, a force sensor and an inertial measurement unit are arranged on the two-wheeled and two-legged robot to collect real-time attitude parameters and motion parameters.

[0027] Step 2: Build a complex terrain model that includes a robot model, obstacles, and environmental constraints.

[0028] In this embodiment, the obstacles are the obstacles caused by the complex terrain, such as structured step terrain, slope terrain, structured wave terrain, rugged terrain, gully terrain, rough gravel terrain, etc. The environmental constraints refer to the constraints generated by the interaction between the robot and the environment during the movement process, such as the interaction with the ground, obstacles, etc., such as friction, contact, collision, and operations such as pushing, kicking, and suddenly applying heavy weight to the robot in the environment.

[0029] Step 3: Build a motion decision controller, and use the reinforcement learning model to perform multi-objective collaborative optimization training on the motion decision controller by adopting an asymmetric training strategy for different complex terrain models.

[0030] Step 3.1: Build a reinforcement learning algorithm model for the control of a two-wheeled foot robot.

[0031] In this embodiment, the reinforcement learning model is a neural network model, including a policy network and an evaluation network. The policy network is used to output the actions of the robot according to the input of the current state observation of the two-wheeled foot robot, and the evaluation network is responsible for evaluating the value of the robot's actions and states and outputting the state-action value function. When building the reinforcement learning model, set the state space, action space, decision-making strategy, and reward function in the reinforcement learning environment.

[0032] Step 3.2: Build a motion decision controller.

[0033] In this embodiment, the motion decision controller includes a multi-source sensor data fusion module and an action decision network module. Among them, the multi-source sensor data fusion module fuses the data collected by the force sensors and inertial measurement units on the two-wheeled foot robot with the environmental terrain information, generates a feature vector and passes it into the action decision network module to perform real-time prediction on the terrain environment. Specifically, it can real-time speculate on the possible terrain environments, such as flat ground, rugged ground, gully terrain, and structured step terrain, pre-train dedicated strategies for different terrains, dynamically adjust the control parameters and reward weight distribution, and use multi-dimensional curriculum learning for dynamic training and parallel training to accelerate and improve the training efficiency, generate corresponding action decisions, and thus control the movement of the two-wheeled foot robot, as Figure 2 、 Figure 3 、 Figure 4 、 Figure 5 、 Figure 6 and Figure 7 shown.

[0034] Step 3.3: Use the reinforcement learning model to perform multi-objective collaborative optimization training on the motion decision controller by adopting an asymmetric training strategy for different complex terrain models.

[0035] Step 3.3.1: Design an asymmetric training strategy based on reinforcement learning.

[0036] In this embodiment, an asymmetric training strategy is adopted for training. Asymmetric training means that different observables are input into the policy network and the evaluation network for training respectively. In this embodiment, considering the acquisition of observables during the subsequent deployment of the training strategy on the real machine, the policy network only uses the observable data that can be conveniently and accurately obtained by the real robot, while the evaluation network uses the observable data with additional external environment data and the observable data that the real robot cannot obtain or is difficult to obtain. The above training method not only does not make the computational complexity of the learning process excessive, but also ensures the accuracy of the reinforcement learning process, improves the robustness of the policy, and facilitates the subsequent sim to sim and sim to real deployment processes.

[0037] In a specific implementation, an asymmetric training strategy is adopted. Different observables are input into the policy network and the evaluation network respectively, and then multiple two-wheeled foot robots are trained in parallel learning and collaborative decision-making mode. A reinforcement learning reward function suitable for two-wheeled foot robots is established. Considering that the two-wheeled foot robot uses the data that can be accurately obtained by itself during subsequent deployment without using the external environment data that cannot be effectively obtained and some difficult-to-obtain self-data, in order to ensure that the robot can make full use of this information and smoothly learn a stable and efficient gait, the expression of the reinforcement learning reward function is designed as follows: .

[0038] Where, is the main reward function, is the speed tracking reward function, including linear velocity and angular velocity tracking. is the height stability reward function, is the touchdown penalty function, is the motion stability reward function, is the energy efficiency reward function, which specifically limits the joint position, velocity, and acceleration. is the reward function for special terrains. Special terrains such as steps and gullies have terrains with vertical angles. is the weight parameter specialized for structured step terrains.

[0039] More specifically: encourages the robot to be able to track the predetermined speed command and give full play to the mobility of the two-wheeled foot robot. The speed tracking reward function is as follows: .

[0040] Where, is the expected linear velocity, is the current linear velocity of the robot, is the expected angular velocity, is the current angular velocity of the robot, and is the scaling factor, and are the weight coefficients.

[0041] The bipedal wheeled robot is encouraged to maintain a predetermined height during movement to avoid unstable movement caused by the body being too high or too low. The height stability reward function is as follows: .

[0042] Where, is the height from the foot end to the body in the z direction, is the predetermined height, are the weight coefficients.

[0043] Penalize the bipedal wheeled robot's base and calf touching the ground to prevent unstable behaviors such as the robot falling or the calf touching the ground to deceive stability, and ensure the most basic balance posture of the robot. The ground contact penalty function is as follows: .

[0044] Where, is the external contact force received by the robot's base, is the external contact force received by the calf of the bipedal wheeled robot, are the weight coefficients.

[0045] Penalize the unstable behaviors of the robot's body to ensure that the robot can move smoothly on flat ground and maintain stability when crossing structured step terrains or gully terrains, and penalize the excessive joint angular velocity to ensure the stability and smoothness of the robot's movement. The motion stability reward function is as follows: .

[0046] Where, is the projection of the gravity vector in the robot coordinate system, is the set gravitational constant, is the acceleration calculated from the velocities at the current moment and the previous moment, and are the weight coefficients.

[0047] Penalize the use of the robot's joint torques and encourage the robot to make full use of the torques to reduce energy consumption. The energy efficiency reward function is as follows: .

[0048] Where, is the The torque of a joint, is the scaling factor, and

[0049] Since the end of the double-wheel foot robot is a wheel joint and it has two legs, it is a great challenge to balance and stably cross structured step terrains and gully terrains. Therefore, a reward function is specialized for special terrains to encourage the robot to actively lift its legs and stably cross the terrain when it comes into contact with structured step terrains or gully terrains. The reward function for special terrains is as follows: .

[0050] Wherein, is the contact condition of whether the wheel part of the double-wheel foot robot touches the step terrain, represents the number of the wheel, is the contact force received by the wheel of the double-wheel foot robot except in the vertical direction, is the contact force received by the wheel in the vertical direction, is the upward speed of the robot, is the upward acceleration of the robot, is the change in the distance from the wheel end of the double-wheel foot robot to the base between the previous moment and the current moment, is the constraint value of the distance change, , and are weight coefficients.

[0051] Step 3.3.2: The policy network generates actions based on the observation, obtains the first batch of trajectories after interacting with the environment, and calculates the return and advantage. According to the actions generated by the policy network, update the robot state, calculate the next state, calculate the new observation, which is used as the input for the next time step of the policy, and calculate the reward based on the current state.

[0052] Specifically, the evaluation network is responsible for evaluating the value of the current action and state, that is, the expected value of the cumulative reward that can be obtained in the future under the current state. The input includes the robot's own state observation information, the body linear velocity, and the perceived external terrain information, and the output is the state value function , which is used to evaluate the value of the current action and state, that is, the expected value of the cumulative reward that can be obtained in the future under the current state. Set the main advantage function , which is used to measure the quality of the current action relative to the average policy, and update the policy network based on the feedback information. The expression of the main advantage function is as follows: .

[0053] Wherein, is the action value function, indicating that when in state take action The cumulative reward after is the state-value function, representing the expected value of the cumulative reward in state .

[0054] During actual calculation, the expression of the advantage function at each time step is as follows: .

[0055] Among them, is the advantage function at time step t, is the action-value function at time step t, is the state-value function at time step t.

[0056] The policy network generates actions based on the current state observations of the two-wheeled foot robot. The goal is to maximize the long-term cumulative reward by optimizing the policy. The expression of the policy objective is as follows: .

[0057] Among them, is the policy objective, represents the expectation (average), is the advantage function calculated by the evaluation network, is the parameterized policy of the policy network, representing the probability of selecting action in state .

[0058] The optimization objective of the evaluation network is to minimize the prediction error of the state value. The expression of the value function loss is as follows: .

[0059] Among them, is the value function loss, represents the expectation (average) of the quantity at time step t, is the target cumulative reward, calculated through the reward and the discount factor.

[0060] Step 3.3.3: To limit the amplitude of policy update, during the update process of the asymmetric training policy, the sample is used to estimate the policy gradient, the clipped objective function is used to update the policy, and the KL divergence is introduced to control the amplitude of policy update. And for unified calculation of the loss, that is, to minimize the loss.

[0061] Define the expression of the policy loss function as follows: .

[0062] Among them, is the policy loss function, Denotes the expectation (average) of the data collected at time step t, is the policy ratio, i.e., the degree of deviation of the updated policy on the existing samples, is the probability that the "new" policy network selects an action in the state, is the probability of the "old" policy network under the same state-action pair, are the parameters of the policy network, is the advantage function at time step t, is the clipping range, is the weight coefficient of KL regularization, which is dynamically adjusted according to the size of the average KL divergence calculated in real time and the target threshold, measures the difference in the policy distribution before and after the update, clip represents the clipping operation, i.e., constraining to and between.

[0063] Step 3.3.4: The asymmetric training policy introduces risk constraints during update to constrain the behavior of the policy.

[0064] Specifically, during the policy update, risk constraints are introduced, and the behavior of the policy is constrained by introducing risk-related value functions and loss functions, so that while optimizing the reward, the policy satisfies the risk constraint threshold condition and avoids dangerous behaviors, excessive motor torques and energy consumption, and unstable actions. Through the risk advantage function and the risk cost penalty term are introduced into the policy optimization objective, and the function expression is as follows: .

[0065] Among them, is the temporal difference error of the risk constraint at time step t, is the risk value function at time step t, is the temporal difference error of the risk constraint at time step t+1, is the immediate risk value at time step t, is the risk value function at time step t+1, is the discount factor, is the GAE parameter, represents the termination time step index of a complete sampling trajectory, is the offset index.

[0066] Following the previous clipping idea, the surrogate loss function of risk is obtained, and the expression is as follows: .

[0067] Among them, represents the proxy loss function.

[0068] For each time step, a risk cost amount is set , which is used to reflect the degree to which the immediate risk exceeds the threshold. The expression of the risk cost penalty term is defined as follows: .

[0069] Among them, is the penalty coefficient.

[0070] Similarly, the risk value loss is obtained: .

[0071] Among them, is the risk value loss function.

[0072] Step 3.3.5: Multi-objective collaborative optimization obtains the globally optimal decision information by setting multiple optimization objectives and setting different weights for weighted combination. The expression of the global loss function designed based on multi-objective collaborative optimization is as follows: In a specific implementation manner, multi-objective collaborative optimization expands the traditional single reward signal into a multi-modal optimization objective, including policy task update, evaluation network value optimization, and safety and stability constraints related to task risks, constructs a composite optimization objective, and obtains the globally optimal decision information through weighted combination with different weights. The expression of the designed global loss function is as follows: .

[0073] Among them, is the global loss function, is the main policy loss, is the risk cost penalty term, is the policy value function loss, is the risk value loss, is the entropy regularization term.

[0074] Step 3.3.6: The training process gradually increases the learning difficulty dynamically.

[0075] During the iterative training process of this embodiment, multi-dimensional curriculum learning is used to dynamically adjust the motion commands, observation noise, terrain difficulty, and external disturbances of the robot during the training process, such as motor errors, random ground friction, random external forces, etc. At the beginning of the training, simple motion commands, observation noise, and simple terrain are given. As the number of training times increases, the difficulty of the motion commands and the complexity of the terrain are gradually increased, and the ranges of the observation noise and external disturbances are enlarged. At the same time, each time the environment is reset during the training, the command range and the terrain difficulty where the robot is located will be dynamically adjusted according to the completion degree and motion performance of the motion commands under the terrain in the previous interaction process of the robot.

[0076] Step 4: Perform policy gradient backpropagation according to the training results, and use an adaptive learning rate adjustment method based on performance feedback to optimize the asymmetric training policy, so as to achieve the adaptive motion control of the two-wheeled foot robot.

[0077] In a specific implementation manner, the policy network samples actions according to the current state observations of the two-wheeled foot robot , the environment executes returns the new state , immediate reward and immediate risk value , and is stored in the experience buffer, and the target comprehensive return and comprehensive advantage estimation are calculated. Among them, the target comprehensive return includes the main reward function and the risk value function, and the comprehensive advantage estimation includes the main advantage function and the risk advantage function.

[0078] Calculate each loss and the global loss to perform policy gradient backpropagation, update the parameters according to the learning rate, and optimize the policy. Specifically, relying on the calculated global loss, apply the chain rule operation from the output layer to the input layer in the reverse direction to obtain the gradient of each parameter, and then update the parameters along the negative gradient direction to make the loss decrease and achieve the purpose of policy optimization.

[0079] However, the traditional policy gradient algorithm uses a fixed learning rate, and the step size of moving forward along the gradient direction is a fixed value each time the parameters are updated. In the case of complex actions, it is easy to have oscillations or divergences caused by the learning rate being too large and crossing the "optimal point", or the training efficiency is low due to the learning rate being too small and the convergence speed being too slow. Therefore, an adaptive learning rate adjustment method based on performance feedback is used to dynamically adjust the learning rate through KL divergence, balance the exploration stage in the initial stage of training and the convergence stage in the later stage of training, and the dynamic adjustment formula of the learning rate is as follows: .

[0080] Among them, represents the learning rate before the k-th update, k is the training round, The average KL value calculated at the end of each epoch, is the scaling factor, and is the KL threshold.

[0081] Example Two: Embodiment Two of the present invention provides a complex terrain adaptive motion control system for a two-wheeled foot robot, including: A parameter acquisition module configured to construct a robot model based on the attitude parameters and motion parameters of the two-wheeled foot robot; A model construction module configured to construct a complex terrain model including the robot model, obstacles, and environmental constraints; A reinforcement learning training module configured to construct a motion decision controller, and use an asymmetric training strategy for multi-objective collaborative optimization training of the motion decision controller for different complex terrain models by using a reinforcement learning model, where the asymmetric training strategy introduces risk constraints during update to constrain the behavior of the strategy; A backpropagation optimization module configured to perform policy gradient backpropagation based on the training results, and optimize the asymmetric training strategy by using an adaptive learning rate adjustment method based on performance feedback, so as to achieve the adaptive motion control of the two-wheeled foot robot.

[0082] Example Three: Embodiment Three of the present invention provides a computer-readable storage medium storing a computer program, which is adapted to be loaded and executed by a processor to perform the steps in the complex terrain adaptive motion control method for a two-wheeled foot robot as described in Embodiment One of the present invention.

[0083] Example Four: Embodiment Four of the present invention provides a computer device, which includes: A processor adapted to execute a computer program; A computer-readable storage medium storing a computer program, and when the computer program is executed by the processor, it implements the steps in the complex terrain adaptive motion control method for a two-wheeled foot robot as described in Embodiment One of the present invention.

[0084] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application. In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data processing device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)), etc. The above description is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any technical person familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A complex terrain adaptive motion control method for a two-wheeled foot robot, characterized in that, Including the following steps: Construct a robot model based on the attitude parameters and motion parameters of the two-wheeled foot robot; Construct a complex terrain model including the robot model, obstacles and environmental constraints; Construct a motion decision controller, and use a reinforcement learning model to perform multi-objective collaborative optimization training on the motion decision controller with an asymmetric training strategy for different complex terrain models. Among them, the asymmetric training strategy introduces a risk constraint during update to constrain the behavior of the strategy; Perform policy gradient backpropagation according to the training results, and use an adaptive learning rate adjustment method based on performance feedback to optimize the asymmetric training strategy, so as to achieve the adaptive motion control of the two-wheeled foot robot.

2. The complex terrain adaptive motion control method of the two-wheeled foot robot according to claim 1, characterized in that The motion decision controller includes a multi-source sensor data fusion module and an action decision network module. Among them, the multi-source sensor data fusion module fuses the data collected on the two-wheeled foot robot with the environmental terrain information, generates a feature vector and passes it into the action decision network module to perform real-time prediction on the terrain environment and generate corresponding action decisions.

3. The complex terrain adaptive motion control method for a two-wheeled foot robot according to claim 1, characterized in that, The reinforcement learning model includes a policy network and an evaluation network. The policy network is used to output the actions of the robot according to the input of the current state observation of the two-wheeled foot robot, and the evaluation network is responsible for evaluating the value of the robot's actions and states and outputting the state-action value function.

4. The complex terrain adaptive motion control method for a two-wheeled foot robot according to claim 1, characterized in that, Adopt an asymmetric training strategy and use the method of parallel learning and collaborative decision-making of multiple two-wheeled foot robots for training.

5. The complex terrain adaptive motion control method of the two-wheeled foot robot according to claim 1, characterized in that The complete reward function of reinforcement learning is: , Among them, is the main reward function, is the speed tracking reward function, is the height stability reward function, is the touchdown penalty function, is the motion stability reward function, is the energy efficiency reward function, is the reward function for special terrains, is the weight parameter specialized for structured step terrains.

6. The complex terrain adaptive motion control method for a two-wheeled foot robot according to claim 1, characterized in that During the update of the asymmetric training strategy, the sample is used to estimate the policy gradient, the policy is updated using the clipped objective function, and the KL divergence is introduced to control the amplitude of the policy update.

7. The complex terrain adaptive motion control method for a two-wheeled foot robot according to claim 1, characterized in that Multi-objective collaborative optimization is achieved by setting multiple optimization objectives and setting different weights for weighted combination to obtain the globally optimal decision information. The expression of the global loss function designed based on multi-objective collaborative optimization is as follows: ; Among them, is the global loss function, is the main policy loss, is the risk cost penalty term, is the policy value function loss, is the risk value loss, is the entropy regularization term.

8. A complex terrain adaptive motion control system for a two-wheeled foot robot, characterized in that, Including: A parameter acquisition module configured to construct a robot model based on the attitude parameters and motion parameters of the two-wheeled foot robot; A model construction module configured to construct a complex terrain model including the robot model, obstacles and environmental constraints; A reinforcement learning training module configured to construct a motion decision controller, and use a reinforcement learning model to perform multi-objective collaborative optimization training on the motion decision controller with an asymmetric training strategy for different complex terrain models. Among them, the asymmetric training strategy introduces a risk constraint during update to constrain the behavior of the strategy; A backpropagation optimization module configured to perform policy gradient backpropagation according to the training results, and use an adaptive learning rate adjustment method based on performance feedback to optimize the asymmetric training strategy, so as to achieve the adaptive motion control of the two-wheeled foot robot.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded and executed by a processor to perform the method for adaptive motion control of a two-wheeled foot robot in complex terrain as described in any one of claims 1-7.

10. A computer device, characterized in that a processor suitable for executing a computer program; A computer-readable storage medium storing a computer program, which when executed by the processor, implements the complex terrain adaptive motion control method of the two-wheeled foot robot according to any one of claims 1-7.

Citation Information

Patent Citations

  • Terrain detection method based on joint motion analysis of wheel-foot robot

    CN115145292A

  • Wheel-leg robot wheel-foot switching control method based on BP neural network

    CN116859975A

  • Intelligent agent control method and system

    CN116991083A

  • Lane keeping assistance method based on risk situation awareness constraint strategy optimization

    CN117302208A

  • Blind hexapod robot motion strategy training method

    CN117340876A

Cited By

  • Robot walking control method, device and equipment and medium

    CN120909328A

  • Motion control strategy network training method and device for foot robot with floating substrate based on reinforcement learning

    CN121179441A

  • Multi-skill integrated control method and system for wheel-foot robot

    CN121918613A

  • Robot motion control method and system in microgravity environment

    CN121957041A

  • Transition control method and system for assisting robot to switch control strategy

    CN122308109A