Dual-wheeled foot robot based on deep reinforcement learning, height obstacle control method and device thereof, and motion control model training method

Through the policy network control method of deep reinforcement learning, the robot performs jumping actions when detecting a height obstacle, solving the problem of poor adaptability of traditional methods in complex environments and achieving more stable obstacle-over-the-blocking ability.

CN120066054BActive Publication Date: 2025-07-18SHENZHEN ZHUJI POWER TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510527007.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-18
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

Traditional robot motion control methods are poor in adaptability when facing complex dynamic environments and high obstacles, and it is difficult to quickly adjust the strategy, resulting in unstable control.

Method used

The control method based on deep reinforcement learning is adopted to output jumping actions through the policy network, including the main body squatting and the rapid forward rotation of the double-wheel foot structure attached to the height obstacle, so as to achieve the robot's obstacle crossing.

Benefits of technology

It improves the performance effect of robots through high obstacles, enhances adaptability to complex environments, reduces dependence on high-precision sensors and complex computing resources, and improves the universality of control strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066054B_ABST
    Figure CN120066054B_ABST
Patent Text Reader

Abstract

The present disclosure provides a height obstacle control method and device for a two-wheeled and legged robot based on deep reinforcement learning, and a motion control model training method. The control method includes: judging whether the robot encounters a height obstacle according to the motion feedback data of the two-wheeled and legged robot; when encountering a height obstacle, controlling the two-wheeled and legged robot to perform a takeoff action, and at the same time controlling the two-wheeled structure to rotate forward quickly while attaching to the height obstacle until the motion feedback data does not meet the preset conditions; wherein, the takeoff action is output by a policy network trained based on a deep reinforcement learning algorithm in a height obstacle terrain. The present disclosure can improve the performance of the two-wheeled and legged robot in passing through height obstacles through the takeoff action of the robot and the quick forward rotation of the two-wheeled structure, thereby enhancing the adaptability of the robot to complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical fields of sensors and robotics, and relates to a method and device for controlling a two-wheeled legged robot to pass through height obstacles based on deep reinforcement learning, and a method for training a motion control model. Background Art

[0002] In the field of robot motion control, traditional control methods usually rely on model predictive control (MPC), proportional-integral-derivative control (PID), or classical optimization algorithms. However, these methods often struggle to achieve stable and precise control in the face of complex dynamic environments, scenarios with drastic speed changes, or high-dimensional control tasks. Especially in the presence of unknown disturbances or partial observation information, traditional methods have poor adaptability and are difficult to quickly adjust strategies, thus affecting the motion performance of robots in complex environments.

[0003] For example, in the Chinese patent application with the publication number CN115892277A, when an obstacle appears on the ground, the hip joint motor and the knee joint motor are controlled by a microcontroller to slowly rotate so that the robot is in a squatting energy storage state, and at the same time, the hub motor is controlled to accelerate and move forward, and then a control signal is sent to make the microcontroller control the hip joint motor and the knee joint motor to rotate quickly in the reverse direction. At this time, the thigh rod and the crank rotate quickly in the reverse direction, and further, the thigh rod can rotate quickly relative to the main body, and the calf can rotate quickly relative to the thigh rod. Finally, the ground reaction force enables the robot to cross the obstacle in a jumping posture. However, in this technical solution, for obstacles of different heights, different controls are required, with poor adaptability and difficulty in quickly adjusting strategies.

[0004] In recent years, control methods based on deep reinforcement learning (DRL) have gradually become an important research direction in the field of robot control. Deep reinforcement learning uses neural networks to extract features from high-dimensional input information and learns optimal control strategies through policy optimization algorithms (such as PPO, DDPG, etc.). For example, the Chinese patent application with the publication number CN118192558A discloses a control algorithm for a wheel-legged robot based on model prediction and deep reinforcement learning. However, regarding the control when passing through height obstacles, specifically: when encountering a height obstacle, dynamic modeling and analysis are performed based on the inverted pendulum model of the wheel-legged robot, and the discrete-time model of the system is obtained through linearization representation and model prediction. The method of solving quadratic programming problems is used to obtain the optimal control input variables and determine the control input at the current moment. This is still a traditional model predictive control-based method, which is easily affected by factors such as the delay of the dynamic modeling algorithm and terrain uncertainty, resulting in unstable action execution.

[0005] Therefore, there is an urgent need to provide a motion control method for a two-wheeled foot robot that performs better in passing over height obstacles, so as not to affect the adaptability of the robot to complex environments. Summary of the Invention

[0006] The present disclosure provides a method and device for controlling a two-wheeled foot robot to pass over height obstacles based on deep reinforcement learning, a method and device for training a motion control model, an electronic device, a robot, and a storage medium, which are used to improve the performance of the two-wheeled foot robot in passing over height obstacles, and further improve the adaptability of the robot to complex environments.

[0007] Additional aspects and advantages of the present disclosure will be partially set forth in the following description, and partially will become apparent from the description, or can be learned through the practice of the present disclosure.

[0008] According to a first aspect of the present disclosure, there is provided a method for controlling a two-wheeled foot robot to pass over height obstacles based on deep reinforcement learning, including:

[0009] When the motion feedback data of the two-wheeled foot robot meets a preset condition, it is determined that the robot encounters a height obstacle;

[0010] Controlling the two-wheeled foot robot to perform a takeoff action, the takeoff action includes the main body squatting down and then rising rapidly, and controlling the two-wheeled foot structure of the two-wheeled foot robot to rotate forward rapidly while attaching to the height obstacle until the motion feedback data of the two-wheeled foot robot does not meet the preset condition;

[0011] Wherein, the takeoff action is output by a policy network trained based on a deep reinforcement learning algorithm in a height obstacle terrain.

[0012] In an exemplary embodiment of the present disclosure, the motion feedback data includes wheel-foot rotation speed data, and the preset condition is that the difference between the target rotation speed and the actual rotation speed of the two wheels exceeds a preset threshold.

[0013] In an exemplary embodiment of the present disclosure, the actual rotation speed is obtained by collecting through a motor encoder installed on the two-wheeled foot robot; the target rotation speed is an expected rotation speed command sent by the controller to the drive motor of the two-wheeled foot robot.

[0014] In an exemplary embodiment of the present disclosure, the difference between the target rotation speed and the actual rotation speed of the two wheels exceeding the preset threshold includes: the difference between the target rotation speed and the actual rotation speed of the two wheels exceeds 30% of the target rotation speed.

[0015] In an exemplary embodiment of the present disclosure, the preset condition further includes: when the duration for which the difference between the target rotation speed and the actual rotation speed of the two wheels exceeds the preset threshold exceeds a preset time length.

[0016] In an exemplary embodiment of the present disclosure, the preset duration is greater than 200 milliseconds.

[0017] In an exemplary embodiment of the present disclosure, the action of the main body quickly rising after squatting includes:

[0018] Controlling the main body of the two-wheeled foot robot to squat briefly by lowering the center of gravity and store energy;

[0019] Providing an upward thrust through the two-wheel drive mechanism to make the robot take off from the ground.

[0020] In an exemplary embodiment of the present disclosure, the quick forward rotation of the two-wheeled foot structure attached to the height obstacle includes:

[0021] Adjusting the driving torque of the two-wheeled foot to actively rotate forward quickly at a speed higher than the normal driving speed to generate effective friction to assist the two-wheeled foot robot to climb.

[0022] In an exemplary embodiment of the present disclosure, the policy network is trained through the following method:

[0023] In a simulation environment with a height obstacle terrain, using the teacher-student model framework, training the policy network for outputting the action policy for controlling the movement of the robot.

[0024] In an exemplary embodiment of the present disclosure, the height obstacle terrain includes various types of height obstacles; there are differences in at least one of the following parameters between different types of height obstacles:

[0025] The overall height of the obstacle, the surface roughness of the obstacle side wall, the material stiffness of the obstacle side wall, the geometric shape of the obstacle top, the surface flatness of the obstacle side wall, the inclination angle between the obstacle side wall and the ground.

[0026] In an exemplary embodiment of the present disclosure, the training process of the policy network specifically includes:

[0027] Encoding the historical motion state information of the two-wheeled foot robot itself through the student encoder to generate a first latent vector, and encoding the privileged state information of the two-wheeled foot robot through the teacher encoder to generate a second latent vector; the privileged state information includes the height obstacle information obtained from the simulation environment;

[0028] Selecting the first latent vector or the second latent vector according to the preset policy and inputting it into the policy network;

[0029] Based on the received latent vector and the current motion state information, the policy network outputs the action policy for controlling the movement of the two-wheeled foot robot; the current motion state information at least includes: attitude sensor data, wheel foot rotation speed data, torque data, and position and speed data;

[0030] Utilize privileged state information through a value network to output a value estimate of the current state;

[0031] Optimize the parameters of the student encoder based on the difference between the first latent vector and the second latent vector;

[0032] Update the parameters of the policy network using a deep reinforcement learning algorithm based on the action policy and the value estimate.

[0033] In an exemplary embodiment of the present disclosure, the deep reinforcement learning algorithm includes a comprehensive reward function for evaluating the action performance of the two-wheeled foot robot in a simulation environment. The reward function includes the success rate of passing through a height obstacle, the takeoff height, and the attitude stability.

[0034] According to a second aspect of the present disclosure, there is provided a method for training a motion control model of a two-wheeled foot robot based on deep reinforcement learning, including:

[0035] Construct a virtual training scenario including a height obstacle terrain, and obtain the motion data of a virtual two-wheeled foot robot in the virtual training scenario;

[0036] Use a deep reinforcement learning algorithm to train a policy network according to the motion data and environmental information; the policy network outputs the following action policy for the two-wheeled foot robot:

[0037] Execute a takeoff action when encountering a height obstacle. The takeoff action includes quickly rising after the main body squats down, and controlling the two-wheeled foot structure of the two-wheeled foot robot to rotate forward quickly while attaching to the height obstacle until the motion feedback data of the two-wheeled foot robot does not meet the preset conditions.

[0038] According to a third aspect of the present disclosure, there is provided a height obstacle control device for a two-wheeled foot robot based on deep reinforcement learning, the device including:

[0039] An obstacle determination module for determining that the robot encounters a height obstacle when the motion feedback data of the two-wheeled foot robot meets the preset conditions;

[0040] An action control module for controlling the two-wheeled foot robot to execute a takeoff action. The takeoff action includes quickly rising after the main body squats down, and controlling the two-wheeled foot structure of the two-wheeled foot robot to rotate forward quickly while attaching to the height obstacle until the motion feedback data of the two-wheeled foot robot does not meet the preset conditions;

[0041] Wherein, the takeoff action is output by a policy network trained based on a deep reinforcement learning algorithm in a height obstacle terrain.

[0042] According to a fourth aspect of the present disclosure, there is provided a device for training a motion control model of a two-wheeled foot robot based on deep reinforcement learning, the device including:

[0043] A training data acquisition module, configured to construct a virtual training scenario including a height obstacle terrain, and acquire motion data of a virtual two-wheeled foot robot in the virtual training scenario;

[0044] A model training module, configured to train a policy network according to the motion data and environmental information by using a deep reinforcement learning algorithm; the policy network outputs the following action policies for the two-wheeled foot robot:

[0045] When encountering a height obstacle, execute a takeoff action, where the takeoff action includes quickly rising after the main body squats down, and controlling the two-wheeled foot structure of the two-wheeled foot robot to rotate forward quickly while attaching to the height obstacle until crossing the height obstacle.

[0046] According to a fifth aspect of the present disclosure, there is provided an electronic device, including:

[0047] A processor; and

[0048] A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the methods of the above embodiments are implemented.

[0049] According to a sixth aspect of the present disclosure, there is provided a robot, including:

[0050] A processor; and

[0051] A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the methods of the above embodiments are implemented.

[0052] In an exemplary embodiment of the present disclosure, the robot is a robot configured with a height obstacle control function for a two-wheeled foot robot based on deep reinforcement learning, and the robot includes any one of a humanoid robot, a cleaning robot, a transportation robot, and a mobile robot.

[0053] According to a seventh aspect of the present disclosure, there is provided a computer-readable storage medium, on which computer program code instructions are stored, and when the computer program code instructions are called by a processor of a robot, the robot is caused to execute the methods of the above embodiments.

[0054] It can be seen from the above technical solutions that the present disclosure has at least one of the following advantages and positive effects:

[0055] In the exemplary embodiment of the present disclosure, the height obstacle control method for a two-wheeled bipedal robot based on deep reinforcement learning enables the two-wheeled bipedal robot to perform a takeoff action when detecting a height obstacle and complete obstacle crossing through the rapid forward rotation of the two-wheeled bipedal structure. On the one hand, based on the real-time judgment of the robot's motion feedback data, when the motion feedback data meets the first condition, the robot can automatically trigger the obstacle-crossing action. Compared with the existing technology that relies on model prediction, the deep reinforcement learning method adopted in the present disclosure enables the robot to autonomously adjust its actions when encountering a height obstacle without relying on complex modeling and prediction. On the other hand, when the robot detects a height obstacle, the present disclosure controls the robot to perform a takeoff action, and this takeoff process includes a rapid rise after the main body squats down, enabling the robot to accumulate sufficient kinetic energy in a short time to complete obstacle crossing. In the existing technology, for obstacles of different heights, different control modes often need to be set, while the present disclosure enables the robot to autonomously adjust the squatting energy storage and rising speed according to the characteristics of different obstacles through the policy network trained by deep reinforcement learning, thereby effectively coping with obstacles of different heights, improving the generality of the control strategy, and avoiding the problem that traditional methods based on preset rules need to frequently adjust parameters when the height changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0057] Figure 1 The system architecture diagram shows the height obstacle control method for a two-wheeled bipedal robot based on deep reinforcement learning in the embodiments of the present disclosure that can be applied.

[0058] Figure 2 The flowchart shows a height obstacle control method for a two-wheeled bipedal robot based on deep reinforcement learning in the embodiments of the present disclosure.

[0059] Figure 3 The schematic diagram shows a scenario of a two-wheeled bipedal robot moving towards a height obstacle in the embodiments of the present disclosure.

[0060] Figure 4 The schematic diagram shows a scenario of a two-wheeled bipedal robot leaping in the embodiments of the present disclosure.

[0061] Figure 5 The schematic diagram shows a scenario of a two-wheeled bipedal robot landing on a height obstacle in the embodiments of the present disclosure.

[0062] Figure 6 Shows a schematic flow diagram of a rapid upward movement after a robot body squats down in an embodiment of the present disclosure.

[0063] Figure 7 Shows a schematic diagram of a two-stage training-inference framework of a policy network in an embodiment of the present disclosure.

[0064] Figure 8 Shows a schematic diagram of another two-stage training-inference framework of a policy network in an embodiment of the present disclosure.

[0065] Figure 9 Shows a schematic flow diagram of a training process of a policy network in an embodiment of the present disclosure.

[0066] Figure 10 Shows a schematic flow diagram of a training method for a motion control model of a two-wheeled foot robot based on deep reinforcement learning in an embodiment of the present disclosure.

[0067] Figure 11 Shows a block diagram of a two-wheeled foot robot based on deep reinforcement learning passing through a height obstacle control device in an embodiment of the present disclosure.

[0068] Figure 12 Shows a block diagram of a training device for a motion control model of a two-wheeled foot robot based on deep reinforcement learning in an embodiment of the present disclosure.

[0069] Figure 13 Shows a schematic diagram of a robot in an embodiment of the present disclosure.

[0070] Figure 14 Shows a schematic structural diagram of an electronic device in an embodiment of the present disclosure.

[0071] Figure 15 Shows a schematic structural diagram of a program product in an embodiment of the present disclosure. Detailed implementation manners

[0072] In the description of the present disclosure, the terms "first" and "second" are only used for description and do not indicate relative importance or imply the number of technical features. Therefore, the features of "first" and "second" may explicitly or implicitly include at least one of such features. The meaning of "a plurality" is at least two, unless otherwise clearly defined.

[0073] Figure 1 Shows a system architecture diagram to which the height obstacle control method for a two-wheeled foot robot based on deep reinforcement learning in an embodiment of the present disclosure can be applied.

[0074] Such as Figure 1As shown in the figure, the system architecture 100 may include a terminal device 101, a two-wheeled foot robot 102, a network 103, and a server 104. Among them, the terminal device 101 includes, but is not limited to, a desktop computer, a portable computer, a smart phone, a tablet computer, and so on. The terminal device 101 can serve as an interactive interface, providing a visualization function to display the running state, motion trajectory, terrain information, etc. of the two-wheeled foot robot 102, and at the same time supporting sending motion control instructions to the two-wheeled foot robot 102.

[0075] The two-wheeled foot robot 102 sends the motion feedback data to the server 104 through the network 103 for centralized processing. When the motion feedback data meets the preset conditions, the server 104 determines that the two-wheeled foot robot 102 has encountered a height obstacle, and the policy network trained in advance based on the deep reinforcement learning algorithm in the height obstacle terrain outputs a takeoff action policy. Subsequently, the takeoff action policy is sent to the two-wheeled foot robot 102 through the network 103 to guide the two-wheeled foot robot 102 to execute a takeoff action including quickly rising after the main body squats down, and an action of quickly rotating straight while the two-wheeled foot structure of the two-wheeled foot robot 102 adheres to the height obstacle until the motion feedback data of the two-wheeled foot robot 102 does not meet the preset conditions.

[0076] Of course, if the two-wheeled foot robot 102 has local inference ability, it can also store the policy network internally, enabling the two-wheeled foot robot 102 to autonomously output a takeoff action policy matching the current height obstacle according to the motion feedback data, thereby realizing edge computing and fast response.

[0077] The network 103 is used to provide a medium for communication links between the terminal device 101, the two-wheeled foot robot 102, and the server 104. The network 103 can include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. It should be understood that Figure 1 the numbers and types of the terminal device, the two-wheeled foot robot, the network, and the server in

[0078] The exemplary embodiment of the present disclosure provides a method for controlling a two-wheeled foot robot to pass through a height obstacle based on deep reinforcement learning. Referring to Figure 2 as shown in the figure, this method may include the following steps S201 and S202:

[0079] Step S201, when the motion feedback data of the two-wheeled foot robot meets the preset conditions, determine that the robot has encountered a height obstacle.

[0080] Among them, the motion feedback data can represent the data information reflecting the current dynamic state of the two-wheeled foot robot, which is collected and transmitted back by its internal sensing components during operation. In the embodiments of the present disclosure, the motion feedback data can include wheel-foot rotation speed data, robot attitude data, torque data, and other parameters that can reflect the motion behavior of the robot. For example, the wheel-foot rotation speed data can represent the motor rotation speed information collected in real time by the motor encoder installed on the drive motor of the two-wheeled foot robot, which can reflect the rotation of the two-wheeled foot structure per unit time. The robot attitude data can include the spatial poses of various parts of the robot body, including position, direction, attitude angle, etc., which are used to describe the overall motion state of the robot and can be obtained by the inertial measurement unit (IMU) set on the robot body. The torque data can represent the drive torque value measured by the current sensor in the robot drive system and calculated in combination with the motor characteristics, which can characterize the force conditions of each drive wheel of the two-wheeled foot robot during motion.

[0081] The height obstacle can represent a terrain structure or object with a vertical height feature encountered in the traveling path of the two-wheeled foot robot, which causes the movement of the robot's wheel feet to be blocked or unable to directly cross. For example, the height obstacle can be a step, a vertical wall, a protruding platform, a ground height difference, or other types of height obstacles. Figure 3 Fig. shows a schematic diagram of the scene during the movement of a two-wheeled foot robot towards a height obstacle in the embodiments of the present disclosure. Among them, during the movement of the robot towards the height obstacle, its internal sensing components obtain motion feedback data in real time. When the motion feedback data meets the preset conditions, it can be determined that the robot has encountered a height obstacle and needs to perform an obstacle-crossing action.

[0082] The preset condition can represent the determination criterion preset in the control system of the two-wheeled foot robot for determining whether there is a height obstacle, and it can be set based on the change of one or more parameters in the motion feedback data. In some exemplary embodiments, the preset condition can be that the difference between the target rotation speed and the actual rotation speed of the two-wheeled feet exceeds a preset threshold. Among them, the actual rotation speed can be collected by the motor encoder installed on the two-wheeled foot robot, and the target rotation speed can be the expected rotation speed command sent by the controller to the drive motor of the two-wheeled foot robot. In the normal driving state, the target rotation speed and the actual rotation speed of the two-wheeled feet are basically the same. However, when the robot encounters a height obstacle during driving, due to the sudden blockage of the height obstacle, the rotation of the wheel feet will be blocked, resulting in a significant decrease in the actual rotation speed, and thus a large deviation from the target rotation speed. Once this deviation exceeds the set preset threshold, it usually indicates that the wheel feet cannot reach the expected motion state per unit time due to encountering a height obstacle. Therefore, the difference between the target rotation speed and the actual rotation speed of the two-wheeled feet exceeding the preset threshold can be used as the trigger judgment condition for the robot to encounter a height obstacle.

[0083] In some exemplary embodiments, the above preset threshold may be 30% of the target rotational speed, that is, when the difference between the actual rotational speed of the two-wheel feet and the target rotational speed exceeds 30% of the target rotational speed, the preset condition is satisfied. Of course, in other embodiments of the present disclosure, the preset threshold may also be other suitable values such as 20%, 24%, 26%, 28%, 32%, 34%, and 36% of the target rotational speed.

[0084] In addition, in some other exemplary embodiments, the preset condition for determining whether the robot encounters a height obstacle may further include: when the duration during which the difference between the target rotational speed and the actual rotational speed of the two-wheel feet exceeds the preset threshold exceeds the preset duration. Exemplarily, the preset duration is greater than 200 milliseconds, which is used to avoid misjudgment caused by short-term fluctuations, so as to more accurately determine that the two-wheel foot robot encounters a height obstacle. Of course, in other embodiments of the present disclosure, the specific value of the preset duration can be flexibly adjusted according to the response speed requirements of the robot application scenario, the sampling frequency of the control system, and the environmental interference characteristics. For example, in a high-dynamic response scenario, the preset duration can be set between 100 milliseconds and 300 milliseconds, so as to improve the determination robustness of whether the robot encounters a height obstacle while ensuring the detection sensitivity.

[0085] Step S202: Control the two-wheel foot robot to perform a takeoff action. The takeoff action includes the main body squatting down and then rising rapidly, and control the two-wheel foot structure of the two-wheel foot robot to rotate forward rapidly while attaching to the height obstacle until the motion feedback data of the two-wheel foot robot does not meet the preset condition; wherein, the takeoff action is output by a policy network trained based on a deep reinforcement learning algorithm in a height obstacle terrain.

[0086] In this step, when the robot detects a height obstacle, it controls the robot to perform a takeoff action. This takeoff process includes the main body squatting down and then rising rapidly, so that the robot can accumulate enough kinetic energy in a short time to complete the obstacle crossing. In addition, during the rising process of the robot main body, the two-wheel foot structure attaches to the obstacle surface and rotates forward rapidly, improving the fitting degree of the robot to the obstacle surface, enabling the robot to maintain a stable obstacle crossing ability under different obstacle forms, and thus reducing the risk of obstacle crossing failure. In this process, the takeoff action is output by a policy network trained based on a deep reinforcement learning algorithm, enabling the robot to dynamically adjust the rotational speed and contact mode of the wheel feet during the motion process, improving the fitting degree of the robot to the obstacle surface, enabling the robot to maintain a stable obstacle crossing ability under different obstacle forms, and thus reducing the risk of obstacle crossing failure, and avoiding the problem of unstable execution caused by environmental changes in the related art. In addition, through the control strategy trained by deep reinforcement learning, the robot can continuously optimize its own motion mode during the training process, form the motion decision-making ability to adapt to complex environments, and reduce the dependence on high-precision sensors and complex computing resources.

[0087] Among them, the takeoff action can represent the action behavior of the two-wheeled foot robot to achieve the purpose of obstacle crossing after detecting a height obstacle. This action includes, but is not limited to, briefly squatting down to store energy by lowering the center of gravity, and then the drive system outputs an upward impact force to complete the rapid rise of the robot body, and combines with the wheel-foot movement to form an overall obstacle-crossing process. The two-wheeled foot structure can represent a motion execution mechanism configured at the bottom of the robot, which has dual functions of rolling driving and attaching support. This structure can include at least two active wheel systems or wheel-foot components with the compound ability of traveling and supporting, and can effectively contact and friction with the obstacle surface during the takeoff and climbing processes. The height obstacle terrain can represent an obstacle area with certain height mutation characteristics. Figure 4 Fig. shows a schematic diagram of the scene of the two-wheeled foot robot during the leaping process. During the leaping process, the two-wheeled foot structure of the two-wheeled foot robot adheres to the height obstacle and rotates forward quickly to enhance the friction between the wheel feet and the obstacle surface. Figure 5 Fig. shows a schematic diagram of the scene of the two-wheeled foot robot stably landing on the height obstacle in an embodiment of the present disclosure. Among them, when the two-wheeled foot robot completes obstacle crossing and is above the height obstacle, the motion feedback data of the two-wheeled foot robot does not meet the preset conditions, and the two-wheeled foot structure of the two-wheeled foot robot stops rotating forward at high speed and returns to the normal driving state to maintain attitude balance and position stability.

[0088] The deep reinforcement learning algorithm can represent an optimal policy search method that combines a deep neural network and a reinforcement learning mechanism. It can continuously update the control policy based on the state of the robot, environmental feedback data, and reward function to improve the decision-making ability of the robot when passing through height obstacles. Among them, deep reinforcement learning includes, but is not limited to, algorithms such as PPO (Proximal Policy Optimization), DDPG (Deep Deterministic Policy Gradient), and SAC (Soft Actor-Critic). The present disclosure does not limit the specific algorithm types of deep reinforcement learning. The policy network can represent a neural network model trained based on deep reinforcement learning. The input of this model is the perception data obtained by the two-wheeled foot robot, and the output is the action decision for controlling the movement of the robot, such as the drive power distribution during takeoff, the rotation speed adjustment of the two-wheeled feet, and the attitude control during the obstacle contact process, so as to achieve stable obstacle crossing in the height obstacle terrain.

[0089] In some exemplary embodiments, refer to Figure 6 As shown, the specific implementation process of the rapid upward movement of the main body of the two-wheeled foot robot after squatting includes the following steps S601 and S602:

[0090] Step S601, control the main body of the two-wheeled foot robot to briefly squat down by lowering the center of gravity and store energy.

[0091] Exemplarily, this step can be implemented in the following manner: By controlling the main body of the two-wheeled foot robot to lower the center of gravity in the vertical direction, a brief squatting process is achieved. This squatting action can be realized by driving the mechanism to reduce the output speed of the wheel feet or adjusting the posture angle of the robot. During the squatting process, the two-wheeled foot robot can apply a preload to the ground by compressing elastic elements or using the reaction force of the wheel foot system, thereby forming a reserve of mechanical energy inside the leg structure. Of course, the above-mentioned squatting and energy storage steps can also be implemented in the following way: By controlling the hip joint motor and the knee joint motor to rotate in coordination, driving the thigh component and the calf component to fold inward, making the lower limb structure of the robot present a bent state, thereby realizing the downward movement of the overall center of gravity of the robot in the vertical direction and completing a brief squatting action. During this process, the robot forms an actively controlled joint bend by driving the leg joints, and combines the contact reaction force between the roller and the ground to apply a preload to the leg structure, thereby forming a reserve of mechanical energy available for takeoff. This energy reserve provides a power basis for subsequent takeoff actions and helps to improve the takeoff height and takeoff stability of the robot.

[0092] Step S602, provide an upward thrust through the two-wheel drive mechanism to make the robot take off from the ground.

[0093] Exemplarily, the takeoff of the two-wheeled foot robot from the ground in this step can be achieved according to the following method: By increasing the rotational speed of the wheel feet through the two-wheel drive mechanism, applying a continuous and concentrated reaction force to the ground, thereby providing an upward thrust for the robot to prompt the main body of the robot to leave the ground and achieve takeoff. This thrust can be precisely controlled by the drive mechanism according to the output instructions of the policy network, providing sufficient vertical kinetic energy to ensure the completion of the takeoff action at the right time and intensity. Of course, in other embodiments of the present disclosure, during the above-mentioned takeoff process, while the rotational speed of the wheel feet is increased, the hip joint motor and the knee joint motor can be controlled to quickly unfold in the reverse direction in coordination, so that the thigh component and the calf component can quickly extend, thereby superimposing the accumulated mechanical energy and the driving acceleration to form a greater takeoff thrust, thereby increasing the success rate of takeoff and obstacle crossing.

[0094] In some exemplary embodiments, the specific process of the two-wheeled foot structure quickly rotating forward while attaching to a high obstacle includes: adjusting the driving torque of the two-wheeled foot to actively rotate forward quickly at a speed higher than the normal driving speed to generate effective friction to assist the two-wheeled foot robot to climb.

[0095] Among them, the driving torque can represent the torque output by the driving motor of the bipedal wheeled robot for driving the wheel feet to rotate, and the magnitude of the driving torque can be dynamically adjusted according to the output of the policy network. The normal driving speed can represent the stable speed maintained by the bipedal wheels of the robot when the bipedal wheeled robot is driving conventionally in an obstacle-free environment. Exemplarily, the bipedal wheel structure can rotate at 30% higher than the normal driving speed. Of course, in other embodiments of the present disclosure, the rotation speed of the bipedal wheel structure can also be dynamically adjusted according to the material, slope, or adhesion difficulty of the high obstacle surface, and its rotation speed can be higher than other appropriate values such as 32%, 34%, 36%, 38%, 40%, 42%, and 45% of the normal driving speed to ensure smooth climbing on different types of high obstacles. In this embodiment, by adjusting the driving torque, the bipedal wheels rotate at a speed higher than the normal driving state when attaching to the high obstacle, thereby enhancing the friction with the obstacle surface. This process not only helps to improve the adhesion effect of the wheel feet but also provides additional propulsion force to enable the robot to climb stably on a vertical or nearly vertical surface. By dynamically regulating the magnitude of the driving torque through the policy network, it is possible to achieve adaptive adjustment to the surface characteristics of different types of obstacles, further enhancing the reliability and success rate of the robot during obstacle crossing.

[0096] In some exemplary embodiments, the policy network can be trained by the following method: In a simulation environment with a high obstacle terrain, the policy network for outputting the action policy for controlling the movement of the robot is trained using the teacher-student model framework.

[0097] Among them, the teacher-student model framework can represent a supervised learning structure for training the policy network. This framework includes two sub-networks, a teacher encoder and a student encoder. Among them, the teacher encoder provides stable feature supervision for the student encoder, guiding the student encoder to learn a higher-quality state representation, so that the policy network has a better convergence speed and generalization ability during training.

[0098] In addition, in some embodiments, the height obstacle terrain may include various types of height obstacles. There are differences in at least one of the parameters among different types of height obstacles, such as the overall height of the obstacle, the surface roughness of the obstacle sidewall, the stiffness of the obstacle sidewall material, the geometry of the obstacle top, the flatness of the obstacle sidewall surface, and the inclination angle between the obstacle sidewall and the ground. Among them, the overall height of the obstacle can represent the vertical distance from the position where the bottom of the obstacle touches the ground to the highest point at the top. The surface roughness of the obstacle sidewall can represent the degree of texture undulation or friction characteristics of the obstacle sidewall surface, which can be quantified by the average deviation of the surface profile or the friction coefficient. The stiffness of the obstacle sidewall material can represent the ability of the material constituting the obstacle to resist deformation under the action of force, which can be quantified by Young's modulus or structural hardness. The geometry of the obstacle top can represent the spatial morphological characteristics of the obstacle top region, including but not limited to other shapes such as flat, arc, and sharp corner. The flatness of the obstacle sidewall surface can represent the degree of obvious unevenness, discontinuity, or structural mutation in the obstacle sidewall. The inclination angle between the obstacle sidewall and the ground can represent the angle between the obstacle sidewall and the horizontal ground.

[0099] Of course, in other embodiments of the present disclosure, there may also be differences in other obstacle-related parameters that affect the stability of the robot's obstacle-crossing action among different types of height obstacles, such as the width of the obstacle, the sharpness of the obstacle top edge, the surface elasticity of the obstacle material, and the slipperiness of the obstacle surface. Among them, the width of the obstacle can represent the lateral extension distance of the obstacle perpendicular to the forward direction of the robot. The sharpness of the obstacle top edge can represent the geometric transition degree of the top boundary of the obstacle. The surface elasticity of the obstacle material can represent the ability of the obstacle to return to its original state after local deformation under the action of force. By training the policy network using the teacher-student model framework in the above simulation environment containing various types of height obstacles, the student model can have stronger policy generalization ability and environmental adaptation ability when facing obstacles with different heights, materials, surface shapes, or slopes, effectively improving the success rate and stability of the bipedal wheeled robot in crossing obstacles in a real and complex environment.

[0100] Furthermore, in some embodiments, the deep reinforcement learning algorithm for training the policy network includes a comprehensive reward function for evaluating the action performance of the bipedal wheeled robot in the simulation environment. This reward function includes the success rate of crossing the height obstacle, the takeoff height, and the attitude stability. This comprehensive reward function is used to evaluate the action performance of the bipedal wheeled robot in completing the obstacle-crossing task in the simulation environment, and its definition is as follows:

[0101]

[0102] In the above formula, represents the comprehensive reward value, which is used to measure the overall performance of the robot in completing the obstacle-crossing task in the current training round. Indicates the success rate reward for passing over the height obstacle, Indicates the takeoff height reward, Indicates the attitude stability reward, respectively represent the weight coefficients of each reward, satisfying , and their values can be adjusted according to the key points of the task concern.

[0103] Among them, if the robot successfully crosses the obstacle and lands stably, the success rate reward for passing over the height obstacle is 1, otherwise it is 0.

[0104] The takeoff height reward can be expressed as:

[0105]

[0106] Among them, represents the maximum vertical height from the ground after the robot takes off, represents the minimum effective takeoff height required for the current obstacle.

[0107] The attitude stability reward can be expressed as:

[0108]

[0109] Among them, represents the duration of the robot's obstacle-crossing process, represents the pitch angle deviation of the robot at time , represents the roll angle deviation of the robot at time .

[0110] In addition, in other embodiments of the present disclosure, the above comprehensive reward function for evaluating the action performance of the two-wheeled bipedal robot in the simulation environment may further include other parameters for evaluating the action efficiency and stability of the two-wheeled bipedal robot during the obstacle-crossing process, such as energy consumption, landing impact magnitude, motion trajectory smoothness, etc., and this embodiment is not specifically limited herein.

[0111] Reference Figure 7As shown in the figure, a schematic diagram of a two-stage training-inference framework for a policy network is presented. Among them, the training stage combines a teacher-student encoder structure with a deep reinforcement learning framework. The teacher-student encoder structure includes a teacher encoder 701 and a student encoder 702. The deep reinforcement learning framework consists of a policy network 703 and a value network 704. Among them, the value network can represent a neural network structure for estimating the value of the current state, and this value estimation can be used to guide the policy optimization in the reinforcement learning process to improve the behavior quality and training stability of the policy network. In the training stage, by introducing privileged state information to guide policy learning, the student encoder 702 can stably output high-quality control actions only relying on observable data in the inference stage. The observable data includes the current motion state information of the two-wheeled foot robot.

[0112] It should be noted that the teacher encoder 701 is only used to extract the high-dimensional semantic representation in the privileged state information during the training stage for subsequent use as the input of the policy to guide learning. The student encoder 702 encodes the historical motion state of the robot itself during the training and inference stages to obtain dynamic features related to action decisions, and then inputs the current motion state information into the policy network 703 to generate control actions. In addition, the student encoder 702 ultimately needs to imitate the representation output by the teacher encoder 701, so distillation training is performed by minimizing the MSE (mean squared error) during the training process.

[0113] The policy network 703 is the core module for action generation. During the training stage, the policy network 703 receives the latent vector from the teacher encoder 701 or the student encoder 702 and the current motion state information, and combines with the value network 704 to update the policy using the PPO algorithm. In the inference stage, the policy network 703 can receive the latent vector from the student encoder 702 and the current motion state information, and output the action policy for controlling the motion of the two-wheeled foot robot.

[0114] Reference Figure 8 As shown in the figure, a schematic diagram of another two-stage training-inference framework for a policy network is presented. Figure 7 and Figure 8 The training framework shown in Figure 8Both the teacher encoder 701 and the student encoder 702 in it are only used in the training phase. After the training of this architecture is completed, the policy network 703 can be trained for different height obstacle terrains. The policy network 703 already has the ability to perform perception and understanding and action output on different height obstacle terrains. Therefore, in the inference phase, there is no need to use the student encoder 702 as an auxiliary module, and the current motion state information can be directly received through the policy network 703, and the action policies corresponding to the robot's takeoff action and the rapid forward rotation action of the dual-wheel foot structure attached to the height obstacle can be output. This example greatly improves the efficiency of the inference phase and the simplicity of deployment, while ensuring the high adaptability and stability of the policy in different height obstacle terrain scenarios.

[0115] Based on Figure 7 the framework schematic diagram shown, refer to Figure 9 as shown, the process of training the policy network for outputting the action policy for controlling the robot's takeoff action may include the following steps S901 to step S906:

[0116] Step S901, encoding the historical motion state information of the dual-wheel foot robot itself through the student encoder to generate a first latent vector, and encoding the privileged state information of the dual-wheel foot robot through the teacher encoder to generate a second latent vector; the privileged state information includes height obstacle information obtained from the simulation environment.

[0117] Among them, the historical motion state information of the robot itself may include attitude sensor data, wheel foot rotation speed data, torque data, and position and speed data, sole contact state, center of gravity trajectory, etc. of the past several frames. For example, the historical motion state sequence of the past 10 frames is input into the student encoder 702 for encoding to obtain a first latent vector. The first latent vector can capture the continuity and dynamic characteristics of the robot's actions.

[0118] The privileged state information may include height obstacle information obtained from the simulation environment. The height obstacle information may represent a set of relevant parameters describing the spatial structure and physical properties of the obstacle in the simulation environment, and the information is used to assist the teacher model in generating a better control strategy. The height obstacle information may include, but is not limited to, the overall height of the obstacle, the surface roughness of the obstacle side wall, the material stiffness of the obstacle side wall, the geometric shape of the obstacle top, the surface flatness of the obstacle side wall, the inclination angle between the obstacle side wall and the ground, the edge shape of the obstacle, and the relative position relationship with the robot, etc. The privileged state information is input into the teacher encoder 701 for encoding to obtain a second latent vector. The second latent vector can highly concentrate the key semantics in the height obstacle information, which helps to construct a more complete high-dimensional feature representation of the height obstacle information.

[0119] It should be noted that the privileged status information can only be obtained during the training phase but is not observable during the testing or deployment phase. Therefore, the teacher encoder 701 can utilize the complete information to learn the optimal latent representation, thereby guiding the student encoder 702 to learn. Both the teacher encoder 701 and the student encoder 702 can compress the high-dimensional, temporal state information into low-dimensional latent representations, providing behavioral semantic representations from different sources for the policy network 703. For example, the teacher encoder 701 and the student encoder 702 can be multi-layer perceptrons, temporal convolutional networks, or recurrent neural networks, with the ability to extract temporal features and compress them into fixed-length semantic vectors. Additionally, the network architectures of the teacher encoder 701 and the student encoder 702 can be the same or different, and the present disclosure does not limit this.

[0120] Step S902, select the first latent vector or the second latent vector according to a preset policy and input it into the policy network.

[0121] According to a preset policy, select to send the first latent vector generated by the student encoder 702 or the second latent vector generated by the teacher encoder 701 into the policy network 703 for decision-making. Among them, the preset policy can be set according to the real-time environmental state, the training phase, or specific performance metrics. The specific performance metrics include action execution error thresholds, simulation and real environment difference degrees, etc.

[0122] For example, one preset policy is to preferentially use the high-quality second latent vector generated by the teacher encoder 701 at the beginning of training to guide the policy network 703 to quickly converge to an approximate optimal solution. When the student encoder 702 is optimized through knowledge distillation, it gradually transitions to only using the first latent vector during the deployment phase to reduce the dependence on privileged information. Another example, another preset policy is to p select the second latent vector in proportion, and p select the first latent vector, and gradually reduce p the value of. Of course, the first latent vector and the second latent vector can also be concatenated and sent into the policy network 703, and dynamically weighted through an attention mechanism or a gating module, enabling the policy network 703 to flexibly combine historical experience and privileged knowledge in complex scenarios.

[0123] The selective input mechanism can not only accelerate the training process by utilizing the ideal state prior knowledge provided by the teacher encoder 701, but also cope with sensor limitations or environmental disturbances through the generalization ability of the student encoder 702 during actual operation. At the same time, through the conditional processing of the latent vector by the policy network 703, smooth switching of motion control and robust decision-making are achieved. For example, when the robot encounters an obstacle of unknown height, it preferentially adjusts the takeoff action and the attachment action of the bipedal wheel structure based on the second latent vector, while relying on the first latent vector to maintain efficiency during the stable walking phase. Finally, the policy selection rule is optimized through closed-loop feedback, enabling the robot to balance obstacle-crossing performance and adaptability under different stages and environmental conditions.

[0124] In addition to the selected latent vector, the current motion state information of the robot can also be input into the policy network 703 for decision-making.

[0125] Step S903, based on the received latent vector and the current motion state information, the policy network outputs an action policy for controlling the motion of the bipedal wheel robot. The current motion state information may include attitude sensor data, wheel-foot rotation speed data, torque data, and position and velocity data.

[0126] Among them, the policy network 703 can be a multi-layer fully connected perceptron, a Transformer structure, etc. The policy network 703 performs multi-modal feature fusion on the first latent vector or the second latent vector and the current motion state information, and outputs an action policy, denoted as .

[0127] For example, the first latent vector or the second latent vector and the current motion state information are concatenated or weighted interacted in the embedding space, and action policies such as joint angle targets, torque commands, or gait phase parameters are generated through non-linear transformation.

[0128] Step S904, the value network utilizes the privileged state information to output an estimated value of the current state.

[0129] During the training process of the policy network 703, the value network 704 is also used to estimate the long-term return of the current state, that is, starting from this state, if the current policy is continuously executed, how much cumulative reward can be obtained in the future, in order to optimize the policy network 703.

[0130] Among them, the value network 704 can be a multi-layer fully connected perceptron. For example, when the value network 704 utilizes the privileged state information to estimate the value of the current state, first the privileged state information is encoded into a high-dimensional feature vector, such as extracting dynamic features through convolutional or fully connected layers, then fused with the current motion state information in the latent space, and subsequently the estimated value representing the current state is output through multi-layer non-linear transformation, denoted as This value estimation is used for policy optimization in reinforcement learning to guide the policy network 703 to learn better behaviors.

[0131] Step S905: Optimize the parameters of the student encoder based on the difference between the first latent vector and the second latent vector.

[0132] This step guides the student encoder 702 to learn to generate latent representations close to those of the teacher encoder 701, achieving teacher knowledge distillation by quantifying the distribution difference between the two in the latent space and backpropagating the difference gradient to update the network weights of the student encoder 702.

[0133] Exemplarily, during the training phase, the parameters of the teacher encoder 701 are fixed. After using the historical motion state information and the corresponding motion state information as parallel inputs to generate latent vectors through the student encoder 702 and the teacher encoder 701 respectively, the distance between the two is minimized using a contrastive learning framework, or the output distribution of the student encoder 702 is made to approximate the latent space characteristics of the teacher encoder 701 through adversarial training. Meanwhile, noise injection or data augmentation is introduced to simulate sensor errors in actual deployment, forcing the student encoder 702 to still extract feature expressions compatible with the privileged information encoding under limited input information conditions.

[0134] For example, to make the output of the student encoder 702 as close as possible to that of the teacher encoder 701, a difference loss function can be constructed, such as the least squares gap or KL divergence. By minimizing this loss, the parameters of the student encoder 702 are optimized, enabling it to learn to extract feature expressions close to the privileged information from the historical motion states, thereby enhancing the generalization ability of the policy network 703, improving the decision-making quality of the policy network 703 in the real environment, and enabling it to approach the teacher level without relying on privileged information during deployment.

[0135] Step S906: Update the parameters of the policy network using a deep reinforcement learning algorithm based on the action policy and the value estimation.

[0136] Taking the PPO algorithm as an example of the deep reinforcement learning algorithm. During the training process of the policy network 703, the current policy network 703 is used to interact with the environment. Based on the current state an action policy is selected , and after executing the action, a reward and the next state are returned. From this, an interaction trajectory sequence ( ) is sampled for subsequent policy optimization. Then, a value network 704 is introduced to evaluate the value of each state . To optimize the policy more stably and efficiently, the policy gradient can be calculated through the advantage function to obtain the parameter update direction of the policy network 703.

[0137] For example, the advantage function is defined as:

[0138]

[0139] where is the advantage function, is the approximate value of the advantage obtained using Generalized Advantage Estimation.

[0140] Then, an optimization objective is constructed according to the clipping objective function of the PPO algorithm, and the policy gradient is calculated through backpropagation to optimize the parameters of the policy network 703, so that the expected cumulative return is continuously improved under the current policy. At the same time, the state-return pair ( ) is used for supervised learning of the value network 704 to minimize the mean square error between its output and the true return , thereby improving the estimation accuracy of the value network 704 for the state value.

[0141] The student encoder 702 performs latent space alignment with the output of the teacher encoder 701 through contrastive learning or knowledge distillation, forcing the student network to still be able to infer a latent representation close to the encoding result of the privileged information through historical states in the actual deployment scenario lacking privileged information, thereby improving the robustness and adaptability of the motion strategy. In the exemplary embodiment of the present disclosure, the robot can simulate the motion decision optimized based on the privileged information only relying on its own sensor data in the real environment, effectively solving the problem of performance degradation caused by unobservable partial states or noise interference during actual deployment. At the same time, the generalization ability of the motion strategy is enhanced through the implicit knowledge transfer of the latent vector.

[0142] In the exemplary embodiment of the present disclosure, a method for training a motion control model of a two-wheeled foot robot based on deep reinforcement learning. Referring to Figure 10 as shown, the method may include the following steps S1001 to step S1002:

[0143] Step S1001, construct a virtual training scenario including a high obstacle terrain, and obtain the motion data of the virtual two-wheeled foot robot in the virtual training scenario.

[0144] In this step, by constructing a virtual training scenario containing various types of height obstacles, the motion environment of the two-wheeled foot robot in a real complex terrain is simulated. The virtual training scenario can include obstacles with various parameter combinations such as height, surface roughness, material stiffness, top geometry, surface flatness, slope angle, etc., to enhance the training diversity and the generalization ability of the strategy. Deploy a virtual two-wheeled foot robot in this environment and record the motion data during its movement. The motion data can include attitude sensor data, wheel-foot rotation speed data, torque data, position and speed data, etc., which are used for input modeling and performance evaluation of the policy network during subsequent training.

[0145] Step S1002: Use a deep reinforcement learning algorithm to train a policy network based on the motion data and environmental information; the policy network outputs the following action strategy for the two-wheeled foot robot: when encountering a height obstacle, execute a takeoff action, and the takeoff action includes quickly rising after the main body squats down, and controlling the two-wheeled foot structure of the two-wheeled foot robot to rotate forward quickly while attaching to the height obstacle until crossing the height obstacle.

[0146] In this step, based on the motion data obtained in step S1001 and the environmental information in the virtual training scenario, a deep reinforcement learning algorithm is used to train the policy network. The deep reinforcement learning can be other suitable deep reinforcement learning algorithms such as the PPO algorithm, DDPG algorithm, SAC algorithm, etc., which are used to optimize the robot control strategy in a high-dimensional continuous action space. During the training process, the policy network takes the current state data of the robot as input, including perception information such as attitude sensor data, wheel-foot rotation speed data, torque data, and position and speed data, and outputs the corresponding action strategy. This action strategy can be used to guide whether the robot executes a takeoff action when detecting a height obstacle and how to control various parameters during the takeoff process. Specifically, the policy network can output a takeoff start signal, the squat amplitude, the rising speed, and the control instructions for the forward rotation speed of the wheel-foot structure, so as to achieve a stable and efficient obstacle-crossing behavior of the robot.

[0147] In the exemplary embodiment of the present disclosure, a height obstacle control device for a two-wheeled foot robot based on deep reinforcement learning is also provided. Refer to Figure 11 As shown, the height obstacle control device 1100 for a two-wheeled foot robot based on deep reinforcement learning may include:

[0148] An obstacle determination module 1101, which can be used to determine that the robot encounters a height obstacle when the motion feedback data of the two-wheeled foot robot meets a preset condition;

[0149] The motion control module 1102 can be used to control the two-wheeled foot robot to perform a takeoff motion. The takeoff motion includes the main body quickly rising after squatting down, and controlling the two-wheeled foot structure of the two-wheeled foot robot to rotate forward quickly while attaching to the height obstacle until the motion feedback data of the two-wheeled foot robot does not meet the preset conditions. Among them, the takeoff motion is output by a policy network trained based on a deep reinforcement learning algorithm in a height obstacle terrain.

[0150] The specific details of the above obstacle determination module 1101 and motion control module 1102 have been described in detail in the corresponding method for controlling a two-wheeled foot robot through height obstacles based on deep reinforcement learning, so they will not be elaborated here.

[0151] In the exemplary embodiment of the present disclosure, a training device for a motion control model of a two-wheeled foot robot based on deep reinforcement learning is also provided. Refer to Figure 12 As shown, the training device 1200 for a motion control model of a two-wheeled foot robot based on deep reinforcement learning may include:

[0152] A training data acquisition module 1201, which can be used to construct a virtual training scene including a height obstacle terrain and acquire the motion data of a virtual two-wheeled foot robot in the virtual training scene.

[0153] A model training module 1202, which can be used to train a policy network according to the motion data and environmental information by using a deep reinforcement learning algorithm. The policy network outputs the following motion strategies for the two-wheeled foot robot: when encountering a height obstacle, perform a takeoff motion, and the takeoff motion includes the main body quickly rising after squatting down, and controlling the two-wheeled foot structure of the two-wheeled foot robot to rotate forward quickly while attaching to the height obstacle until crossing the height obstacle.

[0154] The specific details of the above training data acquisition module 1201 and model training module 1202 have been described in detail in the corresponding method for training a motion control model of a two-wheeled foot robot based on deep reinforcement learning, so they will not be elaborated here.

[0155] In the exemplary embodiment of the present disclosure, a robot is also provided. The robot includes a processor and a memory, and computer-readable instructions are stored on the memory. When the computer-readable instructions are executed by the processor, the above method is implemented. Among them, the robot is a robot configured with a function of controlling a two-wheeled foot robot through height obstacles based on deep reinforcement learning. Exemplarily, the robot includes any one of a humanoid robot, a cleaning robot, a transportation robot, and a mobile robot. Refer to Figure 13 As shown, it shows a schematic diagram of a two-wheeled foot robot.

[0156] Refer to Figure 14As shown, an electronic device capable of implementing the above method is also provided. Among them, the electronic device 1400 includes a processor 1401 and a memory 1402. A computer-readable instruction is stored on the memory 1402. When the computer-readable instruction is executed by the processor 1401, it implements the method for controlling a two-wheeled biped robot through height obstacles or the method for training a motion control model of a two-wheeled biped robot based on deep reinforcement learning in the embodiments of the present disclosure.

[0157] In an exemplary embodiment of the present disclosure, a computer-readable storage medium is also provided, on which computer program code instructions are stored. When the computer program code instructions are called by the processor of the robot, the robot is enabled to execute the method as described in the embodiment.

[0158] Reference Figure 15 As shown, a program product 1500 for implementing the above method for controlling a two-wheeled biped robot through height obstacles or the method for training a motion control model of a two-wheeled biped robot based on deep reinforcement learning according to an embodiment of the present disclosure is described. It can be a portable compact disc read-only memory (CD-ROM) and includes program code and can run on a terminal device. However, the program product of the present disclosure is not limited to this. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or device.

[0159] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0160] Finally, the above preferred embodiments are only used to illustrate the technical solutions of the present application and are not restrictive. Although the present application has been described in detail, those skilled in the art should understand that changes in form and details can be made to it without departing from the scope defined by the claims of the present application. The dimensions of the drawings have nothing to do with the specific physical objects, and the physical dimensions can be arbitrarily changed.

Claims

1. A height obstacle control method for a two-wheeled foot robot based on deep reinforcement learning, characterized in that, Including: When the motion feedback data of the two-wheeled foot robot meets the preset conditions, it is determined that the robot encounters a height obstacle; Controlling the two-wheeled foot robot to perform a takeoff action, the takeoff action includes the main body squatting down and then rising rapidly, and controlling the two-wheeled foot structure of the two-wheeled foot robot to adhere to the height obstacle and rotate forward rapidly at a speed higher than the normal driving speed, thereby generating effective friction to assist the two-wheeled foot robot to climb until the motion feedback data of the two-wheeled foot robot does not meet the preset conditions; Among them, the takeoff action is output by a policy network trained based on a deep reinforcement learning algorithm in a height obstacle terrain.

2. The height obstacle control method for the two-wheeled foot robot according to claim 1, characterized in that, The motion feedback data includes wheel-foot rotation speed data, and the preset conditions are: the difference between the target rotation speed and the actual rotation speed of the two-wheeled feet exceeds a preset threshold.

3. The height obstacle control method for the two-wheeled biped robot according to claim 2, characterized in that, The actual rotation speed is obtained by collecting through a motor encoder installed on the two-wheeled foot robot; the target rotation speed is an expected rotation speed command sent by the controller to the drive motor of the two-wheeled foot robot.

4. The height obstacle control method for the two-wheeled foot robot according to claim 2, characterized in that, The difference between the target rotation speed and the actual rotation speed of the two-wheeled feet exceeding the preset threshold includes: the difference between the target rotation speed and the actual rotation speed of the two-wheeled feet exceeds 30% of the target rotation speed.

5. The method for controlling a two-wheeled foot robot to pass through a height obstacle according to claim 2, wherein The preset conditions also include: when the duration of the difference between the target rotation speed and the actual rotation speed of the two-wheeled feet exceeding the preset threshold exceeds a preset duration.

6. The height obstacle control method for the two-wheeled foot robot according to claim 5, characterized in that, The preset duration is greater than 200 milliseconds.

7. The height obstacle control method for the two-wheeled foot robot according to claim 1, characterized in that The main body squatting down and then rising rapidly includes: Controlling the main body of the two-wheeled foot robot to squat down briefly by lowering the center of gravity and store energy; Providing an upward thrust through a two-wheel drive mechanism to make the robot take off from the ground.

8. The method for controlling a two-wheeled foot robot to pass through a height obstacle according to claim 1, characterized in that, The two-wheeled foot structure adhering to the height obstacle and rotating forward rapidly at a speed higher than the normal driving speed includes: Adjusting the driving torque of the two-wheeled feet to perform active rapid forward rotation at a speed higher than the normal driving speed.

9. The two-wheeled foot robot height obstacle control method according to any one of claims 1 to 8, characterized in that The policy network is trained through the following method: In a simulation environment with a height obstacle terrain, using a teacher-student model framework, training the policy network for outputting an action policy for controlling the movement of the robot.

10. The height obstacle control method for the two-wheeled foot robot according to claim 1, wherein The height obstacle terrain includes various types of height obstacles; there are differences in at least one of the following parameters between different types of the height obstacles: Overall height of the obstacle, surface roughness of the obstacle side wall, material stiffness of the obstacle side wall, geometric shape of the obstacle top, surface flatness of the obstacle side wall, inclination angle between the obstacle side wall and the ground.

11. The height obstacle control method for the two-wheeled foot robot according to claim 9, characterized in that, The training process specifically includes: Encoding the historical motion state information of the two-wheeled foot robot by the student encoder to generate a first latent vector, and encoding the privileged state information of the two-wheeled foot robot by the teacher encoder to generate a second latent vector; the privileged state information includes height obstacle information obtained from the simulation environment; Selecting the first latent vector or the second latent vector according to a preset policy and inputting it into the policy network; Based on the received latent vector and the current motion state information, the policy network outputs an action policy for controlling the movement of the two-wheeled foot robot; the current motion state information at least includes: attitude sensor data, wheel-foot rotation speed data, torque data, and position and speed data; Utilize privileged state information through the value network and output a value estimate of the current state; optimizing parameters of the student encoder based on a difference between the first latent vector and the second latent vector; Based on the action strategy and the value estimate, a deep reinforcement learning algorithm is used to update the parameters of the policy network.

12. The height obstacle control method for the two-wheeled foot robot according to claim 9, characterized in that, The deep reinforcement learning algorithm includes a comprehensive reward function for evaluating the action performance of the two-wheeled leg robot in the simulation environment, and the reward function includes the success rate of passing height obstacles, take-off height and posture stability.

13. A training method for a motion control model of a two-wheeled foot robot based on deep reinforcement learning, which is used to train the policy network in the height obstacle control method of the two-wheeled foot robot described in claim 1, characterized in that, include: Constructing a virtual training scene including a highly obstructed terrain, and obtaining motion data of a virtual two-wheeled legged robot in the virtual training scene; A deep reinforcement learning algorithm is used to train a strategy network based on the motion data and environmental information; the strategy network outputs the following two-wheeled legged robot action strategy: When encountering a height obstacle, a jumping action is performed, which includes the main body squatting and then rising rapidly, and controlling the two-wheeled foot structure of the two-wheeled foot robot to stick to the height obstacle and rotate forward rapidly at a speed higher than the normal driving speed, thereby generating effective friction to assist the two-wheeled foot robot to climb until it crosses the height obstacle.

14. A height obstacle control device for a two-wheeled foot robot based on deep reinforcement learning, characterized in that, include: An obstacle determination module, used for determining that the robot encounters a height obstacle when the motion feedback data of the two-wheeled foot robot meets a preset condition; an action control module, for controlling the two-wheeled foot robot to perform a take-off action, wherein the take-off action includes rapidly rising after the main body squats, and controlling the two-wheeled foot structure of the two-wheeled foot robot to adhere to the height obstacle at a faster forward rotation speed than the normal driving speed, thereby generating effective friction to assist the two-wheeled foot robot in climbing, until the motion feedback data of the two-wheeled foot robot does not meet the preset conditions; The take-off action is output by a strategy network trained based on a deep reinforcement learning algorithm in highly obstacle terrain.

15. A training device for a motion control model of a two-wheeled and legged robot based on deep reinforcement learning, characterized in that, include: A training data acquisition module, used to construct a virtual training scene including a highly obstructed terrain, and to acquire motion data of a virtual two-wheeled legged robot in the virtual training scene; The model training module is used to use a deep reinforcement learning algorithm to train a strategy network based on the motion data and environmental information; the strategy network outputs the following two-wheeled legged robot action strategy: When encountering a height obstacle, a jumping action is performed, which includes the main body squatting and then rising rapidly, and controlling the two-wheeled foot structure of the two-wheeled foot robot to stick to the height obstacle and rotate forward rapidly at a speed higher than the normal driving speed, thereby generating effective friction to assist the two-wheeled foot robot to climb until it crosses the height obstacle.

16. An electronic device, characterized in that, include: processor; as well as A memory having computer-readable instructions stored thereon, wherein the computer-readable instructions, when executed by the processor, implement the method according to any one of claims 1 to 13.

17. A two-wheeled foot robot, characterized in that, include: processor; as well as A memory having computer-readable instructions stored thereon, wherein the computer-readable instructions, when executed by the processor, implement the method according to any one of claims 1 to 13.

18. The robot according to claim 17, characterized in that, The robot is a robot configured with a height obstacle control function for a two-wheeled legged robot based on deep reinforcement learning, and the robot includes any one of a humanoid robot, a cleaning robot, a transportation robot, and a mobile robot.

19. A computer-readable storage medium, characterized in that, Computer program code instructions are stored on the computer-readable storage medium, and when the computer program code instructions are called by a processor of the robot, the robot is caused to execute the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Multi-posture switching two-wheeled robot and working method thereof

    CN115892277A

  • Wheel-legged robot control algorithm based on model prediction and deep reinforcement learning

    CN118192558A

  • Wheel-leg robot wheel-foot switching control method based on BP neural network

    CN116859975A

  • Sleep type obstacle, orientation recognition method thereof and robot control method

    CN119217357A