Method and device for controlling double-wheel-foot robot to pass through height obstacle based on deep reinforcement learning and motion control model training method
By applying deep reinforcement learning algorithms in a double-wheel foot robot, the problem that traditional control methods are difficult to achieve stable and precise control in complex environments is solved, and the robot's autonomy and precise obstacle crossing when there are high obstacles is achieved, and adaptability and control accuracy are improved.
Patent Information
- Application Number
- CN202510527007.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-25
AI Technical Summary
Traditional robot motion control methods are difficult to achieve stable and precise control when facing complex dynamic environments, scenes with severe velocity changes or high-dimensional control tasks, especially in the case of unknown disturbances or partial observation information, and are poor in adaptability.
A double-wheeled foot robot based on deep reinforcement learning uses a high-dimensional barrier control method to extract features of high-dimensional input information through a deep neural network, and learns the optimal control strategy through a strategy optimization algorithm to achieve stable obstacle-surfing of the robot when it is highly barrier-resistant.
It improves the adaptability and control accuracy of the robot in complex environments, can independently adjust its actions to deal with obstacles of different heights, reduces the dependence on complex modeling predictions, and improves the success rate and stability of obstacles.
Smart Images

Figure CN120066054A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical fields of sensors and robotics, and relates to a method and device for controlling a two-wheeled legged robot to pass through height obstacles and a method for training a motion control model based on deep reinforcement learning. Background Art
[0002] In the field of robot motion control, traditional control methods usually rely on model predictive control (MPC), proportional-integral-derivative control (PID), or classical optimization algorithms. However, these methods often struggle to achieve stable and precise control in the face of complex dynamic environments, scenarios with drastic speed changes, or high-dimensional control tasks. Especially in the presence of unknown disturbances or partial observation information, traditional methods have poor adaptability and are difficult to quickly adjust strategies, thus affecting the motion performance of robots in complex environments.
[0003] For example, in the Chinese patent application with publication number CN115892277A, when an obstacle appears on the ground, the hip joint motor and the knee joint motor are controlled by a microcontroller to slowly rotate so that the robot is in a squatting energy storage state, while the hub motor is controlled to accelerate forward, and then a control signal is sent to make the microcontroller control the hip joint motor and the knee joint motor to rotate quickly in the reverse direction. At this time, the thigh rod and the crank rotate quickly in the reverse direction, and further, the thigh rod can rotate quickly relative to the main body, and the calf can rotate quickly relative to the thigh rod. Finally, the reaction force of the ground enables the robot to cross the obstacle in a jumping posture. However, in this technical solution, for obstacles of different heights, different controls are required, with poor adaptability and difficulty in quickly adjusting strategies.
[0004] In recent years, control methods based on deep reinforcement learning (DRL) have gradually become an important research direction in the field of robot control. Deep reinforcement learning uses neural networks to extract features from high-dimensional input information and learns optimal control strategies through policy optimization algorithms (such as PPO, DDPG, etc.). For example, the Chinese patent application with publication number CN118192558A discloses a control algorithm for a wheel-legged robot based on model prediction and deep reinforcement learning. However, regarding the control when passing through height obstacles, specifically: when encountering a height obstacle, dynamic modeling analysis is performed based on the inverted pendulum model of the wheel-legged robot, and the discrete-time model of the system is obtained through linearization representation and model prediction. The method of solving quadratic programming problems is used to obtain the optimal control input variables and determine the control input at the current moment. This is still a traditional model predictive control-based method, which is easily affected by factors such as the delay of the dynamic modeling algorithm and terrain uncertainty, resulting in unstable action execution.
[0005] Therefore, there is an urgent need to provide a motion control method for a two-wheeled foot robot that performs better in passing over height obstacles, so as not to affect the adaptability of the robot to complex environments. Summary of the Invention
[0006] The present disclosure provides a method and device for controlling a two-wheeled foot robot to pass over height obstacles based on deep reinforcement learning, a method and device for training a motion control model, an electronic device, a robot, and a storage medium, which are used to improve the performance of the two-wheeled foot robot in passing over height obstacles, and further improve the adaptability of the robot to complex environments.
[0007] Additional aspects and advantages of the present disclosure will be partly set forth in the description below, and partly will become apparent from the description, or may be learned by practice of the present disclosure.
[0008] According to a first aspect of the present disclosure, there is provided a method for controlling a two-wheeled foot robot to pass over height obstacles based on deep reinforcement learning, including: When the motion feedback data of the two-wheeled foot robot meets a preset condition, it is determined that the robot encounters a height obstacle; Controlling the two-wheeled foot robot to perform a takeoff action, the takeoff action includes the main body squatting down and then rising rapidly, and controlling the two-wheeled foot structure of the two-wheeled foot robot to rotate forward rapidly against the height obstacle until the motion feedback data of the two-wheeled foot robot does not meet the preset condition; Wherein, the takeoff action is output by a policy network trained based on a deep reinforcement learning algorithm in a height obstacle terrain.
[0009] In an exemplary embodiment of the present disclosure, the motion feedback data includes wheel-foot rotation speed data, and the preset condition is that the difference between the target rotation speed and the actual rotation speed of the two wheels exceeds a preset threshold.
[0010] In an exemplary embodiment of the present disclosure, the actual rotation speed is acquired by an encoder installed on the motor of the two-wheeled foot robot; the target rotation speed is an expected rotation speed command sent by the controller to the drive motor of the two-wheeled foot robot.
[0011] In an exemplary embodiment of the present disclosure, the difference between the target rotation speed and the actual rotation speed of the two wheels exceeding the preset threshold includes: the difference between the target rotation speed and the actual rotation speed of the two wheels exceeds 30% of the target rotation speed.
[0012] In an exemplary embodiment of the present disclosure, the preset condition further includes: when the duration for which the difference between the target rotation speed and the actual rotation speed of the two wheels exceeds the preset threshold exceeds a preset duration.
[0013] In an exemplary embodiment of the present disclosure, the preset duration is greater than 200 milliseconds.
[0014] In an exemplary embodiment of the present disclosure, the rapid upward movement after the main body squats down includes: Controlling the main body of the two-wheeled foot robot to briefly squat down by lowering the center of gravity and store energy; Providing an upward thrust through the two-wheel drive mechanism to make the robot take off from the ground.
[0015] In an exemplary embodiment of the present disclosure, the rapid forward rotation of the two-wheeled foot structure attached to the height obstacle includes: Adjusting the driving torque of the two-wheeled foot to actively rotate forward rapidly at a speed higher than the normal driving speed to generate effective friction to assist the two-wheeled foot robot to climb.
[0016] In an exemplary embodiment of the present disclosure, the policy network is trained by the following method: In a simulation environment with a height obstacle terrain, using the teacher-student model framework, training the policy network for outputting the action policy for controlling the movement of the robot.
[0017] In an exemplary embodiment of the present disclosure, the height obstacle terrain includes various types of height obstacles; there are differences in at least one of the following parameters between different types of height obstacles: The overall height of the obstacle, the surface roughness of the obstacle side wall, the material stiffness of the obstacle side wall, the geometric shape of the obstacle top, the surface flatness of the obstacle side wall, the inclination angle between the obstacle side wall and the ground.
[0018] In an exemplary embodiment of the present disclosure, the training process of the policy network specifically includes: Encoding the historical motion state information of the two-wheeled foot robot itself through the student encoder to generate a first latent vector, and encoding the privileged state information of the two-wheeled foot robot through the teacher encoder to generate a second latent vector; the privileged state information includes the height obstacle information obtained from the simulation environment; Selecting the first latent vector or the second latent vector according to a preset policy and inputting it into the policy network; Based on the received latent vector and the current motion state information, the policy network outputs the action policy for controlling the movement of the two-wheeled foot robot; the current motion state information at least includes: attitude sensor data, wheel foot rotation speed data, torque data, and position and speed data; Using the privileged state information through the value network to output the value estimation of the current state; Optimizing the parameters of the student encoder based on the difference between the first latent vector and the second latent vector; Updating the parameters of the policy network using the deep reinforcement learning algorithm based on the action policy and the value estimation.
[0019] In an exemplary embodiment of the present disclosure, a comprehensive reward function for evaluating the action performance of a two-wheeled bipedal robot in a simulation environment is included in the deep reinforcement learning algorithm. The reward function includes the success rate of passing over a height obstacle, the takeoff height, and the attitude stability.
[0020] According to a second aspect of the present disclosure, there is provided a method for training a motion control model of a two-wheeled bipedal robot based on deep reinforcement learning, including: Constructing a virtual training scenario including a height obstacle terrain, and obtaining motion data of a virtual two-wheeled bipedal robot in the virtual training scenario; Training a policy network by using a deep reinforcement learning algorithm according to the motion data and environmental information; The policy network outputs the following action strategy for the two-wheeled bipedal robot: When encountering a height obstacle, execute a takeoff action, which includes quickly rising after the main body squats down, and controlling the two-wheel structure of the two-wheeled bipedal robot to rotate forward quickly while attaching to the height obstacle until the height obstacle is crossed.
[0021] According to a third aspect of the present disclosure, there is provided a height obstacle control device for a two-wheeled bipedal robot based on deep reinforcement learning, the device including: An obstacle determination module, configured to determine that the robot encounters a height obstacle when the motion feedback data of the two-wheeled bipedal robot meets a preset condition; An action control module, configured to control the two-wheeled bipedal robot to execute a takeoff action, which includes quickly rising after the main body squats down, and controlling the two-wheel structure of the two-wheeled bipedal robot to rotate forward quickly while attaching to the height obstacle until the motion feedback data of the two-wheeled bipedal robot does not meet the preset condition; Wherein, the takeoff action is output by a policy network trained based on a deep reinforcement learning algorithm in a height obstacle terrain.
[0022] According to a fourth aspect of the present disclosure, there is provided a device for training a motion control model of a two-wheeled bipedal robot based on deep reinforcement learning, the device including: A training data acquisition module, configured to construct a virtual training scenario including a height obstacle terrain, and obtain motion data of a virtual two-wheeled bipedal robot in the virtual training scenario; A model training module, configured to train a policy network by using a deep reinforcement learning algorithm according to the motion data and environmental information; The policy network outputs the following action strategy for the two-wheeled bipedal robot: When encountering a height obstacle, execute a takeoff action, which includes quickly rising after the main body squats down, and controlling the two-wheel structure of the two-wheeled bipedal robot to rotate forward quickly while attaching to the height obstacle until the height obstacle is crossed.
[0023] According to a fifth aspect of the present disclosure, there is provided an electronic device, including: A processor; and A memory stores computer-readable instructions, which, when executed by a processor, implement the method of the above-described embodiments.
[0024] According to a sixth aspect of the present disclosure, a robot is provided, including: A processor; and A memory stores computer-readable instructions, which, when executed by the processor, implement the method of the above-described embodiments.
[0025] In an exemplary embodiment of the present disclosure, the robot is a robot configured with a height obstacle control function for a two-wheeled bipedal robot based on deep reinforcement learning, and the robot includes any one of a humanoid robot, a cleaning robot, a transportation robot, and a mobile robot.
[0026] According to a seventh aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program code instructions are stored. When the computer program code instructions are called by a processor of a robot, the robot is caused to execute the method of the above-described embodiments.
[0027] It can be seen from the above technical solutions that the present disclosure has at least one of the following advantages and positive effects: In the exemplary embodiment of the present disclosure, for the height obstacle control method of the two-wheeled bipedal robot based on deep reinforcement learning, through the control method based on deep reinforcement learning, the two-wheeled bipedal robot can execute a takeoff action when detecting a height obstacle and complete the obstacle crossing through the rapid forward rotation of the two-wheeled bipedal structure. On the one hand, based on the real-time judgment of the robot motion feedback data, when the motion feedback data meets the first condition, the robot can automatically trigger the obstacle crossing action. Compared with the prior art that relies on model prediction, the deep reinforcement learning method adopted by the present disclosure enables the robot to autonomously adjust its actions when encountering a height obstacle without relying on complex modeling prediction. On the other hand, after the robot detects a height obstacle, the present disclosure controls the robot to execute a takeoff action, and the takeoff process includes a rapid rise after the main body squats down, so that the robot can accumulate enough kinetic energy in a short time to complete the obstacle crossing. In the prior art, for obstacles of different heights, different control modes often need to be set, while the present disclosure enables the robot to autonomously adjust the squatting energy storage and rising speed according to different obstacle characteristics through the policy network trained by deep reinforcement learning, thereby effectively coping with obstacles of different heights, improving the generality of the control strategy, and avoiding the problem of frequent parameter adjustment required by the traditional method based on preset rules when the height changes. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0029] Figure 1 The system architecture diagram of the dual-wheel foot robot based on deep reinforcement learning through height obstacle control method in the embodiments of the present disclosure is shown.
[0030] Figure 2 The flowchart of a method for a dual-wheel foot robot based on deep reinforcement learning to control through height obstacles in the embodiments of the present disclosure is shown.
[0031] Figure 3 The scene diagram of a dual-wheel foot robot moving towards a height obstacle in the embodiments of the present disclosure is shown.
[0032] Figure 4 The scene diagram of a dual-wheel foot robot leaping process in the embodiments of the present disclosure is shown.
[0033] Figure 5 The scene diagram of a dual-wheel foot robot landing on a height obstacle in the embodiments of the present disclosure is shown.
[0034] Figure 6 The flowchart of the process of a robot body quickly rising after squatting in the embodiments of the present disclosure is shown.
[0035] Figure 7 The schematic diagram of the two-stage training-inference framework of a policy network in the embodiments of the present disclosure is shown.
[0036] Figure 8 The schematic diagram of another two-stage training-inference framework of a policy network in the embodiments of the present disclosure is shown.
[0037] Figure 9 The flowchart of the process of training a policy network in the embodiments of the present disclosure is shown.
[0038] Figure 10 The flowchart of a method for training a motion control model of a dual-wheel foot robot based on deep reinforcement learning in the embodiments of the present disclosure is shown.
[0039] Figure 11 The block diagram of a device for a dual-wheel foot robot based on deep reinforcement learning to control through height obstacles in the embodiments of the present disclosure is shown.
[0040] Figure 12The block diagram of a training device for a motion control model of a two-wheeled and two-legged robot based on deep reinforcement learning in an embodiment of the present disclosure is shown.
[0041] Figure 13 A schematic diagram of a robot in an embodiment of the present disclosure is shown.
[0042] Figure 14 The schematic structural diagram of an electronic device in an embodiment of the present disclosure is shown.
[0043] Figure 15 The schematic structural diagram of a program product in an embodiment of the present disclosure is shown. Detailed implementation manners
[0044] In the description of the present disclosure, the terms "first" and "second" are only used for description, and do not indicate relative importance or imply the number of technical features. Therefore, the features of "first" and "second" may explicitly or implicitly include at least one such feature. The meaning of "a plurality" is at least two, unless otherwise clearly defined.
[0045] Figure 1 The system architecture diagram to which the height obstacle control method of the two-wheeled and two-legged robot based on deep reinforcement learning in an embodiment of the present disclosure can be applied is shown.
[0046] As Figure 1 shown, the system architecture 100 may include a terminal device 101, a two-wheeled and two-legged robot 102, a network 103, and a server 104. Among them, the terminal device 101 includes, but is not limited to, a desktop computer, a portable computer, a smart phone, a tablet computer, and the like. The terminal device 101 may serve as an interaction interface, provide a visualization function to display the running state, motion trajectory, terrain information, etc. of the two-wheeled and two-legged robot 102, and at the same time support sending motion control instructions to the two-wheeled and two-legged robot 102.
[0047] The two-wheeled and two-legged robot 102 sends motion feedback data to the server 104 through the network 103 for centralized processing. When the motion feedback data meets the preset conditions, the server 104 determines that the two-wheeled and two-legged robot 102 encounters a height obstacle, and a takeoff action strategy is output by a policy network pre-trained based on a deep reinforcement learning algorithm in a height obstacle terrain. Subsequently, the takeoff action strategy is sent to the two-wheeled and two-legged robot 102 through the network 103 to guide the two-wheeled and two-legged robot 102 to execute a takeoff action including quickly rising after the main body squats down, and an action of quickly rotating straight while the two-wheeled and two-legged structure of the two-wheeled and two-legged robot 102 adheres to the height obstacle, until the motion feedback data of the two-wheeled and two-legged robot 102 does not meet the preset conditions.
[0048] Of course, if the two-wheeled and two-legged robot 102 has local inference capabilities, it can also store the policy network internally, enabling the two-wheeled and two-legged robot 102 to autonomously output a takeoff action policy that matches the current height obstacle based on the motion feedback data, thereby achieving edge computing and fast response.
[0049] The network 103 is a medium for providing a communication link between the terminal device 101, the two-wheeled and two-legged robot 102, and the server 104. The network 103 can include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. It should be understood that Figure 1 the number and type of terminal devices, two-wheeled and two-legged robots, networks, and servers in
[0050] This disclosure's exemplary embodiment provides a method for controlling a two-wheeled and two-legged robot to pass through height obstacles based on deep reinforcement learning. Referring to Figure 2 as shown, this method can include the following steps S201 and S202: Step S201, when the motion feedback data of the two-wheeled and two-legged robot meets a preset condition, determine that the robot encounters a height obstacle.
[0051] Among them, the motion feedback data can represent the data information reflecting the current dynamic state of the two-wheeled and two-legged robot collected and transmitted back by its internal sensing components during operation. In the embodiments of this disclosure, the motion feedback data can include other parameters such as wheel and leg rotation speed data, robot attitude data, torque data, etc. that can reflect the robot's motion behavior. For example, the wheel and leg rotation speed data can represent the motor rotation speed information collected in real time by the motor encoder installed on the drive motor of the two-wheeled and two-legged robot, which can reflect the rotation of the two-wheeled and two-legged structure within a unit time. The robot attitude data can include the spatial poses of various parts of the robot body, including position, direction, attitude angle, etc., which are used to describe the overall motion state of the robot and can be obtained through an inertial measurement unit (IMU) set on the robot body. The torque data can represent the drive torque value measured by the current sensor in the robot drive system and calculated in combination with the motor characteristics, which can characterize the force conditions of each drive wheel of the two-wheeled and two-legged robot during motion.
[0052] The height obstacle can represent a terrain structure or object with a vertical height feature encountered in the traveling path of the two-wheeled and two-legged robot, which causes the robot's wheel and leg movement to be blocked or unable to directly cross. For example, the height obstacle can be a step, a vertical wall, a protruding platform, a ground height difference, or other types of height obstacles. Figure 3Fig. 0 shows a schematic diagram of the scenario during the movement of a two-wheeled foot robot towards a height obstacle in an embodiment of the present disclosure. During the process of the robot moving towards the height obstacle, its internal sensing components acquire motion feedback data in real time. When the motion feedback data meets the preset conditions, it can be determined that the robot has encountered a height obstacle and needs to perform an obstacle-crossing action.
[0053] The preset conditions can represent the judgment criteria preset in the control system of the two-wheeled foot robot for determining whether there is a height obstacle, and can be set based on the change situation of one or more parameters in the motion feedback data. In some exemplary embodiments, the preset conditions can be: the difference between the target speed and the actual speed of the two wheels exceeds a preset threshold. Among them, the actual speed can be obtained by collecting through the motor encoder installed on the two-wheeled foot robot, and the target speed can be the expected speed command sent by the controller to the drive motor of the two-wheeled foot robot. In the normal driving state, the target speed and the actual speed of the two wheels are basically the same. However, when the robot encounters a height obstacle during driving, due to the sudden block of the height obstacle, the rotation of the wheel feet will be blocked, resulting in a significant decrease in the actual speed, and then a large deviation from the target speed. Once this deviation exceeds the set preset threshold, it usually indicates that the wheel feet cannot reach the expected motion state within a unit time due to encountering a height obstacle. Therefore, the difference between the target speed and the actual speed of the two wheels exceeding the preset threshold can be used as the trigger judgment condition for the robot to encounter a height obstacle.
[0054] In some exemplary embodiments, the above preset threshold can be 30% of the target speed, that is, when the difference between the actual speed and the target speed of the two wheels exceeds 30% of the target speed, the preset conditions are met. Of course, in other embodiments of the present disclosure, the preset threshold can also be other suitable values such as 20%, 24%, 26%, 28%, 32%, 34%, and 36% of the target speed.
[0055] In addition, in some other exemplary embodiments, the preset conditions for determining whether the robot encounters a height obstacle can also include: when the duration for which the difference between the target speed and the actual speed of the two wheels exceeds the preset threshold exceeds a preset duration. Exemplarily, the preset duration is greater than 200 milliseconds, which is used to avoid misjudgment caused by short-term fluctuations, so as to more accurately determine that the two-wheeled foot robot has encountered a height obstacle. Of course, in other embodiments of the present disclosure, the specific value of the preset duration can be flexibly adjusted according to the response speed requirements of the robot application scenario, the sampling frequency of the control system, and the environmental interference characteristics. For example, in a high-dynamic response scenario, the preset duration can be set between 100 milliseconds and 300 milliseconds, so as to improve the robustness of the determination of whether the robot encounters a height obstacle while ensuring the detection sensitivity.
[0056] Step S202: Control the two-wheeled foot robot to perform a takeoff action, where the takeoff action includes the main body squatting down and then rising rapidly, and control the two-wheeled foot structure of the two-wheeled foot robot to rotate forward rapidly while attaching to the height obstacle until the motion feedback data of the two-wheeled foot robot does not meet the preset conditions; among them, the takeoff action is output by a policy network trained based on a deep reinforcement learning algorithm in a height obstacle terrain.
[0057] In this step, when the robot detects a height obstacle, it controls the robot to perform a takeoff action. This takeoff process includes the main body squatting down and then rising rapidly, enabling the robot to accumulate sufficient kinetic energy in a short time to complete obstacle crossing. In addition, during the rising process of the robot main body, the two-wheeled foot structure attaches to the obstacle surface and rotates forward rapidly, improving the fitting degree of the robot to the obstacle surface, enabling the robot to maintain a stable obstacle-crossing ability under different obstacle forms, and thus reducing the risk of obstacle-crossing failure. During this process, the takeoff action is output by a policy network trained based on a deep reinforcement learning algorithm, enabling the robot to dynamically adjust the rotation speed and contact mode of the wheel feet during the motion process, improving the fitting degree of the robot to the obstacle surface, enabling the robot to maintain a stable obstacle-crossing ability under different obstacle forms, and thus reducing the risk of obstacle-crossing failure, avoiding the problem of unstable execution caused by environmental changes in the related technology. In addition, through the control strategy trained by deep reinforcement learning, the robot can continuously optimize its own motion mode during the training process, form the ability of motion decision-making adapted to complex environments, and reduce the dependence on high-precision sensors and complex computing resources.
[0058] Among them, the takeoff action can represent the action behavior of the two-wheeled foot robot to achieve the purpose of obstacle crossing after detecting a height obstacle. This action includes, but is not limited to, briefly squatting down by lowering the center of gravity to store energy, and then the drive system outputs an upward impact force to complete the rapid rise of the robot main body, and combines the wheel foot movement to form an overall obstacle-crossing process. The two-wheeled foot structure can represent a motion execution mechanism configured at the bottom of the robot, which has the dual functions of rolling driving and attaching support. This structure can include at least two active wheel systems or wheel foot components with the compound ability of traveling and supporting, and can produce effective contact and friction with the obstacle surface during takeoff and climbing. The height obstacle terrain can represent an obstacle area with certain height mutation characteristics. Figure 4 Fig. shows a schematic diagram of the scenario of the leap process of a two-wheeled foot robot in an embodiment of the present disclosure. During the leap process, the two-wheeled foot structure of the two-wheeled foot robot attaches to the height obstacle and rotates forward rapidly to enhance the friction between the wheel feet and the obstacle surface. Figure 5The figure shows a schematic diagram of the scenario where a two-wheeled foot robot stably lands on a height obstacle in an embodiment of the present disclosure. When the two-wheeled foot robot completes obstacle crossing and is above the height obstacle, the motion feedback data of the two-wheeled foot robot does not meet the preset conditions, and the two-wheeled foot structure of the two-wheeled foot robot stops rotating forward at high speed and returns to the normal driving state to maintain attitude balance and position stability.
[0059] The deep reinforcement learning algorithm can represent an optimal policy search method that combines a deep neural network and a reinforcement learning mechanism. It can continuously update the control policy based on the state of the robot, environmental feedback data, and the reward function to improve the decision-making ability of the robot when passing through height obstacles. Among them, deep reinforcement learning includes, but is not limited to, algorithms such as PPO (Proximal Policy Optimization), DDPG (Deep Deterministic Policy Gradient), and SAC (Soft Actor-Critic). The present disclosure does not limit the specific algorithm type of deep reinforcement learning. The policy network can represent a neural network model trained based on deep reinforcement learning. The input of this model is the perception data obtained by the two-wheeled foot robot, and the output is the action decision for controlling the movement of the robot, such as the drive power distribution during takeoff, the rotation speed adjustment of the two-wheeled feet, and the attitude control during the obstacle contact process, so as to achieve stable obstacle crossing in the height obstacle terrain.
[0060] In some exemplary embodiments, referring to Figure 6 as shown, the specific implementation process of the rapid upward movement of the main body of the two-wheeled foot robot after squatting includes the following steps S601 and S602: Step S601, control the main body of the two-wheeled foot robot to squat briefly by lowering the center of gravity and store energy.
[0061] Exemplarily, this step can be achieved in the following way: by controlling the main body of the two-wheeled foot robot to lower the center of gravity in the vertical direction, a short squatting process is realized. This squatting action can be achieved by driving the mechanism to reduce the output speed of the wheel feet or adjusting the attitude angle of the robot. During the squatting process, the two-wheeled foot robot can apply a preload to the ground by compressing the elastic element or using the reaction force of the wheel foot system, thereby forming a reserve of mechanical energy inside the leg structure. Of course, the above-mentioned squatting and energy storage steps can also be achieved in the following way: by controlling the coordinated rotation of the hip joint motor and the knee joint motor, driving the thigh assembly and the calf assembly to close inward, making the lower limb structure of the robot present a bent state, so as to realize the downward movement of the overall center of gravity of the robot in the vertical direction and complete a short squatting action. During this process, the robot forms an actively controlled joint bend by driving the leg joints, and combines the contact reaction force between the roller and the ground to apply a preload to the leg structure, thereby forming a reserve of mechanical energy available for takeoff. This energy reserve provides a power basis for subsequent takeoff actions and helps to improve the takeoff height and takeoff stability of the robot.
[0062] Step S602, provide an upward thrust through the two-wheel drive mechanism to make the robot take off from the ground.
[0063] Exemplarily, the takeoff of the two-wheeled foot robot from the ground in this step can be achieved according to the following way: by increasing the rotational speed of the wheel feet through the two-wheel drive mechanism, applying a continuous and concentrated reaction force to the ground, thereby providing an upward thrust for the robot and prompting the main body of the robot to leave the ground to achieve takeoff. This thrust can be precisely controlled by the drive mechanism according to the output instructions of the policy network, providing sufficient vertical kinetic energy to ensure the completion of the takeoff action at the right time and intensity. Of course, in other embodiments of the present disclosure, during the above takeoff process, while increasing the rotational speed of the wheel feet, the hip joint motor and the knee joint motor can be controlled to quickly unfold in the reverse direction in coordination, so that the thigh assembly and the calf assembly quickly extend, thereby superimposing the accumulated mechanical energy and the driving acceleration to form a greater takeoff thrust, thereby increasing the success rate of takeoff and obstacle crossing.
[0064] In some exemplary embodiments, the specific process of the two-wheeled foot structure quickly rotating forward while attaching to a high obstacle includes: adjusting the driving torque of the two-wheeled foot to actively rotate forward quickly at a speed higher than the normal driving speed to generate effective friction to assist the two-wheeled foot robot to climb.
[0065] Among them, the driving torque can represent the torque output by the driving motor of the bipedal wheeled robot for driving the wheel feet to rotate, and the magnitude of the driving torque can be dynamically adjusted according to the output of the policy network. The normal driving speed can represent the stable speed maintained by the bipedal wheels of the robot when the bipedal wheeled robot is driving conventionally in an obstacle-free environment. Exemplarily, the bipedal wheel structure can rotate at 30% higher than the normal driving speed. Of course, in other embodiments of the present disclosure, the rotation speed of the bipedal wheel structure can also be dynamically adjusted according to the material, slope or adhesion difficulty of the high obstacle surface, and its rotation speed can be higher than 32%, 34%, 36%, 38%, 40%, 42% and 45% of the normal driving speed and other appropriate values to ensure smooth climbing on different types of high obstacles. In this embodiment, by adjusting the driving torque, the bipedal wheels rotate at a speed higher than the normal driving state when attaching to the high obstacle, thereby enhancing the friction with the obstacle surface. This process not only helps to improve the adhesion effect of the wheel feet, but also provides additional propulsion force to promote the robot to climb stably on a vertical or nearly vertical surface. By dynamically regulating the magnitude of the driving torque through the policy network, the adaptive adjustment of the surface characteristics of different types of obstacles can be realized, further improving the reliability and success rate of the robot during the obstacle-crossing process.
[0066] In some exemplary embodiments, the policy network can be trained by the following method: in a simulation environment with a high obstacle terrain, the policy network for outputting the action policy for controlling the movement of the robot is trained using the teacher-student model framework.
[0067] Among them, the teacher-student model framework can represent a supervised learning structure for training the policy network. The framework includes two sub-networks, a teacher encoder and a student encoder. Among them, the teacher encoder provides stable feature supervision for the student encoder, guiding the student encoder to learn a higher-quality state representation, so that the policy network has a better convergence speed and generalization ability during training.
[0068] In addition, in some embodiments, the height obstacle terrain may include various types of height obstacles. There are differences in at least one of the parameters among different types of height obstacles, such as the overall height of the obstacle, the surface roughness of the obstacle sidewall, the material stiffness of the obstacle sidewall, the geometric shape of the obstacle top, the surface flatness of the obstacle sidewall, and the inclination angle between the obstacle sidewall and the ground. Among them, the overall height of the obstacle can represent the vertical distance from the position where the bottom of the obstacle touches the ground to the highest point at the top. The surface roughness of the obstacle sidewall can represent the degree of texture undulation or friction characteristics of the obstacle sidewall surface, which can be quantified by the mean deviation of the surface profile or the friction coefficient. The material stiffness of the obstacle sidewall can represent the ability of the material constituting the obstacle to resist deformation under the action of force, which can be quantified by Young's modulus or structural hardness. The geometric shape of the obstacle top can represent the spatial morphological characteristics of the obstacle top region, including but not limited to other shapes such as a plane, an arc, a sharp corner, etc. The surface flatness of the obstacle sidewall can represent the degree of obvious concavity and convexity, discontinuity, or structural mutation in the obstacle sidewall. The inclination angle between the obstacle sidewall and the ground can represent the angle between the obstacle sidewall and the horizontal ground.
[0069] Of course, in other embodiments of the present disclosure, there may also be differences in other obstacle-related parameters that affect the stability of the robot's obstacle-crossing action among different types of height obstacles, such as the width of the obstacle, the sharpness of the obstacle top edge, the surface elasticity of the obstacle material, and the slipperiness of the obstacle surface. Among them, the width of the obstacle can represent the lateral extension distance of the obstacle in the direction perpendicular to the forward direction of the robot. The sharpness of the obstacle top edge can represent the geometric transition degree of the top boundary of the obstacle. The surface elasticity of the obstacle material can represent the ability of the obstacle to return to its original state after local deformation under force. By training the policy network using the teacher-student model framework in the above simulation environment containing various types of height obstacles, the student model can have stronger policy generalization ability and environmental adaptation ability when facing obstacles with different heights, materials, surface shapes, or slopes, effectively improving the success rate and stability of the bipedal wheeled robot in crossing obstacles in a real complex environment.
[0070] Furthermore, in some embodiments, the deep reinforcement learning algorithm for training the policy network includes a comprehensive reward function for evaluating the action performance of the bipedal wheeled robot in the simulation environment. This reward function includes the success rate of crossing the height obstacle, the takeoff height, and the attitude stability. This comprehensive reward function is used to evaluate the action performance of the bipedal wheeled robot in completing the obstacle-crossing task in the simulation environment, and its definition is as follows: In the above formula, represents the comprehensive reward value, which is used to measure the overall performance of the robot in completing the obstacle-crossing task in the current training round, represents the success rate reward for crossing the height obstacle, Indicates the takeoff height reward, Indicates the attitude stability reward, respectively represent the weight coefficients of each reward, satisfying , and their values can be adjusted according to the key points of task attention.
[0071] Among them, if the robot successfully crosses the obstacle and lands stably, the success rate reward for crossing the height obstacle is 1, otherwise it is 0.
[0072] Takeoff height reward can be expressed as: Among them, represents the maximum vertical height from the ground after the robot takes off, represents the minimum effective takeoff height required for the current obstacle.
[0073] Attitude stability reward can be expressed as: Among them, represents the duration of the robot's obstacle crossing process, represents the pitch angle deviation of the robot at time , represents the roll angle deviation of the robot at time .
[0074] In addition, in other embodiments of the present disclosure, the above comprehensive reward function for evaluating the action performance of the bipedal wheel robot in the simulation environment may further include other parameters for evaluating the action efficiency and stability of the bipedal wheel robot during the obstacle crossing process, such as energy consumption, landing impact magnitude, motion trajectory smoothness, etc. This embodiment is not specifically limited herein.
[0075] Refer to Figure 7 shown, which shows a schematic diagram of the two-stage training-inference framework of a policy network. Among them, the training stage combines the teacher-student encoder structure and the deep reinforcement learning framework. The teacher-student encoder structure includes a teacher encoder 701 and a student encoder 702. The deep reinforcement learning framework consists of a policy network 703 and a value network 704. Among them, the value network can represent a neural network structure for estimating the value of the current state, and this value estimation can be used to guide the policy optimization in the reinforcement learning process to improve the behavior quality and training stability of the policy network. In the training stage, by introducing privileged state information to guide policy learning, the student encoder 702 can stably output high-quality control actions only relying on observable data during the inference stage. The observable data includes the current motion state information of the bipedal wheel robot.
[0076] It should be noted that the teacher encoder 701 is only used to extract the high-dimensional semantic representation in the privileged state information during the training phase, which is used as the input of the policy for subsequent learning guidance. The student encoder 702 encodes the historical motion state of the robot itself during the training and inference phases to obtain dynamic features related to action decisions, and then inputs the current motion state information into the policy network 703 together to generate control actions. In addition, the student encoder 702 ultimately needs to imitate the representation output by the teacher encoder 701, so during the training process, distillation training is carried out by minimizing the MSE (mean square error).
[0077] The policy network 703 is the core module for action generation. During the training phase, the policy network 703 receives the latent vector from the teacher encoder 701 or the student encoder 702 and the current motion state information, and combines with the value network 704 to update the policy using the PPO algorithm. During the inference phase, the policy network 703 can receive the latent vector from the student encoder 702 and the current motion state information, and output the action policy for controlling the movement of the two-wheeled foot robot.
[0078] Reference Figure 8 As shown, a schematic diagram of a two-stage training-inference framework of another policy network is shown. Figure 7 And Figure 8 is the same as the training framework shown in, but the inference framework is different. Specifically,[[]]END]] Figure 8 In, both the teacher encoder 701 and the student encoder 702 are only used in the training phase. After the training is completed through this architecture, the policy network 703 can be trained for different height obstacle terrains, and the policy network 703 already has the ability to perform perception and understanding and action output on different height obstacle terrains. Therefore, in the inference phase, there is no need to use the student encoder 702 as an auxiliary module, and the policy network 703 can directly receive the current motion state information and output the action policy corresponding to the control robot's takeoff action and the fast forward rotation action of the two-wheeled foot structure attached to the height obstacle. This example greatly improves the efficiency of the inference phase and the simplicity of deployment, while ensuring the high adaptability and stability of the policy in different height obstacle terrain scenarios.
[0079] Based on Figure 7 the schematic diagram of the framework shown, with reference to Figure 9 shown, the process of training the policy network for outputting the action policy for controlling the robot's takeoff action may include the following steps S901 to step S906: Step S901, encoding the historical motion state information of the two-wheeled foot robot itself through the student encoder to generate a first latent vector, and encoding the privileged state information of the two-wheeled foot robot through the teacher encoder to generate a second latent vector; the privileged state information includes the height obstacle information obtained from the simulation environment.
[0080] Among them, the historical motion state information of the robot itself may include attitude sensor data, wheel and foot rotation speed data, torque data, and position and speed data, sole contact state, center of gravity trajectory, etc. of several past frames. For example, the historical motion state sequence of the past 10 frames is input into the student encoder 702 for encoding to obtain a first latent vector. The first latent vector can capture the continuity and dynamic characteristics of the robot's actions.
[0081] The privileged state information may include height obstacle information obtained from the simulation environment. The height obstacle information may represent a set of relevant parameters describing the spatial structure and physical properties of the obstacle in the simulation environment, and the information is used to assist the teacher model in generating a better control strategy. The height obstacle information may include, but is not limited to, the overall height of the obstacle, the surface roughness of the obstacle sidewall, the material stiffness of the obstacle sidewall, the geometric shape of the obstacle top, the surface flatness of the obstacle sidewall, the inclination angle between the obstacle sidewall and the ground, the edge shape of the obstacle, and the relative position relationship with the robot, etc. The privileged state information is input into the teacher encoder 701 for encoding to obtain a second latent vector. The second latent vector can highly concentrate the key semantics in the height obstacle information, which helps to construct a more complete high-dimensional feature representation for expressing the height obstacle information.
[0082] It should be noted that the privileged state information can only be obtained during the training phase but is not observable during the testing or deployment phase. Therefore, the teacher encoder 701 can use the complete information to learn the best latent representation, thereby guiding the student encoder 702 to learn. Both the teacher encoder 701 and the student encoder 702 can compress the high-dimensional, time-series state information into a low-dimensional latent representation, providing behavioral semantic representations from different sources for the policy network 703. For example, the teacher encoder 701 and the student encoder 702 can be multi-layer perceptrons, temporal convolutional networks, or recurrent neural networks, and have the ability to extract temporal features and compress them into fixed-length semantic vectors. In addition, the network architectures of the teacher encoder 701 and the student encoder 702 can be the same or different, and the present disclosure does not limit this.
[0083] Step S902, select the first latent vector or the second latent vector according to a preset policy and input it into the policy network.
[0084] According to the preset policy, select to send the first latent vector generated by the student encoder 702 or the second latent vector generated by the teacher encoder 701 into the policy network 703 for decision-making. Among them, the preset policy can be set according to the real-time environment state, training phase, or specific performance metrics. The specific performance metrics include action execution error thresholds, simulation and real environment difference degrees, etc.
[0085] For example, one preset strategy is to preferentially use the high-quality second latent vectors generated by the teacher encoder 701 at the initial stage of training to guide the policy network 703 to quickly converge to an approximate optimal solution. When the student encoder 702 is optimized through knowledge distillation, it gradually transitions to using only the first latent vectors during the deployment stage to reduce the dependence on privileged information. Another example is that another preset strategy is to select the second latent vectors in proportion p and select the first latent vectors according to 1 - p and gradually reduce the p value. Of course, the first latent vectors and the second latent vectors can also be concatenated and fed into the policy network 703, and dynamically weighted through an attention mechanism or a gating module, enabling the policy network 703 to flexibly combine historical experience and privileged knowledge in complex scenarios.
[0086] The selective input mechanism can not only utilize the prior knowledge of the ideal state provided by the teacher encoder 701 to accelerate the training process, but also cope with sensor limitations or environmental disturbances through the generalization ability of the student encoder 702 during actual operation. At the same time, through the conditional processing of the latent vectors by the policy network 703, smooth switching and robust decision-making of motion control are achieved. For example, when the robot encounters an obstacle of unknown height, it preferentially adjusts the takeoff action and the attachment action of the bipedal wheel structure based on the second latent vectors, and relies on the first latent vectors to maintain efficiency during the stable walking stage. Finally, the policy selection rule is optimized through closed-loop feedback, enabling the robot to balance the obstacle-crossing performance and adaptability under different stages and environmental conditions.
[0087] In addition, in addition to the selected latent vectors, the current motion state information of the robot can also be input into the policy network 703 for decision-making.
[0088] Step S903, based on the received latent vectors and the current motion state information, the policy network outputs an action policy for controlling the motion of the bipedal wheel robot. The current motion state information may include attitude sensor data, wheel-foot rotation speed data, torque data, and position and speed data.
[0089] Among them, the policy network 703 can be a multi-layer fully connected perceptron, a Transformer structure, etc. The policy network 703 performs multi-modal feature fusion on the first latent vectors or the second latent vectors and the current motion state information, and outputs an action policy, denoted as .
[0090] For example, the first latent vectors or the second latent vectors and the current motion state information are concatenated or weighted and interacted in the embedding space, and action policies such as joint angle targets, torque commands, or gait phase parameters are generated through non-linear transformation.
[0091] Step S904, utilize the privileged state information through the value network to output the value estimate of the current state.
[0092] During the training process of the policy network 703, the long-term return of the current state is also estimated through the value network 704, that is, starting from this state, if the current policy is continuously executed, how much cumulative reward can be obtained in the future, so as to optimize the policy network 703.
[0093] Among them, the value network 704 can be a multi-layer fully connected perceptron. For example, when using the privileged state information through the value network 704 to estimate the value of the current state, first encode the privileged state information into a high-dimensional feature vector, such as extracting dynamic features through a convolutional or fully connected layer, then fuse it with the current motion state information in the latent space, and then output the value estimate representing the current state through multi-layer non-linear transformation, denoted as . This value estimate is used for policy optimization in reinforcement learning to guide the policy network 703 to learn better behaviors.
[0094] Step S905, optimize the parameters of the student encoder based on the difference between the first latent vector and the second latent vector.
[0095] This step quantifies the distribution difference between the two in the latent space and backpropagates the difference gradient to update the network weights of the student encoder 702, so as to guide the student encoder 702 to learn to generate a latent representation close to that of the teacher encoder 701, realizing teacher knowledge distillation.
[0096] Exemplarily, during the training phase, fix the parameters of the teacher encoder 701, take the historical motion state information and the corresponding motion state information as parallel inputs, generate latent vectors through the student encoder 702 and the teacher encoder 701 respectively, then use the contrast learning framework to minimize the distance between the two, or make the output distribution of the student encoder 702 approximate the latent space characteristics of the teacher encoder 701 through adversarial training. At the same time, introduce noise injection or data augmentation to simulate sensor errors in actual deployment, and force the student encoder 702 to still extract feature expressions compatible with the privileged information encoding under the condition of limited input information.
[0097] For example, to make the output of the student encoder 702 as close as possible to that of the teacher encoder 701, a difference loss function can be constructed, such as the least squares gap or KL divergence. By minimizing this loss, optimize the parameters of the student encoder 702, so that it can learn to extract feature expressions close to the privileged information from the historical motion state, thereby enhancing the generalization ability of the policy network 703, improving the decision-making quality of the policy network 703 in the real environment, and enabling it to approach the teacher level without relying on privileged information during deployment.
[0098] Step S906: Update the parameters of the policy network using a deep reinforcement learning algorithm based on the action policy and value estimation.
[0099] Taking the Proximal Policy Optimization (PPO) algorithm as an example of the deep reinforcement learning algorithm. During the training process of the policy network 703, the current policy network 703 is used to interact with the environment. Based on the current state select the action policy , and after executing the action, return the reward and the next state . Thus, an interaction trajectory sequence ( ) is sampled for subsequent policy optimization. Then, a value network 704 is introduced to evaluate the value of each state . To optimize the policy more stably and efficiently, the policy gradient can be calculated through the advantage function to obtain the parameter update direction of the policy network 703.
[0100] For example, the advantage function is defined as: where is the advantage function, and is the advantage approximation obtained using Generalized Advantage Estimation (GAE).
[0101] Then, an optimization objective is constructed according to the clipping objective function of the PPO algorithm, and the policy gradient is calculated through backpropagation to optimize the parameters of the policy network 703, so that it continuously improves the expected cumulative return under the current policy. At the same time, supervised learning is performed on the value network 704 using the state - return pairs ( ) to minimize the mean square error between its output and the true return , thereby improving the estimation accuracy of the value network 704 for the state value.
[0102] The student encoder 702 aligns the latent space with the output of the teacher encoder 701 through contrastive learning or knowledge distillation, forcing the student network to infer a latent representation close to the encoding result of the privileged information through historical states in the actual deployment scenario lacking privileged information, thereby improving the robustness and adaptability of the motion policy. In the exemplary embodiment of the present disclosure, the robot can simulate the motion decision optimized based on privileged information only relying on its own sensor data in the real environment, effectively solving the problem of performance degradation caused by unobservable partial states or noise interference during actual deployment. At the same time, the generalization ability of the motion policy is enhanced through the implicit knowledge transfer of the latent vector.
[0103] In the exemplary embodiment of the present disclosure, a method for training a motion control model of a two - wheeled foot robot based on deep reinforcement learning. Refer to Figure 10As shown, the method may include the following steps S1001 to S1002: Step S1001: Construct a virtual training scenario including a height obstacle terrain, and obtain the motion data of the virtual two-wheeled biped robot in the virtual training scenario.
[0104] In this step, by constructing a virtual training scenario containing various types of height obstacles, the motion environment of the two-wheeled biped robot in a real complex terrain is simulated. The virtual training scenario may include obstacles with various parameter combinations such as height, surface roughness, material stiffness, top geometry, surface flatness, slope angle, etc., to enhance the training diversity and the generalization ability of the strategy. Deploy the virtual two-wheeled biped robot in this environment, and record its motion data during the movement. The motion data may include attitude sensor data, wheel-foot rotation speed data, torque data, position and speed data, etc., which are used for input modeling and performance evaluation of the policy network during the subsequent training process.
[0105] Step S1002: Use a deep reinforcement learning algorithm to train a policy network based on the motion data and environmental information; the policy network outputs the following action strategy for the two-wheeled biped robot: when encountering a height obstacle, execute a takeoff action, and the takeoff action includes quickly rising after the main body squats down, and controlling the two-wheeled foot structure of the two-wheeled biped robot to rotate forward quickly while attaching to the height obstacle until crossing the height obstacle.
[0106] In this step, based on the motion data obtained in step S1001 and the environmental information in the virtual training scenario, a deep reinforcement learning algorithm is used to train the policy network. The deep reinforcement learning can be other suitable deep reinforcement learning algorithms such as the PPO algorithm, DDPG algorithm, SAC algorithm, etc., which are used to optimize the robot control strategy in a high-dimensional continuous action space. During the training process, the policy network takes the current state data of the robot as input, including perception information such as attitude sensor data, wheel-foot rotation speed data, torque data, and position and speed data, and outputs the corresponding action strategy. This action strategy can be used to guide the robot to determine whether to execute a takeoff action and how to control various parameters during the takeoff process when detecting a height obstacle. Specifically, the policy network can output control instructions for the takeoff start signal, squatting amplitude, rising speed, and forward rotation speed of the wheel-foot structure, so as to achieve a stable and efficient obstacle-crossing behavior of the robot.
[0107] In the exemplary embodiment of the present disclosure, a height obstacle control device for a two-wheeled biped robot based on deep reinforcement learning is also provided. Refer to Figure 11 As shown, the height obstacle control device 1100 for a two-wheeled biped robot based on deep reinforcement learning may include: An obstacle determination module 1101, which can be used to determine that the robot encounters a height obstacle when the motion feedback data of the two-wheeled biped robot meets a preset condition; An action control module 1102, which can be used to control the two-wheeled foot robot to perform a takeoff action. The takeoff action includes the main body quickly rising after squatting, and controlling the two-wheeled foot structure of the two-wheeled foot robot to rotate forward quickly while attaching to the height obstacle until the motion feedback data of the two-wheeled foot robot does not meet the preset conditions. Among them, the takeoff action is output by a policy network trained based on a deep reinforcement learning algorithm in a height obstacle terrain.
[0108] The specific details of the above obstacle determination module 1101 and action control module 1102 have been described in detail in the corresponding method for controlling a two-wheeled foot robot through height obstacles based on deep reinforcement learning, so they will not be elaborated here.
[0109] In an exemplary embodiment of the present disclosure, a training device for a motion control model of a two-wheeled foot robot based on deep reinforcement learning is also provided. Refer to Figure 12 As shown, the training device 1200 for a motion control model of a two-wheeled foot robot based on deep reinforcement learning may include: A training data acquisition module 1201, which can be used to construct a virtual training scenario including a height obstacle terrain and acquire the motion data of a virtual two-wheeled foot robot in the virtual training scenario. A model training module 1202, which can be used to train a policy network according to the motion data and environmental information by using a deep reinforcement learning algorithm. The policy network outputs the following action strategy for the two-wheeled foot robot: when encountering a height obstacle, perform a takeoff action, and the takeoff action includes the main body quickly rising after squatting, and controlling the two-wheeled foot structure of the two-wheeled foot robot to rotate forward quickly while attaching to the height obstacle until the height obstacle is crossed.
[0110] The specific details of the above training data acquisition module 1201 and model training module 1202 have been described in detail in the corresponding method for training a motion control model of a two-wheeled foot robot based on deep reinforcement learning, so they will not be elaborated here.
[0111] In an exemplary embodiment of the present disclosure, a robot is also provided. The robot includes a processor and a memory, and computer-readable instructions are stored on the memory. When the computer-readable instructions are executed by the processor, the above method is implemented. Among them, the robot is a robot configured with a function of controlling a two-wheeled foot robot through height obstacles based on deep reinforcement learning. Exemplarily, the robot includes any one of a humanoid robot, a cleaning robot, a transportation robot, and a mobile robot. Refer to Figure 13 As shown, it shows a schematic diagram of a two-wheeled foot robot.
[0112] Refer to Figure 14As shown, an electronic device capable of implementing the above method is also provided. Among them, the electronic device 1400 includes a processor 1401 and a memory 1402. A computer-readable instruction is stored on the memory 1402. When the computer-readable instruction is executed by the processor 1401, the method for controlling a two-wheeled bipedal robot through a height obstacle based on deep reinforcement learning or the method for training a motion control model of a two-wheeled bipedal robot based on deep reinforcement learning in the implementation of the present disclosure is realized.
[0113] In an exemplary embodiment of the present disclosure, a computer-readable storage medium is also provided, on which computer program code instructions are stored. When the computer program code instructions are called by the processor of the robot, the robot is enabled to execute the method as described in the embodiment.
[0114] Referring to Figure 15 As shown, a program product 1500 for implementing the above method for controlling a two-wheeled bipedal robot through a height obstacle based on deep reinforcement learning or the method for training a motion control model of a two-wheeled bipedal robot based on deep reinforcement learning according to an embodiment of the present disclosure is described. It can adopt a portable compact disc read-only memory (CD-ROM) and include program code, and can run on a terminal device. However, the program product of the present disclosure is not limited to this. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or device.
[0115] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described here can be implemented by software or in a manner of software combined with necessary hardware. Therefore, the technical solution according to the embodiment of the present disclosure can be embodied in the form of a software product. This software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on the network, including several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiment of the present disclosure.
[0116] Finally, the above preferred embodiments are only used to illustrate the technical solution of the present application and are not restrictive. Although the present application has been described in detail, those skilled in the art should understand that changes in form and details can be made to it without departing from the scope defined by the claims of the present application. The dimensions of the drawings have nothing to do with the specific physical object, and the physical dimensions can be arbitrarily changed.
Claims
1. A method for controlling a two-wheeled robot to pass through high obstacles based on deep reinforcement learning, characterized in that: include: When the motion feedback data of the two-wheeled foot robot meets the preset conditions, it is determined that the robot encounters a height obstacle; Controlling the two-wheeled foot robot to perform a take-off action, wherein the take-off action includes rapidly rising after the main body squats, and controlling the two-wheeled foot structure of the two-wheeled foot robot to adhere to the height obstacle and rapidly rotate forward until the motion feedback data of the two-wheeled foot robot does not meet the preset conditions; The take-off action is output by a strategy network trained based on a deep reinforcement learning algorithm in highly obstacle terrain.
2. The method for controlling a two-wheeled legged robot to pass through a high obstacle according to claim 1, characterized in that: The motion feedback data includes wheel-foot rotation speed data, and the preset condition is that a difference between a target rotation speed and an actual rotation speed of the dual-wheel-foot exceeds a preset threshold.
3. The method for controlling a two-wheeled legged robot to pass through a high obstacle according to claim 2, characterized in that: The actual rotation speed is obtained by collecting the motor encoder installed on the two-wheeled leg robot; the target rotation speed is the expected rotation speed instruction issued by the controller to the driving motor of the two-wheeled leg robot.
4. The method for controlling a two-wheeled legged robot to pass through a high obstacle according to claim 2, characterized in that: The difference between the target rotational speed and the actual rotational speed of the two-wheeled foot exceeds a preset threshold value, including: the difference between the target rotational speed and the actual rotational speed of the two-wheeled foot exceeds 30% of the target rotational speed.
5. The method for controlling a two-wheeled legged robot to pass through a high obstacle according to claim 2, characterized in that: The preset condition also includes: when the duration of the difference between the target rotation speed and the actual rotation speed of the two-wheeled foot exceeding the preset threshold exceeds a preset time length.
6. The method for controlling a two-wheeled legged robot to pass through a high obstacle according to claim 5, characterized in that: The preset duration is greater than 200 milliseconds.
7. The method for controlling a two-wheeled legged robot to pass through a high obstacle according to claim 1, characterized in that: The subject squats and then quickly rises, including: Controlling the two-wheeled foot robot body to squat briefly by lowering the center of gravity and storing energy; The dual-wheel drive mechanism provides an upward thrust, allowing the robot to jump off the ground.
8. The method for controlling a two-wheeled legged robot to pass through a high obstacle according to claim 1, characterized in that: The dual-wheeled foot structure attaches to a high obstacle and quickly rotates forward, comprising: The driving torque of the two-wheeled foot is adjusted to actively and quickly rotate in the forward direction at a speed higher than the normal driving speed to generate effective friction to assist the two-wheeled foot robot in climbing.
9. The method for controlling a two-wheeled legged robot to pass through a height obstacle according to any one of claims 1 to 8, characterized in that: The policy network is trained by the following method: In a simulated environment with highly obstacle-ridden terrain, a policy network that outputs action policies for controlling robot motion is trained using a teacher-student model framework.
10. The method for controlling a two-wheeled legged robot to pass through a high obstacle according to claim 1, characterized in that: The high obstacle terrain includes multiple types of high obstacles; different types of high obstacles differ in at least one of the following parameters: The overall height of the obstacle, the surface roughness of the obstacle side wall, the stiffness of the obstacle side wall material, the geometric shape of the obstacle top, the surface flatness of the obstacle side wall, and the inclination angle between the obstacle side wall and the ground.
11. The method for controlling a two-wheeled legged robot to pass through a high obstacle according to claim 9, characterized in that: The training process specifically includes: The student encoder encodes the historical motion state information of the two-wheeled legged robot and generates a first latent vector, and the teacher encoder encodes the privileged state information of the two-wheeled legged robot and generates a second latent vector; the privileged state information includes height obstacle information obtained from the simulation environment; Selecting a first latent vector or a second latent vector according to a preset strategy and inputting the first latent vector into the strategy network; Outputting, through the strategy network, a motion strategy for controlling the movement of the two-wheeled foot robot based on the received potential vector and current motion state information; the current motion state information at least includes: posture sensor data, wheel foot speed data, torque data, and position and speed data; Utilize privileged state information through the value network and output a value estimate of the current state; optimizing parameters of the student encoder based on a difference between the first latent vector and the second latent vector; Based on the action strategy and the value estimate, the parameters of the policy network are updated using a deep reinforcement learning algorithm.
12. The method for controlling a two-wheeled legged robot to pass through a high obstacle according to claim 9, characterized in that: The deep reinforcement learning algorithm includes a comprehensive reward function for evaluating the action performance of the two-wheeled leg robot in the simulation environment, and the reward function includes the success rate of passing height obstacles, take-off height and posture stability.
13. A method for training a motion control model of a two-wheeled legged robot based on deep reinforcement learning, used for training the control model in the method for controlling a two-wheeled legged robot passing through a high obstacle as claimed in claim 1, characterized in that: include: Constructing a virtual training scene including a highly obstructed terrain, and acquiring motion data of a virtual two-wheeled legged robot in the virtual training scene; A deep reinforcement learning algorithm is used to train a strategy network based on the motion data and environmental information; the strategy network outputs the following two-wheeled legged robot action strategy: When encountering a height obstacle, a take-off action is performed, wherein the take-off action includes the main body squatting and then rising rapidly, and the two-wheeled foot structure of the two-wheeled foot robot is controlled to adhere to the height obstacle and rotate rapidly forward until the height obstacle is crossed.
14. A control device for a two-wheeled robot passing through high obstacles based on deep reinforcement learning, characterized in that: include: An obstacle determination module, used for determining that the robot encounters a height obstacle when the motion feedback data of the two-wheeled foot robot meets a preset condition; an action control module, used to control the two-wheeled foot robot to perform a take-off action, wherein the take-off action includes a main body squatting and then rising rapidly, and controlling the two-wheeled foot structure of the two-wheeled foot robot to adhere to the height obstacle and rotate rapidly forward until the motion feedback data of the two-wheeled foot robot does not meet the preset conditions; The take-off action is output by a strategy network trained based on a deep reinforcement learning algorithm in highly obstacle terrain.
15. A two-wheeled foot robot motion control model training device based on deep reinforcement learning, characterized in that: include: A training data acquisition module, used to construct a virtual training scene including a highly obstructed terrain, and to acquire motion data of a virtual two-wheeled legged robot in the virtual training scene; The model training module is used to use a deep reinforcement learning algorithm to train a strategy network based on the motion data and environmental information; the strategy network outputs the following two-wheeled legged robot action strategy: When encountering a height obstacle, a take-off action is performed, wherein the take-off action includes the main body squatting and then rising rapidly, and the two-wheeled foot structure of the two-wheeled foot robot is controlled to adhere to the height obstacle and rotate rapidly forward until the height obstacle is crossed.
16. An electronic device, characterized in that: include: processor; as well as A memory having computer-readable instructions stored thereon, wherein the computer-readable instructions, when executed by the processor, implement the method according to any one of claims 1 to 13.
17. A two-wheeled foot robot, characterized in that: include: processor; as well as A memory having computer-readable instructions stored thereon, wherein the computer-readable instructions, when executed by the processor, implement the method according to any one of claims 1 to 13.
18. The robot according to claim 17, characterized in that: The robot is a robot equipped with a two-wheeled and footed robot control function for passing high obstacles based on deep reinforcement learning, and the robot includes any one of a humanoid robot, a cleaning robot, a transport robot, and a mobile robot.
19. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program code instructions, and when the computer program code instructions are called by a processor of the robot, the robot executes the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Multi-posture switching two-wheeled robot and working method thereof
CN115892277A
Wheel-legged robot control algorithm based on model prediction and deep reinforcement learning
CN118192558A
Wheel-leg robot wheel-foot switching control method based on BP neural network
CN116859975A
Sleep type obstacle, orientation recognition method thereof and robot control method
CN119217357A
High power passive obstacle crossing robot
CN2673583Y
Cited By
Motion control strategy network training method and device for foot robot with floating substrate based on reinforcement learning
CN121179441A