Torque control foot type robot with rotatable waist and control method thereof

Through the torque control method of waist rotatable, the gait dataset is optimized by imitation learning framework and GAIL training, which solves the problems of high control complexity and insufficient axial rotation of the four-legged robot, and achieves flexible steering and stability improvement, adapting to complex terrain and external disturbances.

CN120386181AInactive Publication Date: 2025-07-29FUDAN UNIVERSITY

Patent Information

Application Number
CN202510299739.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-29
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing four-legged robot has high control complexity and insufficient axial rotation ability, making it difficult to achieve efficient and dynamic response in complex terrain. It relies on the hip joint when steering, and has poor scene adaptability.

Method used

Using a torque control method with rotatable waist, a simulation learning framework based on asymmetric actor-critic structure is constructed, combined with Generative Adversarial Imitation Learning (GAIL) is used for iterative training, and the waist adaptive gait data set is optimized, and the waist motor output torque is used to achieve full-body control.

Benefits of technology

It reduces control complexity, improves steering ability and scene adaptability, and achieves flexible steering and stability, especially excellent performance under complex terrain and external disturbances.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386181A_ABST
    Figure CN120386181A_ABST
Patent Text Reader

Abstract

The invention relates to a torque control foot type robot with a rotatable waist and a control method of the torque control foot type robot. The control method comprises the following steps that an initial gait data set of the foot type robot is obtained; a simulation learning framework based on an asymmetric act-critic structure is constructed; performing iterative training on a strategy based on the simulation reward and the task reward by utilizing the simulation learning framework to obtain a gait data set after waist self-adaptive training; based on the gait data set after waist self-adaptive training, motor parameters and environment parameters of the foot type robot are trained through course learning, and a final gait data set is obtained; based on the final gait data set, the waist of the foot type robot is made to rotate by controlling the output torque of a waist motor, and whole-body control over the foot type robot is achieved. Compared with the prior art, the method has the advantages of low control complexity, high steering capability, good scene adaptability and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of legged robots, and more particularly to a torque-controlled legged robot with a rotatable waist and a control method thereof. Background Art

[0002] In nature, the flexible and agile movements of many animals rely on the waist as the core power support. When a cheetah chases prey at high speed, it needs to rely on the waist strength to maintain stability during the high-speed movement. A cat can always rely on rotating its waist in the air to ensure that its four feet can land smoothly. Inspired by this, some work has focused on improving and optimizing the torso structure of quadruped robots to achieve more agile motion performance. For example, some solutions add an inertial tail with 3 degrees of freedom to the robot torso to enhance the agility of movement and improve landing safety. Some solutions introduce a 4-degree-of-freedom spine to achieve good speed following of the robot dog under various gaits. Although these works have improved the motion performance of quadruped robots in different aspects, it cannot be ignored that the introduction of complex mechanisms with multiple degrees of freedom often brings huge challenges in control.

[0003] Chinese Patent Application Publication No. CN118907264A provides a leg structure including a rotating component and a telescopic foot component, and realizes complex rotational movements between the thigh and the calf through six rotating motors, aiming to improve the adaptability and movement flexibility of the robot in complex terrains. It mainly provides a modular design of the leg structure and a safe recycling mechanism for the telescopic foot. However, on the one hand, this application focuses on the optimization of the leg structure and does not involve the design of the axial rotation function of the spine. Therefore, the problem of dependence on the hip joint during turning cannot be solved, and the robot still needs to achieve turning through the refined control of the leg joints. On the other hand, although the leg structure design of this application improves flexibility, the introduction of multiple rotating motors increases the complexity of the control system and may exacerbate the energy consumption problem.

[0004] In summary, the prior art has the following disadvantages:

[0005] (1) High control complexity: Existing quadruped robots improve their motion performance by introducing multi-degree-of-freedom spine or inertial tail structures. However, the complex mechanism design significantly increases the control difficulty. Coordinating the movements of multiple degrees of freedom requires high-precision algorithms and sensor support, increasing the system energy consumption and development cost.

[0006] (2) Insufficient axial rotation ability: The research on low-degree-of-freedom spines mainly focuses on radial motion, and there is less research on the practical application of spinal axial rotation, still remaining at the simulation level. This design limits the stability and maneuverability of the robot during turning, resulting in the need to rely on the adduction / abduction movements of the hip joint during turning and unable to fully utilize the potential of the bionic spine.

[0007] (3) Poor adaptability to actual scenarios: The existing control framework has insufficient adaptability to the low-degree-of-freedom spine and is difficult to achieve efficient dynamic response in complex terrains. For example, it is easy to lose balance when making a quick turn or sudden stop. Summary of the Invention

[0008] The purpose of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a torque-controlled legged robot with a rotatable waist and its control method to solve or partially solve the problems of insufficient axial rotation ability and high control complexity of legged robots.

[0009] The purpose of the present invention can be achieved through the following technical solutions:

[0010] In one aspect of the present invention, a control method for a torque-controlled legged robot with a rotatable waist is provided, including the following steps:

[0011] Obtain the initial gait data set of the legged robot;

[0012] Construct an imitation learning framework based on an asymmetric actor-critic structure;

[0013] Using the imitation learning framework, iteratively train the policy based on imitation rewards and task rewards to obtain a gait data set after waist adaptation training;

[0014] Based on the gait data set after waist adaptation training, train the motor parameters and environmental parameters of the legged robot through curriculum learning to obtain the final gait data set;

[0015] Based on the final gait data set, rotate the waist of the legged robot by controlling the output torque of the waist motor to achieve the overall control of the legged robot.

[0016] As a preferred technical solution, in the imitation learning framework, the action space of the legged robot includes the output torques of the waist motor and the leg motors, and the state space of the legged robot includes the base quaternion, forward speed, yaw angular velocity, joint positions, joint speeds, and the joint positions at the previous moment.

[0017] As a preferred technical solution, in the imitation learning framework based on an asymmetric actor-critic structure, the input of the actor network is data matching the state space, and the input of the critic network includes data matching the state space and the base linear velocity.

[0018] As a preferred technical solution, in the imitation learning framework, the discriminator takes the states corresponding to the trajectory segments in a historical period of time as inputs and calculates the least-squares loss. Among them, the states input to the discriminator include the base linear velocity, the base angular velocity, the base normal vector, the base height, the joint positions, and the joint velocities.

[0019] As a preferred technical solution, during the process of iteratively training the policy, it also includes a yaw angular velocity tracking reward in the z-axis direction for encouraging the robot to use the waist to turn and track in a timely manner, a foot-lifting reward for avoiding frequent contact between the feet and the ground, and an action change constraint reward for restricting the movement frequency of the robot.

[0020] As a preferred technical solution, after obtaining the gait dataset after the waist adaptive training, it also includes:

[0021] Performing domain randomization processing on the robot joint friction, the base mass, the initial joint angles, and the centroid position, etc.

[0022] As a preferred technical solution, the process of training the motor parameters and the environmental parameters of the legged robot through curriculum learning includes the following steps:

[0023] Judge whether the walking distance of the robot under the current motor parameters and environmental parameters reaches the preset requirements;

[0024] If the walking distance of the robot reaches the preset requirements, increase the curriculum difficulty and move under a more difficult curriculum during the next reset;

[0025] If the moving distance of the robot is less than half of the target distance at the end of a preset period of time, reduce the curriculum difficulty;

[0026] If the highest-level curriculum is completed, loop to a randomly selected curriculum level.

[0027] Another aspect of the present invention provides a torque control legged robot with a rotatable waist for implementing the foregoing control method. The legged robot includes:

[0028] A robot main body, the base is a front-rear symmetric structure that can rotate along the axis. Among them, both symmetric sides include drive motors, the rotation directions of the two drive motors are opposite, and each drive motor is respectively connected to a bipolar belt transmission. The connecting shaft inside the base is hollow for accommodating power supply and communication cables.

[0029] As a preferred technical solution, the base includes a front half-base and a rear half-base. The front half-base is connected to the two front legs, and the rear half-base is connected to the two rear legs.

[0030] As a preferred technical solution, it further includes:

[0031] Main board;

[0032] Multiple FOC control boards, connected to the main board and respectively connected to the motors of the waist, front legs and hind legs, to achieve field-oriented control of the motors;

[0033] IMU unit, connected to the main board.

[0034] Compared with the prior art, the present invention has at least one of the following beneficial effects:

[0035] (1) Low control complexity: By only adding 1 degree of freedom of waist rotation, through data set migration and collaborative optimization, the present invention avoids the design of complex reward functions, improves the training efficiency, decouples the mechanical structure and the control algorithm, simplifies the hardware implementation, and reduces the manufacturing cost and maintenance difficulty. Through the combination of a low-degree-of-freedom waist mechanism and Generative Adversarial Imitation Learning (GAIL) migration optimization, complex motion performance is achieved with a simple structure.

[0036] (2) Strong steering ability: The present invention realizes flexible steering through the axial rotation of the waist. Compared with the traditional method that relies on the hip joint, the yaw angular velocity tracking of Solo9 is more stable, the turning radius is smaller, and it is even better than Solo12 with a higher degree of freedom. High-difficulty tasks can be completed in actual machine tests, overcoming the uncontrollable offset problem of Solo8 caused by ground friction.

[0037] (3) Good scene adaptability: The waist mechanism of the present invention allows dynamic adjustment of the relative positions of the front and rear bases. In rough terrains, the leg support stability is significantly optimized. In anti-disturbance tests, the survival rate of Solo9 under large-amplitude random speed disturbances has been significantly improved, and the waist adaptive adjustment effectively absorbs external impacts. The combination of domain randomization and curriculum learning improves the generalization ability of the strategy to unknown environments, verifying the practical value of the waist mechanism in complex scenarios. Description of the Drawings

[0038] Figure 1 It is a flowchart of the control method of the torque control legged robot with a rotatable waist in the embodiment;

[0039] Figure 2 It is a schematic diagram of the whole-body control - data set optimization of the robot in the embodiment;

[0040] Figure 3 It is a schematic diagram of the torque control legged robot with a rotatable waist in the embodiment;

[0041] Figure 4For the comparison of the turning effects of the Solo8, Solo9, and Solo12 robots in the leap and trot motions under the conditions of a forward speed of 0.6 m / s and a yaw angular velocity of -0.4 rad / s in the embodiment;

[0042] Figure 5 It is a schematic diagram of the electronic device in the embodiment. Specific implementation manners

[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0044] Embodiment 1

[0045] Aiming at the problems existing in the foregoing prior art, this embodiment provides a control method for a torque-controlled legged robot with a rotatable waist, which realizes the full-body control of the quadruped robot Solo9 with a waist based on Generative Adversarial Imitation Learning (GAIL). During the training process, the discriminator part input is used to optimize the policy and the data set in multiple rounds, so as to realize the flexible turning motion of Solo9 in multiple gaits. Finally, extensive tests are carried out on the turning ability, terrain adaptability, and robustness of Solo9 in simulation and on the real machine, and detailed comparisons are made with Solo8 and Solo12, which proves the effectiveness of the control algorithm and the advantages of the new waist mechanism.

[0046] See Figure 1 and Figure 2 , this method includes the following steps:

[0047] Step S1, obtain the initial gait data set of the legged robot.

[0048] In this embodiment, a data set of multiple high-quality gaits open-sourced by the Solo8 robot is used.

[0049] Step S2, construct an imitation learning framework based on an asymmetric actor-critic structure.

[0050] Reinforcement Learning RL Architecture: The action space of Solo9 is defined as the torques output by 9 motors, including a waist motor and 8 leg motors, with 2 motors on each leg to drive the thigh and calf. The basic state space is constructed based on the robot's proprioception and has a total of 33 dimensions, namely the base quaternion (4 dimensions), two-dimensional control commands: including forward speed and yaw angular velocity, the positions (9 dimensions), speeds (9 dimensions), and positions at the previous moment (9 dimensions) of 9 joints. In the physical machine, it is obtained through the sensor readings of brushless motor driver boards and inertial measurement units (IMUs). Considering that it is difficult to accurately obtain the base linear velocity in the actual environment, this embodiment adopts an asymmetric actor-critic structure. The input dimension of the actor network is the 33-dimensional basic state space, and the critic network adds this privileged observation of the base linear velocity on this basis (which can be directly obtained in the simulation environment), so the state space is 36 dimensions.

[0051] Imitation Learning Architecture: The framework adopts generative adversarial imitation learning, considering the imitation observation space O I , mapping the complete state space S of the underlying Markov decision process to the imitation observation space O through the function f I , using s t = (s t-HI+1 ,..., s t ) to represent the state corresponding to the trajectory segment of length HI at time t. The discriminator is optimized using the least squares GAN (LSGAN) loss, and the optimization goal is to make the expert data and the generated data (data generated by the policy) as close as possible in the discriminator output, thereby guiding the optimization of the policy. Use HI-step input and gradient penalty. The discriminator can consider the input observations including the following parts, namely the base position (3 dimensions), base quaternion (4 dimensions), base linear velocity (3 dimensions), base angular velocity (3 dimensions), base normal vector (3 dimensions), base height (1 dimension), joint positions (9 dimensions), and joint speeds (9 dimensions). Considering that the base position and base yaw angular velocity will interfere with the robot's turning training, they are ignored in this embodiment. The base quaternion and the normal vector have similar representation capabilities here, and the normal vector is used instead of the quaternion during actual training. In this way, the actual observations of the discriminator are the base linear velocity, base angular velocity, base normal vector, base height, joint positions, and joint speeds, a total of 27 dimensions.

[0052] Step S3, using the imitation learning framework, iteratively train the policy based on the imitation reward and the task reward to obtain the gait dataset after the waist adaptive training.

[0053] First, based on the Isaac Gym environment, using the open-source initial gait dataset, after preliminary training, various robust gaits of Solo8 are obtained, such as trot, leap, etc. Then, the selected gaits are stored, and waist information data is supplemented in the sequence. In this embodiment, the waist data is all set to 0 here as the initial gait dataset for Solo9 training.

[0054] Discriminator partial observation: Considering that the initial dataset of Solo9 lacks useful waist information, it will be difficult to play the role of the waist directly in imitation learning. To address this issue, in this embodiment, the discriminator is not allowed to input waist and other turning-related information, allowing the waist to adaptively train and encouraging the robot to use the waist to learn turning actions. If all observations are directly input into the discriminator, it will lead to a conflict between the imitation reward and the turning reward in the task reward, affecting the training effect. Additionally, in this stage, additional gait rewards are added for fine-tuning, such as adding foot-lifting rewards and slip penalties to promote the sim2real work.

[0055] Dataset and gait co-optimization: Considering that the imitation reward largely standardizes the robot's behavioral actions, only a small number of additional reward item weights need to be adjusted, and the action sequences generated by the new policy are stored as a new reference dataset for iterative training. This embodiment plays a significant role in the Solo9 gait learning process based on the co-optimization idea of policy & dataset. It can decouple the connections between different reward functions, steadily improve the quality of the dataset gait, and gradually stimulate the potential of the waist in aspects such as robot turning and robustness. As the iteration progresses, while increasing the weight of the imitation reward, the weight of the task reward is continuously reduced to steadily improve the quality of the robot's movement in a relatively small number of iteration rounds.

[0056] Reward engineering: During the training process, a yaw angular velocity tracking reward for the z-axis direction is introduced to encourage the robot to use the waist to perform turning tracking in a timely manner. Specifically, foot-lifting rewards and foot slip penalties are introduced to avoid frequent contact between the feet and the ground. In addition, considering that the robot's movement frequency has a significant impact on sim2real, too high or too low movement frequencies will both lead to failure in the actual deployment process. To address this issue, this embodiment restricts it by adding an action change constraint reward item action_rate. The key reward items in the experimental process are presented in groups in Table 1.

[0057] Table 1 Training reward items

[0058]

[0059] Step S4: Based on the gait dataset after waist adaptive training, the motor parameters and environmental parameters of the legged robot are trained through curriculum learning to obtain the final gait dataset.

[0060] Domain randomization: To improve the robustness of the robot's policy, extensive domain randomization is carried out during training. Based on the open-source legged_gym project, in this embodiment, extensive domain randomization processing is performed on the robot's joint friction, base mass, initial joint angles, center of mass position, etc. It should be noted that domain randomization needs to be carried out after determining the trot, leap, and crawl gait datasets of Solo9 to reduce the impact of domain randomization on the Solo9 gait. See Table 2 for the results after domain randomization processing.

[0061] Table 2 Results of domain randomization processing

[0062]

[0063] Curriculum learning: Considering that the PD parameters of the robot's motors and the rugged terrain height parameters in the environment have a great impact on the training complexity, directly using large parameter settings at the beginning of training is likely to cause the policy to fall into a local optimum, resulting in performance degradation or even inability to learn. In this embodiment, the above parameters are trained through the method of curriculum learning. Specifically, if the robot walks a distance that meets the requirements under the current parameters, the curriculum difficulty is increased, and it will move under a more difficult curriculum during the next reset. However, if the distance the robot moves is less than half of the target distance at the end of a period of time, the curriculum difficulty will be reduced. The robot that completes the highest-level curriculum is cycled back to a randomly selected level to increase diversity and avoid catastrophic forgetting.

[0064] Step S5: Based on the final gait dataset, the waist of the legged robot is rotated by controlling the output torque of the waist motor to achieve the full-body control of the legged robot.

[0065] Embodiment 2

[0066] Based on Embodiment 1, this embodiment provides a torque-controlled legged robot with a rotatable waist, namely Solo9. See Figure 3 (a), (b), (c), and (d) are respectively the schematic diagram of the Solo9 waist structure, the schematic diagram of the Solo9 whole-body structure, the top view of the Solo9 waist structure, and the top view of the Solo9 whole-body structure.

[0067] Solo9 is based on Solo8, with the pedestal modified into an axially rotatable and front-back symmetric structure. Specifically, the waist of Solo9 consists of two high-torque brushless motors and a low-gear ratio transmission. The pedestal is divided into two front-back symmetric parts, and rotation is achieved by the drive motors (T-MOTOR MN5008 340KV) near the waist on each part. The two motors rotate in opposite directions, and FOC and a high-precision encoder are used to ensure that the rotational speed of each motor remains consistent. Each motor is connected to a 9:1 bipolar belt transmission. The connecting shaft between the front and rear pedestals is designed to be hollow to allow power supply and communication cables to pass through. We do not impose mechanical structural limitations on the rotation angle of the robot's waist. Solo9 is 46.5 cm in body length and weighs 2.3 kg (Solo8 is about 43 cm and weighs 1.8 kg), with a width of 31 cm, the same as Solo8.

[0068] By dividing the pedestal into two front-back symmetric parts, with a rotary motor near the waist on each part, they jointly drive to achieve powerful axial movement of the waist. This enables Solo9 not only to have a flexible steering ability that Solo8 does not possess but also to enhance the robot's robustness in dealing with complex terrains and disturbances.

[0069] Solo9 uses the same brushless motor driver boards as Solo8 for Field Oriented Control (FOC) of the motors. In this embodiment, an additional control board is added to drive the two motors of the waist. The main board controls 5 FOC driver boards, with 3 placed on the rear pedestal to drive the motors of the two rear legs and the waist, and 2 placed on the front pedestal to drive the motors of the two front legs. The main board and the IMU are installed at the front end. The main board can be connected to a control computer through an Ethernet or WiFi network interface. The robot senses the pose of the front end of the body through the IMU and jointly calculates the pose of the rear end of the body through the angles of the waist motors and the IMU.

[0070] Construction of the simulation model: In this embodiment, a urdf file of Solo9 is constructed, with the front half of the pedestal as the root tag, connecting the two front legs and the rear half of the pedestal, and the rear half of the pedestal is then connected to the two rear legs. Since the single-joint multi-motor structure is not supported in isaac gym, only a single motor is used to drive the waist rotation in the simulation environment. To depict the relative angular positions of the front and rear pedestals, the relative joint angle between the pedestals is defined. Similar to the leg joints, the relative joint angle and relative angular velocity between the front and rear pedestals are added to the observations. Since two motors are used to drive the waist joint in the real machine, the output torque of the motor in the simulation model is twice that of a single motor in the real machine.

[0071] To demonstrate the rationality and effectiveness of the Solo9's movable waist mechanism, we conducted steering, terrain, and disturbance tolerance tests in both simulated and real-world environments. The experimental results show that, despite only adding one degree of freedom compared to the Solo8, the Solo9 possesses more flexible and stable steering capabilities, even outperforming the Solo12 robot, which has even more degrees of freedom. Furthermore, proper waist control can significantly improve the robot's motion robustness. The specific testing process is as follows:

[0072] (1) Steering ability test

[0073] Simulation test: We tested the steering capability of Solo9 in three gaits: trot, leap, and crawl. We recorded the yaw angular velocity command and the actual yaw angular velocity of the robot in each gait, and kept the forward linear velocity command at 0.6m / s. We tested Solo8, Solo9, and Solo12 respectively. Figure 4 As shown, the left side is a schematic diagram of the trot gait of the three robots Solo12, Solo8, and Solo9 from top to bottom, and the right side is the yaw angular velocity of the corresponding gait. When the angular velocity command is -0.4rad / s, the yaw angle tracking ability of Solo8 is significantly weaker than that of Solo9 and Solo12. Compared with Solo12, Solo9 has more stable yaw angular velocity tracking and a smaller turning radius. This also reflects the advantage of waist steering over hip steering. According to Figure 4 , YOLO9 has the smallest turning radius and the most stable yaw rate.

[0074] Actual machine test: Solo9 is controlled to complete the steering tasks of the specified path, including small 90-degree steering movements (the diagonal of the black carpet in the figure), large 180-degree steering movements, circular movements, "S"-shaped steering tasks, and even using its waist to complete in-situ steering movements on a carpet with a high friction coefficient. Solo9 can complete these tasks very well. Compared with Solo9, the Solo8 robot lacks the freedom of hip adduction and abduction. Although it also shows a certain steering tracking ability in the simulation environment, in the actual machine deployment, due to the influence of ground friction, Solo8 will often have irregular, small and uncontrollable steering offsets when walking, resulting in the basic loss of yaw angular velocity tracking ability. Due to hardware conditions, this embodiment did not test the steering ability of the Solo12 actual machine, but the simulation shows that Solo9 has a steering ability comparable to or even better than Solo12, and has fewer joint degrees of freedom.

[0075] (2) Terrain test

[0076] The motion stability of Solo8 and Solo9 in complex terrains was compared in simulation and real environments to verify the terrain adaptability of the waist mechanism. In the simulation environment, two types of rugged terrains were generated for testing, including an irregular ground with a height range of ±3.5 cm and irregular steps with the initial heights of the irregular steps randomized at 2.5 cm, 2.75 cm, and 2.9 cm. Both robots were trained on flat terrains. By counting the survival rates of the robots after running a given number of time steps, their robustness in complex terrains was characterized. The experimental results are shown in Table 3. It can be seen that whether it is the irregular ground or the irregular step terrain, thanks to the flexible rotation of the waist mechanism, Solo9 has a higher survival rate, and its robustness has an obvious advantage compared to Solo8. As shown in the table, the walking success rate of Solo9 on irregular steps is nearly 40% higher than that of Solo8, indicating that the reasonable torsion of the waist has a significant effect on improving the lateral stability of the robot.

[0077] Table 3 Terrain Test Results of Solo8 and Solo9

[0078]

[0079] In the real machine test environment, a ramp was used to test the robustness of Solo8 and Solo9 in complex terrains. Two groups of experiments were designed, including walking in an inclined state and terrain crossing tasks, and both of these two experimental scenarios were unknown environments that had never been encountered in the simulation. The experimental results show that Solo9 can dynamically adjust the relative positions of the front and rear trunks by using the rotatable waist, so as to ensure that the legs can strongly support the ground and absorb stress, and obtain higher stability on complex terrains. In contrast, since the trunk of Solo8 is a whole, there are often moments when the legs miss the ground when it crosses an inclined plane, resulting in a higher probability of tipping over and even damage to the robot.

[0080] (3) Anti-Disturbance Test

[0081] The adaptability of Solo8 and Solo9 robots to external disturbances was tested in a simulation environment. Both robots were trained in the same environmental configuration. In the test environment, we applied random velocity disturbances of ±0.5 m / s, ±0.7 m / s, and ±1 m / s in the x and y directions to the robot's base at each time step, and counted the survival rate of the robot after running for a given time. Table 4 summarizes the survival rates of Solo8 and Solo9. The results show that as the disturbance increases, the success rates of both robots decrease. However, due to the adaptability of the waist to disturbances, the anti-disturbance performance of Solo9 is always better than that of Solo8. Especially under large disturbances, the success rate can be increased by 15.6%. In addition, an ablation experiment was conducted on the waist control method of Solo9. Among them, Solo9 fixed means fixing the waist angle at 0, and Solo9 free means not actively controlling the waist and allowing it to move freely. The results show that when the waist is fixed, Solo9 has the same structure as Solo8 in the simulation, so the success rates of the two are similar. It should be noted that if the waist is not controlled and allowed to move freely, Solo9 cannot walk normally at all. This also indicates that after adding a waist mechanism, an appropriate motion control strategy is needed to maximize the advantages of the waist.

[0082] Table 4 Test results of anti-disturbance of Solo8 and Solo9

[0083]

[0084] In summary, the present invention provides a novel quadruped robot Solo9 with a simple and efficient axially rotatable waist. And a whole-body control algorithm for a quadruped robot with a waist based on the idea of collaborative optimization of dataset-strategy is proposed. Only by using the linear motion dataset of the Solo8 robot, the GAIL method, combined with instruction guidance and through multiple rounds of dataset and strategy iteration, can Solo9 flexibly turn and walk steadily in various gaits. Extensive experiments were carried out in simulation and real environments. The results show that although Solo9 only has one more degree of freedom than Solo8, it has higher mobility and can track a larger steering angular velocity, even better than the Solo12 robot with a higher degree of freedom. Under rough terrain and external disturbances, the survival rate of Solo9 is significantly better than that of Solo8 robot, indicating that the rotatable waist structure is very helpful for legged robots to adapt to complex environments.

[0085] Embodiment 3

[0086] This embodiment provides an electronic device, including: one or more processors and a memory. The memory stores one or more programs, and the one or more programs include instructions for executing the control method of the torque-controlled legged robot with a rotatable waist as described in Embodiment 1.

[0087] As Figure 5 described above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 described method. Of course, in addition to the software implementation, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but may also be hardware or a logic device.

[0088] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0089] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0090] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A control method for a torque-controlled legged robot with a rotatable waist, characterized in that, It includes the following steps: Obtain the initial gait dataset of the legged robot; Construct an imitation learning framework based on an asymmetric actor-critic structure; Utilize the imitation learning framework to iteratively train the policy based on imitation rewards and task rewards to obtain a gait dataset after lumbar adaptative training; Based on the gait dataset after lumbar adaptative training, train the motor parameters and environmental parameters of the legged robot through curriculum learning to obtain the final gait dataset; Based on the final gait dataset, rotate the waist of the legged robot by controlling the output torque of the waist motor to achieve the full-body control of the legged robot.

2. The control method of a torque-controlled foot-type robot with a rotatable waist according to claim 1, characterized in that In the imitation learning framework, the action space of the legged robot includes the output torques of the waist motor and the leg motors, and the state space of the legged robot includes the base quaternion, forward speed, yaw angular velocity, joint positions, joint velocities, and the joint positions at the previous moment.

3. A control method for a torque-controlled foot-type robot with a rotatable waist according to claim 2, characterized in that, In the imitation learning framework based on the asymmetric actor-critic structure, the input of the actor network is data matching the state space, and the input of the critic network includes data matching the state space and the base linear velocity.

4. A control method for a torque-controlled foot-type robot with a rotatable waist, as claimed in claim 1, wherein In the imitation learning framework, the discriminator takes the states corresponding to a trajectory segment in a historical period as input and calculates the least squares loss. Among them, the states input to the discriminator include the base linear velocity, base angular velocity, base normal vector, base height, joint positions, and joint velocities.

5. A control method for a torque-controlled foot-type robot with a rotatable waist according to claim 1, characterized in that During the process of iteratively training the policy, it also includes a yaw angular velocity tracking reward in the z-axis direction for encouraging the robot to use the waist to turn and track in time, a foot-lifting reward for avoiding frequent contact between the feet and the ground, and an action change constraint reward for restricting the movement frequency of the robot.

6. The control method of a torque-controlled foot-type robot with a rotatable waist according to claim 1, wherein, After obtaining the gait dataset after lumbar adaptative training, it also includes: Perform domain randomization processing on the robot joint friction, base mass, initial joint angles, and centroid position, etc.

7. A control method for a torque-controlled foot-type robot with a rotatable waist, as claimed in claim 1, characterized in that The process of training the motor parameters and environmental parameters of the legged robot through curriculum learning includes the following steps: Judge whether the walking distance of the robot reaches the preset requirement under the current motor parameters and environmental parameters; If the walking distance of the robot reaches the preset requirement, increase the curriculum difficulty and move under a more difficult curriculum at the next reset; If the moving distance of the robot is less than half of the target distance at the end of a preset period of time, decrease the curriculum difficulty; If the highest-level curriculum is completed, loop to a randomly selected curriculum level.

8. A torque-controlled legged robot with a rotatable waist, characterized in that, For implementing the control method as described in any one of claims 1-7, the legged robot includes: A robot main body, the base is a front-back symmetric structure that can rotate along the axis. Among them, both symmetric sides include drive motors, the rotation directions of the two drive motors are opposite, and each drive motor is respectively connected to a bipolar belt transmission. The connecting shaft inside the base is hollow and used to accommodate power supply and communication cables.

9. A torque-controlled legged robot with a rotatable waist according to claim 8, characterized in that, The base includes a front half base and a rear half base. The front half base is connected to two front legs, and the rear half base is connected to two rear legs.

10. A torque-controlled legged robot with a rotatable waist according to claim 8, characterized in that, It also includes: A main board; Multiple FOC control boards, connected to the main board and respectively connected to the motors of the waist, front legs, and hind legs, to achieve field-oriented control of the motors; An IMU unit, connected to the main board.

Citation Information

Patent Citations

  • Control method of quadruped robot with waist degree of freedom

    CN109871018A

  • Humanoid robot gait planning deep reinforcement learning new method

    CN111546349A

  • Robot walking control method and system based on deep reinforcement learning and medium

    CN111580385A

Cited By

  • Quadruped robot low-noise gait control method and system based on soft landing reward function

    CN121209388A

  • Leg mechanism optimization method, device and equipment of leg-foot robot, medium and program product

    CN121257129A