High degree of freedom dual-leg wheel robot device, control device and control method
Patent Information
- Application Number
- CN202411772812.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-04
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2044-12-04
AI Technical Summary
[0016] The beneficial technical effects of this application are as follows: The bipedal wheeled robot constructed using the high-degree-of-freedom bipedal wheeled robot device provided in this application has three degrees of freedom in each hip, similar to humans, thus possessing high flexibility. Simultaneously, by utilizing optimal control theory and reinforcement learning control, a control mode and method are proposed, enabling the robot to perform multimodal motion. Compared to other robots (including legged and wheeled robots) and existing bipedal wheeled robots (such as Ascento from ETH Zurich), this robot has more hip degrees of freedom, a more complex structure, and enhanced motion stability and flexibility.
Smart Images

Figure CN119898418B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent robots, and more particularly to a high-degree-of-freedom bipedal wheeled robot device, control device, and control method. Background Technology
[0002] With the rapid development of artificial intelligence (AI) technology, it is infusing industrial production and daily life with increasingly vibrant machine intelligence. In the process of automation and intelligentization in human society, there is an urgent need for general-purpose robots with high mobility and operational performance, serving as algorithm carriers to assist humans in completing tasks. The combination of general artificial intelligence and general-purpose robots will greatly expand the application areas of robots, improving their efficiency and value in industrial production, healthcare, and daily services. The robot world envisioned by Asimov will gradually become a reality.
[0003] Human production and living scenarios are highly complex and unstructured. Bipedal robots, which resemble human forms, have the main design goal of "imitating human bipedal movement, enabling robots to imitate, approach, and even surpass humans in terms of speed, stability, flexibility, and energy efficiency." They represent a possible solution for future general-purpose robots.
[0004] In recent years, wheeled-legged robots have received widespread research attention. As a novel mobile platform, wheeled-legged robots combine the efficiency of wheels with the flexibility of legs, integrating the advantages of both. They can move quickly and flexibly on flat terrain and effectively overcome obstacles. Their superior mobility has the potential to solve important problems such as movement, transportation, and navigation in complex real-world environments (e.g., home scenarios, assistive devices for the disabled), thus making them a promising research subject with broad application prospects. However, existing wheeled-legged robots typically have only one or two degrees of freedom at the hip joint, a structure that limits tasks such as lateral movement and obstacle avoidance. For example, ETH Zurich's Ascento robot has only 4 degrees of freedom, and its legs (2DOFs) can only perform curved movements in the sagittal plane, not movements around the sagittal axis, making it difficult to move on lateral slopes. For special movement commands in the horizontal plane (such as moving to the left and forward), Ascento must achieve this through differential turning, which requires additional body rotation and a change in body orientation, thus failing to meet certain special obstacle avoidance and photography requirements. The Ollie robot, on the other hand, uses a five-bar linkage, with each leg having 3 degrees of freedom, enabling free movement in the sagittal plane within the feasible region, making its leg movements more flexible. However, because the movement plane of its legs is orthogonal to the coronal and transverse planes, it still has limitations in more complex movement tasks. Summary of the Invention
[0005] The purpose of this application is to provide a high-degree-of-freedom bipedal wheeled robot device, control device, and control method. Through innovation in mechanism, control, and algorithm, a high-performance control algorithm for high-degree-of-freedom wheeled robots is developed to achieve human-like movements such as gliding and jumping, which can be applied to scenarios such as logistics transportation, special operations, and life services.
[0006] To achieve the above objectives, this application provides a high-degree-of-freedom bipedal wheeled robot device, comprising a robot base and at least two leg components; each leg component includes a hip joint mechanism and mechanical wheels, one side of the robot base is connected to the mechanical wheels via the hip joint mechanism, for driving the robot body mounted on the other side of the robot base to move via the mechanical wheels; wherein, the hip joint mechanism comprises a ball-joint-like structure formed by multiple rotating components, the rotating components being used to drive the mechanical wheels to rotate relative to the robot base along yaw angle, roll angle, and pitch angle according to received control commands.
[0007] In the aforementioned high-degree-of-freedom bipedal wheeled robot device, optionally, the hip joint mechanism includes a yaw link, a rolling link, a thigh link, a first drive motor, a second drive motor, and a third drive motor; the first drive motor is fixedly disposed on the robot base, and its rotating side is connected to one side of the yaw link, for driving the yaw link to rotate along the yaw angle of the robot base according to the received control command; the second drive motor is fixedly disposed on the other side of the yaw link, and its rotating side is connected to one side of the rolling link, for driving the rolling link to rotate along the rolling angle of the robot base according to the received control command; the third drive motor is fixedly disposed on the other side of the rolling link, and its rotating side is connected to one side of the thigh link, for driving the thigh link to rotate along the pitch angle of the robot base according to the received control command.
[0008] In the aforementioned high-degree-of-freedom bipedal wheeled robot device, optionally, the mechanical wheel includes a lower leg link, a foot wheel, a fourth drive motor, and a fifth drive motor; the fourth drive motor is fixedly located on the other side of the upper leg link and its rotating side is connected to one side of the lower leg link, and is used to drive the lower leg link to rotate along the pitch angle of the robot base through a parallel mechanism according to the received control command; the fifth drive motor is fixedly located on the other side of the lower leg link and its rotating side is connected to the fixed axis of the foot wheel, and is used to drive the foot wheel to rotate clockwise or counterclockwise according to the received control command.
[0009] In the above-mentioned high-degree-of-freedom bipedal wheeled robot device, optionally, the fourth drive motor and the thigh link form a parallelogram linkage mechanism; the first drive motor, the second drive motor, the third drive motor, the fourth drive motor and the fifth drive motor all include dual encoders, the dual encoders are used to receive externally provided control commands, and to provide feedback on the position, speed and torque of the corresponding drive motors.
[0010] This application also provides a control device for the aforementioned high-degree-of-freedom bipedal wheeled robot device. The control device includes an inertial measurement module and a microcontroller module. The inertial measurement module is used to collect the operating status of the high-degree-of-freedom bipedal wheeled robot device. The microcontroller module is used to provide a communication channel between external devices and motors through a hybrid communication method. The microcontroller module includes a first microcontroller and a second microcontroller. The first microcontroller is connected to a first drive motor, a second drive motor, and a fourth drive motor via a controller area network (CLAN) bus. The second microcontroller is connected to a third drive motor and a fifth drive motor via a CLAN bus. The first microcontroller, the second microcontroller, and the external devices communicate via Ethernet for control automation technology, and the first microcontroller and the second microcontroller are connected in parallel for data transmission.
[0011] This application also provides a control method for the aforementioned high-degree-of-freedom bipedal wheeled robot device. The method includes: acquiring control requirements; parsing the control requirements to obtain an action flow; decomposing the action flow to obtain one or more control actions; calculating control parameters based on the action type of the control actions using an optimal control algorithm and reinforcement learning; and controlling the high-degree-of-freedom bipedal wheeled robot device to complete the corresponding action based on the control parameters.
[0012] In the above control method, optionally, the action type includes balancing gliding and walking actions; when the control action is a balancing gliding action, control parameters are calculated using an optimal control algorithm; when the control action is a walking action, control parameters are calculated using a reinforcement learning method; wherein, calculating control parameters using an optimal control algorithm includes: obtaining target data based on the control action analysis, calculating a weighting matrix based on the target data using a preset linear quadratic regulator, and calculating control parameters based on the weighting matrix using a preset control law; calculating control parameters using a reinforcement learning method includes: obtaining target data based on the control action analysis, and calculating corresponding control parameters based on the target data using an objective function constructed by a proximal policy optimization algorithm.
[0013] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method.
[0014] This application also provides a computer-readable storage medium storing a computer program that performs the above-described methods.
[0015] This application also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.
[0016] The beneficial technical effects of this application are as follows: The bipedal wheeled robot constructed using the high-degree-of-freedom bipedal wheeled robot device provided in this application has three degrees of freedom in each hip, similar to humans, thus possessing high flexibility. Simultaneously, by utilizing optimal control theory and reinforcement learning control, a control mode and method are proposed, enabling the robot to perform multimodal motion. Compared to other robots (including legged and wheeled robots) and existing bipedal wheeled robots (such as Ascento from ETH Zurich), this robot has more hip degrees of freedom, a more complex structure, and enhanced motion stability and flexibility. Attached Figure Description
[0017] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, do not constitute a limitation thereof. In the drawings:
[0018] Figure 1A This is a schematic diagram of the structure of a high-degree-of-freedom bipedal wheeled robot device provided in an embodiment of this application;
[0019] Figures 1B to 1G This is a schematic diagram of the structure of each component in a high-degree-of-freedom bipedal wheeled robot device provided in an embodiment of this application;
[0020] Figure 2 This is a schematic diagram of the degrees of freedom of a high-degree-of-freedom wheeled robot device provided in an embodiment of this application;
[0021] Figure 3 This is a schematic diagram of a control method provided in an embodiment of this application;
[0022] Figure 4 This is a schematic diagram of the robot software architecture provided in an embodiment of this application;
[0023] Figure 5 This is a schematic diagram of a dynamic model provided in an embodiment of this application;
[0024] Figure 6This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0025] Attached icon number
[0026] 100: Robot Base
[0027] 101: Yaw Series Robot Hip Joint Motors
[0028] 102: Robot Link Yaw Series
[0029] 103: Robot Hip Joint Motor Roll Series
[0030] 104: Pitch Series Robot Hip Joint Motors
[0031] 105: Robot Linkage Roll Series
[0032] 106: Robot Knee Motor
[0033] 107: Robot's Thigh Link
[0034] 108: Robot's lower leg linkage
[0035] 109: Robot Wheel Hub
[0036] 110: Robot hub motor Detailed Implementation
[0037] The following will describe in detail the implementation methods of this application with reference to the accompanying drawings and embodiments, so as to fully understand how this application uses technical means to solve technical problems and achieve technical effects, and to implement it accordingly. It should be noted that, as long as there is no conflict, the various embodiments and features in each embodiment of this application can be combined with each other, and the resulting technical solutions are all within the protection scope of this application.
[0038] Furthermore, the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0039] This application provides a high-degree-of-freedom bipedal wheeled robot device, comprising a robot base and at least two leg components; each leg component includes a hip joint mechanism and mechanical wheels, one side of the robot base is connected to the mechanical wheels via the hip joint mechanism, for driving the robot body mounted on the other side of the robot base to move via the mechanical wheels; wherein, the hip joint mechanism comprises a ball-joint-like structure formed by multiple rotating components, the rotating components being used to drive the mechanical wheels to rotate relative to the robot base along yaw angle, roll angle and pitch angle according to received control commands.
[0040] Specifically, from a biological perspective, the human hip joint has three degrees of freedom to ensure stable walking. A significant difference between wheeled robots and humans is that wheeled robots use wheels to contact the ground. The robot's structural system should allow the wheels to have a wide range of postures relative to the ground, thus providing more diverse foot placement options. Therefore, in the high-degree-of-freedom bipedal wheeled robot device provided in this application, the leg components simulate the function of two feet. Technically, the overall improvement is achieved by increasing the degrees of freedom of the hip joint to enhance the stability and flexibility of the wheeled robot, while maintaining its wheeled and legged movement capabilities. To facilitate a clearer understanding of the implementation principle of the degrees of freedom provided by the hip joint mechanism in this application, the specific structures of the hip joint mechanism and the mechanical wheels will be described in detail in subsequent embodiments, and will not be elaborated upon here.
[0041] In one embodiment of this application, the hip joint mechanism includes a yaw link, a rolling link, a thigh link, a first drive motor, a second drive motor, and a third drive motor. The first drive motor is fixedly disposed on the robot base and its rotating side is connected to one side of the yaw link, and is used to drive the yaw link to rotate along the yaw angle of the robot base according to a received control command. The second drive motor is fixedly disposed on the other side of the yaw link and its rotating side is connected to one side of the rolling link, and is used to drive the rolling link to rotate along the rolling angle of the robot base according to a received control command. The third drive motor is fixedly disposed on the other side of the rolling link and its rotating side is connected to one side of the thigh link, and is used to drive the thigh link to rotate along the pitch angle of the robot base according to a received control command.
[0042] In one embodiment of this application, the mechanical wheel includes a lower leg link, a foot wheel, a fourth drive motor, and a fifth drive motor. The fourth drive motor is fixedly located on the other side of the upper leg link and its rotating side is connected to one side of the lower leg link. It is used to drive the lower leg link to rotate along the pitch angle of the robot base through a parallel mechanism with the third drive motor according to the received control command. The fifth drive motor is fixedly located on the other side of the lower leg link and its rotating side is connected to the fixed axis of the foot wheel. It is used to drive the foot wheel to rotate clockwise or counterclockwise according to the received control command.
[0043] For details, please refer to Figure 1A As shown, the high-degree-of-freedom bipedal wheeled robot device generally comprises: a robot base 100, robot hip joint motors (Yaw series) 101, robot linkages (Yaw series) 102, robot hip joint motors (Roll series) 103, robot hip joint motors (Pitch series) 104, robot linkages (Roll series) 105, robot knee motors 106, robot thigh linkages 107, robot lower leg linkages 108, robot wheel hubs 109, and robot hub motors 110. Considering the contact between the wheel and the ground, due to its axisymmetry along the pitch (along the direction of wheel rotation), its posture can be described by two coordinates (i.e., roll angle (Roll) and yaw angle (Yaw)). In this application, the design of the hip pitch and knee pitch degrees of freedom allows the robot to maintain contact with the ground and improves its adaptability to uneven terrain. The wheel's posture is achieved by increasing the degrees of freedom of hip-roll and hip-yaw, while hip pitch and knee pitch ensure walking motion.
[0044] Please refer to Figures 1B to 1GAs shown, referring to the lower body of a human, the robot device provided in this application consists of a robot base, yaw links for the left and right legs, roll links, pitch links for the thighs, lower leg links, and foot wheels. This structure gives the robot the flexibility of human legs while using wheels to replace human feet, making the robot move more efficiently on the ground. The base connects the two legs, carries the power supply and control board, and has two motors mounted underneath, connected to the yaw links. The yaw links connect the base and the thigh links, can rotate around the yaw axis, and are fixed to two motors located on the rear of the body, connected to the roll links. The roll links connect the base and the thigh links, can rotate around the roll axis, and are fixed to two motors located on the inner side of the body, connected to the thigh links. The thigh links act as the robot's thighs, with motors on the outer side that drive the lower leg links to rotate around the knee joint via a parallel mechanism. A motor is fixed to the end of the lower leg links, driving the foot wheels to rotate.
[0045] Please refer to Figure 2 As shown in the above embodiment, the robot has joint motors installed at each of its 10 joints, thus giving each leg 5 degrees of freedom (3 degrees of freedom for the hip joint, 1 degree of freedom for the knee joint, and 1 degree of freedom for the foot), and the entire robot 10 degrees of freedom. Taking the left leg as an example, each motor corresponds to one degree of freedom: the Hip-Yaw Motor corresponds to the yaw angle of the hip joint, the Hip-Roll Motor corresponds to the roll angle of the hip joint, the Hip-Pitch Motor corresponds to the pitch angle of the hip joint, the Knee Motor corresponds to the pitch angle of the knee joint, and the Wheel Motor corresponds to the rotation angle of the foot. The right leg is similar. The Hip-Yaw Motor, Hip-Roll Motor, and Hip-Pitch Motor together form a ball-joint-like structure, allowing each leg to rotate around three axes, simulating the human hip joint. The Knee Motor connects to the lower leg via a linkage, forming a parallelogram linkage mechanism with the thigh, driving the lower leg to rotate, simulating the human knee joint. Wheel Motor can control the robot to move at high speed on the ground, and can control the rotation speed of the left and right wheels through differential speed control to achieve the robot's rotation, replacing human feet with wheels. Therefore, the robot can perform simple movements such as forward movement and rotation at high speed, as well as complex movements such as swaying, squatting, and jumping.
[0046] In the above embodiment, the fourth drive motor and the thigh linkage form a parallelogram linkage mechanism; the first drive motor, the second drive motor, the third drive motor, the fourth drive motor and the fifth drive motor all include dual encoders, which are used to receive externally provided control commands and to provide feedback on the position, speed and torque of the corresponding drive motors.
[0047] Specifically, in the high-degree-of-freedom bipedal wheeled robot device provided in this application, the eight motors are concentrated near the base. While this increases the center of gravity height and control difficulty, it reduces the weight of the robot's legs, making leg movements more flexible. Furthermore, a hybrid series-parallel mechanical connection structure is designed for the legs. The parallelogram mechanism reduces knee joint weight, lowers leg inertia, and enables precise leg control. It also effectively transfers torque from the hip to the knee joint. Therefore, the overall leg structure is kinematically equivalent to a serial manipulator, a configuration beneficial for kinematic and dynamic solutions. Finally, a single motor is directly connected to drive the wheel's rotation. From a practical perspective, hip yaw and hip rolling movements play a crucial role in establishing the wheel's ground contact posture, while hip pitching movements facilitate thigh lifting. Knee movements ensure continuous contact between the wheel and the ground, satisfying the required constraints.
[0048] This application also provides a control device for the aforementioned high-degree-of-freedom bipedal wheeled robot device. The control device includes an inertial measurement module and a microcontroller module. The inertial measurement module is used to collect the operating status of the high-degree-of-freedom bipedal wheeled robot device. The microcontroller module is used to provide a communication channel between external devices and motors through a hybrid communication method. The microcontroller module includes a first microcontroller and a second microcontroller. The first microcontroller is connected to a first drive motor, a second drive motor, and a fourth drive motor via a controller area network (CLAN) bus. The second microcontroller is connected to a third drive motor and a fifth drive motor via a CLAN bus. The first microcontroller, the second microcontroller, and the external devices communicate via Ethernet for control automation technology, and the first microcontroller and the second microcontroller are connected in parallel for data transmission.
[0049] For details, please refer to Figure 2 As shown, in response to the unique mechanical structure of the high-degree-of-freedom bipedal wheeled robot device provided in this application, this application provides a corresponding control device for the high-degree-of-freedom bipedal wheeled robot device. In actual operation, due to the concentration of motors and aluminum structures in the base and hip, the robot's center of gravity (COG) is relatively high. The team effectively controls 8 joints (all joints except the Hip-Pitch Motor) by using 8 motors, providing preset peak torque.
[0050] The Hip-Pitch Motor is chosen to withstand the enormous load torque at the knee, thus its peak torque is higher than that of the other eight joints. Each motor is equipped with dual encoders for precise feedback of position, velocity, and torque, adaptable to various closed-loop control methods. To capture the overall state of the wheeled robot, a MicroStrain 3DM-GX5-AHRS is used as an inertial measurement unit (IMU). It features a high-precision accelerometer and a stable gyroscope with a sampling rate of 1000Hz. For communication, two microcontrollers act as a bridge between the motors and the computer. To improve real-time performance, a hybrid communication method is used between the computer (host computer), microcontrollers (slave stations), and motors. Communication between the host computer and slave stations, and between slave stations themselves, is via Ethercat (Ethernet for Control Automation Technology). The four slave stations are connected in parallel, allowing them to exchange data. The host computer only needs to connect directly to one slave station to control all slave stations. The slave stations communicate with the motors via CAN (Controller Area Network), with each slave station connecting to and controlling two or three motors. The host computer communicates directly with the inertial navigation module (IMU+Gyro) and the game controller via a serial port.
[0051] In terms of power supply, a lithium-ion polymer (LiPo) battery pack can be used to provide stable power to the entire system. After passing through the power module, the output voltage powers 10 parallel motors and converts the output voltage to supply 4 parallel slave stations. The host computer uses the user computer to power the inertial navigation module (IMU+Gyro) and the motion control module. The controller is powered by a built-in power supply, which is equipped with an air switch to ensure electrical safety.
[0052] Please refer to Figure 3 As shown, this application also provides a control method applied to the aforementioned high-degree-of-freedom bipedal wheeled robot device, the method comprising:
[0053] S301 Obtains control requirements and parses the action flow based on the control requirements;
[0054] S302 Decomposes the action flow to obtain one or more control actions, and calculates control parameters according to the action type of the control actions using the optimal control algorithm and reinforcement learning method;
[0055] S303 controls the high-degree-of-freedom bipedal wheeled robot device to complete the corresponding actions according to the control parameters.
[0056] In practical applications, given the unique mechanical structure of high-degree-of-freedom bipedal wheeled robot devices, this application establishes a distributed simulation control platform based on IsaacGym+ROS. This distributed simulation control platform mainly consists of three parts: a simulation simulator (Sim) for importing robot models and configuring parameters; a controller for implementing the robot's motion control strategy; and a communication platform (ROS) that acts as a communication channel to bridge the controller and the simulation platform.
[0057] The aforementioned simulation platform can be used to conduct experiments to verify that the wheeled robot has better performance.
[0058] The components include: 1) Simulator: simulates the real environment and robot movement; 2) Controller: executes the robot motion control algorithm; 3) Communicator: connects the controller and simulator via the Robot Operating System (ROS).
[0059] This framework decouples the simulation environment, control algorithms, and communication protocols. This facilitates the porting of these algorithms to practical applications. Utilizing IsaacGym as the simulation environment allows for massive parallelism when training the robot on GPUs, followed by PhysX as the physics engine and deployment of LQR and PPO algorithms. Based on this framework, multiple environments can be created via APIs to enable reinforcement learning on GPUs. The simulation platform can switch between traditional control algorithms and perform reinforcement learning training. Furthermore, thanks to its ROS communication mechanism, the platform boasts a robust communication architecture and debugging APIs, meeting the demands of high real-time control simulations and allowing for rapid addition of new controllers, demonstrating strong scalability. The program portion of the simulation environment deployed by the team consists of three parallel components: RosMaster—a ROS-based communication layer responsible for multi-threaded information interaction and management; Simulation—a simulation module developed based on IsaacGym for Whleaper (the robot's name), responsible for running and modifying the physical simulation environment; and Control Algorithm—a relatively decoupled distributed control algorithm layer carrying the main control strategies.
[0060] Specifically, the control method provided in this application can be referenced in terms of software architecture. Figure 4As shown, its control architecture can be divided into three layers: the command layer, the control layer, and the hardware layer. The command layer is responsible for user interaction and command transmission. Users can input the desired posture and speed into the control layer by manipulating the joystick, while simultaneously observing real-time feedback on the overall posture and joint states through the user interface (UI). As a key component in feedback control, the control layer is responsible for processing information and issuing commands. It calculates the estimated state vector by integrating data from the IMU and motors. Subsequently, a designated controller (e.g., LQR) calculates the motor torque command. These commands are then transmitted to the hardware layer, smoothly adjusting the robot to the desired state. The hardware layer is responsible for low-level data acquisition and the control of hardware components (including the IMU and motors). A PID controller is used in the low-level control of the motors. The hardware layer also forwards overall posture information from the IMU and joint states from the motor encoders to the control layer.
[0061] In one embodiment of this application, the action type includes balancing gliding and walking actions; when the control action is a balancing gliding action, control parameters are calculated using an optimal control algorithm; when the control action is a walking action, control parameters are calculated using a reinforcement learning method; wherein, calculating control parameters using the optimal control algorithm includes: obtaining target data based on the control action parsing, calculating a weighted matrix based on the target data using a preset linear quadratic regulator, and calculating control parameters based on the weighted matrix using a preset control law; calculating control parameters using the reinforcement learning method includes: obtaining target data based on the control action parsing, and calculating corresponding control parameters based on the target data using an objective function constructed by a proximal policy optimization algorithm.
[0062] In practical applications, to achieve multiple motion modes, namely gliding and walking, this invention designs a dedicated control algorithm for each motion mode. Balanced gliding is achieved using the Lowest Quantity Optimization (LQR) algorithm, while walking, jumping, and other actions are implemented using a reinforcement learning method based on proximal policy optimization (PPO).
[0063] Specifically, sliding: the balance control given to LQR includes:
[0064] 1) Coordinate system and symbols: Coordinate system as follows Figure 5 As shown in the diagram.
[0065] The team represents the LQR state vector as: The team denoted the rotation of the hip along the yaw, pitch, and roll axes as γ, β, and α, respectively. Knee rotation was represented by φ, wheel rotation by θ, and the subscripts r and l by the right and left sides, respectively. Considering the limited range of motion of the mechanism, the base, hip, and connected thigh were assumed to have instantaneously constant mass. Furthermore, considering that the angles γ and α vary within a small range, the leg was assumed to be effectively approximated as a moving planar rigid body. The system was abstracted as a first-order inverted pendulum system (IPS) and a system was established as follows: Figure 5 The coordinate system is shown, and clear labels are assigned to the symbols. Based on the robot's posture, the equivalent mass, inertia, and length matrices can be calculated:
[0066] m,I,l=m,I,l pendulum (γ l ,γ r ,β l ,β r ,α l ,α r ,φ l ,φ r (1)
[0067] 2) Derivation of the dynamic equations: For a given model, the Lagrange equations of motion are given by the following equations (2) and (3), where q i Represents each generalized coordinate.
[0068] L = TV (2)
[0069]
[0070]
[0071]
[0072] 3) Stability Control: To achieve stable sliding motion, a robust stabilization algorithm is required. The goal of this algorithm is to enable the robot to handle external disturbances to maintain balance while using as little time and energy as possible. However, in most cases, saving time and energy is contradictory. To address this issue, the wheeled robot employs an LQR (Linear Quadratic Regulator), which is an optimal control method for controlling linear systems at minimal cost. It has been previously demonstrated that using an LQR controller can effectively ensure strong reliability and robustness in stabilization scenarios involving two-wheeled inverted pendulums. The state equation of the LTI (Linear Time Invariant) discrete system is as follows: Equation (6)
[0073] x(k+1)=Ax(k)+Bu(k). (6)
[0074] In the formula, x is the state vector and u is the input vector. The quadratic performance index J is expressed as the discrete-time state-space representation of the system defined by the control matrices A and B, which originates from the robot model around the state vector x =
[0075] Linearization of [0,0,0,0]T. In the given equations, Q and R denote the weighted matrices of the state covariance, while K denotes an unspecified matrix. Equation (7)
[0076]
[0077] Q and R are determined based on the required stability for each state. Experiments have shown that a diagonal weight matrix is sufficient for good control, so the team simplified the model, reducing the number of adjustable weight parameters to the order of the weight matrix. Experiments were conducted under various conditions to obtain suitable weight parameters.
[0078] The optimal feedback gain matrix L can be determined by solving the discrete-time algebraic Riccati equations (DARE). Equations (8)(9)
[0079] -K+Q+A T KA-A T KB(R+B T AK) -1 B T KA=0 (8)
[0080] L = (R + B) T KB) -1 B T KA (9)
[0081] The robot makes dynamic adjustments using the following control law. Equation (10)
[0082] u=-Kx (10)
[0083] The state vector x can be derived from the IMU and kinematics. For a linear system, controllability and observability are separable. Therefore, the system can maintain stable operation, resist external disturbances, and reach the target position. Furthermore, for the main posture control of the legs, position control is adopted, corresponding to the modification of the equivalent mass parameters in the inverted pendulum model.
[0084] Walking: Reinforcement Learning Control:
[0085] To achieve robust walking and jumping control for a 10-DOF robot, the team implemented the PPO algorithm in simulations. PPO (Proximal Policy Optimization) is a reinforcement learning algorithm that ensures stability by limiting the magnitude of policy updates and uses a pruning mechanism to control the degree of updates. The objective function is given by the following equation: Equation (11).
[0086]
[0087] The first term of the minimal function is represented by L. CLIP The second item Used to limit the probability ratio r t The range of (θ) is restricted to between 1-ε and 1+ε. Regarding rewards, the team uses a fine-tuned reward function to adjust the robot's movements, enabling it to walk and jump. The team sets the termination and joint position constraints to negative values with large absolute values. The aim is to accelerate the robot's convergence to a steady state, maintain balance, and adhere to joint constraints. Taking jumping as an example, the team rewards the time its feet leave the ground, its velocity along the positive Z-axis, and its base height. Simultaneously, the team penalizes rotation and angular velocity. These lower reward functions ensure that the robot's jumps occur in place.
[0088] It is worth noting that, in addition to the flexibility of walking and gliding, in the simulation environment, the team can freely switch the robot to either LQR-controlled gliding mode or RL-controlled walking mode by changing the strategy. This switching of movement modes allows the robot to adopt different movement modes when dealing with various scenarios, further enhancing the flexibility of the wheeled robot.
[0089] The beneficial technical effects of this application are as follows: The bipedal wheeled robot constructed using the high-degree-of-freedom bipedal wheeled robot device provided in this application has three degrees of freedom in each hip, similar to humans, thus possessing high flexibility. Simultaneously, by utilizing optimal control theory and reinforcement learning control, a control mode and method are proposed, enabling the robot to perform multimodal motion. Compared to other robots (including legged and wheeled robots) and existing bipedal wheeled robots (such as Ascento from ETH Zurich), this robot has more hip degrees of freedom, a more complex structure, and enhanced motion stability and flexibility.
[0090] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method.
[0091] Figure 6 This is a schematic diagram of the physical structure of an electronic device provided in an embodiment of the present invention, such as... Figure 6 As shown, the electronic device includes: a processor 501, a memory 502, and a bus 503.
[0092] The processor 501 and the memory 502 communicate with each other via the bus 503.
[0093] The processor 501 is used to call program instructions in the memory 502 to execute the methods provided in the above-described method embodiments.
[0094] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described control method.
[0095] This invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-described control method.
[0096] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0097] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.
[0098] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.
[0099] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.
[0100] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A control method for a high-degree-of-freedom bilegged wheeled robot device, characterized in that, The method includes: Obtain control requirements, and parse the action flow based on the control requirements; One or more control actions are obtained by decomposing the action flow, and control parameters are calculated by using the optimal control algorithm and reinforcement learning method according to the action type of the control actions. The high-degree-of-freedom bipedal wheeled robot device is controlled to complete the corresponding actions according to the control parameters. The types of movements include balance gliding and walking. When the control action is a balancing sliding action, the control parameters are calculated using the optimal control algorithm. When the controlled action is a walking action, the control parameters are calculated using reinforcement learning. The process of obtaining control parameters through the optimal control algorithm includes: obtaining target data by parsing the control action; obtaining a weighting matrix by calculating the target data using a preset linear quadratic regulator; and obtaining control parameters by calculating the weighting matrix using a preset control law. The control parameters are obtained by reinforcement learning, which includes: parsing the control action to obtain target data, and calculating the corresponding control parameters based on the target data using a target function constructed by a proximal policy optimization algorithm. The high-degree-of-freedom bipedal wheeled robot device includes a robot base and at least two leg components; The leg assembly includes a hip joint mechanism and mechanical wheels. One side of the robot base is connected to the mechanical wheels via the hip joint mechanism, which is used to drive the robot body mounted on the other side of the robot base to move via the mechanical wheels. The hip joint mechanism includes a ball joint-like structure formed by multiple rotating components, which are used to drive the mechanical wheel foot to rotate relative to the robot base along yaw angle, roll angle and pitch angle according to the received control commands.
2. A high-degree-of-freedom bipedal wheeled robot device suitable for the control method described in claim 1, characterized in that, The device includes a robot base and at least two leg components; The leg assembly includes a hip joint mechanism and mechanical wheels. One side of the robot base is connected to the mechanical wheels via the hip joint mechanism, which is used to drive the robot body mounted on the other side of the robot base to move via the mechanical wheels. The hip joint mechanism includes a ball joint-like structure formed by multiple rotating components. The rotating components are used to drive the mechanical wheel foot to rotate relative to the robot base along yaw angle, roll angle and pitch angle according to the received control commands.
3. The high-degree-of-freedom bipedal wheeled robot device according to claim 2, characterized in that, The hip joint mechanism includes a yaw link, a rolling link, a thigh link, a first drive motor, a second drive motor, and a third drive motor; The first drive motor is fixed on the robot base and rotates on one side of the yaw link, and is used to drive the yaw link to rotate along the yaw angle of the robot base according to the received control command; The second drive motor is fixed on the other side of the yaw link and rotates on one side of the rolling link, and is used to drive the rolling link to rotate along the rolling angle of the robot base according to the received control command; The third drive motor is fixed on the other side of the rolling link and rotates on one side of the thigh link, and is used to drive the thigh link to rotate along the pitch angle of the robot base according to the received control command.
4. The high-degree-of-freedom bipedal wheeled robot device according to claim 3, characterized in that, The mechanical wheel includes a lower leg connecting rod, a foot wheel, a fourth drive motor, and a fifth drive motor; The fourth drive motor is fixed on the other side of the thigh link and rotates on one side of the lower leg link. It is used to drive the lower leg link to rotate along the pitch angle of the robot base through a parallel mechanism according to the received control command. The fifth drive motor is fixed on the other side of the lower leg connecting rod, and its rotating side is connected to the fixed axis of the foot wheel. It is used to drive the foot wheel to rotate clockwise or counterclockwise according to the received control command.
5. The high-degree-of-freedom bipedal wheeled robot device according to claim 4, characterized in that, The fourth drive motor and the thigh linkage form a parallelogram linkage mechanism. The first drive motor, the second drive motor, the third drive motor, the fourth drive motor, and the fifth drive motor all include dual encoders. The dual encoders are used to receive externally provided control commands and to provide feedback on the position, speed, and torque of the corresponding drive motors.
6. A control device applied to the high-degree-of-freedom bipedal wheeled robot device as described in claim 4 or 5, characterized in that, The control device includes an inertial measurement module and a microcontroller module; The inertial measurement module is used to collect the operating status of the high-degree-of-freedom bipedal wheeled robot device; The microcontroller module is used to provide a communication channel between external devices and the motor through a hybrid communication method; The microcontroller module includes a first microcontroller and a second microcontroller. The first microcontroller is connected to the first drive motor, the second drive motor and the fourth drive motor via a controller local area network bus, and the second microcontroller is connected to the third drive motor and the fifth drive motor via a controller local area network bus. The first microcontroller, the second microcontroller, and the external device communicate via Ethernet for control automation technology, and the first microcontroller and the second microcontroller are connected in parallel for data transmission.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of claim 1.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of claim 1.
9. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method of claim 1.
Citation Information
Patent Citations
High-flexibility seven-degree-of-freedom wheel-foot robot leg structure
CN114454980A