PID (Proportion Integration Differentiation) control and gait planning method of wheel-foot composite structure and electronic equipment

Through the optimization of the PID control model with coupling compensation terms and the PPO algorithm, combined with IMU and lidar, the control accuracy and mode switching impact problems of the wheel-leg composite structure robot in complex environments were solved, and the robot's motion coordination efficiency and environmental adaptability were improved.

CN120779708AActive Publication Date: 2025-10-14FUDAN UNIVERSITY

Patent Information

Application Number
CN202511202969.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-10-14
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

Under traditional PID control and gait planning methods, wheel-leg composite structure robots face problems such as insufficient control accuracy due to joint dynamics coupling, large and uneven impact of wheel-foot mode switching, synchronization errors and excessive energy consumption in multi-machine collaboration, and poor parameter adaptability in complex terrain.

Method used

A PID control model with coupling compensation terms is adopted, combined with IMU feedback and lidar prediction, to adjust PID parameters in real time. The parameters are optimized through the PPO algorithm to achieve coordinated motion between joints, and linear interpolation and sliding mode control are used to reduce impact when switching modes.

Benefits of technology

It achieves high-precision motion control of robots in complex environments, reduces joint synchronization errors and mode switching impacts, and improves multi-machine collaboration efficiency and environmental adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120779708A_ABST
    Figure CN120779708A_ABST
Patent Text Reader

Abstract

The invention relates to the field of robot control, and particularly discloses a PID control and gait planning method of a wheel-foot composite structure and electronic equipment, and the method comprises the steps: building a PID control model with a coupling compensation item; angle errors of all joints are calculated, motion interference of adjacent joints is compensated according to a PID control model with a coupling compensation item, in a wheel type mode, an outer ring takes a second threshold value as a target to control the speed, an inner ring corrects the posture according to IMU feedback, and the posture control weight is enhanced when the inclination angle exceeds a third threshold value; in the foot mode, PID parameters are adjusted in advance according to laser radar prediction, and the complex terrain adaptability is improved; terrain, gait, control precision and energy consumption data are collected in real time, PID parameters are automatically adjusted based on a PPO algorithm, and the control precision of the robot in the unknown terrain is improved; according to the invention, through a dynamic coupling compensation mechanism, the motion interference between the joints is accurately solved, so that the joints keep high coordination in motion, and the joints can stably keep accurate motion postures no matter what complex terrains face.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of robot control, in particular to a PID control and gait planning method of a wheel-foot composite structure and an electronic device. BACKGROUND

[0002] In the robot control, the joint movement of the wheel-leg composite structure has a dynamic coupling (such as the mutual interference of the hip joint and the knee joint movement), and when the traditional wheel-foot composite structure robot adopts the conventional PID control and gait planning mode, it faces the problems of insufficient control accuracy caused by the joint dynamic coupling, large and non-smooth switching impact of the wheel-foot mode, synchronization error and high energy consumption in multi-robot cooperation, and poor parameter adaptability in complex terrain, which jointly restrict the improvement of the motion performance, cooperation efficiency and environmental adaptability in complex environments. SUMMARY

[0003] The purpose of the present application is to overcome the defects of the prior art, and to provide a PID control and gait planning method of a wheel-foot composite structure, and to provide a PID control and gait planning electronic device of a wheel-foot composite structure to solve the problems raised in the background.

[0004] The PID control and gait planning method of the wheel-foot composite structure of the present application comprises the following steps: Establish a PID control model with a coupling compensation term; Calculate the joint angle error, compensate for the adjacent joint movement interference according to the PID control model with the coupling compensation term, and correct the historical control amount difference for the control delay of the first threshold value; In the wheeled mode, the outer ring controls the speed with the second threshold value as the target, and the inner ring corrects the attitude according to the IMU feedback, and when the inclination exceeds the third threshold value, the attitude control weight is enhanced; In the foot mode, the PID parameters are adjusted in advance according to the laser radar prediction to improve the adaptability to complex terrain; Real-time collection of terrain, gait, control accuracy and energy consumption data, automatic adjustment of PID parameters based on PPO algorithm to improve the control accuracy of the robot in unknown terrain; Each robot module has independent PID control, and for the fourth threshold value communication delay, the parameters of the robot module are adjusted according to the control amount of the neighbor node, the PID parameters are linearly interpolated when the wheel-foot mode is switched, and the control amount is adjusted according to the joint angular velocity and angular acceleration direction during the foot landing stage, so as to reduce the switching impact.

[0005] Further, the control equation of the PID control model with the coupling compensation term is: ; The basic PID control is calculated, and the preliminary control amount is generated based on the current error, the cumulative error and the error change rate respectively; To couple the compensation term, the preliminary control quantity is modified, u i (t) is the control quantity of the i-th joint, e i (t) is the joint angle error, K p,i (t) is the time-varying proportional gain, K i,i is the integral gain, K d,i is the differential gain; C i,j is the joint coupling coefficient matrix, used to describe the degree of mutual influence of the motion of different joints, u j (t) is the control quantity of the j-th joint at time t, u j (t-T d ) is the control quantity of the j-th joint at time t-T d .

[0006] Further, the joint angle error is calculated, and the motion interference of adjacent joints is compensated according to the PID control model with coupling compensation term, including: the actual angle value of each joint is obtained in real time through the encoder, and the target angle value in the preset trajectory is subtracted to obtain the angle error; based on the PID control model with coupling compensation term, the control quantity of each joint is calculated, the motion interference degree between joints is quantified through the predefined coupling coefficient matrix, and the control delay caused by mechanical transmission is compensated by using the difference between the control quantity at the current time and the historical time.

[0007] Further, the control delay of the first threshold value is modified by the historical control quantity difference, including: the control quantity of the adjacent joint at the current time and the historical time is obtained, and the difference between the two is calculated as the delay compensation signal; the difference is multiplied by the predefined joint coupling coefficient matrix to generate the coupling compensation term, which is superimposed into the PID control quantity of the current joint.

[0008] Further, the outer ring targets the second threshold value for speed control, including: the wheel set speed is collected in real time by the encoder and compared with the target value to generate a speed error, and a preliminary control quantity is calculated through a PID algorithm; the inner ring is modified by the IMU feedback for the attitude, including: the attitude inner ring takes the real-time attitude data provided by the IMU as the feedback, obtains the three-axis angular velocity and acceleration through high-frequency sampling, and calculates the pitch angle and roll angle after fusing the data using the complementary filtering algorithm, and generates an attitude error by comparing with the preset attitude reference; when the inclination angle exceeds the third threshold value, the attitude control weight is increased to 0.6 from the default value, and the correction force of the vehicle body inclination is enhanced.

[0009] Further, in the foot mode, the PID parameters are adjusted in advance according to the prediction of the laser radar, including: acquiring terrain elevation data in real time through continuous scanning of the laser radar and constructing a three-dimensional terrain model; analyzing the acquired terrain data to identify complex terrain features such as gullies, slopes and steps; adjusting the PID control parameters in advance based on a set prediction time window; dynamically optimizing the proportional gain, integral gain and differential gain of each joint according to the identified terrain type and trend, and introducing a second derivative compensation term to correct the control quantity according to the predicted terrain change trend.

[0010] Further, real-time terrain, gait, control accuracy and energy consumption data are collected, and the PID parameters are automatically adjusted based on the PPO algorithm, including: acquiring terrain point cloud data through real-time scanning of the laser radar to calculate the terrain roughness index, identifying the current gait based on encoder and IMU data, calculating error entropy values based on joint angle error sequences to represent control accuracy, and monitoring energy consumption data to construct a four-dimensional state space containing terrain, gait, control accuracy and energy consumption data indicators; inputting the real-time state into a pre-trained PPO algorithm model, the pre-trained PPO algorithm model is a reinforcement learning model based on a proximal policy optimization framework, the reinforcement learning model includes a policy network and a value network, the policy network is used to output the adjustment amount of the PID parameters, and the value network is used to evaluate the value of the current state to assist policy optimization, the reinforcement learning model has initial decision-making capability through training of various terrain and gait scenarios; a pre-set reward function integrates control error, energy efficiency and smoothness through a weighting method, the control error is calculated by the deviation between the actual joint angle and the target angle, the smaller the deviation, the higher the reward value; the energy efficiency is calculated based on the energy consumption per unit time, the lower the energy consumption, the higher the reward value; the smoothness is calculated by the fluctuation of the joint angle change rate, the smaller the fluctuation, the higher the reward value, the weights of the three are set in the proportion of control error being the highest, energy efficiency being the second, and smoothness being the third; after the real-time state is input into the reinforcement learning model, the policy network outputs the adjustment strategy of the proportional, integral and differential gains according to the current state, and the adjustment strategy of the proportional, integral and differential gains is directly used to update the parameters of the current PID controller.

[0011] Further, each robot module has independent PID control, and for the fourth threshold communication delay, the parameters of each robot module are adjusted according to the control quantity of the neighbor node, including: each robot module independently performs PID calculation using local sensor data, generates an initial control quantity, and for the communication delay between modules, acquires the control quantity data of the neighbor node at the delay time, calculates the difference between the current control quantity of the robot module and the historical control quantity of the neighbor node, weights the difference with a coordination weight factor, generates a communication delay compensation term and adds it to the local control quantity.

[0012] Further, during wheel-foot mode switching, linearly interpolate the PID parameters, and during the foot landing phase, adjust the control amount according to the joint angular velocity and angular acceleration direction, including: after the wheel-foot mode switching is started, define the current time in the switching process as t based on the preset total switching time T, calculate the transition coefficient a(t) through a quintic polynomial function, wherein the value of the transition coefficient a(t) increases smoothly from 0 to 1 with t, which is used to quantify the transition degree of the wheel mode to the foot mode; for the proportional gain, integral gain and differential gain, respectively, linear interpolation formula is used for parameter transition; during the transition process, the pre-trained PPO algorithm model is used to optimize the dynamic characteristics of the transition coefficient a(t), and the state space of the model includes the current transition coefficient a(t), the joint angle error, the parameter change rate and the energy consumption rate; the action space is the adjustment amount of the transition coefficient.

[0013] Further disclosed is an electronic device, comprising: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above-mentioned PID control and gait planning method for a wheel-foot composite structure.

[0014] The present application has the following advantages: The PID control and gait planning method and system for a wheel-foot composite structure of the present application can accurately solve the motion interference between joints through a dynamic coupling compensation mechanism, so that each joint can maintain high coordination during motion, and can stably maintain accurate motion posture regardless of complex terrain.

[0015] During wheel-foot mode switching, the optimized parameter transition method realizes natural and smooth conversion between the two modes, completely eliminates the impact and jerk during the switching process, and makes the motion of the robot more coherent.

[0016] In terms of multi-machine cooperation, through an efficient cooperation strategy and power distribution scheme, the modular robot can maintain consistent motion rhythm after being combined, the carrying capacity is significantly enhanced, and energy consumption is reduced, solving the problem of low efficiency in the traditional cooperation mode.

[0017] The robot can automatically optimize the control parameters according to different terrain conditions and motion states, and can flexibly adapt to various unknown environments without manual operation, greatly improving the environmental adaptability. BRIEF DESCRIPTION OF DRAWINGS

[0018] Figure 1 The workflow diagram of the PID control and gait planning method for a wheel-foot composite structure of the present application. DETAILED DESCRIPTION

[0019] The application discloses a PID control and gait planning method for a wheel-foot composite structure, which comprises the following steps: Figure 1 The PID control and gait planning method for the wheel-foot composite structure comprises the following steps: S100. Establishing a PID control model with a coupling compensation term; S101. Calculating joint angle errors, compensating for adjacent joint motion interference according to the PID control model with the coupling compensation term, correcting a historical control amount difference for control delay of a first threshold value, ensuring synchronous motion of the joints, and reducing joint synchronization errors; S102. In a wheeled mode, an outer ring controls speed with a second threshold value as a target, and an inner ring corrects a posture according to IMU feedback, and when a tilt angle exceeds a third threshold value, the weight of the posture control is increased; S103. In a foot mode, PID parameters are adjusted in advance according to laser radar prediction, and adaptability to complex terrains is improved; S104. Terrain, gait, control accuracy and energy consumption data are collected in real time, PID parameters are automatically adjusted based on a PPO algorithm, and control accuracy of the robot in unknown terrains is improved; S105. Each robot module is independently PID controlled, and according to a fourth threshold value communication delay, the parameters of the module are adjusted according to control amounts of neighbor nodes, power distribution is optimized through a coordination matrix, joint synchronization of the quadruped assembly is realized, load capacity is improved, and energy consumption is reduced; and S106. When the wheel-foot mode is switched, PID parameters are linearly interpolated by using a quintic polynomial function, and in the foot landing stage, the control amount is adjusted according to the direction of the joint angular velocity and the angular acceleration, so that the switching impact is reduced.

[0020] The PID control and gait planning method for the wheel-foot composite structure can be realized by a computer program, and a computer program for implementing the method of the application can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing apparatus, so that the computer programs make the functions / operations specified in the flowcharts and / or block diagrams be implemented when executed by the processor. The computer programs can be executed entirely on a machine, partially on a machine, partially on a machine as a separate software package and partially on a remote machine, or entirely on a remote machine or server.

[0021] For S100. Establishing a PID control model with a coupling compensation term, the control equation of the PID control model with the coupling compensation term is: ; For basic PID control calculation, a preliminary control amount is generated based on the current error, the cumulative error and the error change rate respectively; For the coupling compensation term, the preliminary control amount is corrected by calculating the change and delay influence of the control amount of other joints, and finally the accurate joint control amount u is obtained i(t), enabling coordinated motion between joints, reducing joint synchronization errors, and improving the motion control accuracy of robots in complex environments. The basic PID parameters include: u i (t) : Control variable of the i-th joint (e.g., hip joint, knee joint) to drive the joint motor, determining the joint's angular position and velocity.

[0022] e i (t) : Joint angle error, calculated by the difference between the target angle and the current actual angle. This error value is the basis for the PID controller's adjustment.

[0023] K p,i (t) : Time-varying proportional gain. The proportional gain determines the response speed and intensity of the controller to the current error. Here K p,i (t) is dynamically adjusted in real time according to terrain characteristics, such as maintaining a low value for smooth motion on flat surfaces and automatically increasing to speed up joint angle adjustment when crossing obstacles.

[0024] K i,i : Integral gain. Its role is to accumulate past errors and eliminate the system's static error. For example, when the robot is climbing a slope, if there is a persistent angle deviation, the integral term will continuously accumulate, gradually increasing the control variable until the deviation is eliminated.

[0025] K d,i : Differential gain. Based on the rate of change of error, it can predict the trend of error change, adjust the control variable in advance, and suppress the system's overshoot and oscillation. For example, when the joint is moving quickly and approaching the target angle, the differential action will slow down the joint motion speed to avoid overshooting the target angle.

[0026] Coupling compensation parameters include: C i,j : Joint coupling coefficient matrix. Used to describe the degree of mutual influence between different joints. For example, C 2,1 =0.3 indicates that the motion of the knee joint (2nd joint) has a 0.3 coefficient influence on the hip joint (1st joint). During the robot's walking process, when the knee joint is bent, the compensation value for the hip joint control is calculated through this coefficient to correct the hip joint control and offset the interference caused by the knee joint motion, ensuring the overall motion coordination.

[0027] u j (t) : Control variable of the j-th joint at time t, used to calculate the coupling influence between joints.

[0028] u j (t-T d ) : Control variable of the j-th joint at time t-T d . Td =0.05s is the adjacent joint control delay compensation time. Due to mechanical structure and signal transmission factors, there is a certain delay in joint control. By introducing the control amount at the historical time, the delay can be compensated, and the coupling compensation is more accurate.

[0029] For S101. Calculate the joint angle error, compensate the adjacent joint motion interference according to the PID control model with coupling compensation term, correct the control delay for the first threshold value with the historical control amount difference, ensure the joint synchronous motion, reduce the joint synchronization error, which in implementation is, through the high-precision encoder, the actual angle value of each joint is obtained, which is subtracted from the target angle value in the preset trajectory to obtain the angle error. Based on the PID control model with coupling compensation term, the control amount of each joint is calculated, not only considering the proportional, integral and differential control effect of the current joint, but also introducing the coupling compensation mechanism between adjacent joints. Through the predefined coupling coefficient matrix, the motion interference degree between each joint is quantified, and the control delay caused by mechanical transmission is compensated by the control amount difference between the current time and the historical time. In complex terrain motion, the control parameters are dynamically adjusted according to the terrain characteristics identified by the environment perception system.

[0030] For the control delay of the first threshold value, the historical control amount difference is corrected to ensure the joint synchronous motion and reduce the joint synchronization error. In implementation, for the 0.05 second control delay caused by mechanical transmission and signal transmission, the control amount u j (t) and u j (t-0.05s) of adjacent joints at the current time t and the historical time t-0.05s are obtained, and the difference between the two is calculated as the delay compensation signal. Multiply the difference by the predefined joint coupling coefficient matrix C i,j , generate the coupling compensation term, and superimpose it into the PID control amount of the current joint. By real-time monitoring of the motion asynchronization phenomenon caused by the control delay of each joint, the compensation strength is dynamically adjusted by using the historical control amount difference, so that the control signals of adjacent joints are synchronized and calibrated in time dimension, effectively eliminating the motion interference caused by delay, ensuring the attitude consistency in multi-joint collaborative motion, reducing the joint synchronization error, and improving the motion stability of the robot in complex terrain.

[0031] For S102. In wheeled mode, the outer ring controls the speed with the second threshold value as the target, and the inner ring corrects the attitude according to the IMU feedback. When the inclination exceeds the third threshold value, the attitude control weight is enhanced.

[0032] In implementation, specifically, when the wheeled mode is running, the speed outer loop takes the target speed of 1.2 m / s as the control threshold, the encoder is used to collect the wheel group speed in real time, and the speed error is generated by comparing with the target value, and the preliminary control quantity is calculated through the PID algorithm. The attitude inner loop takes the real-time attitude data provided by the IMU as the feedback, and the three-axis angular velocity and acceleration are obtained through high-frequency sampling, and the pitch angle and roll angle are calculated after the data are fused by the complementary filtering algorithm. The attitude error is generated by comparing with the preset attitude reference. When the IMU detects that the inclination angle exceeds 10°, the system automatically increases the attitude control weight from the default value to 0.6, and enhances the correction strength of the body inclination. The double closed-loop control quantity is output to the wheel group motor after weighted fusion, realizing the cooperative control of speed and attitude, ensuring the stability of the robot in high-speed running on flat road, and avoiding the risk of rollover caused by uneven ground or turning.

[0033] For S103. When the foot mode is running, the PID parameters are adjusted in advance based on the prediction of the laser radar, and the adaptability to complex terrain is improved. In implementation, specifically, the laser radar is used to continuously scan the terrain in front of a certain range, and the terrain elevation data is obtained in real time and a three-dimensional terrain model is constructed. The system analyzes the obtained terrain data, identifies complex terrain features such as gullies, slopes, and steps, adjusts the PID control parameters in advance based on the set prediction time window. According to the identified terrain type and trend, the proportional gain, integral gain and differential gain of each joint are dynamically optimized, and a second derivative compensation term is introduced. The control quantity is corrected according to the predicted terrain change trend, so that the robot can adjust the joint motion trajectory in advance according to the terrain change during foot walking, improve the adaptability to complex terrain, and ensure smooth walking.

[0034] For S104. Real-time collection of terrain, gait, control accuracy, and energy consumption data, automatically adjust the PID parameters based on the PPO algorithm to improve the control accuracy of the robot in unknown terrain. In implementation, the terrain roughness index is calculated by real-time scanning of the laser radar to obtain the terrain point cloud data, the current gait mode is identified using the encoder and IMU data, the error entropy value is calculated based on the joint angle error sequence to represent the control accuracy, and the energy consumption rate is monitored to construct a four-dimensional state space containing terrain, gait, control accuracy, and energy consumption data indicators. The real-time state is input into the pre-trained PPO algorithm model, which is a reinforcement learning model based on the proximal policy optimization framework, containing a policy network and a value network. The policy network is used to output the adjustment amount of the PID parameters, and the value network is used to evaluate the value of the current state to assist policy optimization. The model has initial decision-making ability through training on various terrain and gait scenarios. The pre-set reward function integrates control error, energy efficiency, and smoothness through weighting, where control error is calculated by the deviation between actual and target joint angles, with smaller deviation corresponding to higher reward value; energy efficiency is calculated based on energy consumption per unit time, with lower energy consumption corresponding to higher reward value; smoothness is calculated by the fluctuation of joint angle change rate, with smaller fluctuation corresponding to higher reward value. The weights of the three indicators are set according to the proportion of control error being the highest, energy efficiency being the second, and smoothness being the third. When the real-time state is input into the model, the policy network outputs the adjustment strategy for proportional, integral, and derivative gains based on the current state. These adjustment strategies are directly used to update the parameters of the current PID controller.

[0035] For S105. Independent PID control of each robot module, for the fourth threshold communication delay, adjust its own parameters according to the neighbor node control amount, optimize power distribution through the cooperation matrix, realize joint synchronization of the quadruped assembly, improve load capacity, and reduce energy consumption; In implementation, each robot module independently performs PID calculation using local sensor data, generates an initial control amount, and for a communication delay of 0.1 seconds between modules, obtains the control amount data of the neighbor node at the delay time, calculates the difference between the current control amount and the historical control amount of the neighbor node, weights the difference by a cooperation weight factor of 0.2, generates a communication delay compensation term, and adds it to the local control amount. At the same time, based on the cooperative coupling matrix of the quadruped assembly, the joint control amount is optimized, the energy and driving capacity are shared and distributed through the power bus, the joint motion of the multi-machine assembly tends to be synchronized, the overall load capacity is improved, and the energy consumption is reduced.

[0036] For S106. When the wheel-foot mode is switched, the PID parameters are linearly interpolated by a quintic polynomial function, and during the foot landing phase, the control amount is adjusted according to the joint angular velocity and angular acceleration direction to reduce the switching impact. In practice, after the wheel-foot mode switching is started, the total switching time T (T = 0.6 s, i.e. the total duration of the switching process) is preset as the reference, and the current time t (t ∈ [0, T]) in the switching process is defined. The transition coefficient α(t) is calculated by a quintic polynomial function α(t) = 10(t / T)³−15(t / T) 4 +6(t / T) 5 The transition coefficient α(t) is calculated, where the value of α(t) smoothly increases from 0 to 1 with t, which is used to quantify the transition degree from wheel mode to foot mode.

[0037] For the proportional gain Kp, the integral gain Ki, and the derivative gain Kd, linear interpolation formulas are used for parameter transition respectively: Proportional gain transition: Kp(t) = Kp_wheel⋅(1−α(t))+Kp_leg⋅α(t), where Kp_wheel is the preset value of the proportional gain in wheel mode, and Kp_leg is the preset value of the proportional gain in foot mode; Integral gain transition: Ki(t) = Ki_wheel⋅(1−α(t))+Ki_leg⋅α(t), where Ki_wheel is the preset value of the integral gain in wheel mode, and Ki_leg is the preset value of the integral gain in foot mode; Derivative gain transition: Kd(t) = Kd_wheel⋅(1−α(t))+Kd_leg⋅α(t), where Kd_wheel is the preset value of the derivative gain in wheel mode, and Kd_leg is the preset value of the derivative gain in foot mode.

[0038] During the transition process, a pre-trained PPO algorithm model is used to optimize the dynamic characteristics of α(t). The state space of the model includes the current transition coefficient α(t), the joint angle error, the parameter change rate, and the energy consumption rate; the action space is the adjustment amount of the transition coefficient. The preset reward function is R = w1⋅err_norm + w2⋅(1−energy_rate) + w3⋅smoothness, where err_norm is the normalized joint angle error (reflecting control accuracy), energy_rate is the proportion of energy consumption per unit time (reflecting energy efficiency), and smoothness is the sum of the squares of the PID parameter change rate (reflecting transition smoothness). w1 = 0.5, w2 = 0.3, and w3 = 0.2 are the weight coefficients of each index. Through this reward function, the model optimizes the transition process, so that the proportional, integral, and derivative gains have no sudden changes during the switching process, achieving smooth transition between wheel and foot modes.

[0039] Preferably, during the wheel-foot mode switching process, the dynamic coupling between joints (such as the motion interference between the hip and knee joints) will exacerbate the switching impact and posture fluctuation. This application designs a sliding surface containing joint coupling terms and combines it with a high-frequency switching control law to offset the coupling interference in real time. The sliding surface design containing joint coupling terms is as follows: For the i-th joint (such as the hip joint), the sliding surface is defined as: ;

[0040] Parameter meaning: e i (t): Angle error of the i-th joint (the difference between the target angle and the actual angle, consistent with the error definition in basic PID); : rate of change of angle error (reflecting the dynamic trend of joint movement); k i : Sliding surface convergence coefficient (positive number, ranging from 1.2 to 2.5, determines the speed at which the system state converges to the sliding surface; the larger the value, the faster the convergence); C i,j : Joint coupling coefficient matrix (completely consistent with the coupling compensation parameters in the basic PID control model, such as C2,1=0.3 means that the coupling influence coefficient of the knee joint on the hip joint is 0.3); : The rate of change of the angle error of the jth adjacent joint (such as the knee joint), which is used to quantify its dynamic interference on the i-th joint in the sliding surface The dynamic interference of adjacent joints (such as the impact of rapid swing of the knee joint on the hip joint) is captured in real time through the coupling coefficient matrix, so that the sliding surface can perceive the coupling effect in advance, providing an accurate interference quantification basis for subsequent control law design.

[0041] The control law actively offsets the coupling interference through the synergistic effect of "equivalent control + switching control". The mathematical relationship is: ; As an equivalent control item, the PID parameter interpolation result in the wheel-foot mode switching is reused, that is, the real-time PID control quantity generated by the quintic polynomial transition function: , where K p (t), K i (t), K d (t) is the time-varying PID parameter after quintic polynomial interpolation (consistent with the parameter transition logic in mode switching), ensuring that the sliding mode control is seamlessly integrated with the existing PID framework. is a switching control item used to offset coupling interference at high frequencies. Its mathematical expression is: ; Parameter meaning: K sw : Switching gain (positive number, ranging from 5 to 15, dynamically adjusted by the PPO algorithm according to the real-time terrain roughness, the greater the interference, the greater the value); sign(s i(t)): sign function (output 1 when s i (t)>0, output -1 when s i (t)<0), determines the direction of switching control; : coupling strength coefficient (the sum of the absolute values of the total coupling influence of adjacent joints on the i th joint), used to dynamically adjust the amplitude of switching control. The stronger the coupling, the greater the compensation.

[0042] Core logic: when the sliding surface s i (t) is not equal to 0 (i.e. the system deviates from the ideal state), the switching control term outputs a compensation quantity opposite to the direction of the coupling disturbance through high-frequency switching, directly offsetting the disturbance caused by the coupling, forcing the system state to quickly converge to the sliding surface s i (t)=0 state. Through the sliding surface with coupling term and the high-frequency switching control law, the switching impact caused by joint coupling disturbance and joint synchronization error are reduced.

[0043] It can be understood that the present application discloses an electronic device, comprising: at least one processor; and a memory in communication connection with the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above-mentioned PID control and gait planning method for a multi-legged composite structure.

[0044] The present application discloses a PID control and gait planning system for a multi-legged composite structure, which is a hardware and software integrated system for high-precision control and efficient gait planning of a multi-legged composite structure robot. Through the collaborative work of multiple modules, problems such as joint coupling disturbance, mode switching impact, low efficiency of multi-machine cooperation, and poor environmental adaptability are solved. The specific composition and functions are as follows: The perception layer module acquires real-time robot motion state and environmental information, providing data support for control decision-making.

[0045] The perception layer module comprises: The joint state perception unit is composed of high-precision encoders, which can obtain the actual angle values of each joint (such as the hip joint and the knee joint) in real time, calculate the angle error and the error rate, and provide basic feedback data for PID control.

[0046] The attitude and motion perception unit integrates an IMU (Inertial Measurement Unit), which can obtain three-axis angular velocity and acceleration through high-frequency sampling, and calculate the attitude information of the robot, such as the pitch angle and the roll angle, through a complementary filtering algorithm, which is used for attitude closed-loop control in wheeled mode.

[0047] The terrain environment perception unit is used to: contain a laser radar, continuously scan the front terrain and build a three-dimensional terrain model, identify complex terrain features such as gullies, slopes, steps, etc., and provide environmental data for PID parameter pre-adjustment and terrain prediction in foot mode.

[0048] The control layer module realizes high-precision PID control based on perception data, and the core is a control model with coupling compensation terms to offset joint dynamics coupling and delay interference.

[0049] The control layer module includes: The coupling compensation PID control unit is used to: build a PID control model with joint coupling coefficient matrix, when calculating the joint control amount, not only the current joint proportion, integral, and differential terms are included, but also the coupling compensation terms are generated through the difference between the current and historical control amounts of adjacent joints, to correct the control delay caused by mechanical transmission and signal transmission (such as 0.05 seconds), and to ensure synchronous movement of multiple joints.

[0050] The mode dual-closed-loop control unit is used to: in wheel mode, a double closed loop is composed of a speed outer loop (with target speed as threshold) and an attitude inner loop (based on IMU feedback), and when the inclination angle exceeds the threshold, the attitude control weight is automatically increased to balance the speed and attitude stability.

[0051] In foot mode, combined with the terrain prediction data of the laser radar, the PID parameters (such as proportional gain and differential gain) of each joint are adjusted in advance, and a second derivative compensation term is introduced to adapt to the movement requirements of complex terrain.

[0052] The cooperative control module realizes efficient cooperation of multiple robot modules and solves the problems of communication delay and power distribution.

[0053] The cooperative control module includes: The distributed cooperative unit is used to: each robot module runs PID control independently, while obtaining the historical control amount of the neighbor node (considering 0.1 second communication delay) through the communication bus, calculating the difference between the local control amount and the neighbor node control amount, and generating a delay compensation term through a cooperative weight factor weighting, and adding it to the local control amount.

[0054] The power distribution optimization unit is used to: based on the cooperative coupling matrix, the joint control amount of the multi-machine combination (such as quadruped mode) is optimized, the energy and driving ability are shared and distributed through the power bus, the overall load capacity is improved, and the energy consumption is reduced.

[0055] The learning and optimization module realizes autonomous optimization of PID parameters based on reinforcement learning, and improves the adaptability to unknown environment.

[0056] The learning and optimization module includes: The state perception and data fusion unit is used for collecting in real time the terrain roughness (based on the laser radar point cloud), the current gait mode (based on the encoder and IMU data), the control accuracy (based on the joint angle error entropy value), and the energy consumption data, and constructing a four-dimensional state space.

[0057] The PPO algorithm optimization unit is used for embedding a pre-trained PPO reinforcement learning model (containing a policy network and a value network), the policy network outputs a PID parameter adjustment strategy according to a real-time state, the value network evaluates a state value to assist optimization, and the model is guided to learn autonomously through a reward function (comprehensively considering the control error, energy consumption efficiency, and motion smoothness), so that parameter adaptive adjustment under unknown terrain is realized without human intervention.

[0058] The mode switching module realizes smooth transition of the wheeled mode and the foot mode, and reduces switching impact.

[0059] The mode switching module comprises: The parameter transition unit is used for calculating a transition coefficient a(t) by using a quintic polynomial function, performing linear interpolation on proportional, integral, and differential gains, and smoothly transitioning parameters from the wheeled mode to the foot mode; during the transition process, the dynamic characteristics of a(t) are optimized by the PPO algorithm, and the control accuracy, energy consumption, and smoothness are balanced.

[0060] The impact suppression unit is used for adjusting a control amount in real time in combination with the joint angular velocity and the angular acceleration direction during the foot landing stage; in an optimal scheme, a sliding mode surface containing a joint coupling term and a high-frequency switching control law are introduced, the coupling disturbance is perceived through the sliding mode surface (containing a joint coupling coefficient), the switching control term outputs a compensation amount at a high frequency, and the motion interference of the hip joint and the knee joint is actively offset, so that the switching impact is further reduced.

[0061] Finally, it should be noted that the above-described preferred embodiments of the present application are merely used for limiting the present application, and although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or equivalently replace some technical features, any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

[0062] The above are preferred embodiments of the present application, but not for limiting the protection scope of the present application, therefore: any equivalent change made according to the structure, shape, principle of the present application shall be included in the protection scope of the present application.

Claims

1. PID control and gait planning method for wheel-foot composite structure, characterized by: Including steps: Establish a PID control model with coupling compensation terms; Calculate the angle error of each joint, compensate for the adjacent joint motion interference based on the PID control model with coupling compensation terms, and use the historical control value difference to correct the control delay of the first threshold; In wheeled mode, the outer loop uses the second threshold as the target speed control, and the inner loop corrects the attitude based on IMU feedback. When the inclination angle exceeds the third threshold, the attitude control weight is increased. In foot mode, the PID parameters are adjusted in advance based on the lidar prediction to improve adaptability to complex terrain; Real-time collection of terrain, gait, control accuracy, and energy consumption data, and automatic adjustment of PID parameters based on the PPO algorithm, to improve the robot's control accuracy in unknown terrains; Each robot module is independently PID controlled. To address the fourth threshold communication delay, the robot adjusts its own parameters according to the control amount of neighboring nodes. When switching between wheel-foot mode, the PID parameters are linearly interpolated. During the foot-stance landing phase, the control amount is adjusted according to the direction of the joint angular velocity and angular acceleration to reduce the switching impact.

2. The PID control and gait planning method of the wheel-foot composite structure according to claim 1 is characterized in that: The control equation of the PID control model with coupling compensation term is: ; For basic PID control calculation, the preliminary control quantity is generated based on the current error, cumulative error and error change rate; is the coupling compensation term, which corrects the initial control quantity, u i (t) is the control value of the i-th joint, e i (t) is the joint angle error, K p,i (t) is the time-varying proportional gain, K i,i is the integral gain, K d,i is the differential gain; C i,j is the joint coupling coefficient matrix, which is used to describe the degree of mutual influence between the movements of different joints, u j (t) is the control value of the j-th joint at time t, u j (tT d ) is the jth joint at tT d The amount of control at any moment.

3. The PID control and gait planning method of the wheel-foot composite structure according to claim 1, characterized in that: The angle error of each joint is calculated, and the motion interference of adjacent joints is compensated according to the PID control model with coupling compensation terms. This includes: obtaining the actual angle value of each joint in real time through the encoder, and subtracting it from the target angle value in the preset trajectory to obtain the angle error; based on the PID control model with coupling compensation terms, the control amount of each joint is calculated, and the degree of motion interference between the joints is quantified through the predefined coupling coefficient matrix. The control delay caused by mechanical transmission is compensated by the difference between the control amount at the current moment and the historical moment.

4. The PID control and gait planning method of the wheel-foot composite structure according to claim 1, characterized in that: The control delay of the first threshold is corrected using the difference in historical control quantities, including: obtaining the control quantities of adjacent joints at the current moment and the historical moment, calculating the difference between the two as a delay compensation signal; multiplying the difference by a predefined joint coupling coefficient matrix to generate a coupling compensation term, which is superimposed on the PID control quantity of the current joint.

5. The PID control and gait planning method of the wheel-foot composite structure according to claim 1, characterized in that: The outer loop uses the second threshold as the target speed control, including: collecting the wheel speed in real time through the encoder and comparing it with the target value to generate a speed error, and obtaining the preliminary control quantity through the PID algorithm; the inner loop corrects the attitude based on the IMU feedback, including: the attitude inner loop uses the real-time attitude data provided by the IMU as feedback, obtains the three-axis angular velocity and acceleration through high-frequency sampling, and uses the complementary filtering algorithm to fuse the data and then calculates the pitch angle and roll angle, which are compared with the preset attitude reference to generate the attitude error; enhancing the attitude control weight when the inclination angle exceeds the third threshold, including: when the IMU detects that the inclination angle exceeds the third threshold, it automatically increases the attitude control weight from the default value to 0.6 to enhance the correction force for the vehicle body tilt.

6. The PID control and gait planning method of the wheel-foot composite structure according to claim 1, characterized in that: In foot-operated mode, the PID parameters are adjusted in advance based on the lidar prediction, including: obtaining terrain elevation data in real time through continuous lidar scanning and building a three-dimensional terrain model; analyzing the acquired terrain data to identify complex terrain features such as gullies, slopes, and steps, and adjusting the PID control parameters in advance based on the set prediction time window; dynamically optimizing the proportional gain, integral gain, and differential gain of each joint based on the identified terrain type and trend, and introducing a second-order derivative compensation term to correct the control amount based on the predicted terrain change trend.

7. The PID control and gait planning method of the wheel-foot composite structure according to claim 1, characterized in that: Real-time collection of terrain, gait, control accuracy, and energy consumption data, and automatic adjustment of PID parameters based on the PPO algorithm include: obtaining terrain point cloud data through real-time scanning of the lidar to calculate the terrain roughness index, using encoder and IMU data to identify the current gait, and calculating the error entropy value based on the joint angle error sequence to characterize the control accuracy. At the same time, energy consumption data is monitored to construct a four-dimensional state space containing terrain, gait, control accuracy, and energy consumption data indicators; the real-time state is input into the pre-trained PPO algorithm model. The pre-trained PPO algorithm model used is a reinforcement learning model built based on the proximal policy optimization framework. The reinforcement learning model includes a policy network and a value network. The policy network is used to output the adjustment amount of the PID parameters, and the value network is used to evaluate the value of the current state to assist in policy optimization. The model has initial decision-making capabilities through training in various terrain and gait scenarios; the preset reward function comprehensively considers the three indicators of control error, energy efficiency and smoothness in a weighted manner, among which the control error is calculated by the deviation between the actual joint angle and the target angle. The smaller the deviation, the higher the corresponding reward value; the energy efficiency is calculated based on the energy consumption per unit time, and the lower the energy consumption, the higher the corresponding reward value; the smoothness is calculated based on the fluctuation of the joint angle change rate, and the smaller the fluctuation, the higher the corresponding reward value. The weights of the three are set according to the ratio of the highest proportion of control error, the second highest proportion of energy efficiency, and the third highest proportion of smoothness; when the real-time state is input into the reinforcement learning model, the policy network outputs the adjustment strategy of the proportional, integral and differential gains according to the current state, and the adjustment strategy of the proportional, integral and differential gains is directly used to update the parameters of the current PID controller.

8. The PID control and gait planning method of the wheel-foot composite structure according to claim 1, characterized in that: Each robot module is independently PID controlled, and for the fourth threshold communication delay, its own parameters are adjusted according to the control quantity of the neighboring node. The following includes: each robot module uses local sensor data to independently perform PID calculation, generates an initial control quantity, and then obtains the control quantity data of the neighboring node at the time of delay for the communication delay between modules, calculates the difference between its own current control quantity and the historical control quantity of the neighboring node, weights the difference through the collaborative weight factor, generates a communication delay compensation item and adds it to the local control quantity.

9. The PID control and gait planning method of the wheel-foot composite structure according to claim 1, characterized in that: When the wheel-foot mode is switched, the PID parameters are linearly interpolated, and in the foot-stance landing stage, the control quantity is adjusted according to the direction of the joint angular velocity and angular acceleration. The following is included: after the wheel-foot mode switch is started, the preset total switching time T is used as the benchmark, and the current moment in the switching process is defined as t. The transition coefficient α(t) is calculated by a quintic polynomial function, where the value of the transition coefficient α(t) increases smoothly from 0 to 1 with t, which is used to quantify the degree of transition from the wheel mode to the foot mode; for the proportional gain, integral gain, and differential gain, linear interpolation formulas are used for parameter transition respectively; during the transition process, the pre-trained PPO algorithm model is used to optimize the dynamic characteristics of the transition coefficient α(t), and the state space of the model contains the current transition coefficient α(t), joint angle error, parameter change rate, and energy consumption rate; the action space is the adjustment amount of the transition coefficient.

10. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the PID control and gait planning method of the wheel-foot composite structure described in claim 1.

Citation Information

Patent Citations

  • Predictive control method and device for wheeled robot and wheeled robot

    CN116679706A

  • Wheel-leg robot wheel-foot switching control method based on BP neural network

    CN116859975A

  • Control method and system applied to jumping motion of single-wheel-leg robot

    CN118061197A

  • Balance control method, device and equipment for wheel-legged robot and storage medium

    CN119024875A

  • Robot control method and apparatus, robot, computer-readable storage medium, and computer program product

    US20240176365A1

Cited By

  • Wheel-foot composite robot, motion control method thereof, terminal and storage medium

    CN121209567A