Four-rotor unmanned aerial vehicle suspension swing suppression flight control method based on DDPG
By using the DDPG-based reinforcement learning algorithm, fuzzy PID, and auto-disturbance rejection control technology, an inner and outer loop cascade control system was designed to solve the load swing problem of the quadrotor UAV suspension system and achieve improved stability and trajectory tracking accuracy.
Patent Information
- Application Number
- CN202510720846.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-05
AI Technical Summary
Traditional control methods are unable to effectively cope with the strong nonlinear characteristics and dynamic interference of the quadrotor UAV suspension system, which causes load swing to affect the stability and trajectory tracking performance of the UAV.
An inner and outer loop cascade control system is designed by combining a DDPG-based reinforcement learning algorithm with fuzzy PID and active disturbance rejection control technology. The ZVD input shaper is used to predict load swing disturbances and generate offset commands. DDPG is used to automatically adjust the control strategy to suppress load swing and accurately track the trajectory.
It improves the stability and trajectory tracking capability of the UAV hanging system, reduces the difficulty of parameter adjustment, and enhances the system's anti-interference ability and trajectory tracking accuracy.
Smart Images

Figure CN120595845A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of deep reinforcement learning and automated control, and relates to the technical field of flight control of a quadrotor UAV with suspension swing angle suppression, and in particular to a flight control method for a quadrotor UAV with suspension swing suppression based on DDPG (deep deterministic policy gradient algorithm). Background Art
[0002] With technological advancements and innovations, drones are rapidly developing. Quadrotor drones with hanging wings have been widely used in various missions and scenarios, such as logistics and transportation, aerial reconnaissance, agricultural spraying, and emergency rescue, demonstrating their unique value and promising future. Quadrotors have attracted considerable attention for their exceptional flight performance, including vertical takeoff and landing, precise hovering, and flexible control. However, the periodic oscillation of hanging objects during flight can affect the stability and safety of quadrotors.
[0003] Input shaping control technology is an effective strategy for reducing load swing, particularly in precisely controlled systems like quadrotor drones. By designing the input signal format, input shaping can effectively reduce load swing without affecting the system's ultimate state, thereby enhancing the stability of the quadrotor's suspension system. The advantages of input shaping lie in its simplicity and effectiveness. By preprocessing the input signal, the system's dynamic response can be improved without complicating the system.
[0004] The basic principle of input shaping is to preemptively offset potential system oscillations by controlling the command signal, thereby reducing system vibration. This is a form of feedforward control. By solving a set of constraint equations, a series of pulse signals with varying intensities and time delays are generated. These pulse signals are then convolved with the desired signal to generate a shaped signal, which serves as the system input.
[0005] During flight missions, the drone's suspended payload exhibits complex nonlinear characteristics. The suspended payload swing system model cannot be fully modeled as a second-order oscillation system, and accurate system parameters are difficult to obtain. Input shaping has limited sway control effectiveness, so a suspended payload sway angle suppressor is also required. By combining these two control methods, the payload can be quickly stabilized, ensuring the overall stability of the drone's lifting system. However, suspended payloads are more susceptible to interference in mid-air, so a sway angle suppression method based on deep deterministic policy gradients is employed.
[0006] Traditional control methods currently used to suppress the swing angle of hoisting drones typically require manual adjustment of control parameters and rely on precise mathematical model parameters. In complex flight environments, the suspended load can experience large swings, severely impacting the drone's path tracking performance and stability. Conventional control methods, however, require frequent parameter adjustments. DDPG, on the other hand, is a reinforcement learning-based algorithm with adaptive capabilities. Through interaction with the environment, it can automatically learn and adjust control strategies to adapt to dynamic and highly uncertain environments. This enables DDPG to make flexible decisions when tackling complex tasks. Compared to other policy gradient algorithms, DDPG is particularly well-suited for problems in continuous action spaces, outputting specific action values rather than probability distributions.
[0007] For complex, multi-objective tasks, traditional control methods typically require multiple control modules or carefully designed controllers. The DDPG algorithm, however, considers multiple objectives and constraints simultaneously and, through the design of a properly designed reward function, can fundamentally optimize strategies for complex tasks. For nonlinear systems, the performance of traditional control methods often declines significantly. DDPG, through its deep learning model, is able to capture the nonlinear characteristics of the system, enabling more effective control of complex control systems with highly nonlinear responses. Traditional control methods struggle to effectively handle high-dimensional state and action spaces, while DDPG leverages deep neural networks to approximate complex strategies and value functions to address these challenges. Traditional control methods often rely on immediate state decisions and lack effective use of historical data. DDPG uses an experience replay mechanism to store interaction experience with the environment, enabling reuse of historical data during the learning process, thereby improving learning efficiency and stability. Summary of the Invention
[0008] In view of the technical problems that traditional nonlinear control technology relies on precise mathematical models, is difficult to cope with the strong nonlinear characteristics and dynamic interference of the suspension system, is sensitive to noise and lumped interference, easily leads to trajectory tracking deviation and swing angle oscillation, and cannot completely eliminate the system steady-state error through a single control loop, affecting the accuracy of the transportation mission, the present invention proposes a DDPG-based flight control method for suppressing the suspension swing of a quad-rotor UAV. The UAV flight control system is designed based on fuzzy PID and active disturbance rejection control technology. By using the DDPG reinforcement learning algorithm to automatically adjust the UAV attitude angle, a closed-loop control system is formed to replace the traditional control method, reduce manual intervention, improve the degree of automation, and effectively suppress the swing of the suspended load while meeting a certain trajectory tracking accuracy.
[0009] In order to achieve the above object, the technical solution of the present invention is implemented as follows: a DDPG-based quadrotor UAV suspension swing suppression flight control method, the steps of which are as follows:
[0010] Step 1: Use the Newton-Euler method to establish the dynamic model of the quadrotor UAV suspension system;
[0011] Step 2: Use the ZVD input shaper to perform input shaping on the desired reference trajectory signal to obtain a shaped trajectory signal;
[0012] Step 3: A cascade trajectory tracking control strategy is used to decouple the quadrotor UAV suspension system into position control of the position loop and attitude control of the attitude loop. A fuzzy PID controller for the position loop is designed based on fuzzy theory. The shaped trajectory signal is input into the fuzzy PID controller to obtain the PID controller parameter correction value, and the PID controller parameters are adjusted online.
[0013] Step 4: Perform attitude settlement on the virtual control quantity output by the fuzzy PID controller and input it into the active disturbance rejection controller to obtain the virtual control quantity of the attitude loop;
[0014] Step 5: Design a DDPG-based swing angle controller to generate roll, pitch, and yaw rotation moment control variables based on the virtual control variables of the attitude loop and the control variables generated by the DDPG-based swing angle controller.
[0015] Preferably, the dynamic model of the quadrotor UAV suspension system is:
[0016]
[0017] Where u1 represents the total thrust control input, u2, u3, and u4 represent the moment control inputs for controlling the pitch, roll, and yaw of the drone, M and m represent the masses of the drone and the payload, L is the length of the rope, l represents the arm length of the quadrotor drone, φ, θ, and ψ represent the roll angle, pitch angle, and yaw angle of the drone relative to the inertial coordinate system {W}, respectively. denote the angular velocities of roll, pitch and yaw respectively, are the angular accelerations of roll, pitch and yaw, I x , I y , I z is the moment of inertia of the quadrotor; g is the acceleration due to gravity, is the acceleration of the drone, are the angular velocities of the load swing angles α and β, respectively, They are the load swing angle and the load swing angle, respectively; the load swing angle α represents the rope in the body coordinate system {X B , Y B , Z B Plane Y B O B Z B The projection on Z B The angle between the axes and the load swing angle β represents the angle between the rope and the plane X.B O B Z B The projection on Z B The angle between the axes; a1, a2, a3, b1, b2, b3, c1, and c2 are all intermediate variables.
[0018] Preferably, the method for constructing the dynamic model of the quadrotor drone suspension system is: converting the total thrust F in the body coordinate system {B} to the inertial coordinate system {W} through the rotation matrix R from the body coordinate system {B} to the inertial coordinate system {W} T Decomposed into the force components [F x ,F y ,F z ] T ;
[0019] According to the load swing angles α and β, the positional relationship between the UAV and the load is determined, and the second-order derivative of the UAV position (x, y, z) is obtained to obtain the relationship between the acceleration of the load and the acceleration of the UAV; according to Newton's second law, the component force [F x ,F y ,F z ] T The relationship between the acceleration of the load and the acceleration of the drone is obtained by the equation I of the acceleration of the drone. Similarly, the components of the load in the x, y, and z axis directions under the rope tension T in the body coordinate system {B} are obtained [T x ,T y ,T z ] T ; According to the rotation matrix R from the load coordinate system {l} to the inertial coordinate system {W} l The suspended load is subjected to the total thrust F of the UAV T Transform from the load reference system {l} to the inertial coordinate system {W} and obtain the total force F acting on the suspended load in the body coordinate system {B} l ;
[0020] The Newton-Lagrange method is used to analyze the forces acting on the drone and the load, and the equation II for the relationship between the acceleration of the load and the acceleration of the drone is obtained. The rope tension T is calculated based on the relationship between the acceleration of the load and the acceleration of the drone.
[0021] Without considering the resistance and ignoring the torque generated by the load, the angular acceleration of the UAV's roll angle, pitch angle and yaw angle are obtained. According to the UAV acceleration equation I, the UAV acceleration equation II and the rope tension T, the component force [F x ,F y ,F z ] T Determine the acceleration of the drone
[0022] The intermediate variable Intermediate variables
[0023]
[0024] The high-level control instructions are mapped to the speeds of the four motors through control distribution:
[0025]
[0026] Among them, c t is the lift coefficient, c m is the counter-torque coefficient, ω1, ω2, ω3, and ω4 are the speeds of the four motors.
[0027] Preferably, the ZVD input shaper performs input shaping processing on the input reference trajectory signal, which is achieved by using three pulse signals, and the constraint equation is:
[0028]
[0029] Among them, A1, A2, and A3 represent the amplitudes of the three pulse signals input to the ZVD shaper, t1 represents the initial time; T(ξ,ω n ) is the residual oscillation percentage, ω n is the natural frequency of the system, ξ is the damping ratio of the system;
[0030] According to the constraint equation, the amplitude and time delay t of the ZVD input shaper can be obtained. i respectively:
[0031]
[0032] Where k is the damping correction factor, τ d is the pulse time point.
[0033] Preferably, the reference trajectory signal (x d ,y d ,z d ) through the three pulse trains of ZVD input shaper and ZVD input shaper Perform convolution to obtain the shaped trajectory signal which is input into the fuzzy PID controller;
[0034] The residual oscillation percentage T(ξ,ω n ) is the ratio of the amplitude of the system response after input shaping to the amplitude of the total system response without input shaping, and ω d is the damped oscillation frequency of the system,
[0035] Preferably, the fuzzy PID controller is a two-input three-output structure, the input of the fuzzy PID controller is the trajectory tracking error e and the corresponding change rate ec, and the output is the PID controller parameter correction ΔK p , ΔK i , ΔK d The trajectory tracking error e is the difference between the shaped trajectory signal r(t) and the current state. The trajectory tracking error e and the rate of change ec are fuzzified, fuzzy inference and defuzzified, and the output is the parameter correction ΔK p , ΔK i , ΔK d , input into the PID controller, and adjust the control parameter K of the PID controller in real time p , K i , K d , ensuring that the suspension system accurately tracks the reference trajectory while effectively suppressing the load swing angle;
[0036] The fuzzy control rules of the fuzzy PID controller adopt the standardized design created by Mamdani, and the output fuzzy control rule table is established as follows:
[0037]
[0038] The membership function of the fuzzy subset of the trajectory tracking error e adopts trigonometric function, and the seven-level fuzzy subset {NB, NM, NS, ZO, PS, PM, PB} is selected, where N represents negative, P represents positive, B represents large, M represents medium, S represents small, and ZO represents zero;
[0039] The centroid method is used to obtain the precise value of the parameter correction.
[0040] Preferably, the fuzzy PID controller is based on the trajectory tracking error e of the reference trajectory and the current position. p =[e x e y e z ] T Get the virtual control quantity u of the position loop p =[u x u y u z ] T , and the virtual control quantity is: is the trajectory tracking error e p The derivative of
[0041] The attitude solver calculates the desired attitude angle φ of the drone d and θ d And the thrust control quantity u1 of the UAV is:
[0042]
[0043] Among them, ψ d is the reference yaw angle.
[0044] Preferably, an attitude loop auto-disturbance rejection controller is designed based on the dynamic model of the quadrotor UAV suspension system. The auto-disturbance rejection controller obtains the virtual control quantity of the attitude loop according to the error between the current attitude angle and the desired attitude angle after attitude settlement. The auto-disturbance rejection controller is divided into three parts: a tracking differentiator, an extended state observer, and a nonlinear feedback error. All uncertainties affecting the UAV are regarded as a lumped disturbance. The extended state observer estimates and compensates for the lumped disturbance in real time, and approximately transforms the uncertain nonlinear system into an easy-to-control linear integral series system. The tracking differentiator is the desired attitude angle φ. d and θ d Arrange a smooth, non-overshoot transition process and extract the differential signal; the nonlinear feedback error uses the transition signal generated by the tracking differentiator and the state estimated by the extended state observer and the lumped disturbance to dynamically compensate the control quantity and generate the torque control quantity (u φ ,u θ ,u ψ ).
[0045] Preferably, according to the virtual control amount (u φ ,u θ ,u ψ ) and the control variables D1, D2, and D3 generated by the DDPG-based angle controller generate the moment control inputs for roll, pitch, and yaw:
[0046]
[0047] The state space of the DDPG-based swing angle inhibitor is S(t)=[x,y,z,e x ,e y ,e z ,φ,θ,ψ,e φ ,e θ ,e ψ ,α,β], (x,y,z) is the position of the UAV, (φ,θ,ψ) is the attitude angle of the UAV, (e x ,e y ,e z ) is the position error of the UAV, (e φ ,e θ ,e ψ ) is the attitude angle error of the UAV, (α, β) is the load swing angle;
[0048] Control action a t =[D1,D2,D3] TThese are the three torque control signals of the quadrotor, and the action value range of the control quantities D1, D2, and D3 is [-20, 20].
[0049] Preferably, the DDPG-based swing angle inhibitor adopts a deterministic strategy. The Actor current network directly outputs the specific action value rather than the probability distribution of the action; the Critic current network evaluates the Q value of the state-action pair Q = (s t ,a t |θ Q ), guides the optimization strategy of the Actor's current network; the Actor's target network generates the action a of the next state t+1 , used to calculate the target Q value; the critic target network is used to calculate the target Q value and maintain training stability; the target network uses a soft update method to stably approach the current network:
[0050] θ Q′ ←τθ Q +(1-τ)θ Q′ ,θ μ′ ←τθ μ +(1-τ)θ μ′ ;
[0051] Among them, τ<<1; θ μ is the parameter of the Actor's current network, θ Q is the parameter of the current Critic network, θ μ' is the parameter of the Actor target network, θ Q' is the parameter of the Critic target network; s t Indicates the state at the current time t;
[0052] Randomly initialize the Actor current network and the Critic current network, copy the parameters to the target network; perform action a t , and store the transferred samples in the experience replay pool, and calculate the target Q value of the Critic target network: y t = r + γQ'(s', μ'(s')); where s' and μ'(s) are the next state and target action, respectively. Represents expectation; the optimization goal of the Critic target network is to minimize the loss function
[0053] The Actor target network optimizes the strategy by maximizing the Q value of the Critic target network. The objective function is: N represents the number of samples;
[0054] Update the actor's current network parameters θ by applying the chain rule with respect to the actor's parameters μ :
[0055]
[0056] Among them, θ μ is the weight of the function approximator of the Actor's current network, Respectively represent the Q value of μ θ , a finds the gradient;
[0057] The input of the Actor network is the state space S(t) of the system. Two fully connected layers are used to extract high-order state features. The ReLU activation function is used to enhance the nonlinear expression capability of the hidden layer. The output layer uses the Tanh activation function to normalize the output to [-1,1]. A scaling layer is used to scale the actual action range to ensure consistency with the upper and lower limits of the action space. The output action μ(s t |θ μ ); Perform L2 regularization on the Actor network;
[0058] The critic network has a dual input path of state + action followed by a fusion layer, ultimately outputting a Q value. The state path inputs the state vector S(t), compressing the high-dimensional state information into low-dimensional features. Two fully connected layers are used to extract high-order state features. A normalization layer is added between the two fully connected layers to accelerate convergence. The dynamic path inputs the action vector to extract action features. The dual fully connected layers are designed to enhance nonlinear expression capabilities. The state features and action features are merged through the splicing layer to output the Q value.
[0059] The critic network estimates the Q value of the state-action pair by minimizing the temporal difference error, and the TD error δ t The calculation of depends on the reward value, and
[0060] Q target (s t+1 ,a t+1 )=r t +γQ target (s t ,a t );
[0061] δ t =Q target (s t ,a t )-Q(s t ,a t );
[0062] Among them, r t is the current reward, Q target represents the target Q value, Q(s t ,a t ) represents the Q-value function of the current state-action pair, and γ is the discount factor;
[0063] The reward function is Among them, when the drone ends the current training early due to unexpected circumstances, it will be punished: The termination condition is
[0064] The reward function for trajectory tracking and payload swing angle is: Where Ω∈[x, y, z, α, β], ε is a constant;
[0065] When the attitude angle error is always less than a certain value, a reward is given, and when it exceeds a certain value, a penalty is given. The reward function is:
[0066]
[0067] Compared with the existing technology, the beneficial effects of the present invention are as follows: strong nonlinear coupling and load swing interference in the UAV suspension flight system, a UAV control strategy based on auto-disturbance rejection, and feedforward control of zero vibration first-order differential shaper (ZVD) are used to suppress residual vibration. At the same time, a dual closed-loop control system based on fuzzy PID and auto-disturbance rejection is designed, and the feedback control of the DDPG reinforcement learning algorithm is used to control the torque control signal of the inner loop attitude controller to suppress the problem of load residual oscillation. The use of a fuzzy PID controller in the position loop makes it possible to accurately track the desired trajectory while effectively suppressing the load swing angle, thereby enhancing the stability of the UAV suspension system, reducing the difficulty of parameter adjustment, and improving the trajectory tracking capability.
[0068] This invention aims to provide a novel solution for a quadrotor unmanned aerial vehicle (UAV) suspended payload system to successfully complete transport missions. It employs inner and outer loop cascade control technology and a hierarchical control architecture to decompose complex problems, reduce the design complexity of individual loops, and improve system stability. To address the problems of load swing during flight and residual load oscillation during hovering, an input shaping technique, namely a ZVD shaper, is proposed to predict load swing disturbances and generate offsetting commands to suppress residual oscillations. A dynamic fuzzy-PID parameter adjustment mechanism is employed to enhance the robustness of the outer loop trajectory tracking. An end-to-end reinforcement learning scheme, designed using the DDPG algorithm, combines feedforward and feedback control, eliminating the need for an accurate dynamic model. This scheme balances swing angle suppression, trajectory tracking, and energy optimization through a multimodal reward function. Simulation results demonstrate that the system successfully ensures reliable transport mission execution under complex external interference conditions, significantly improving system control accuracy and anti-interference capabilities. By accurately tracking the preset trajectory and effectively suppressing load swing, this invention provides a complete and efficient control strategy for quadrotor suspended payload systems in transport scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0070] Figure 1 This is a structural diagram of the UAV hanging system of the present invention.
[0071] Figure 2 This is the overall control architecture of the drone suspension system of the present invention.
[0072] Figure 3 This is a principle block diagram of the active disturbance rejection control (ADRC) of the present invention.
[0073] Figure 4 This is the principle diagram of the position loop fuzzy PID controller of the present invention.
[0074] Figure 5 This is an example diagram of the fuzzy subset membership function of the present invention.
[0075] Figure 6 This is the framework diagram of the DDPG algorithm of the present invention.
[0076] Figure 7 for Figure 6 The network update flow chart in .
[0077] Figure 8 for Figure 6 The Actor network structure diagram shown.
[0078] Figure 9 for Figure 6 The Critic network structure diagram shown in Figure 2 is as follows:
[0079] Figure 10 A trajectory tracking performance comparison chart showing the error comparison of different control algorithms (traditional PID, backstepping sliding mode, and the present invention) in spiral trajectory tracking.
[0080] Figure 11 This is a graph showing the change of the attitude angle, roll angle, and pitch angle of the drone over time.
[0081] Figure 12 Schematic diagram of the convergence process of load swing angles α, β under different control strategies in the present invention.
[0082] Figure 13 This is a comparison diagram of the actual flight trajectory of the UAV in the present invention and the expected spiral trajectory. DETAILED DESCRIPTION
[0083] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without creative work are within the scope of protection of the present invention.
[0084] like Figure 1 As shown in the figure, a DDPG-based flight control method for quadrotor UAV suspension swing suppression is proposed. First, the Newton-Euler method is used to establish the dynamic model of the quadrotor UAV suspension system. Then, a cascade trajectory tracking strategy is used to decouple the system into an inner loop attitude control and an outer loop position control. Appropriate controllers are designed for the two loops. To address the problems of load swing and severe coupling and mutual interference during UAV flight, a nonlinear controller is designed based on a reinforcement learning algorithm to achieve decoupling control under underactuated constraints. The specific implementation steps of the present invention are as follows:
[0085] Step 1: Use the Newton-Euler method to establish the dynamic model of the quadrotor UAV suspension system.
[0086] like Figure 1 As shown in Figure 1, consider a quadrotor drone connected to a point load via a massless and inextensible rope. Assume the rope is connected to the center of mass of the quadrotor and air resistance is negligible. Assume the inertial coordinate system {W}, the drone's body coordinate system {B}, and the load coordinate system {l}, where {X W , Y W , Z W} and {X B , Y B , Z B} and {X l , Y l , Z l} respectively represent the positive directions of the inertial coordinate system {W}, the body coordinate system {B}, and the load coordinate system {l}. α represents the positive direction of the rope in plane Y B O B Z B The projection on Z B The angle between the axes, β represents the angle between the rope and the plane X. B O B Z B The projection on Z B The angle between the axes, O B is the origin coordinate of the aircraft coordinate system {B}. M and m represent the masses of the drone and the payload, respectively, and L is the length of the rope.
[0087] Compared with the controller based on Lagrangian modeling, the nonlinear controller based on the Newton-Euler dynamics model has a simpler controller structure, avoids complex matrix calculations, and greatly reduces the computational burden. The dynamic model establishment process of the quadcopter UAV hanging load system is as follows:
[0088] The control input vector of the quadrotor is defined as:
[0089] U=[u1,u2,u3,u4] T =[F T ,τ1,τ2,τ3] T (1)
[0090] Among them, u1 represents the total thrust control input, u2, u3, and u4 represent the torque control inputs for controlling the attitude (pitch, roll, and yaw) of the drone, and F T represents the total thrust generated by the four rotors; τ1, τ2, and τ3 represent the torques generated by the rotors along the X, Y, and Z axes in the body coordinate system {B}.
[0091] First, the total thrust F in the body coordinate system {B} is converted to T Decomposed into the force components [F x ,F y ,F z ] T for:
[0092]
[0093] Where φ, θ, and ψ are the roll, pitch, and yaw angles of the quadrotor drone relative to the inertial coordinate system {W}, respectively. According to the load swing angle, the positional relationship between the drone and the load satisfies:
[0094]
[0095] The load position is represented by p l =(x l ,y l ,z l ) T The position of the drone is represented by p = (x, y, z) T Expressed. Taking the second-order derivative of x, y, and z from equation (3) yields:
[0096]
[0097] Among them, the intermediate variable According to Newton's second law, we can get:
[0098]
[0099] Substituting equations (2), (3) and (4) into equation (5) yields:
[0100]
[0101] Where g is the acceleration due to gravity, is the acceleration of the quadrotor drone. are the angular velocities of the load swing angles α and β, respectively, are the angular accelerations of the load swing angles respectively. And Intermediate variables
[0102] Similarly, according to formula (3), the load components in the x-, y-, and z-axis directions under the rope tension T in the body coordinate system {B} can be obtained as follows:
[0103]
[0104] The rotation matrix from the load coordinate system {l} to the inertial coordinate system {W} is:
[0105]
[0106] According to the above rotation matrix R l The suspended load is subjected to the total thrust F of the UAV T Transformed from the load reference frame {l} to the inertial coordinate system {W}, the total force F acting on the suspended load l for:
[0107]
[0108] Among them, (F lx ,F ly ,F lz ) are the thrust F l The components of force in the x, y, and z axes in the body coordinate system {B}.
[0109] The Newton-Lagrange method is used to analyze the stress on the drone and the load and obtain the following:
[0110]
[0111] Combining equations (4), (11), and (12), we can obtain the rope tension T as follows:
[0112]
[0113] Without considering the resistance and ignoring the torque generated by the load, the attitude dynamics model of the quadrotor drone is obtained:
[0114]
[0115] Where l represents the arm length of the quadrotor, φ, θ, and ψ represent the roll, pitch, and yaw angles of the quadrotor relative to the inertial coordinate system {W} (rad), are the angular velocities of roll, pitch, and yaw (rad / s), The angular accelerations of roll, pitch, and yaw (rad / s 2 ). I x , I y , I z is the moment of inertia of the quadrotor.
[0116] Arranging equations (6), (11), and (14) yields the quadrotor UAV dynamics model:
[0117]
[0118] The total thrust and torque of the quadcopter is determined by the rotational speed of the four rotors. For the "X" type UAV control, the high-level control instructions are mapped to the rotational speeds of the four motors through control distribution:
[0119]
[0120] Among them, c t is the lift coefficient, c m is the counter-torque coefficient, ω1, ω2, ω3, and ω4 are the speeds of the four motors.
[0121] Step 2: Use the ZVD input shaper to perform input shaping processing on the desired reference trajectory signal to obtain a shaped trajectory signal.
[0122] Aiming at the load swing during flight and the residual oscillation of the load when the UAV is hovering, a strategy based on the ZVD input shaper is designed to suppress the residual oscillation of the quadrotor UAV suspension system, thereby suppressing the load swing when the UAV is transporting the load.
[0123] Input shaping is a control strategy that pre-processes the control input to adjust the system response in advance to reduce vibration or oscillation. The Zero Vibration First-Order Differential (ZVD) shaper is a commonly used input shaping tool that effectively suppresses residual vibration in the system's dynamic response.
[0124] In the presence of resistance, the system model of the suspended load is approximately regarded as a second-order oscillation system, and its transfer function is:
[0125]
[0126] Among them, ωn is the natural frequency of the system, ξ is the damping ratio of the system, s represents the complex frequency variable, and the unit impulse response of the suspended load swing angle system is:
[0127]
[0128] Among them, ω d is the damped oscillation frequency of the system,
[0129] After input shaping, the reference trajectory signal generates three pulse signals. The first one acts on the quadcopter's suspension system at the initial time, causing it to swing. After a certain delay time, the second signal acts on the quadcopter's suspension system, causing it to swing in the opposite direction to the swing caused by the first signal, but with the same amplitude. In this way, the swings generated by the quadcopter's suspension system will cancel each other out, thereby stabilizing the load. Therefore, the core of input shaping technology lies in determining the strength and delay time of each pulse signal.
[0130] The ZVD input shaper contains a series of pulse signals with different intensities and delay times, whose amplitudes and delays can be expressed as a i and t i , the system response caused by the signal is:
[0131]
[0132] Among them, t i represents the time delay, t represents the current time, and i represents the i-th pulse signal.
[0133] The total response of the quadrotor drone's suspended load system is the sum of the responses caused by all signals:
[0134]
[0135] in,
[0136] Compare the amplitude of the system response after input shaping to the amplitude of the total system response without input shaping, and define this ratio as the residual oscillation percentage T(ξ,ω n ),and:
[0137]
[0138] The influence of the input shaper on the system can be judged by observing the residual oscillation percentage. Therefore, in order to eliminate the residual vibration generated by the system, the residual oscillation percentage needs to be equal to zero. From formula (21), it can be seen that the constraint needs to be satisfied:
[0139]
[0140] The ZVD input shaper is designed to perform input shaping on the input reference trajectory signal. This is achieved using three pulse signals. The constraint equation is:
[0141]
[0142] Among them, A1, A2, and A3 represent the amplitudes of the three pulse signals of the ZVD input shaper, and t1 represents the initial time. According to the above constraints, the amplitude and time delay of the ZVD input shaper can be obtained as follows:
[0143]
[0144] Where k is the damping correction factor, τ d is the pulse time point.
[0145] Reference trajectory signal (x d ,y d ,z d ) through the three pulse trains of ZVD input shaper and ZVD input shaper Perform convolution to obtain the shaped trajectory signal which is input into the fuzzy PID controller.
[0146] Input shaping is a feedforward technology that eliminates residual vibration of flexible systems. It suppresses terminal vibration by shaping the input signal. The ZVD input shaper is designed and solved according to the characteristics of the system. The reference trajectory signal (x d ,y d ,z d ) is convolved with the pulse train of the shaper, and the control signal after passing through the shaper is used to control the system.
[0147] Step 3: Use the cascade trajectory tracking control strategy to decouple the quadrotor UAV suspension system into the position control of the position loop and the attitude control of the attitude loop, and design the fuzzy PID controller of the position loop based on fuzzy theory; input the shaped trajectory signal into the fuzzy PID controller to obtain the PID controller parameter correction, and adjust the PID controller parameters online.
[0148] The present invention designs an outer-loop fuzzy PID controller and an inner-loop anti-disturbance control controller to address the external interference of the hanging load on the UAV and the internal interference due to model accuracy. While reducing the impact of the hanging load on the UAV body, the trajectory is accurately tracked, the robustness and anti-interference ability of the UAV are improved, and the UAV hanging system is ensured to maintain stable flight in various environments.
[0149] Based on fuzzy theory, on the basis of theoretical analysis, control rules are designed according to control experience knowledge, and control coefficients are designed. The principle block diagram of fuzzy PID controller is as follows Figure 4As shown, fuzzification involves the range of values of the input error of the quadcopter UAV and the corresponding mapping. The position loop fuzzy PID controller designed in this invention has a two-input and three-output structure. The input of the fuzzy PID controller is the trajectory tracking error e and the corresponding change rate ec, and the output is the PID controller parameter correction value ΔK p , ΔK i , ΔK d The trajectory tracking error e is the difference between the shaped trajectory signal r(t) and the current state c(t). When the control system is running, the input trajectory tracking error e and the rate of change ec are fuzzified, fuzzy inference and defuzzified to obtain the output parameter correction ΔK p , ΔK i , ΔK d Input into the PID controller and adjust the control parameter K of the PID controller in real time p , K i , K d , ensuring that the suspension system accurately tracks the reference trajectory while effectively suppressing the load swing angle.
[0150] Establishing fuzzy control rules is the most important step in designing a fuzzy PID controller. The fuzzy rules in this invention adopt the standardized design created by Mamdani. The output fuzzy control rule table is established as shown in Table 1. All the fuzzified membership degrees are applied one by one to the IF-THEN rules in the fuzzy control rule table to calculate the triggering conditions of each rule. For example, "IF e is PB, and ec is NB THEN ΔK p =ZO,ΔK i =ZO,ΔK d =PB”
[0151] Fuzzification is the process of determining the membership of the precise input value based on the membership function of the fuzzy set, normalizing the input quantity, discretizing the domain of the normalized input quantity, and converting it into the corresponding variable value on the domain. Figure 5As shown, taking the membership function of the input trajectory tracking error e as an example, the membership function of the fuzzy subset adopts trigonometric function. The trigonometric membership function only requires three parameters (left endpoint, vertex, right endpoint) to be defined, and the mathematical expression is simple. This structure facilitates the rapid calculation of membership values, especially in fuzzy PID control, which can significantly improve the computational efficiency. Define the range of input and output as the domain [-3, 3] on the fuzzy set, and select the seven-level fuzzy subset {NB, NM, NS, ZO, PS, PM, PB}, where N represents negative, P represents positive, B represents large, M represents medium, S represents small, and ZO represents zero, which is used to fuzzily describe the physical range of input / output. Input the precise error value and map each input variable to the membership of the PID parameter of the corresponding fuzzy set. According to the membership of the quantized value of the currently input trajectory tracking error e and the rate of change ec, all activated rules are screened out. For example, if the membership of e to "PB" is a and the membership of ec to "NM" is b, then ΔK p , ΔK i , ΔK d The membership degree of ZO / ZO / PM is a*b, and the activation of the above rules includes the combination of PB and NM. The product of the membership degrees is used as the rule activation strength. The subsequent outputs of all activated rules are superimposed to form a comprehensive fuzzy output. The maximum membership degree of each rule output is taken to improve computational efficiency.
[0152] According to the centroid of the area enclosed by the fuzzy set membership curve and the coordinate axis, its horizontal coordinate is used as the accurate output value. The centroid method is used to obtain the accurate value. Let A be a non-empty fuzzy set, and in x1, x2…x m Form m vertical slices at the position, and take the weighted average value of the membership degree corresponding to each element in the fuzzy control quantity to obtain the exact value x A , ΔK p For example:
[0153]
[0154] Where ΔK p,i is the fuzzy variable element, μ(ΔK p,i ) is the degree of membership of the element. In summary, we can get the control quantity K of fuzzy PID p , K i , K d .
[0155]
[0156] Among them, k p , k i , k d are the initial parameters of the PID controller.
[0157] Table 1 Parameter correction value ΔKp , ΔK i , ΔK d Fuzzy control rules
[0158]
[0159] Step 4: The virtual control quantity output by the fuzzy PID controller is input into the active disturbance rejection controller after attitude settlement.
[0160] The outer loop fuzzy PID controller is based on the deviation e between the target (reference) trajectory and the current position. p Get the virtual control quantity u of the position loop p =[u x u y u z ] T Let the trajectory tracking error e p =[e x e y e z ] T , error e p The derivative of Then the virtual control quantity is:
[0161]
[0162] The attitude solver calculates the desired attitude angle φ of the drone d and θ d And the thrust control quantity u1 of the UAV is:
[0163]
[0164] Among them, ψ d is the reference yaw angle.
[0165] The attitude loop's ADRC controller is designed based on the dynamic model of the quadrotor UAV's suspension system. The ADRC obtains the virtual control variable of the attitude loop based on the error between the current attitude angle and the desired attitude angle after attitude settlement. The ADRC consists of three parts: tracking differentiator (TD), extended state observer (ESO), and nonlinear feedback error (NLSEF). Figure 3 As shown in Figure 2. All uncertainties affecting the UAV (model parameter errors, unmodeled dynamics, external disturbances) are considered as a "sum disturbance". This disturbance is estimated and compensated in real time by the extended state observer to obtain the lumped disturbance z3, thereby approximately transforming the uncertain nonlinear system into an easy-to-control linear integral series system. The control quantity of the position loop is the reference attitude angle φ obtained by attitude angle solution. d ,θ dThe tracking differentiator TD arranges a smooth, non-overshoot transition process v1 for the reference attitude angle and extracts the differential signal v2 of the transition process v1. NLSEF uses the transition signals v1, v2 generated by TD and the states z1, z2 estimated by the extended state observer ESO, as well as the lumped disturbance z3 to dynamically compensate the control variable. Finally, the attitude angle control variable (u φ ,u θ ,u ψ ).
[0166] Step 5: Design a DDPG-based swing angle controller to generate roll, pitch, and yaw rotation moment control variables based on the virtual control variables of the attitude loop and the control variables generated by the DDPG-based swing angle controller.
[0167] According to the virtual control amount of the attitude loop (u φ ,u θ ,u ψ ) is combined with the control quantities D1, D2, and D3 generated by the DDPG-based swing angle controller to generate the roll, pitch, and yaw rotation torque control quantities u2, u3, and u4.
[0168]
[0169] DDPG is a deep reinforcement learning algorithm based on the Actor-Critic architecture, which is specially designed to solve the control problem of continuous action space. x ,e y ,e z ), attitude angle error (e φ ,e θ ,e ψ ) and the load swing angle (α, β) constitute the state space S(t), and
[0170] S(t)=[x,y,z,e x ,e y ,e z ,φ,θ,ψ,e φ ,e θ ,e ψ ,α,β] (30)
[0171] Among them, (x, y, z) is the coordinate of the drone’s position, (e x ,e y ,e z ) is the coordinate of the position error of the UAV, (φ,θ,ψ) is the attitude angle of the UAV, and the attitude angle error (e φ ,e θ ,e ψ) is the attitude angle reference value minus the current attitude angle, and (α, β) is the swing angle of the load.
[0172] The control action a(t) is the three torque control signals of the quadrotor, and the action value range of the control quantities D1, D2, and D3 is [-20, 20].
[0173] a t =[D1,D2,D3] T (31)
[0174] like Figure 6 As shown, DDPG contains a total of 4 networks, namely Actor current network for outputting action a t =μ(s t |θ μ ), the Critic current network is used to evaluate the current state action value Q = (s t ,a t |θ Q ), the Actor target network is used to generate the next state-action pair, and the Critic target network is used to calculate the target Q value. θ μ is the parameter of the Actor's current network, θ Q is the parameter of the current network of Critic, θ μ' is the parameter of the Actor target network, θ Q' is the parameter of the Critic target network, s t Represents the state at the current time t, μ represents the action. And set the experience playback buffer of size D to store the interaction experience tuple (s t ,a t ,r t ,s t+1 ), where s t+1 is the next state. t For the current state s t Reward for the action.
[0175] DDPG uses a deterministic strategy, where the policy network (Actor) directly outputs specific action values rather than the probability distribution of actions. This design avoids the computational burden of sampling in a high-dimensional continuous action space and is suitable for tasks such as drone control that require precise actions. The Actor network is used to input states and output deterministic actions. The Critic network evaluates the Q value of the state-action pair and guides the optimization strategy of the Actor network. The Actor target network generates the action a for the next state. t+1 , used to calculate the target Q value. The critic target network is used to calculate the target Q value and maintain training stability. The target network uses a soft update method to steadily approach the current network instead of directly copying the current network parameters.
[0176] θ Q′ ←τθ Q +(1-τ)θ Q′ ,θ μ′ ←τθ μ +(1-τ)θ μ′ (32)
[0177] Wherein, τ<<1 (e.g., 0.005).
[0178] The network update process is as follows Figure 7 As shown, first randomly initialize the Actor current network and the Critic current network, and copy the parameters to the target network. Execute action a t And store the transferred samples in the experience replay pool. Then calculate the target Q value of the Critic target network:
[0179] y t =r+γQ'(s',μ'(s')) (33)
[0180] s' and μ'(s) are the next state and target action respectively, Represents expectation. The optimization goal of the Critic target network is to minimize the loss function:
[0181]
[0182] Similarly, the Actor target network optimizes the strategy by maximizing the Q value of the Critic target network. The objective function is:
[0183]
[0184] N represents the number of samples. Update the Actor's current network parameters θ by applying the chain rule to the actor parameters μ :
[0185]
[0186] Among them, θ μ is the weight of the function approximator of the Actor's current network. Respectively represent the Q value of μ θ , a finds the gradient.
[0187] The DDPG algorithm interacts with the environment through the control agent. t , the Actor network generates a deterministic action a t =μ(s t ∣θ μ ), by interacting with the environment, we get reward r t and the next state s t+1, and then the interaction experience tuple (s t ,a t ,r t ,s t+1 ) stores the experience back into the buffer. The Actor network directly outputs continuous actions (torque commands), while the Critic network provides action value feedback. The two work together to achieve model-free reinforcement learning, eliminating the reliance on precise dynamic models.
[0188] The deep network design of the DDPG algorithm includes two deep networks: the Actor strategy network and the Critic value network. The Actor network structure in this invention is as follows: Figure 8 As shown in Figure 1. The input of the Actor network is the state space S(t) of the system, so the input has 14 nodes. Two fully connected layers (256→128) are used to extract high-order state features. The ReLU activation function is used to enhance the nonlinear expression capability of the hidden layer. The output layer uses the Tanh activation function to normalize the output to [-1,1]. A scaling layer is used to scale the actual action range to ensure consistency with the upper and lower limits of the action space. The output action μ(s t |θ μ The Actor network is set to 8 layers, and the number of nodes in each layer is [14, 256, 256, 128, 128, 3, 3, 3]. To prevent overfitting or gradient vanishing during training, the present invention performs L2 regularization on the Actor network.
[0189] The critic network structure is as follows Figure 9 As shown, the architecture of the dual input path (state + action) is followed by a fusion layer, and the Q value is finally output. The input in the state path is the state vector S(t), the purpose of which is to compress the high-dimensional state information into low-dimensional features. There are 14 nodes, and two fully connected layers are used to extract high-order state features. A normalization layer is added to the two fully connected layers to accelerate convergence, and the dimension transformation is (14→256→256→128). The input in the dynamic path is the action vector, the purpose of which is to extract action features. A dual fully connected layer is designed to enhance the nonlinear expression ability, and the dimension transformation is (3→128→128→128). Next, the state and action features are merged through the splicing layer to output the Q value. The dimension is gradually reduced (256→64→1), balancing nonlinear modeling and preventing overfitting. In the network designed by the present invention, the state and action are encoded independently to avoid early feature confusion, and a normalization layer is used to improve training stability. ReLU activation is followed by each layer to introduce nonlinearity. The deep structure is suitable for approximating complex Q-value functions.
[0190] The reward function is the interactive feedback mechanism between the agent and the environment, and is essentially a mathematical expression of the task goal. The reward function quantifies the task goals such as more accurate trajectory tracking, smaller load swing angle, and more stable posture into specific values, enabling the agent to learn to achieve the goal by maximizing the cumulative reward. The reward function guides the agent to choose specific actions through a reward and punishment mechanism. The reward function not only focuses on immediate rewards, but also needs to balance short-term and long-term returns through the discount factor γ to ensure that the agent chooses actions that are better for long-term cumulative benefits. The Critic network estimates the Q value of the state-action pair by minimizing the temporal difference error, and the TD error δ t The calculation of depends on the reward value reward, and
[0191] Q target (s t+1 ,a t+1 )=r t +γQ target (s t ,a t ) (37)
[0192] δ t =Q target (s t ,a t )-Q(s t ,a t ) (38)
[0193] Among them, r t is the current reward, which is the measure of the immediate task completion fed back by the environment, Q target Represents the target Q value, calculated by the Critic target network. Q(s t ,a t ) represents the Q-value function of the current state-action pair.
[0194] The Actor network updates its strategy by maximizing the Q value, and the accuracy of the Q value depends on the rationality of the reward function. Reasonable reward function settings can significantly reduce training time. During the design process, it is necessary to balance sparse rewards and dense rewards. The size of the reward needs to be controlled within a reasonable range. Too large may lead to unstable gradients, while too small will slow learning. Reward components of different dimensions (such as speed, distance, and swing angle) are normalized to prevent certain components from dominating the learning process. When designing rewards, it is necessary to prevent the agent from falling into a local optimal trap. The reward function for trajectory tracking and swing angle designed by this invention is shown in Equation (39). When the error is less than the threshold, a reward is given, otherwise a penalty is given. Where Ω∈[x, y, z, α, β], in order to prevent the reward value from fluctuating violently when the error approaches zero, a minimum value ε is added to prevent division by zero errors.
[0195]
[0196] At the same time, rewards are given when the attitude angle error is always less than a certain value, and penalties are given when it exceeds the certain value, as shown in formula (40), j = [α, β]
[0197]
[0198] In reinforcement learning, the termination condition is a key rule that defines when a task ends during the interaction between the environment and the agent. It directly determines how the agent explores the environment, how it accumulates experience, and ultimately how it learns the optimal strategy. For example, when the drone's flight altitude, trajectory tracking error, or attitude angle exceed a certain threshold, the current round of training is terminated. A reasonable termination condition can adjust the agent's balance between exploring unknown states and utilizing known strategies. When the threshold is too strict, the agent may be unable to explore the environment due to premature termination and unable to learn complex behaviors. When the threshold is too loose, the agent will linger in the invalid area for a long time, reducing learning efficiency. The termination condition designed by the present invention is shown in Equation (41).
[0199]
[0200] When the drone ends the current training early due to unexpected circumstances, it will be punished:
[0201]
[0202] In summary, the reward function of the present invention is:
[0203]
[0204] The desired position and yaw angle of the UAV are set, and a simulation experiment is conducted through Simulink to evaluate its effect on suppressing the swing of the suspended load.
[0205] To verify the tracking performance and payload sway suppression of the proposed algorithm (DDPG) for suspended drones, a comparative analysis was conducted on the tracking accuracy and payload sway suppression capabilities of the algorithm before modification (no control), with the shaper alone (shaper), and with the addition of a backstepping sliding mode angle controller (backstepping sliding mode) in spiral trajectory tracking experiments. Table 2 shows the model parameters.
[0206] Table 2 System model parameters
[0207]
[0208] The hyperparameter settings of the DDPG algorithm are shown in Table 3.
[0209] Table 3 DDPG algorithm hyperparameters
[0210]
[0211] The simulation time is 100s, the initial position of the quadrotor is (0, 0, 0)m, the yaw angle is set to 0.2rad, and the reference trajectory is:
[0212]
[0213] The signal obtained after the desired trajectory (desired position and desired yaw angle) of the drone passes through the ZVD input shaper is used as the input of the fuzzy PID controller to drive the drone to move along the reference spiral trajectory for 100 seconds to obtain data, such as Figure 10 From the trajectory tracking structure diagram, it can be seen that the controller proposed in this invention has better performance in trajectory tracking in terms of overshoot and steady-state error than other algorithms. It can closely follow the reference trajectory in all dimensions and stably track the reference trajectory after 2s. The maximum deviation after stabilization is less than 0.05m. Compared with using only the shaper and backstepping sliding mode, the controller proposed in this article does not cause trajectory tracking delay, and has smaller overshoot and steady-state error.
[0214] like Figure 11 The attitude angle can be seen in the controller proposed by the present invention in controlling the attitude angle data. Figure 11 It can be seen that in the early stages of system operation, the UAV's attitude fluctuates significantly to quickly suppress the payload's swing angle. After about 20 seconds, the UAV's attitude stabilizes. Table 4 provides quantitative data on the attitude angle control performance, which demonstrates that the UAV has good stability during routine flight missions.
[0215] Table 4 Quantitative evaluation of attitude angle control performance
[0216]
[0217] from Figure 12 It can be seen that the control algorithm proposed in this paper brings the curve closer to the zero axis, providing stronger suppression of steady-state errors. Backstepping sliding mode control exhibits periodic small oscillations due to the inherent chattering phenomenon of sliding mode control. While the control algorithm proposed in this paper shows increased fluctuations in all curves, the proposed algorithm still maintains fast and efficient suppression capabilities. The proposed controller exhibits superior performance in suppressing load swing angles. Figure 13This is the trajectory of the UAV suspension system under the proposed controller. The swing amplitude of the swing angle α is consistently limited to ±2°, while other methods fluctuate by up to ±5°. The swing amplitude of the swing angle β is consistently limited to ±5°, while other methods fluctuate by up to ±8°. DDPG learns the optimal control strategy through interaction with the environment. Combined with the timing optimization of the input signal by the ZVD (input shaper), it can actively compensate for the phase lag and energy accumulation of the load swing, thereby suppressing oscillations. Experimental results verify the effectiveness of the DDPG-ZVD controller in promoting load swing attenuation.
[0218] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A DDPG-based flight control method for quadrotor UAV suspension swing suppression, characterized in that: The steps are: Step 1: Use the Newton-Euler method to establish the dynamic model of the quadrotor UAV suspension system; Step 2: Use the ZVD input shaper to perform input shaping on the desired reference trajectory signal to obtain a shaped trajectory signal; Step 3: A cascade trajectory tracking control strategy is used to decouple the quadrotor UAV suspension system into position control of the position loop and attitude control of the attitude loop. A fuzzy PID controller for the position loop is designed based on fuzzy theory. The shaped trajectory signal is input into the fuzzy PID controller to obtain the PID controller parameter correction value, and the PID controller parameters are adjusted online. Step 4: Perform attitude settlement on the virtual control quantity output by the fuzzy PID controller and input it into the active disturbance rejection controller to obtain the virtual control quantity of the attitude loop; Step 5: Design a DDPG-based swing angle controller to generate roll, pitch, and yaw rotation moment control variables based on the virtual control variables of the attitude loop and the control variables generated by the DDPG-based swing angle controller.
2. The DDPG-based quadrotor UAV suspension swing suppression flight control method according to claim 1 is characterized in that: The dynamic model of the quadrotor UAV suspension system is: Where u1 represents the total thrust control input, u2, u3, and u4 represent the moment control inputs for controlling the pitch, roll, and yaw of the drone, M and m represent the masses of the drone and the payload, L is the length of the rope, l represents the arm length of the quadrotor drone, φ, θ, and ψ represent the roll angle, pitch angle, and yaw angle of the drone relative to the inertial coordinate system {W}, respectively. denote the angular velocities of roll, pitch and yaw respectively, are the angular accelerations of roll, pitch and yaw, I x , I y , I z is the moment of inertia of the quadrotor; g is the acceleration due to gravity, is the acceleration of the drone, are the angular velocities of the load swing angles α and β, respectively, They are the load swing angle and the load swing angle, respectively; the load swing angle α represents the rope in the body coordinate system {X B , Y B , Z B Plane Y B O B Z B The projection on Z B The angle between the axes and the load swing angle β represents the angle between the rope and the plane X. B O B Z B The projection on Z B The angle between the axes; a1, a2, a3, b1, b2, b3, c1, and c2 are all intermediate variables.
3. The DDPG-based quadrotor UAV suspension swing suppression flight control method according to claim 1 is characterized in that: The method for constructing the dynamic model of the quadrotor UAV suspension system is as follows: the total thrust F in the body coordinate system {B} is converted to the inertial coordinate system {W} by the rotation matrix R from the body coordinate system {B} to the inertial coordinate system {W}. T Decomposed into the force components [F x ,F y ,F z ] T ; According to the load swing angles α and β, the positional relationship between the UAV and the load is determined, and the second-order derivative of the UAV position (x, y, z) is obtained to obtain the relationship between the acceleration of the load and the acceleration of the UAV; according to Newton's second law, the component force [F x ,F y ,F z ] T The relationship between the acceleration of the load and the acceleration of the UAV is obtained by equation I of the acceleration of the UAV; Similarly, the load is subjected to the rope tension T in the x, y, and z axis directions in the body coordinate system {B} [T x ,T y ,T z ] T ; According to the rotation matrix R from the load coordinate system {l} to the inertial coordinate system {W} l The suspended load is subjected to the total thrust F of the UAV T Transform from the load reference system {l} to the inertial coordinate system {W} and obtain the total force F acting on the suspended load in the body coordinate system {B} l ; The Newton-Lagrange method is used to analyze the forces acting on the drone and the load, and the equation II for the relationship between the acceleration of the load and the acceleration of the drone is obtained. The rope tension T is calculated based on the relationship between the acceleration of the load and the acceleration of the drone. Without considering the resistance and ignoring the torque generated by the load, the angular acceleration of the UAV's roll angle, pitch angle and yaw angle are obtained. According to the UAV acceleration equation I, the UAV acceleration equation II and the rope tension T, the component force [F x ,F y ,F z ] T Determine the acceleration of the drone The intermediate variable Intermediate variables The high-level control instructions are mapped to the speeds of the four motors through control distribution: Among them, c t is the lift coefficient, c m is the counter-torque coefficient, ω1, ω2, ω3, and ω4 are the speeds of the four motors.
4. The DDPG-based quadrotor UAV suspension swing suppression flight control method according to any one of claims 1 to 3, characterized in that: The ZVD input shaper performs input shaping processing on the input reference trajectory signal using three pulse signals. The constraint equation is: Among them, A1, A2, and A3 represent the amplitudes of the three pulse signals input to the ZVD shaper, t1 represents the initial time; T(ξ,ω n ) is the residual oscillation percentage, ω n is the natural frequency of the system, ξ is the damping ratio of the system; According to the constraint equation, the amplitude and time delay t of the ZVD input shaper can be obtained. i respectively: Where k is the damping correction factor, τ d is the pulse time point.
5. The DDPG-based quadrotor UAV suspension swing suppression flight control method according to claim 4 is characterized in that: Reference trajectory signal (x d ,y d ,z d ) through the three pulse trains of ZVD input shaper and ZVD input shaper Perform convolution to obtain the shaped trajectory signal which is input into the fuzzy PID controller; The residual oscillation percentage T(ξ,ω n ) is the ratio of the amplitude of the system response after input shaping to the amplitude of the total system response without input shaping, and ω d is the damped oscillation frequency of the system, 6. The DDPG-based flight control method for suppressing hanging swing of a quadrotor UAV according to any one of claims 1, 2, 3, and 5, characterized in that: The fuzzy PID controller has a two-input and three-output structure. The input of the fuzzy PID controller is the trajectory tracking error e and the corresponding change rate ec, and the output is the PID controller parameter correction value ΔK p , ΔK i , ΔK d The trajectory tracking error e is the difference between the shaped trajectory signal r(t) and the current state. The trajectory tracking error e and the rate of change ec are fuzzified, fuzzy inference and defuzzified, and the output is the parameter correction ΔK p , ΔK i , ΔK d , input into the PID controller, and adjust the control parameter K of the PID controller in real time p , K i , K d , ensuring that the suspension system accurately tracks the reference trajectory while effectively suppressing the load swing angle; The fuzzy control rules of the fuzzy PID controller adopt the standardized design created by Mamdani, and the output fuzzy control rule table is established as follows: The membership function of the fuzzy subset of the trajectory tracking error e adopts trigonometric function, and the seven-level fuzzy subset {NB, NM, NS, ZO, PS, PM, PB} is selected, where N represents negative, P represents positive, B represents large, M represents medium, S represents small, and ZO represents zero; The centroid method is used to obtain the precise value of the parameter correction.
7. The DDPG-based quadrotor UAV suspension swing suppression flight control method according to claim 6, characterized in that: The fuzzy PID controller is based on the trajectory tracking error e of the reference trajectory and the current position. p =[e x e y e z ] T Get the virtual control quantity u of the position loop p =[u x u y u z ] T , and the virtual control quantity is: is the trajectory tracking error e p The derivative of The attitude solver calculates the desired attitude angle φ of the drone d and θ d And the thrust control quantity u1 of the UAV is: Among them, ψ d is the reference yaw angle.
8. The DDPG-based quadrotor UAV suspension swing suppression flight control method according to claim 7, characterized in that: The attitude loop's active disturbance rejection controller is designed based on the dynamic model of the quadrotor UAV's suspension system. The active disturbance rejection controller obtains the virtual control quantity of the attitude loop based on the error between the current attitude angle and the desired attitude angle after attitude settlement. The active disturbance rejection controller is divided into three parts: a tracking differentiator, an extended state observer, and a nonlinear feedback error. All uncertainties affecting the UAV are regarded as a lumped disturbance. The extended state observer estimates and compensates for the lumped disturbance in real time, and approximately transforms the uncertain nonlinear system into an easily controllable linear integral series system. The tracking differentiator is the desired attitude angle φ d and θ d Arrange a smooth transition process without overshoot and extract the differential signal; The nonlinear feedback error uses the transient signal generated by the tracking differentiator and the state estimated by the extended state observer and the lumped disturbance to dynamically compensate the control quantity and generate the torque control quantity (u φ ,u θ ,u ψ ).
9. The DDPG-based quadrotor UAV suspension swing suppression flight control method according to claim 8, characterized in that: According to the virtual control amount of the attitude loop (u φ ,u θ ,u ψ ) and the control variables D1, D2, and D3 generated by the DDPG-based angle controller generate the moment control inputs for roll, pitch, and yaw: The state space of the DDPG-based swing angle inhibitor is S(t)=[x,y,z,e x ,e y ,e z ,φ,θ,ψ,e φ ,e θ ,e ψ ,α,β], (x,y,z) is the position of the UAV, (φ,θ,ψ) is the attitude angle of the UAV, (e x ,e y ,e z ) is the position error of the UAV, (e φ ,e θ ,e ψ ) is the attitude angle error of the UAV, (α, β) is the load swing angle; Control action a t =[D1,D2,D3] T These are the three torque control signals of the quadrotor, and the action value range of the control quantities D1, D2, and D3 is [-20, 20].
10. The DDPG-based quadrotor UAV suspension swing suppression flight control method according to claim 9, characterized in that: The DDPG-based swing angle suppressor adopts a deterministic strategy. The Actor current network directly outputs the specific action value rather than the probability distribution of the action; the Critic current network evaluates the Q value of the state-action pair Q = (s t ,a t |θ Q ), guiding the optimization strategy of the Actor's current network; Actor target network generates action a for the next state t+1 , used to calculate the target Q value; the critic target network is used to calculate the target Q value and maintain training stability; the target network uses a soft update method to stably approach the current network: i Q′ ←tth Q +(1-τ)θ Q′ ,i μ′ ←tth μ +(1-τ)θ μ′ ; Among them, τ<<1; θ μ is the parameter of the Actor's current network, θ Q is the parameter of the current Critic network, θ μ' is the parameter of the Actor target network, θ Q' is the parameter of the Critic target network; s t Indicates the state at the current time t; Randomly initialize the Actor current network and the Critic current network, copy the parameters to the target network; perform action a t , and store the transferred samples in the experience replay pool, and calculate the target Q value of the Critic target network: y t = r + γQ'(s', μ'(s')); where s' and μ'(s) are the next state and target action, respectively. Represents expectation; the optimization goal of the Critic target network is to minimize the loss function The Actor target network optimizes the strategy by maximizing the Q value of the Critic target network. The objective function is: N represents the number of samples; Update the actor's current network parameters θ by applying the chain rule with respect to the actor's parameters μ : Among them, θ μ is the weight of the function approximator of the Actor's current network, ▽ a Respectively represent the Q value of μ θ , a finds the gradient; The input of the Actor network is the state space S(t) of the system. Two fully connected layers are used to extract high-order state features. The ReLU activation function is used to enhance the nonlinear expression capability of the hidden layer. The output layer uses the Tanh activation function to normalize the output to [-1,1]. A scaling layer is used to scale the actual action range to ensure consistency with the upper and lower limits of the action space. The output action μ(s t |θ μ ); Perform L2 regularization on the Actor network; The critic network has a dual input path of state + action followed by a fusion layer, ultimately outputting a Q value. The state path inputs the state vector S(t), compressing the high-dimensional state information into low-dimensional features. Two fully connected layers are used to extract high-order state features. A normalization layer is added between the two fully connected layers to accelerate convergence. The dynamic path inputs the action vector to extract action features. The dual fully connected layers are designed to enhance nonlinear expression capabilities. The state features and action features are merged through the splicing layer to output the Q value. The critic network estimates the Q value of the state-action pair by minimizing the temporal difference error, and the TD error δ t The calculation of depends on the reward value, and Q target (s t+1 ,a t+1 )=r t +γQ target (s t ,a t ); δ t =Q target (s t ,a t )-Q(s t ,a t ); Among them, r t is the current reward, Q target Indicates the target Q value, Q(s t ,a t ) represents the Q-value function of the current state-action pair, and γ is the discount factor; The reward function is Among them, when the drone ends the current training early due to unexpected circumstances, it will be punished: The termination condition is The reward function for trajectory tracking and payload swing angle is: Where, Ω∈[x, y, z, α, β], ε is a constant; When the attitude angle error is always less than a certain value, a reward is given, and when it exceeds a certain value, a penalty is given. The reward function is:
Citation Information
Cited By
Flight control and load swing suppression method of four-rotor unmanned aerial vehicle hanging load system
CN116449867A
A flight control and load swing suppression method for a four-rotor unmanned aerial vehicle hanging load system
CN116449867B
Small robot dynamic balance control system based on DDPG
CN121179442A
Unmanned aerial vehicle sling delivery dynamic tracking control system and method
CN121277214A
Control system for power transmission line monitoring device transporting and loading unmanned aerial vehicle
CN121634984A