Four-rotor unmanned aerial vehicle attitude preset performance control method based on reinforcement learning

Through the reinforcement learning-based channel dynamic decoupling architecture and multi-objective reward function, the attitude control problem of the quadrotor drone under time-varying inertial parameters and external interference is solved, high-precision and stable tracking of the drone's attitude is achieved, and the robustness and control effect of the system are improved.

CN120821291AActive Publication Date: 2025-10-21NANJING UNIV OF INFORMATION SCI & TECH

Patent Information

Application Number
CN202511301110.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-10-21
Estimated Expiration
2045-09-12

AI Technical Summary

Technical Problem

Quadrotors have difficulty achieving high-precision attitude control in complex dynamic environments, especially under time-varying inertial parameters and external interference. Traditional methods have the problem of difficulty in balancing robustness and dynamic performance, and reinforcement learning methods have difficulty balancing multi-objective conflicts, resulting in fluctuations in control variables and system instability.

Method used

A reinforcement learning-based channel dynamic decoupling architecture is adopted. By building a deep reinforcement learning parameter generator and designing a multi-objective reward function, preset performance control of the UAV attitude is achieved. The Actor network is used to generate adaptive control parameters. Combined with non-singular fast terminal sliding mode control technology, the performance of each channel is independently optimized to improve anti-interference robustness and parameter stability.

Benefits of technology

It achieves smooth tracking of the UAV's attitude in a time-varying environment, reduces the interference amplification between channels, improves the controller's adaptability and system stability, and meets the needs of high-precision attitude tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120821291A_ABST
    Figure CN120821291A_ABST
Patent Text Reader

Abstract

The invention discloses a reinforcement learning-based quadrotor unmanned aerial vehicle attitude preset performance control method, which comprises the following steps of: establishing an unmanned aerial vehicle attitude kinetic equation by considering a time-varying inertial parameter and external unknown interference; a sub-channel preset performance controller is constructed, and a robust control quantity is generated in combination with preset performance and a non-singular fast terminal sliding mode control technology; designing a reinforcement learning parameter generator, and dynamically optimizing 12 time-varying parameter estimated values through a time sequence feature extraction network and a residual network; and establishing an online reinforcement learning training mechanism, constructing a multi-target reward function, and optimizing the output of the parameter generator through the multi-target reward function. According to the method, a reinforcement learning method is adopted to replace a traditional self-adaptive method to estimate time-varying parameters, multi-degree-of-freedom decoupling optimization is achieved through a sub-channel control architecture, tracking error preset performance constraint is achieved under the working condition of time-varying inertial parameters, and the dynamic adjustment capacity and anti-interference robustness of an attitude system are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for controlling the attitude preset performance of a quad-rotor unmanned aerial vehicle (UAV) based on reinforcement learning, and belongs to the technical field of UAV control. Background Art

[0002] In recent years, high-precision attitude control of quadrotor drones in complex dynamic environments has faced significant challenges. Traditional preset performance control methods, such as adaptive control based on sliding mode variable structures, can achieve limited-precision tracking. However, their fixed-parameter strategies present significant limitations when faced with time-varying inertial parameters, such as sudden changes in the inertia matrix caused by load changes, and complex disturbances such as aerodynamic disturbances and sensor noise coupling. Existing literature suggests that model-based adaptive control methods rely heavily on precise dynamic models and struggle to cope with practical scenarios such as time-varying inertia matrices and unknown upper bounds on external disturbances. Furthermore, traditional methods rely on expert experience to tune control parameters, and once fixed, they cannot be adaptively adjusted online. This makes it difficult to balance control robustness and dynamic performance. This is particularly true in attitude control problems involving multiple degrees of freedom coupling and strong nonlinear characteristics, which can easily lead to inter-channel interference amplification.

[0003] Deep reinforcement learning technology offers a new approach to UAV attitude control. However, existing methods still face numerous technical bottlenecks. Most reinforcement learning schemes employ a single-objective reward function, focusing solely on tracking error penalties while neglecting the comprehensive optimization of control energy consumption, parameter stability, and transient performance (such as overshoot and convergence rate). This makes it difficult for the algorithm to balance multiple conflicting objectives during training, and in practical applications, it is prone to significant fluctuations in the control variable or actuator saturation. Regarding network architecture design, traditional fully connected networks have limited feature extraction capabilities and cannot effectively handle time-varying parameter estimation and multi-channel coupling interference, resulting in large standard deviations of steady-state errors and difficulty meeting high-precision control requirements. Furthermore, the fixed exploration rate and static experience replay strategy used in training mechanisms are prone to slow convergence and local optima. This is especially true when subjected to sudden torque disturbances or parameter mutations, which significantly increase the fluctuation of the control variable and threaten system stability. Therefore, to avoid attitude tracking failures in quadrotors caused by time-varying inertial parameters and external disturbances, and to comprehensively consider multiple performance indicators during attitude tracking, it is necessary to design a control method that enables stable attitude tracking for quadrotors. Summary of the Invention

[0004] The technical problem to be solved by the present invention is: to provide a preset performance control method for the attitude of a quadrotor unmanned aerial vehicle based on reinforcement learning, to solve the attitude control problem under the coupling of time-varying inertial parameters and external unknown interference, to realize independent optimization of multiple degrees of freedom through a channel-based dynamic decoupling architecture, and to generate adaptive control parameters in real time based on reinforcement learning, to improve the anti-interference robustness and parameter stability while meeting the preset performance constraints.

[0005] The present invention adopts the following technical solutions to solve the above technical problems: A method for controlling the attitude preset performance of a quadrotor drone based on reinforcement learning comprises the following steps: Step 1: Considering time-varying parameters and external unknown interference, a mathematical model of the quadrotor drone is established, and state variables are defined to establish the control-oriented state equation of the quadrotor drone. Step 2: Select a time-varying preset performance boundary function, introduce an error transformation function, and design a channel-by-channel preset performance attitude controller based on the quadrotor UAV state equation. The error transformation function is used to ensure that the tracking error of the UAV attitude system meets the preset performance constraints. Step 3: Build a deep reinforcement learning parameter generator, namely the Actor network. Determine the input and output of the Actor network based on the channel-by-channel preset performance attitude controller. The input is the state vector of the UAV, and the output is the time-varying parameter adjustment. Step 4: Train the Actor network and the Critic network corresponding to the Actor network based on the online reinforcement learning training mechanism of deep deterministic policy gradient, define a multi-objective reward function, and optimize the output of the Actor network through the multi-objective reward function; Step 5: Use the Actor network to output the time-varying parameter adjustment at the current moment, fuse it with the time-varying parameters at the previous moment through the smooth update mechanism, obtain the time-varying parameters at the current moment, substitute them into the channel preset performance attitude controller, obtain the control input of each channel of the drone at the current moment, and realize the preset performance control of the drone attitude.

[0006] Compared with the prior art, the present invention adopts the above technical solution and has the following technical effects: 1. The present invention constructs a quadrotor UAV attitude model by considering time-varying inertial parameters and external unknown interference. By decoupling state variables, a control-oriented channel model is constructed, which improves the controller's adaptability to time-varying environments.

[0007] 2. The present invention adopts channel preset performance and non-singular fast mid-terminal sliding mode control technology to design the attitude controller, so that the channels can independently meet the preset performance constraints, avoid the problem of inter-channel interference amplification in traditional coupling control, and improve the attitude tracking stability.

[0008] 3. The present invention innovatively designs a reinforcement learning parameter generator, which replaces the traditional adaptive parameter calculation method and improves the control effect of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 This is a flow chart of a method for controlling the attitude preset performance of a quadrotor drone based on reinforcement learning according to the present invention; Figure 2 is a network structure diagram of the reinforcement learning parameter generator of the present invention; Figure 3 The rolling angle of the quadrotor UAV of the present invention is compared with the tracking effect of the traditional method; Figure 4 The rolling angle of the quadrotor UAV of the present invention is compared with the tracking error of the traditional method; Figure 5 The pitch angle of the quadrotor UAV of the present invention is compared with the tracking effect of the traditional method; Figure 6 The pitch angle of the quadrotor UAV of the present invention is compared with the tracking error of the traditional method; Figure 7 The yaw angle of the quadrotor UAV of the present invention is compared with the tracking effect of the traditional method; Figure 8 The yaw angle of the quadrotor UAV of the present invention is compared with the tracking error of the traditional method. DETAILED DESCRIPTION

[0010] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be interpreted as limiting the present invention.

[0011] like Figure 1 As shown, this paper proposes a reinforcement learning-based method for controlling the attitude of a quadrotor drone with preset performance. This method uses a channel-by-channel independent control architecture. Each attitude channel (roll, pitch, and yaw) generates a dynamic error bound based on preset performance constraints. A reinforcement learning parameter generator is used to optimize controller parameters, thereby achieving the attitude tracking task of the quadrotor drone. The specific steps are as follows: Step 1: Considering time-varying parameters and external unknown interference, establish a mathematical model of the quadrotor drone, define state variables, and establish the state equation for the quadrotor drone control. The specific process is as follows: S101. For a quadrotor drone with time-varying parameters, its mathematical model is described as: , in, Represents the attitude angle of the UAV in the inertial system, 、 、 is the inertia matrix of the system, is the air resistance coefficient, is the external disturbance acting on each channel, , Represents the control torque for the roll, pitch, and yaw motion of the drone.

[0012] S102. Establish a control-oriented design model Select state variables , the control-oriented design model is as follows: , in, , , , , , , , , .

[0013] Step 2: Select the preset performance function, design the channel-specific non-singular fast terminal sliding mode function and the preset performance attitude controller, and independently calculate the control variable for each channel subsystem. The error transformation function is used to ensure that the tracking error meets the preset performance constraints. The specific process is as follows: S201. Select a preset performance function and introduce error transformation Use the performance function to quantitatively describe the expected performance indicators, and then design the controller to ensure that the full-time performance constraints are met. Select a smooth decreasing positive function Preset the performance boundary function for time-varying performance, considering the following performance constraints: , in, 、 、 is a pre-given positive constant, Indicates the preset lower limit of steady-state error, represents the upper limit of steady-state error, Limit the minimum convergence speed of the system, Indicates time. , is the overshoot suppression parameter, satisfying .

[0014] To handle time-varying performance constraints, the tracking error of the UAV attitude system with time-varying parameters is defined as , , ,in, 、 、 They are 、 、 The reference trajectory of . , In order to normalize the tracking error, the error transformation function is introduced: , Error transformation function The derivative is: , in, .

[0015] S202. Design of a three-channel non-singular fast terminal sliding mode function and a three-channel attitude control law for a quadrotor UAV The following takes the roll angle subsystem as an example to illustrate the control design process. The non-singular fast terminal sliding mode function of the roll angle subsystem is designed as follows: , in, 、 are all constants greater than zero, and are all positive odd numbers and satisfy , ,right Taking the derivative, we can get: , Substituting the control-oriented design model into the equation: , To handle time-varying parameters 、 、 ,Will 、 、 Separate the constant terms 、 、 and time-varying error terms 、 、 ,Right now , , , and use the reinforcement learning method to estimate the above constant term, and design a sliding mode controller to suppress the above time-varying error term.

[0016] The design of the roll angle subsystem controller is as follows: , , .

[0017] The non-singular fast terminal sliding mode functions of the pitch angle and yaw angle subsystems are as follows: , , in, 、 、 and All parameters are greater than zero.

[0018] The design of the pitch angle subsystem controller is as follows: , , ; The design of the yaw angle subsystem controller is as follows: , , , in, 、 、 They are The estimated value of , , is the control gain to be designed, 、 、 They are 、 、 The estimated value of 、 、 They are 、 、 The estimated value of 、 、 They are 、 、 The estimated value of yes The separated constant term, is a positive time-varying integrable function that satisfies: , , and They are The upper and lower bounds of express: .

[0019] Step 3: Build a deep reinforcement learning parameter generator, design a 27-dimensional input state vector, and dynamically optimize 12 time-varying parameter estimates through a temporal feature extraction network and a residual network.

[0020] The deep reinforcement learning parameter generator is the policy network (Actor network) structure as follows Figure 2As shown, the input of the network is a 27-dimensional state vector, which is composed of the normalized real-time attitude angle , angular velocity , three-channel tracking error and its differential , historical error moving average integral , and the current parameter estimates constitute.

[0021] The normalized 27-dimensional state vector is then normalized by a learnable layer normalization module. The layer normalization module independently calculates the mean and variance for each feature dimension of the input data and achieves normalization through affine transformation. The normalized data is input into the temporal feature extraction network, which consists of a fully connected layer (input dimension 27, output dimension 64), a GELU activation function, a layer normalization layer, a second fully connected layer (input and output dimensions are both 64), a GELU activation function, and a layer normalization layer connected in series. The output of the temporal feature extraction network is connected to a residual module consisting of three residual blocks. Each residual block contains a fully connected layer (input dimension 64, output dimension 128), a GELU activation function, a layer normalization layer, a fully connected layer (input dimension 128, output dimension 64), a GELU activation function, and a layer normalization layer. The output of the residual module is mapped by a fully connected layer (input dimension 64, output dimension 12) to generate an action vector, i.e., a 12-dimensional parameter adjustment quantity. , parameter adjustment amount The new parameter estimates are integrated into the traditional controller's parameter estimates through a smooth update mechanism. The specific formula is: , Among them, the smoothing coefficient Set independently by control channel.

[0022] Step 4: Establish an online reinforcement learning training mechanism based on DDPG (Deep Deterministic Policy Gradient) and define a multi-objective reward function.

[0023] The online reinforcement learning training mechanism is based on the Markov decision process, whose state space is a 27-dimensional enhanced state vector , which is the input of the deep reinforcement learning parameter generator. The action space is the 12-dimensional parameter adjustment of the deep reinforcement learning parameter generator The parameter generator network output is optimized through a multi-objective reward function. The multi-objective reward function consists of a tracking error penalty term, an error derivative penalty term, a control energy penalty term, and a control smoothness penalty term. Its mathematical expression is: , Among them, the tracking error , error derivative and control quantity change rate All are processed by limiting. , ,in, is the squared error penalty weight, is the penalty term weight of the absolute value of the error derivative, To control the input penalty weight, is the penalty weight of the control quantity change law. Represents the sum of all rewards.

[0024] Step 5: Optimize network parameters by combining experience replay, construct the target Q-value function and loss function, and establish the Actor and Critic network update process.

[0025] The experience replay buffer stores the latest 5000 transfer data ( ), during training, 512 data are randomly sampled to batch update the network parameters. The initial 511 steps are the warm-up phase, which only collects data without performing network updates. After the buffer reaches the minimum batch size (512), the network starts to perform parameter updates. The system randomly samples batches of data from the buffer. , the state data is input into the parameter generation network, or policy network (Actor network), for calculation. The 12-dimensional vector output by the Actor network represents the parameter adjustment. The Critic network evaluates the state-action pair (where the action is the parameter adjustment output by the Actor network) and outputs a Q-value representing the value of the current state-action pair. The target Q-value is constructed as follows: First, the Actor target network generates a target action based on the next state. Then, the next state and target action are input into the Critic target network to obtain the target Q-value. Its mathematical expression is: , in, is the discount factor, which highlights the importance of recent rewards. The critic target network evaluates the value of the next state-target action pair. The action strategy generated by the Actor target network according to the next state, Indicates the next state the environment transitions to after the current action is executed.

[0026] For the critic network, the update of its neural network parameters adopts the time difference (TD) error method, which is implemented by the mean square error. Its loss function is defined as follows: , in, , Represents the parameters of the Critic network, It is the Q-value estimate of the current state-action pair by the Critic network.

[0027] The gradient descent method is used to solve the gradient of the loss function, and then the critic network parameters are updated according to the calculation results: , , in, Represents the loss function gradient of the Critic network, is the updated critic network parameter, is the learning rate of the Critic network.

[0028] Similarly, the update formula of the Actor network is as follows: , , in, Represents the parameters of the Actor network, represents the objective function gradient of the Actor network, is the policy function about The gradient, Represents the Q-value function of the Critic network About Action The gradient of Calculated under the conditions, To update the Actor network parameters, is the learning rate of the Actor network.

[0029] The Actor Target Network and the Actor Network have exactly the same network structure, but the Actor Network is a real-time updated network. The Actor Target Network does not update in real time, but is periodically "soft-updated" from the Actor Main Network. By choosing an appropriate soft update rate, the Actor Target Network and the Critic Target Network can be soft-updated by their respective Actor and Critic Networks: , in, is the soft update rate, is the parameter of the Critic target network, Parameters for the Actor target network.

[0030] In order to evaluate the effectiveness and innovation of the designed control strategy, a simulation verification was carried out based on the embodiment. The parameters of the quadrotor UAV system are set to , , initial posture , , , given a reference trajectory , , , air resistance coefficient , the external disturbance is set to , , The control objective of the present invention is to design the control input so that the output signal 、 、 Asymptotic tracking reference instructions 、 、 , and the transient convergence speed is not less than .

[0031] In order to meet the above preset performance requirements, the preset performance function is selected as: , the controller parameters are selected as: , , , , , , controller gain , a time-varying integrable function: , , , , , .

[0032] Figure 3-8 This is the simulation result of the system using the three-channel attitude control law established by the present invention. Figure 3 、 Figure 4 、 Figure 5 It shows that the control strategy designed by the present invention can track the control target faster than the traditional adaptive law calculation method, reduce overshoot, and have a smoother tracking effect. Figure 6 、 Figure 7 、 Figure 8 The results show that the control strategy designed by the present invention reduces the error and improves the system performance compared with the traditional adaptive law calculation method.

[0033] Based on the same inventive concept, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, the steps of the aforementioned reinforcement learning-based quadcopter attitude preset performance control method are implemented.

[0034] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of the aforementioned reinforcement learning-based quadrotor drone attitude preset performance control method.

[0035] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0036] The present invention is described with reference to flowcharts of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process in the flowcharts, as well as combinations of processes in the flowcharts, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts. Figure 1 A device that specifies functions in a process or multiple processes.

[0037] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A function specified in a process or multiple processes.

[0038] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 The steps of a specified function in a process or multiple processes.

[0039] The above embodiments are only for illustrating the technical idea of ​​the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution in accordance with the technical idea proposed by the present invention shall fall within the protection scope of the present invention.

Claims

1. A method for controlling the attitude preset performance of a quadrotor drone based on reinforcement learning, characterized in that: The steps include: Step 1: Considering time-varying parameters and external unknown interference, a mathematical model of the quadrotor drone is established, and state variables are defined to establish the control-oriented state equation of the quadrotor drone. Step 2: Select a time-varying preset performance boundary function, introduce an error transformation function, and design a channel-by-channel preset performance attitude controller based on the quadrotor UAV state equation. The error transformation function is used to ensure that the tracking error of the UAV attitude system meets the preset performance constraints. Step 3: Build a deep reinforcement learning parameter generator, namely the Actor network. Determine the input and output of the Actor network based on the channel-by-channel preset performance attitude controller. The input is the state vector of the UAV, and the output is the time-varying parameter adjustment. Step 4: Train the Actor network and the Critic network corresponding to the Actor network based on the online reinforcement learning training mechanism of deep deterministic policy gradient, define a multi-objective reward function, and optimize the output of the Actor network through the multi-objective reward function; Step 5: Use the Actor network to output the time-varying parameter adjustment at the current moment, fuse it with the time-varying parameters at the previous moment through the smooth update mechanism, obtain the time-varying parameters at the current moment, substitute them into the channel preset performance attitude controller, obtain the control input of each channel of the drone at the current moment, and realize the preset performance control of the drone attitude.

2. The method for controlling the attitude preset performance of a quadrotor drone based on reinforcement learning according to claim 1, characterized in that: The specific process of step 1 is as follows: Considering time-varying parameters and external unknown interference, the mathematical model of the quadrotor drone is described as: , in, and Respectively represent the roll, pitch and yaw attitude angles of the UAV in the inertial system, and are the air resistance coefficients of the roll, pitch and yaw channels respectively, 、 and are the inertia matrices of the UAV attitude system. and They represent the control torques for the roll, pitch and yaw motions of the UAV, and are the external disturbances acting on the roll, pitch and yaw channels respectively; Defining state variables , then the state equation of the quadrotor drone for control is designed as follows: , in, , , , , , , , , , and They represent the control inputs of the roll, pitch, and yaw channels of the quadrotor drone, respectively.

3. The method for controlling the attitude preset performance of a quadrotor drone based on reinforcement learning according to claim 2, characterized in that: The specific process of step 2 is as follows: Step 2.1, select a smoothly decreasing positive function Preset performance boundary function for time variation, 、 Respectively represent the preset lower and upper limits of the steady-state error, is the base of natural logarithms, is a preset positive constant, Indicates time, represent the roll, pitch and yaw channels respectively; Tracking error of UAV attitude system The following performance constraints are met: , in, and are all overshoot suppression parameters, , ,satisfy ; Introducing error transformation function as follows: , in, is the normalized tracking error, ; Step 2.2, design a three-channel non-singular fast terminal sliding mode function : , in, 、 、 、 、 and are all parameters greater than zero, and are all positive odd numbers, is a positive real number that satisfies , ; Step 2.3: Establish the three-channel attitude control law of the quadrotor drone based on the sliding mode function: , in, for The estimated value of for The separated constant term, , , , , , , , , and are the control gains to be designed, , , , 、 and They are 、 and The estimated value of and They are and The separated constant term, 、 and They are 、 and The estimated value of and They are and The separated constant term, 、 and They are 、 and The estimated value of and They are and The separated constant term, and They are and The reference trajectory of is a positive time-varying integrable function that satisfies: , , is the integration variable, and They are The upper and lower bounds of .

4. The method for controlling the attitude preset performance of a quadrotor drone based on reinforcement learning according to claim 3, characterized in that: In step 3, the deep reinforcement learning parameter generator includes a layer normalization module, a temporal feature extraction network, a residual module and a fifth fully connected layer connected in sequence, wherein the temporal feature extraction network includes a first fully connected layer, a first GELU activation function, a first normalization layer, a second fully connected layer, a second GELU activation function and a second normalization layer connected in sequence; the residual module includes three residual blocks with the same structure and connected in sequence, each residual block includes a third fully connected layer, a third GELU activation function, a third normalization layer, a fourth fully connected layer, a fourth GELU activation function and a fourth normalization layer connected in sequence; The input of the deep reinforcement learning parameter generator is a 27-dimensional state vector, including the normalized real-time posture angle , angular velocity , three-channel tracking error and its differential , historical error moving average integral and parameter estimates ; The 27-dimensional state vector is normalized by the layer normalization module, and then passes through the temporal feature extraction network, the residual module and the fifth fully connected layer in sequence to output the action vector, which is the 12-dimensional time-varying parameter adjustment. .

5. The method for controlling the attitude preset performance of a quadrotor drone based on reinforcement learning according to claim 1, characterized in that: In step 4, the multi-objective reward function is composed of a tracking error penalty term, an error derivative penalty term, a control energy penalty term, and a control smoothness penalty term. The expression is: , in, represents the time step corresponding to the current moment, is the squared error penalty weight, is the penalty term weight of the absolute value of the error derivative, is the weight of the control penalty term, is the penalty term weight for the control variable change rate, is the tracking error, is the tracking error derivative, To control the amount, , is the rate of change of the controlled quantity; Starting from the initial moment, collect the data corresponding to each moment Put it into the experience playback buffer, each moment corresponds to a time step, Represents the time step The corresponding state, action and reward, the state is the state vector input by the Actor network, and the action is the time-varying parameter adjustment output by the Actor network. Represents the time step The corresponding state; the 1st to 511th time steps are the warm-up phase, where only data is collected without updating the Actor network and Critic network parameters; when the experience replay buffer accumulates to the minimum batch size, i.e. 512 data items, the network parameter update begins; Randomly sample batches of data from the experience replay buffer , the current state Input the Actor network for calculation, and the Actor network outputs the corresponding action , the Critic network evaluates the current state-action pair and outputs the Q value of the current state-action pair; the Actor target network evaluates the next state Generate the target action and then the next state The target action is input into the Critic target network to obtain the target Q value; the Actor target network has the same structure as the Actor network, and the Critic target network has the same structure as the Critic network; the target Q value expression is as follows: , in, is the target Q value, is the multi-objective reward function corresponding to the current state-action pair, is the discount factor, The critic target network evaluates the value of the next state-target action pair. The target action generated by the Actor target network according to the next state, Indicates the next state the environment will transition to after the current action is executed; The Critic network parameters are updated through the gradient descent method, and the Actor network parameters are updated using the gradient information of the Critic network.

6. A computer device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: When the processor executes the computer program, the steps of the method for controlling the preset performance of the attitude of a quadrotor drone based on reinforcement learning are implemented.

7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for controlling the preset performance of the attitude of a quadrotor drone based on reinforcement learning are implemented.

Citation Information

Patent Citations

  • Reinforced learning optimization control method for attitude of quad-rotor unmanned aerial vehicle

    CN114859704A

  • Four-rotor unmanned aerial vehicle preset performance tracking control method based on reinforcement learning

    CN116661478A

  • Four-rotor unmanned aerial vehicle optimal robust fault-tolerant control method considering external disturbance

    CN116974203A

  • Four-rotor unmanned aerial vehicle active-disturbance-rejection control method based on DDPG

    CN118795918A

  • Four-rotor unmanned aerial vehicle model prediction control method based on DQN

    CN120215553A

Cited By

  • Sparse reward environment optimization learning identification method and system based on demonstration data enhancement

    CN121157051A

  • Method and system for learning identification in sparse reward environment based on demonstration data augmentation

    CN121157051B

  • Finite time sliding mode attitude control method for quad-rotor unmanned aerial vehicle based on reinforcement learning

    CN121300437A