Four-rotor unmanned aerial vehicle end-to-end attitude control method based on reinforcement learning
By employing the Actor-Critic framework and "base+bias" control method on a quadcopter UAV, combined with a network structure featuring long short-term memory and residual connections, the attitude control challenge of quadcopter UAVs in complex environments was solved, achieving fast convergence and high-precision attitude control.
Patent Information
- Application Number
- CN202510925553.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-05
- Publication Date
- 2025-12-02
AI Technical Summary
Existing quadcopter UAV control methods suffer from low steady-state accuracy, long learning time, and difficulty in convergence when facing complex and ever-changing flight environments. In particular, when using reinforcement learning agents for end-to-end control, the exploration space is huge and it is difficult to achieve precise attitude control.
A reinforcement learning algorithm based on the Actor-Critic framework is adopted, combined with the "base+bias" control method and the "long short-term memory + residual connection" network structure to design a reinforcement learning agent. By constructing a quadcopter drone model and a rotor motor model, the interactive training process of the reinforcement learning agent is optimized, the invalid exploration space is compressed, and the gradient stability and steady-state accuracy are improved.
Precise attitude control of quadcopter UAVs was achieved, shortening training time and improving model convergence speed. Furthermore, oscillation problems were suppressed through long short-term memory networks, enhancing steady-state accuracy and control response capabilities.
Smart Images

Figure CN121050433A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of unmanned aerial vehicle (UAV) control technology, specifically relating to an end-to-end attitude control method for a quadrotor UAV based on reinforcement learning. Background Technology
[0002] Quadrotor drones are widely used in military, aerial photography, and agricultural applications due to their simple structure and flexible operation. However, a quadcopter drone is a highly nonlinear and strongly coupled system, making its control complex. Attitude control is one of the fundamental problems in quadcopter drone control and a prerequisite for position control. In simulations, almost all parameters are controllable, but the actual flight environment of a drone is volatile and subject to disturbances, significantly impacting its control. Therefore, the drone itself needs to possess strong autonomous decision-making capabilities to make appropriate adjustments in different situations.
[0003] Currently, flight controllers for quadcopter drones are mainly divided into linear controllers, model-based nonlinear controllers, and learning-based controllers. Linear controllers, such as PID controllers, are simple in structure and easy to implement, but they have certain limitations when facing complex and ever-changing flight environments and mission requirements. Model-based nonlinear controllers, such as sliding mode control, can better handle nonlinear problems, but they have drawbacks such as chattering or sensitivity to external disturbances. Learning-based controllers have high adaptability, but they are rarely used as attitude controllers due to potential problems such as low steady-state accuracy or long learning time, or they are often combined with traditional algorithms to adjust some parameter values.
[0004] Using reinforcement learning agents for end-to-end control of UAV attitude is a challenging task. Due to its vast exploration space, it is often difficult to converge and requires tens of thousands of training sessions. Furthermore, reinforcement learning-based controllers face certain challenges in terms of steady-state control accuracy. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide an end-to-end attitude control method for quadrotor UAVs based on reinforcement learning, so as to accurately control the flight attitude of the UAV.
[0006] To address the aforementioned technical problems, this invention provides an end-to-end attitude control method for a quadrotor unmanned aerial vehicle based on reinforcement learning, comprising the following process:
[0007] S1. Establish a model of a quadcopter drone, including the airframe model and the rotor motor model;
[0008] S2. Construct a reinforcement learning interactive environment, including reward policies and state space;
[0009] S3. Construct a reinforcement learning agent, including an Actor network and a Critic network. The state space is used as the input to the Actor network and the Critic network, respectively. The Actor network outputs an action prediction, which is a 5-dimensional vector including one Base value and four Bias values. The Critic network outputs a value estimate.
[0010] S4. Using a reinforcement learning algorithm and objective function based on the Actor-Critic framework and policy gradient, the reinforcement learning agent is interactively trained using the "base + bias" method.
[0011] S5. Use the Actor network of trained reinforcement learning agents to control the attitude of the quadcopter drone.
[0012] As an improvement to the reinforcement learning-based end-to-end attitude control method for quadrotor UAVs of the present invention:
[0013] The machine model is:
[0014]
[0015] Wherein, φ, θ, and ψ represent the roll angle, pitch angle, and yaw angle, respectively; These represent the actual rotational speeds of rotor motors 1-4, respectively. This represents the acceleration in the drone's body coordinate system; K represents the angular acceleration in the UAV's body coordinate system; g represents the gravitational acceleration; K represents the acceleration due to gravity. T The lift coefficient is represented by d; the distance between the centers of the two diagonally opposite rotor motors of the quadcopter UAV is represented by J. r The moment of inertia of the rotor and its motor rotor is represented by p, q, and r, respectively, which represent the attitude angular velocities in the UAV body coordinate system; I x I y I z K represents the moments of inertia of the UAV along the x, y, and z axes of the body coordinate system, respectively; Q Indicates the anti-torque coefficient;
[0016] Velocity information of a quadcopter drone in Earth coordinate system Angular velocity information in Earth coordinate system as follows:
[0017]
[0018] in, V represents the rotation matrix of the body coordinate system relative to the Earth coordinate system; W represents the transformation matrix from the angular velocity in the body coordinate system to the angular velocity in the Earth coordinate system; Bω represents the velocity of the UAV in the body coordinate system, and ω represents the attitude angular velocity of the UAV in the body coordinate system.
[0019] As a further improvement to the reinforcement learning-based end-to-end attitude control method for quadrotor UAVs of the present invention:
[0020] The rotor motor model is
[0021]
[0022] in, T l =L / R are both time constants, where R and L represent the total resistance and total inductance in the equivalent circuit of the rotor motor, respectively. C m This represents the torque coefficient of a rotor motor under rated excitation. C e I is the electromotive force coefficient of the motor under rated magnetic flux; d Indicates electromagnetic torque current; I dL represents the load torque current; s represents the complex frequency in the Laplace transform.
[0023] As a further improvement to the reinforcement learning-based end-to-end attitude control method for quadrotor UAVs of the present invention:
[0024] The state space includes, in the Earth coordinate system, the velocity, acceleration, roll angle, pitch angle, yaw angle, roll angle angular velocity, pitch angle angular velocity, yaw angle angular velocity, yaw angle error, rotational speed and duty cycle of the four rotor motors of the quadcopter UAV along the x, y, and z axes. The rotational speed and duty cycle of the four rotor motors are used as motor states, and the rest are used as main physical quantity states.
[0025] As a further improvement to the reinforcement learning-based end-to-end attitude control method for quadrotor UAVs of the present invention:
[0026] The input layer network structures of the Actor network and the Critic network are the same, both including three parallel branches: a motor state preprocessing network, a main physical quantity state preprocessing network, and a residual connection channel. The motor state vector is used as the input of the motor state preprocessing network, the main physical quantity state vector is used as the input of the main physical quantity state preprocessing network, and the state space vector is used as the input of the residual connection channel.
[0027] The backbone networks of the Actor network and the Critic network have the same structure. After the output of the motor state preprocessing network and the output of the main physical quantity state preprocessing network are spliced, they pass through the fully connected layer, the ReLU nonlinear activation layer, the long short-term memory network layer, and the fully connected layer in sequence. Then, they are residually connected to the output of the residual connection channel of the input layer network, and then input to the output layer network through the ReLU nonlinear activation layer.
[0028] The output layer network of the Actor network includes a parallel expected value output network and a standard deviation output network. The expected value output network includes a fully connected layer, a ReLU nonlinear activation layer, and a fully connected layer connected in sequence. The standard deviation output network includes two consecutive fully connected layers. Then, the output of the expected value output network corresponds one-to-one with the output of the standard deviation output network, and outputs a 5-dimensional vector of action prediction through a normal distribution.
[0029] The output layer of the Critic network is a fully connected layer that outputs a value estimate of the current state.
[0030] As a further improvement to the reinforcement learning-based end-to-end attitude control method for quadrotor UAVs of the present invention:
[0031] The "base+bias" method is as follows: four bias values correspond to the speed adjustment control signals of the four rotors, and the duty cycle input to each rotor motor is:
[0032] PWM i =Base+Bias i (12)
[0033] Where i represents the serial number of the four rotor motors.
[0034] As a further improvement to the reinforcement learning-based end-to-end attitude control method for quadrotor UAVs of the present invention:
[0035] The motor state preprocessing network includes a fully connected layer, a long short-term memory layer, and a fully connected layer connected in sequence. The main physical quantity state preprocessing network includes a fully connected layer, a ReLU nonlinear activation layer, a fully connected layer, a ReLU nonlinear activation layer, a fully connected layer, and a ReLU nonlinear activation layer connected in sequence.
[0036] The residual connection channel is a fully connected layer.
[0037] As a further improvement to the reinforcement learning-based end-to-end attitude control method for quadrotor UAVs of the present invention:
[0038] The interactive training process is as follows:
[0039] (1) During each training cycle, the target yaw angle is randomly initialized, the state space vector of the quadcopter is periodically collected, and then input into the reinforcement learning agent; the action prediction output by the Actor network calculates the duty cycle of the four rotor motors using the "base+bias" method, and is used to input the quadcopter model; the Critic network outputs the value estimate based on the current state, calculates the immediate reward of the current state according to the reward strategy, and then stores the immediate reward, the current state space vector, the action prediction, and the value estimate of the current state into the experience replay pool as training data;
[0040] (2) After a predetermined number of training cycles, check whether the accumulated amount of training data in the experience replay pool meets the minimum batch size. If it does, extract data from the experience replay pool and update the network parameters in batches: the Actor network updates according to the objective function, and the Critic network uses the estimated reward and the actual reward for regression training.
[0041] As a further improvement to the reinforcement learning-based end-to-end attitude control method for quadrotor UAVs of the present invention:
[0042] The reward strategy is as follows:
[0043]
[0044] Among them, f reward The immediate reward for the current state is given, and the stopping condition is that the speed of the quadcopter UAV along the X, Y, and Z axes in the Earth coordinate system is greater than 5 m / s or the roll angle and pitch angle are greater than 32°.
[0045] f penalty_stop =-2000 (9)
[0046] f reward_alive =30 (8)
[0047] f penalty_velocity =4·[tanh(abs(V x ))+tanh(abs(V y ))+tanh(abs(V z (7)
[0048] Among them, V x V y V z This represents the velocity of the quadcopter UAV along the x, y, and z axes in the Earth coordinate system;
[0049] f penalty_posture=4·[tanh(2·abs(roll))+tanh(2·abs(pitch))]+10·tanh(2·abs(error_yaw)) (6)
[0050] Where tanh is the hyperbolic tangent function, tanh(x) = (e^(-tanh / x))^(-tanh / x ... x -e -x ) / (e x +e -x ); roll and pitch represent the roll and pitch angles of the quadcopter drone in the Earth coordinate system; error_yaw represents the yaw error of the quadcopter drone in the Earth coordinate system; abs is an absolute value function, i.e., abs(x) = |x|.
[0051] The beneficial effects of this invention are mainly reflected in:
[0052] 1. This invention proposes a novel end-to-end control method for reinforcement learning agents by replacing the direct output of four duty cycle signals with a "base + bias" approach. Using "bias" to adjust actions significantly reduces the ineffective exploration space, ensures gradient stability, and accelerates model convergence.
[0053] 2. This invention innovatively designs a policy network structure for reinforcement learning agents. It effectively suppresses oscillations by using a long short-term memory network, while introducing residual connections to avoid the loss of details during the propagation of minute numerical changes in the deep network, thereby improving steady-state accuracy.
[0054] 3. In this invention, by adopting the "base+bias" control method and combining the "long short-term memory + residual connection" network structure, end-to-end attitude control response of a quadcopter UAV to a given target attitude can be achieved. Attached Figure Description
[0055] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0056] Figure 1 This is a flowchart illustrating an end-to-end attitude control method for a quadrotor UAV based on reinforcement learning according to the present invention.
[0057] Figure 2 This refers to the Actor network structure in Embodiment 1 of the present invention;
[0058] Figure 3 This refers to the Critic network structure in Embodiment 1 of the present invention;
[0059] Figure 4 This is a graph showing the reward curve obtained by the agent in the early training stage of Example 1.
[0060] Figure 5 This is a graph showing the reward curve obtained by the agent during the training phase in Example 1;
[0061] Figure 6 This is a graph showing the reward curve obtained by the agent in the later stage of training in Example 1;
[0062] Figure 7 This is a graph showing the control effect of the reinforcement learning agent trained in Example 1 when the target yaw angle is 0°.
[0063] Figure 8 The graph shows the control effect of the reinforcement learning agent trained in Example 1 on a target yaw angle of 10°.
[0064] Figure 9 The graph shows the control effect of the reinforcement learning agent trained in Example 1 on a target yaw angle of -10°.
[0065] Figure 10 The graph shows the control effect of the reinforcement learning agent trained in Example 1 on a target yaw angle of 20°.
[0066] Figure 11 This is a graph showing the control effect of the reinforcement learning agent trained in Example 1 on a target yaw angle of -20°.
[0067] Figure 12 This is a graph showing the control effect of the reinforcement learning agent trained in Example 1 on the step signal;
[0068] Figure 13 This is a graph showing the control effect of the reinforcement learning agent trained in Example 1 on the sinusoidal signal.
[0069] Figure 14 The training results of the agent without using the "base+biase" control method are shown in the figure.
[0070] Figure 15 The training effect of an agent without long short-term memory and residual connections is shown in the figure.
[0071] Figure 16 This is a diagram illustrating the control effect of an agent without long short-term memory and residual connections. Detailed Implementation
[0072] The present invention will be further described below with reference to specific embodiments, but the scope of protection of the present invention is not limited thereto:
[0073] Example 1: An end-to-end attitude control method for a quadcopter UAV based on reinforcement learning, such as... Figure 1-3As shown, taking the Proximal Policy Optimization (PPO) reinforcement learning algorithm framework for training a reinforcement learning agent to control yaw angle as an example, the end-to-end attitude controller based on reinforcement learning designed in this invention includes the following steps:
[0074] Step 1: Build a quadcopter drone model
[0075] 1) Body model:
[0076] The rotor rotation provides lift and generates corresponding torque. The magnitude of lift and torque is related to the rotor speed. The resulting lift and torque act on the UAV, causing it to acquire acceleration and angular acceleration, which in turn produces corresponding changes in velocity, angular velocity, attitude angle, and position. Based on dynamics and kinematic analysis, the airframe model of a quadcopter UAV can be obtained:
[0077]
[0078] Wherein, φ, θ, and ψ represent the roll angle, pitch angle, and yaw angle, respectively; These represent the actual rotational speeds of rotor motors 1-4, respectively. This represents the acceleration in the drone's body coordinate system; K represents the angular acceleration in the UAV's body coordinate system; g represents the gravitational acceleration; K represents the acceleration due to gravity. T The lift coefficient is represented by d; the distance between the centers of the two diagonally opposite rotor motors of the quadcopter UAV is represented by J. r The moment of inertia of the rotor and its motor rotor is represented by p, q, and r, respectively, which represent the attitude angular velocities in the UAV body coordinate system; I x I y I z K represents the moments of inertia of the UAV along the x, y, and z axes of the body coordinate system, respectively; Q This represents the anti-torque coefficient.
[0079] Velocity information of a quadcopter drone in Earth coordinate system Angular velocity information in Earth coordinate system as follows:
[0080]
[0081] in, V represents the rotation matrix of the body coordinate system relative to the Earth coordinate system; W represents the transformation matrix from the angular velocity in the body coordinate system to the angular velocity in the Earth coordinate system; B ω represents the velocity of the UAV in the body coordinate system, and ω represents the attitude angular velocity of the UAV in the body coordinate system.
[0082] 2) Rotor motor model:
[0083] The motion expression of a rotor motor can be simplified as follows:
[0084]
[0085] Among them, T e T represents the electromagnetic torque generated by the rotor motor. L GD represents the load torque. 2 This represents the flywheel torque, and n represents the rotor's mechanical speed. This represents the mechanical rotational speed differentiated with respect to time.
[0086] The rotor motor circuit equation can be expressed as:
[0087]
[0088] Among them, R, L, E, I d These represent the total resistance, total inductance, back electromotive force, power supply voltage, and current in the circuit (i.e., the electromagnetic torque current of the rotor motor) in the equivalent circuit of the rotor motor, respectively. This represents the derivative of current with respect to time.
[0089] Applying the Laplace transform to equations (3) and (4) yields the following:
[0090]
[0091] in, For electromechanical time constant, T l =L / R is the electromagnetic time constant; C m C represents the torque coefficient of a rotor motor under rated excitation. e This is the electromotive force coefficient of the motor under rated magnetic flux. I d Indicates electromagnetic torque current; I dL E(s) represents the load torque current; E(s) is the Laplace transform of the motor back electromotive force, and s represents the complex frequency in the Laplace transform.
[0092] 3) Quadcopter drone model:
[0093] The above-mentioned airframe model is combined with the rotor motor model to obtain a quadcopter drone model.
[0094] Step 2: Set up a reinforcement learning interactive environment
[0095] Step 2.1 Reward Strategy:
[0096] Taking a relatively simple reward strategy as an example, it includes:
[0097] Attitude error penalty:
[0098] fpenalty_posture =4·[tanh(2·abs(roll))+tanh(2·abs(pitch))]+10·tanh(2·abs(error_yaw)) (6)
[0099] Where tanh is the hyperbolic tangent function, tanh(x) = (e^(-tanh / x))^(-tanh / x ... x -e -x ) / (e x +e -x ); roll and pitch represent the roll and pitch angles of the UAV in the Earth coordinate system; error_yaw represents the yaw angle error of the UAV in the Earth coordinate system; abs is an absolute value function, i.e., abs(x) = |x|.
[0100] Speed penalty:
[0101] f penalty_velocity =4·[tanh(abs(V x ))+tanh(abs(V y ))+tanh(abs(V z (7)
[0102] Among them, V x V y V z This represents the velocity of the drone along the x, y, and z axes in the Earth coordinate system.
[0103] Survival Rewards:
[0104] If only punishment is applied, the agent is likely to converge to a point where the environment reaches a stopping condition. To allow the agent to maintain control in the long term, it should be encouraged to continue interacting.
[0105] f reward_alive =30 (8)
[0106] Stop the punishment:
[0107] This design incorporates undesirable states as stopping conditions to prevent unwanted actions, such as excessive drone attitude angles or speeds. If these conditions occur, further training is unnecessary, as they would realistically result in a drone crash.
[0108] f penalty_stop =-2000 (9)
[0109] In summary, the reward strategy is as follows:
[0110]
[0111] Among them, f reward The immediate reward for the current state is set as follows: the drone's movement speed along the X, Y, and Z axes in the Earth coordinate system is greater than 5 m / s or the roll angle and pitch angle are greater than 32°.
[0112] Step 2.2 Training Strategy:
[0113] Taking a simple yaw angle control training strategy as an example: the reinforcement learning agent randomly resets the target yaw angle after each round of interaction.
[0114] This invention uses a reinforcement learning algorithm based on the Actor-Critic framework and policy gradient to determine the target function to be optimized based on the selected reinforcement learning algorithm:
[0115] When using the Actor-Critic framework, two neural networks are constructed: one as the Actor network and the other as the Critic network. The Actor network is responsible for the actual interactions in the reinforcement learning environment; the Critic network can be divided into "state-value functions".
[0116] There are two types of functions: "state-action value function" (based on the discounted reward that the Actor can earn according to the state of the reinforcement learning environment) and "state-action value function" (based on the action output by the Actor network and the discounted reward that the Actor can earn according to the state of the reinforcement learning environment). This example uses the "state value function". The two are used together to guide the Actor network to achieve the optimal policy through the Critic network.
[0117] The training objective is to maximize the sum of immediate rewards obtained by the reinforcement learning agent, thereby learning an optimal control policy for a quadcopter drone. The Actor network of the reinforcement learning agent is updated by policy gradients. The reinforcement learning agent obtains immediate rewards by interacting with the environment, and then obtains the final discounted reward for each round of interaction. The expected value of the discounted reward is obtained through multiple rounds of interaction, and the derivative of the expected value is used to obtain the gradient of the network update. After obtaining the gradient, gradient ascent is performed to find the optimal control policy.
[0118] Step 2.3 Interaction Objects:
[0119] This refers to a quadcopter drone model where the actions output by the reinforcement learning agent act on the quadcopter drone's motors, causing changes in the environmental conditions (i.e., changes in rotor speed, drone attitude, speed, air density, air pressure, etc.).
[0120] Step 2.4 Determine the state space:
[0121] This invention selects a portion of the environmental states from the motors of a quadcopter drone to construct the spatial state of a quadcopter drone model as input to a reinforcement learning agent. This spatial state comprises 21 dimensions, specifically including: velocity (velocity along the x, y, and z axes of the Earth coordinate system), acceleration (acceleration along the x, y, and z axes of the Earth coordinate system), attitude (roll, pitch, and yaw angles) in the Earth coordinate system, and angular velocity (angular velocity in the Earth coordinate system). The system uses eight state variables—roll angle angular velocity, pitch angle angular velocity, yaw angle angular velocity, yaw angle error, rotational speeds of rotor motors 1-4 (the rotational speeds of the four rotors), and duty cycles of the control rotor motors (the actual duty cycles input to the four rotor motors)—as motor states. In the current state, the remaining 13 state variables (i.e., the speed, acceleration, attitude, angular velocity, and yaw angle error of the UAV in the Earth coordinate system) are used as the main physical quantities and are used as inputs to the reinforcement learning agent.
[0122] Step 3: Constructing a reinforcement learning agent:
[0123] Taking the PPO algorithm framework as an example, two networks, Actor and Critic, need to be constructed to form the reinforcement learning agent of the UAV of this invention. The network structure is constructed based on the "long short-term memory + residual connection" structure. At the same time, attention should be paid to preprocessing the motor state. The Actor network structure is as follows: Figure 2 As shown, the Critic network structure is as follows: Figure 3As shown. In this embodiment, the input layer network and backbone network of the Actor network and Critic network are the same, while the output layer network is slightly different. Specifically, the input layer network of both the Actor network and Critic network includes three parallel branches: a motor state preprocessing network, a main physical quantity state preprocessing network, and a residual connection channel (a fully connected layer). The motor state preprocessing network is mainly a "fully connected layer---long short-term memory layer---fully connected layer" structure; the main physical quantity state preprocessing network is mainly a "fully connected layer---ReLU nonlinear activation layer---fully connected layer---ReLU nonlinear activation layer---fully connected layer---ReLU nonlinear activation layer" structure. The initial state determined by the 21-dimensional input in step 2.4 after processing by the fully connected layer, the motor state vector after processing by the motor state preprocessing network, and the main physical quantity state vector after processing by the main physical quantity state preprocessing network are all used as the input of the backbone network. The first layer of the backbone network first connects the outputs from the motor state preprocessing network and the main physical quantity state preprocessing network in series. Then, it passes through a fully connected layer, a ReLU nonlinear activation layer, and then a long short-term memory network layer. After passing through another fully connected layer, it forms a residual connection with the initial state processed by the residual connection channel (fully connected layer) of the input layer network, and then passes through another ReLU nonlinear activation layer. For the output layer network, since the Critic network only needs to predict one value, the output of the backbone network is passed through a fully connected layer and then directly predicts a scalar. The Actor network's output layer consists of parallel expectation value output networks and standard deviation output networks. The expectation value output network comprises a fully connected layer, a ReLU nonlinear activation layer, and another fully connected layer connected in sequence. The standard deviation output network comprises two consecutive fully connected layers. For the Actor network, a 5-dimensional vector needs to be output as the action prediction for controlling the drone. The action prediction includes one base value and four bias values. The base value serves as the basic duty cycle for the four rotors, and the four bias values represent the adjustment amount of the basic duty cycle based on the current state of the drone. To explore randomness during training, the expected action output by the expectation value output network corresponds one-to-one with the standard deviation output by the standard deviation output network, and the action prediction is output through a normal distribution.
[0124] Based on the general approximation theorem and the characteristics of recurrent networks, a recurrent neural network can theoretically be constructed to approximate any nonlinear dynamic system. This design introduces a long short-term memory network layer, which helps the network to better learn the complex timing and dynamic characteristics in the control of a quadcopter UAV, achieving the goals of fast response and reduced overshoot. In addition, since the sensitivity of deep neural networks to small numerical changes is difficult to control, the numerical changes in the state are very small when the actual attitude of the UAV is close to the target attitude. These small changes are easily lost during multi-layer propagation. Introducing residual connections helps to preserve the original information and achieve more accurate control.
[0125] Reinforcement learning agents are used to interact with the environment in a reinforcement learning environment; in this design, this means controlling the state of a drone. Specifically, the control works as follows: the initial state of the drone is input to the reinforcement learning agent, which then outputs control actions through the Actor. The drone's state changes under the Actor's actions. After the new drone state is input back to the reinforcement learning agent, the agent generates a new action based on the Actor, causing the drone's state to change again, thus forming a closed-loop control.
[0126] Step 4: Conduct interactive training using the "base + bias" control method:
[0127] Step 4.1 Determine the objective function to be optimized:
[0128] In this example, the PPO algorithm is used to train the reinforcement learning agent, and its objective function is:
[0129]
[0130] in, ε is the shearing factor. Represents the generalized dominance function for state s t The advantage estimate. P θ This indicates that the current Actor network input is state s. t Output action a t The probability, P θ′ This indicates that the input to the old Actor network is state s. t Output action a t The probability of.
[0131] Step 4.2 Calculate the rotor motor duty cycle using the "base+bias" method:
[0132] The motion prediction output by the Actor network of the reinforcement learning agent is a 5-dimensional vector. The first element of the 5-dimensional vector is used as the Base value, and the remaining 4 elements are used as Bias values, which are then used as the speed adjustment control signals for the four rotors. When inputting to the rotor motors, the duty cycle input to each rotor motor is:
[0133] PWM i =Base+Bias i (12)
[0134] Where i represents the serial number of the four rotor motors.
[0135] Step 4.3 Training:
[0136] The constructed reinforcement learning agent is interactively trained according to the selected reinforcement learning algorithm until it converges or the training results are saved in stages for transfer learning. Specifically:
[0137] Within each training cycle, the target yaw angle is randomly initialized according to the preset training strategy. The reinforcement learning agent interacts with the environment every 2.5 milliseconds and reads the environmental state (i.e., the UAV state) every 2.5 milliseconds to generate a 21-dimensional state space vector. The state space vector is input to the reinforcement learning agent (i.e., the state space vector is input to the Actor network and the Critic network respectively) to obtain action prediction and the agent's value estimate of the current state. During each interaction, the duty cycle of the four rotor motors is calculated using the "base+bias" method based on the action output by the current Actor network. This duty cycle is input to the quadcopter UAV model. The Critic network outputs the value estimate based on the current state. At the same time, the immediate reward for transitioning from the state at the time of the last interaction to the current state is calculated according to the reward strategy in step 2.1. Then, the immediate reward, the current state space vector, the action prediction output by the Actor network based on the current UAV state, and the value estimate output by the Critic network based on the current state are stored in the experience replay pool as training data.
[0138] The above steps are repeated during the next interaction (i.e., after 2.5 milliseconds). After a certain number of training rounds, the accumulated amount of training data in the experience replay pool is checked to see if it meets the batch size (in this embodiment, it is checked every 6 rounds of training in the early stage, with a batch size of 1024; in subsequent training, it is checked every 12 rounds, with a batch size of 4096). If it meets the requirements, the data is used in batches to update the network parameters. The Actor network uses the obtained data to update according to the objective function preset in step 4.1 to optimize its action strategy. The Critic network uses the estimated reward in the training data and the final discounted reward to perform regression training, thereby improving its accuracy in evaluating state value.
[0139] Performance evaluation and training termination:
[0140] After each training round, the sum of the immediate rewards obtained in that round is calculated. If the total reward obtained in this round reaches the preset saving threshold, the agent's current model will be saved for later use. In addition, after each training round, the average total reward of the last 200 rounds is calculated (the average is taken based on the maximum number of training rounds if the training does not exceed 200 rounds), and compared with the set training termination threshold. Once this threshold is reached, it is considered that the agent has learned a sufficiently good policy, and the training process will terminate.
[0141]
[0142] Step 5: Use the trained reinforcement learning agent to control the quadcopter drone model:
[0143] Step 5.1: Load the model
[0144] Load the Actor network, which has been trained offline in step 4, into the reinforcement learning environment you have built (training components such as the Critic network and experience replay pool will no longer be needed).
[0145] Step 5.2: Set the target attitude
[0146] A target yaw angle is set for the quadcopter drone through programmatic configuration. This target attitude is compared with the drone's actual attitude to derive the attitude error, which is then provided to the Actor network.
[0147] Step 5.3: Real-time control loop (every 2.5 milliseconds)
[0148] The quadcopter drone enters a high-speed, cyclical control state, consistent with the interaction frequency during training, that is, it executes once every 2.5 milliseconds.
[0149] (1) State perception
[0150] The relevant variables in the reference program include variables that store the drone's speed, acceleration, angular velocity, attitude angle, target attitude, motor speed, and duty cycle.
[0151] (2) State space construction
[0152] The real-time attitude data of the quadcopter UAV model simulation is combined with the preset target yaw angle to form a state space, which includes:
[0153] Speed: The speed of a quadcopter drone along the x, y, and z axes in the Earth coordinate system;
[0154] Acceleration: The acceleration of a quadcopter UAV along the x, y, and z axes in the Earth coordinate system;
[0155] Angular velocity: the roll rate, pitch rate, and yaw rate of a quadcopter drone;
[0156] Attitude angles: the roll angle, pitch angle, and preset target yaw angle of the quadcopter UAV;
[0157] Attitude angle error: The error between the attitude of the quadcopter UAV and the desired attitude;
[0158] Motor speed: The rotational speed of the four rotor motors of the quadcopter drone;
[0159] Duty cycle: The duty cycle of the four rotor motors of a quadcopter drone;
[0160] (3) Actor Network Decision
[0161] The state vector is obtained from the state space (i.e., the values of each variable in the state space are read every 2.5 milliseconds, and a set of state vectors are formed after calculating the attitude angle error). It is input into the loaded Actor network to calculate and output a 5-dimensional vector representing the "action", including a base lift PWM value and four attitude adjustment PWM values.
[0162] (4) Calculate the duty cycle signal PWM using the "base+bias" rule. i The signals are sent to the electronic speed controllers (ESCs) of the four rotor motors, and the ESCs (ESCs) process the received duty cycle signals according to the PWM signals. i By changing the voltage across the motor terminals, the rotational speed of each rotor is altered. The differentiated combination of the four rotor speeds generates a specific torque, driving the quadcopter drone's fuselage to yaw towards the target.
[0163] During this process, the Actor network continuously senses new states, inputs them into the network, outputs new actions, and adjusts motor speeds. This high-speed closed-loop control process enables the UAV to respond smoothly, quickly, and accurately to commands, ultimately stabilizing at the target attitude.
[0164] experiment:
[0165] The quadrotor drone control simulation experiment was conducted using the reinforcement learning-based end-to-end attitude controller for quadrotor drones described in Example 1. This verified the feasibility and effectiveness of the controller based on the "base+bias" and "long short-term memory+residual connection" network structures for controlling the quadrotor drone. The simulation results are as follows: Figures 4 to 13 As shown.
[0166] Figures 4-6 These are all training effect graphs. In the graphs, the yellow curve represents the agent Critic's expected total reward for a single round of interaction, the light blue curve represents the total reward actually obtained by the agent in a single round of interaction, and the dark blue curve represents the average reward obtained in the past 200 rounds of interaction. Figure 4 As shown, after approximately 630 rounds of training, the drone was able to avoid premature termination of training due to tipping over or excessive speed. Figure 5 Shown in Figure 4 The training curve for transfer learning based on the training results further improves the agent's control performance. Figure 6 Yes Figure 5 Further training, as shown in the training results, eventually leads to a convergence in the total reward earned by the agent.
[0167] Figures 7-13 All figures are curves showing the control effect of the reinforcement learning agent. The first column shows the velocity change curve, with red, green and blue representing the velocity curves along the X, Y and Z axes of the Earth coordinate system, respectively. The second column shows the attitude angle change curve, with red, green and blue representing the roll angle, pitch angle and yaw angle of the UAV relative to the Earth coordinate system, respectively. The third column shows the actual speed change curve of the motor, with blue, black, red and green representing the actual speed of rotor motors 1 to 4, respectively. Figures 7-13 The results shown demonstrate that the reinforcement learning agent enabled the UAV to track the target's yaw angle.
[0168] exist Figure 12 In the step signal test shown, a step signal with an amplitude of 10° is applied to the target yaw angle at the 10th second, before which the target yaw angle is 0°.
[0169] exist Figure 13 In the sinusoidal signal test shown, a sinusoidal signal with an amplitude of 10° and a period of 12 seconds is applied to the target yaw angle at the 10th second. Before this, the target yaw angle is 0°.
[0170] It is worth noting that only a simple training strategy was used during the training process, while Figure 12 and Figure 13 The simulation results shown indicate that the reinforcement learning agent is still able to respond to the tested step and sinusoidal signals and achieve good control results.
[0171] Figure 14 To achieve the same network structure as the reinforcement learning agent in this embodiment, but without employing the "base+bias" control method, the Actor's output is only 4 dimensions directly used to control the drone's duty cycle. In the experiment, its training was extremely difficult, showing no signs of convergence even after tens of thousands of training rounds. In contrast, without the "base+bias" control method, the agent's overall exploration space is extremely large and the gradient is very unstable, with most of it being an ineffective exploration space, thus making convergence difficult. However, the "base+bias" control method effectively compresses the ineffective exploration space, thereby reducing the overall exploration space, ensuring gradient stability, and enabling effective convergence. This significantly reduces training resource overhead, resulting in more stable and smoother drone control, avoiding the problem of excessively large differences in the rotational speeds of the four wings leading to loss of control.
[0172] For a feedforward neural network consisting only of fully connected layers and non-linear activation layers, its input and output can be viewed as a function:
[0173] y k =f(x) k (13)
[0174] Where f is the input-output mapping function of the feedforward neural network, x k y represents the input of the feedforward neural network at time t=k. k This indicates that the feedforward neural network is for x k The output.
[0175] For a neural network with long short-term memory layers and residual connections, its input and output can be viewed as:
[0176] y k =h(x k ,x k-1 )+r(x k (14)
[0177] Where h represents the long short-term memory network in this structure for input x. k The input-output relationship mapping, r represents the residual connection in the network with respect to input x under this structure. k The input and output relationship mapping.
[0178] Clearly, a feedforward neural network's output only considers the input at the current moment. However, the attitude adjustment of a quadcopter drone is a process, and using a feedforward neural network as an attitude controller would lack awareness of the temporal sequence of this process. In contrast, a neural network with long short-term memory layers and residual connections considers not only the current input but also previous inputs. Furthermore, the residual connections retain the original information from previous inputs, allowing this network structure to not only perceive the quadcopter drone's attitude adjustment process but also retain sensitivity to the original information.
[0179] Figure 15 To achieve a network structure that is roughly the same as that in this embodiment, but using a network structure with no residual connections, all long short-term memory network layers replaced by fully connected layers, and employing a "base+bias" control method, the agent training effect is comparable to that in this embodiment. Its performance at a target yaw angle of 0° is as follows: Figure 16 As shown, under the control of this network structure, the agent was initially able to adjust the drone's attitude appropriately, but subsequently lost its adjustment ability and caused large oscillations. As can be seen from the training effect diagram, after a considerable number of training rounds, compared with the structure used in Example 1, the agent under this structure is still at a certain local extreme point and requires a lot of training. In contrast, the network structure in Example 1 has high training efficiency and can grasp the timing characteristics of the drone control process more quickly.
[0180] The experimental results show that the end-to-end attitude controller design for a quadcopter UAV based on reinforcement learning described in this invention has the characteristics of fast convergence speed, good control effect, reasonable network structure and good generalization ability. It has certain innovation and practicality in the field of reinforcement learning control of UAVs.
[0181] Finally, it should be noted that the above examples are merely some specific embodiments of the present invention. Obviously, the present invention is not limited to the above embodiments and many variations are possible. All variations that can be directly derived or conceived by those skilled in the art from the disclosure of the present invention should be considered within the scope of protection of the present invention.
Claims
1. An end-to-end attitude control method for a quadrotor unmanned aerial vehicle based on reinforcement learning, characterized in that: The process includes, S1. Establish a model of a quadcopter drone, including the airframe model and the rotor motor model; S2. Construct a reinforcement learning interactive environment, including reward policies and state space; S3. Construct a reinforcement learning agent, including an Actor network and a Critic network. The state space is used as the input to the Actor network and the Critic network, respectively. The Actor network outputs an action prediction, which is a 5-dimensional vector including one Base value and four Bias values. The Critic network outputs a value estimate. S4. Using a reinforcement learning algorithm and objective function based on the Actor-Critic framework and policy gradient, The "base+bias" method is used for interactive training of reinforcement learning agents; S5. Use the Actor network of trained reinforcement learning agents to control the attitude of the quadcopter drone.
2. The end-to-end attitude control method for a quadrotor UAV based on reinforcement learning according to claim 1, characterized in that: The machine model is: Wherein, φ, θ, and ψ represent the roll angle, pitch angle, and yaw angle, respectively; These represent the actual rotational speeds of rotor motors 1-4, respectively. This represents the acceleration in the drone's body coordinate system; K represents the angular acceleration in the UAV's body coordinate system; g represents the gravitational acceleration; K represents the acceleration due to gravity. T The lift coefficient is represented by d; the distance between the centers of the two diagonally opposite rotor motors of the quadcopter UAV is represented by J. r The moment of inertia of the rotor and its motor rotor is represented by p, q, and r, respectively, which represent the attitude angular velocities in the UAV body coordinate system; I x I y I z K represents the moments of inertia of the UAV along the x, y, and z axes of the body coordinate system, respectively; Q Indicates the anti-torque coefficient; Velocity information of a quadcopter drone in Earth coordinate system Angular velocity information in Earth coordinate system as follows: in, V represents the rotation matrix of the body coordinate system relative to the Earth coordinate system; W represents the transformation matrix from the angular velocity in the body coordinate system to the angular velocity in the Earth coordinate system; B ω represents the velocity of the UAV in the body coordinate system, and ω represents the attitude angular velocity of the UAV in the body coordinate system.
3. The end-to-end attitude control method for a quadrotor UAV based on reinforcement learning according to claim 2, characterized in that: The rotor motor model is in, Both are time constants, R and L represent the total resistance and total inductance in the equivalent circuit of the rotor motor, respectively, and C m This represents the torque coefficient of a rotor motor under rated excitation. C e I is the electromotive force coefficient of the motor under rated magnetic flux; d Indicates electromagnetic torque current; I dL represents the load torque current; s represents the complex frequency in the Laplace transform.
4. The end-to-end attitude control method for a quadrotor UAV based on reinforcement learning according to claim 3, characterized in that: The state space includes, in the Earth coordinate system, the velocity, acceleration, roll angle, pitch angle, yaw angle, roll angle angular velocity, pitch angle angular velocity, yaw angle angular velocity, yaw angle error, rotational speed and duty cycle of the four rotor motors of the quadcopter UAV along the x, y, and z axes. The rotational speed and duty cycle of the four rotor motors are used as motor states, and the rest are used as main physical quantity states.
5. The end-to-end attitude control method for a quadrotor UAV based on reinforcement learning according to claim 4, characterized in that: The input layer network structures of the Actor network and the Critic network are the same, both including three parallel branches: a motor state preprocessing network, a main physical quantity state preprocessing network, and a residual connection channel. The motor state vector is used as the input of the motor state preprocessing network, the main physical quantity state vector is used as the input of the main physical quantity state preprocessing network, and the state space vector is used as the input of the residual connection channel. The backbone networks of the Actor network and the Critic network have the same structure. After the output of the motor state preprocessing network and the output of the main physical quantity state preprocessing network are spliced, they pass through the fully connected layer, the ReLU nonlinear activation layer, the long short-term memory network layer, and the fully connected layer in sequence. Then, they are residually connected to the output of the residual connection channel of the input layer network, and then input to the output layer network through the ReLU nonlinear activation layer. The output layer network of the Actor network includes a parallel expected value output network and a standard deviation output network. The expected value output network includes a fully connected layer, a ReLU nonlinear activation layer, and a fully connected layer connected in sequence. The standard deviation output network includes two consecutive fully connected layers. Then, the output of the expected value output network corresponds one-to-one with the output of the standard deviation output network, and outputs a 5-dimensional vector of action prediction through a normal distribution. The output layer of the Critic network is a fully connected layer that outputs a value estimate of the current state.
6. The end-to-end attitude control method for a quadrotor UAV based on reinforcement learning according to claim 5, characterized in that: The "base+bias" method is as follows: four bias values correspond to the speed adjustment control signals of the four rotors, and the duty cycle input to each rotor motor is: PWM i =Base+Bias i (12) Where i represents the serial number of the four rotor motors.
7. The end-to-end attitude control method for a quadrotor UAV based on reinforcement learning according to claim 6, characterized in that: The motor state preprocessing network includes a fully connected layer, a long short-term memory layer, and a fully connected layer connected in sequence. The main physical quantity state preprocessing network includes a fully connected layer, a ReLU nonlinear activation layer, a fully connected layer, a ReLU nonlinear activation layer, a fully connected layer, and a ReLU nonlinear activation layer connected in sequence. The residual connection channel is a fully connected layer.
8. The end-to-end attitude control method for a quadrotor UAV based on reinforcement learning according to claim 7, characterized in that: The interactive training process is as follows: (1) During each training cycle, the target yaw angle is randomly initialized, the state space vector of the quadcopter is periodically collected, and then input into the reinforcement learning agent; the action prediction output by the Actor network calculates the duty cycle of the four rotor motors using the "base+bias" method, and is used to input the quadcopter model; the Critic network outputs the value estimate based on the current state, calculates the immediate reward of the current state according to the reward strategy, and then stores the immediate reward, the current state space vector, the action prediction, and the value estimate of the current state into the experience replay pool as training data; (2) After a predetermined number of training cycles, check whether the accumulated amount of training data in the experience replay pool meets the minimum batch size. If it does, extract data from the experience replay pool and update the network parameters in batches: the Actor network updates according to the objective function, and the Critic network uses the estimated reward and the actual reward for regression training.
9. The end-to-end attitude control method for a quadrotor UAV based on reinforcement learning according to claim 8, characterized in that: The reward strategy is as follows: Among them, f reward The immediate reward for the current state is given, and the stopping condition is that the speed of the quadcopter UAV along the X, Y, and Z axes in the Earth coordinate system is greater than 5 m / s or the roll angle and pitch angle are greater than 32°. f penalty_stop =-2000 (9) f reward_alive =30 (8) f penalty_velocity =4·[tanh(abs(V x ))+tanh(abs(V y ))+tanh(abs(V z ))] (7) Among them, V x V y V z This represents the velocity of the quadcopter UAV along the x, y, and z axes in the Earth coordinate system; f penalty_posture =4·[tanh(2·abs(roll))+tanh(2·abs(pitch))]+10·tanh(2·abs(error_yaw)) (6) Where tanh is the hyperbolic tangent function, tanh(x) = (e^(-tanh / x))^(-tanh / x ... x -e -x ) / (e x +e -x ); roll and pitch represent the roll and pitch angles of the quadcopter drone in the Earth coordinate system; error_yaw represents the yaw error of the quadcopter drone in the Earth coordinate system; abs is an absolute value function, i.e., abs(x) = |x|.