Six-degree-of-freedom rapid cooperative fine formation method for quadrotor unmanned aerial vehicle
By constructing a six-degree-of-freedom dynamic model and a collision box model for quadrotor UAVs, and training the agent using the Kronecker factor trust threshold executor/commentator deep reinforcement learning algorithm, the problems of slow and poor trajectory planning in UAV formation flight are solved, and fast and accurate UAV swarm collaborative formation is achieved.
Patent Information
- Application Number
- CN202310582466.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-23
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2043-05-23
AI Technical Summary
Existing drone formation flight methods suffer from slow and inaccurate trajectory planning, difficulty in four-dimensional trajectory planning, and low control accuracy in complex environments.
A six-degree-of-freedom dynamics model and a collision box model of a quadrotor UAV are used, combined with a Kronecker factor trust threshold executor/commentator deep reinforcement learning algorithm to train the agent for cooperative formation trajectory planning.
It achieves safe, efficient, fast and accurate collaborative formation trajectory planning for UAV swarms, reducing formation time to 0.753 seconds, and the trajectory planning results include timestamps and four-dimensional tracks with a granularity of 1 second.
Smart Images

Figure CN116679746B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of collaborative flight path planning for unmanned aerial vehicles (UAVs), and more particularly to a method for rapid collaborative fine formation of six-degree-of-freedom quadrotor UAVs. Background Technology
[0002] Drone formation flying has matured in military reconnaissance, raids, search and capture and other collaborative command and combat applications. In recent years, with the relaxation of low-altitude airspace control policies and the maturity and cost reduction of drone technology, drone applications have gradually expanded from military to civilian fields. In particular, drone formation flying is widely used in general aviation fields such as civilian flight performances, aerial photography, and emergency transportation.
[0003] Due to the forward-looking nature and complexity of UAV formation flight research, as well as its multidisciplinary nature, there has been some progress and achievements in UAV formation flight methods both domestically and internationally. Yan Junkun et al. disclosed a UAV formation trajectory optimization method based on reconnaissance and positioning tasks, establishing a formation trajectory optimization model based on a two-dimensional plane coordinate system. Its shortcoming lies in not considering the UAV's motion trajectory along the z-axis, making the motion model overly idealized. Huangfu Yafan et al. disclosed a multi-UAV trajectory planning method based on virtual navigation, achieving collaborative trajectory planning for multiple UAVs through the artificial potential field method. Its advantages include smooth and safe paths and ease of calculation; however, the artificial potential field method generally suffers from local optima problems, resulting in low reliability in complex environments and poor control accuracy when communication capabilities are insufficient.
[0004] Therefore, existing UAV swarm flight methods suffer from problems such as slow trajectory planning, poor accuracy of planned trajectories, and difficulty in four-dimensional trajectory planning based on longitude, latitude, altitude, and time. There is an urgent need to propose an efficient, flexible, and precise method for rapid collaborative formation of UAV swarms. Summary of the Invention
[0005] The purpose of this invention is to overcome the above-mentioned defects in the background technology and provide a method for rapid and precise coordinated formation of quadcopter UAVs with six degrees of freedom, so as to achieve safe, efficient, fast and accurate UAV swarm coordinated formation trajectory planning.
[0006] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0007] A method for rapid, coordinated, and precise formation of six-degree-of-freedom quadcopter UAVs, characterized in that the method specifically includes the following steps:
[0008] Step S1: Construct a six-degree-of-freedom dynamic model of a quadcopter UAV;
[0009] Step S2: Construct a collision box model of a quadcopter drone;
[0010] Step S3: Agent training based on Kronecker factor trust threshold enforcer / commentator deep reinforcement learning;
[0011] Step S4: Perform collaborative formation trajectory planning based on artificial intelligence agents.
[0012] Furthermore, the specific content of step S1, which involves constructing a six-degree-of-freedom dynamic model of a quadcopter UAV, is as follows:
[0013] If we consider a quadcopter drone as a centrally symmetric rigid body, then its motion in space mainly consists of translation and rotation. Considering only the drone's translation, we have Newton's second law:
[0014]
[0015] Among them, F 合 For the net force acting on the drone, F T For the lift force on the drone, F G For the weight of the drone, F f Let be the air resistance experienced by the drone, a be the acceleration of the drone, x be the displacement of the drone, and m be the mass of the drone.
[0016] Considering the gimbal lock in Euler rotation and the motion characteristics of the UAV, it is necessary to avoid pitch angles of ±90 degrees during UAV operation. The UAV rotation sequence is defined as Z-axis → Y-axis → X-axis, corresponding to yaw angle ψ, pitch angle θ, and roll angle φ. Therefore, the Euler rotation matrix R of the quadcopter UAV is... E :
[0017]
[0018] For a drone control system, the method to control the drone's translation and rotation is to adjust the rotational speed of the four rotors, which yields the following force balance equation for the drone:
[0019]
[0020] Where x” represents the drone's acceleration along the x-axis, y” represents the drone's acceleration along the y-axis, and z” represents the drone's acceleration along the z-axis; x' represents the drone's velocity along the x-axis, y' represents the drone's velocity along the y-axis, and z' represents the drone's velocity along the z-axis; k x Let k be the air drag coefficient of the UAV along the x-axis. y k is the air drag coefficient in the y-axis direction. z R is the air drag coefficient in the z-axis direction; E Here, b is the Euler rotation matrix of the quadcopter UAV, b is the rotor lift coefficient of the UAV, and w is the Eulerian rotation matrix of the quadcopter UAV. iLet be the rotational speed of the i-th propeller of the UAV, g be the acceleration due to gravity, and U1 be the resultant lift force generated by the UAV rotor.
[0021] Therefore, the state transition equation for a quadcopter UAV is:
[0022]
[0023]
[0024]
[0025] Where x is the drone displacement in the x-axis direction, y is the drone displacement in the y-axis direction, and z is the drone displacement in the z-axis direction.
[0026] Furthermore, the specific content of step S2, which involves constructing the quadcopter drone collision box model, is as follows:
[0027] The three-dimensional coordinates of the key coordinate points of the UAV shape in the geodetic coordinate system are described as follows:
[0028]
[0029] Where ψ is the UAV yaw angle, θ is the UAV pitch angle, and φ is the UAV roll angle; [x',y',z'] T Let [x0, y0, z0] be the vector coordinates of the UAV in the coordinate system of the UAV's rotation center. T This represents the displacement of the UAV's rotation center relative to the origin of the geodetic coordinate system. This represents the matrix cross product; in this case, x, y, and z are the coordinates of the key points of the UAV after displacement [x0, y0, z0]. T At this point, [x,y,z] T These are the geodetic coordinates of the UAV after rotating around the yaw axis, pitch axis, and roll axis [ψ,θ,φ].
[0030] Furthermore, the specific steps of the agent training based on Kronecker factor trust threshold executor / commentator deep reinforcement learning in step S3 can be divided into:
[0031] S31: Environmental design;
[0032] The reward function for the drone approaching the target formation position is designed as follows:
[0033]
[0034] Among them, S i,t Let DP be the state of the i-th UAV at time t. i S represents the position of the i-th UAV target formation; i,t and DP iEach of these is a column vector consisting of the UAV's longitude x, latitude y, and altitude h.
[0035] The reward function for maintaining a safe distance between two drones is designed as follows:
[0036] in
[0037] Among them, P it Let drone i form the key drone location set at second t, and
[0038] P it ={(x it ,y it ,z it )1,(x it ,y it ,z it )2,(x it ,y it ,z it )3,(x it ,y it ,z it )4,(x it ,y it ,z it )5}
[0039] Among them, (x it ,y it ,z it These are the five key collision detection points for quadcopter drones: the positions of the left front rotor, right front rotor, left rear rotor, and right rear rotor.
[0040] Ultimately, the two quadcopter drones received a single-step reward R at time t. t for:
[0041]
[0042] S32: Agent Design;
[0043] Design an internal neural network, and denote the neural network q(S) of the fully connected backpropagation error state-action pair q inside the UAV. t A t ,W);
[0044] The action selection strategy is designed so that the agent selects the action corresponding to the maximum q-value based on the e-greedy strategy, that is:
[0045]
[0046] Where e is a small positive number in the interval 0-1;
[0047] S33: Environment-Agent Interaction Mode;
[0048] For each environment-agent interaction, the agent is based on a q-valued neural network q(S). t A t The e-greedy strategy selects action A. t Two quadcopter drones performed action A t A state transition is performed, and both drones enter the next state S. t+1 At the same time, the environment provides the agent with a state S. t Next, execute action A t Feedback reward value R t This is used by the agent to update the hidden parameters of its internal q-value neural network.
[0049] S34: Iteration of agent action value;
[0050] Initialize the agent's internal neural network as an arbitrary weight matrix W and the optimal policy estimation neural network, with the optimal policy estimation neural network having a weight matrix θ, and denoted as the optimal policy estimate π(θ); initialize the implicit learning rate α. (θ) α (W) Initialize the discount factor γ;
[0051] intelligent agents according to Two drones in status S t Make action decision A t Afterwards, you will receive an environmental reward value R. t Let U←R+γq(S) t+1 A t ,W);
[0052] Mapping Then, based on the Kronecker factor approximation curvature expression of these two fully connected neural networks, the corresponding Fisher matrix F can be derived. θ and F W :
[0053]
[0054]
[0055] Then, update To reduce -q(A) t |S t ;W)·π(A t |S t ;θ) value;
[0056] renew To reduce [Uq(A)] t |S t;W)]·π(A t |S t The value of θ);
[0057] Update drone status, S t =S t+1 Then proceed to the next iteration.
[0058] Furthermore, the specific steps of step S4, which involves collaborative formation trajectory planning based on an artificial intelligence agent, can be divided into:
[0059] S41: Initialize the set of aircraft K = {uav1, uav2, uav3, ..., uav4} in the airspace. n};
[0060] S42: Priority selection for drone pairs;
[0061] Find the UAV of the drone i With UAV drones j Euclidean distance Distance(uav) i ,uav j ):
[0062] Distance(uav i ,uav j )=d ij =||[x i ,y i ]-[x j ,y j ]||2
[0063] Select the UAV pair i, j with the closest Euclidean distance:
[0064]
[0065] S43: Unmanned aerial vehicle (UAV) collaborative trajectory planning and UAV swarm traversal;
[0066] After selecting the UAV pair i and j with the closest Euclidean distance, if uav i ∪uav j If ∈K, then the agent is uav i and UAV j Assign actions and remove uav from set K. i and UAV j ;like The intelligent agent is based on UAV. j The assigned action is UAV. i Assign actions and remove uav from set K. i Repeat the above steps until... Then, the first round of UAV collaborative trajectory planning at this time stamp is completed, and the process of UAV collaborative trajectory planning at the next time stamp begins;
[0067] S44: The UAV determines its flight path and intent;
[0068] After t seconds of collaborative trajectory planning and drone swarm traversal, the determination of whether the drones have reached the expected target formation position is made by the following formula:
[0069]
[0070] Among them, S i,t DP represents the state of drone i at time t. i S represents the target formation position of drone i. i,t and DP i Each vector consists of the UAV's longitude x, latitude y, and altitude h.
[0071] Compared with the prior art, the present invention, employing the above technical solution, has the following beneficial effects:
[0072] 1. This invention is based on the six degrees of freedom flexibility of a quadcopter UAV and the spatial position tracking of key collision nodes. It uses a deep reinforcement learning algorithm based on the Kronecker factor trust threshold executor / commentator to train the artificial intelligence agent, and reduces the complexity of UAV cooperative formation by designing heuristic algorithms.
[0073] 2. As per the instruction manual Figure 4 , 5 As shown, this invention can achieve coordinated formation of 25 aircraft in just 0.753 seconds, and the formation trajectory planning result is a four-dimensional trajectory with timestamp and longitude, latitude and altitude. At the same time, the trajectory granularity reaches 1 second, which can realize safe, efficient, fast and accurate UAV swarm coordinated formation trajectory planning. Attached Figure Description
[0074] Figure 1 A flowchart illustrating the overall steps of a rapid, coordinated, and precise formation method for a six-DOF quadcopter UAV.
[0075] Figure 2 A detailed technical roadmap for a six-degree-of-freedom rapid cooperative fine formation method for quadrotor UAVs;
[0076] Figure 3 A schematic diagram of the spatial rigid body coordinate rotation of a rotary-wing unmanned aerial vehicle;
[0077] Figure 4 A schematic diagram of a UAV's six-degree-of-freedom attitude flight;
[0078] Figure 5 Demonstrates a scenario for drone collaborative formation;
[0079] Figure 6 This is a diagram showing the effect of the UAV collaborative formation at the 5th second in this invention;
[0080] Figure 7 This is a rendering of the UAV collaborative formation at the 10th second in this invention;
[0081] Figure 8 This is a rendering of the UAV collaborative formation at 25 seconds in this invention;
[0082] Figure 9 This is a rendering of the UAV collaborative formation at 30 seconds in this invention;
[0083] Figure 10 This is a diagram showing the effect of the UAV collaborative formation at 50 seconds in this invention. Detailed Implementation
[0084] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.
[0085] like Figure 1 As shown, a method for rapid, coordinated, and precise formation of six-DOF quadcopter UAVs includes the following steps:
[0086] Step S1: Construct a six-degree-of-freedom dynamic model of a quadcopter UAV
[0087] A six-degree-of-freedom dynamic model of a quadrotor UAV is constructed. The quadrotor UAV is regarded as a centrally symmetric rigid body. Then the motion of the quadrotor UAV in space is mainly translation and rotation. The motion of the quadrotor UAV can be regarded as a combination of the three-dimensional spatial motion of the UAV center in the geographic coordinate system and the six-degree-of-freedom rotational motion of the UAV itself.
[0088] Considering only the translational motion of the drone, Newton's second law applies:
[0089]
[0090] Among them, F 合 For the net force acting on the drone, F T For the lift force on the drone, F G For the weight of the drone, F f Let be the air resistance experienced by the drone, a be the acceleration of the drone, x be the displacement of the drone, and m be the mass of the drone.
[0091] For the spatial motion of a UAV that involves both translation and rotation, it can be described by a combination of the motion of the UAV's geometric center in the geographic coordinate system and the Euler angles of the UAV's flight state, such as yaw, pitch, and roll angles. Figure 2 As shown;
[0092] Considering the gimbal lock in Euler rotation and the motion characteristics of the UAV, it is necessary to avoid pitch angles of ±90 degrees during UAV operation. The UAV rotation sequence is defined as Z-axis → Y-axis → X-axis, corresponding to yaw angle ψ, pitch angle θ, and roll angle φ. Therefore, the Euler rotation matrix R for the quadcopter UAV is... E :
[0093]
[0094] For a drone control system, the method to control the drone's translation and rotation is to adjust the rotational speed (angular velocity) of the four rotors, which yields the following force balance equation for the drone:
[0095]
[0096] Where x” represents the drone's acceleration along the x-axis, y” represents the drone's acceleration along the y-axis, and z” represents the drone's acceleration along the z-axis; x' represents the drone's velocity along the x-axis, y' represents the drone's velocity along the y-axis, and z' represents the drone's velocity along the z-axis; k x Let k be the air drag coefficient of the UAV along the x-axis. y k is the air drag coefficient in the y-axis direction. z R is the air drag coefficient in the z-axis direction; E Here, b is the Euler rotation matrix of the quadcopter UAV, b is the rotor lift coefficient of the UAV, and w is the Eulerian rotation matrix of the quadcopter UAV. i Let be the rotational speed of the i-th propeller of the UAV, g be the acceleration due to gravity, and U1 be the resultant lift force generated by the UAV rotor.
[0097] Therefore, the state transition equation for a quadcopter UAV is:
[0098]
[0099]
[0100]
[0101] Where x is the drone displacement in the x-axis direction, y is the drone displacement in the y-axis direction, and z is the drone displacement in the z-axis direction.
[0102] Step S2: Construct a quadcopter drone collision box model
[0103] like Figure 3As shown, since the geometric center of the centrally symmetric rotary-wing UAV is the Euler rotation center and the UAV is a rigid body structure, the shape characteristics of the UAV remain unchanged during the Euler rotation, meaning that the relative positions of the coordinates of each key point of the UAV with respect to the geometric center of the UAV remain unchanged. Therefore, the three-dimensional coordinates of the key coordinate points of the UAV shape in the geodetic coordinate system (old reference system) can be described as follows:
[0104]
[0105] Where ψ is the UAV yaw angle, θ is the UAV pitch angle, and φ is the UAV roll angle; [x',y',z'] T Let [x0, y0, z0] be the vector coordinates of the UAV in the coordinate system of the UAV's rotation center. T This represents the displacement of the UAV's rotation center relative to the origin of the geodetic coordinate system. This represents the matrix cross product. At this point, x, y, and z are the coordinates of the UAV's key points after displacement [x0, y0, z0]. T At this point, [x,y,z] T These are the geodetic coordinates of the UAV after rotating around the yaw axis, pitch axis, and roll axis [ψ,θ,φ].
[0106] Step S3: Agent training based on Kronecker factor trust threshold enforcer / commentator deep reinforcement learning; specifically including the following steps:
[0107] S31: Environmental Design
[0108] The reward function for the drone approaching the target formation position is designed as follows:
[0109]
[0110] Among them, S i,t Let DP be the state of the i-th UAV at time t. i S represents the position of the i-th UAV target formation; i,t and DP i Each of these is a column vector consisting of the UAV's longitude x, latitude y, and altitude h.
[0111] The reward function for maintaining a safe distance between two drones is designed as follows:
[0112] in
[0113] Among them, P it Let drone i form the key drone location set at second t, and
[0114] P it ={(x it ,yit ,z it )1,(x it ,y it ,z it )2,(x it ,y it ,z it )3,(x it ,y it ,z it )4,(x it ,y it ,z it )5}
[0115] Among them, (x it ,y it ,z it These are the five key collision detection points for quadcopter drones: the positions of the left front rotor, right front rotor, left rear rotor, and right rear rotor.
[0116] Ultimately, the two quadcopter drones received a single-step reward R at time t. t for:
[0117]
[0118] S32: Agent Design
[0119] Design an internal neural network, and denote the backpropagation error q-value (state-action pair value) of the fully connected internal UAV neural network q(S). t A t ,W);
[0120] The action selection strategy is designed so that the agent selects the action corresponding to the maximum q-value based on the e-greedy strategy, that is:
[0121]
[0122] Where e is a small positive number in the interval 0-1;
[0123] S33: Environment-Agent Interaction Mode
[0124] For each environment-agent interaction, the agent is based on a q-valued neural network q(S). t A t The e-greedy strategy selects action A. t Two quadcopter drones performed action A t A state transition is performed, and both drones enter the next state S. t+1 At the same time, the environment provides the agent with a state S. t Next, execute action A t Feedback reward value R tThis is used by the agent to update the hidden parameters of its internal q-value neural network.
[0125] S34: Iteration of Agent Action Value
[0126] Initialize the agent's internal neural network as an arbitrary weight matrix W and the optimal policy estimation neural network, with the optimal policy estimation neural network having a weight matrix θ, and denoted as the optimal policy estimate π(θ); initialize the implicit learning rate α. (θ) α (W) Initialize the discount factor γ;
[0127] intelligent agents according to (S t A t ;W) Status of the two drones S t Make action decision A t Afterwards, you will receive an environmental reward value R. t Let U←R+γq(S) t+1 A t ,W);
[0128] Now, we have two fully connected neural network weight matrices W and θ, denoted as mapping... Then, based on the Kronecker factor approximation curvature expression of these two fully connected neural networks, the corresponding Fisher matrix F can be derived. θ and F W :
[0129]
[0130]
[0131] Then, update To reduce -q(A) t |S t ;W)·π(A t |S t ;θ) value;
[0132] renew To reduce [Uq(A)] t |S t ;W)]·π(A t |S t The value of θ);
[0133] Update drone status, S t =S t+1 Then, proceed to the next iteration.
[0134] Step S4: Cooperative formation path planning, specifically including:
[0135] S41: Initialize the set of aircraft K = {uav1,,uav2,,uav3,,...,uav...} in the airspace. n};
[0136] S42: Prioritization of UAV pairs
[0137] Find the UAV of the drone i With UAV drones j Euclidean distance Distance(uav) i ,uav j ):
[0138] Distance(uav i ,uav j )=d ij =||[x i ,y i ]-[x j ,y j ]||2
[0139] Select the UAV pair i, j with the closest Euclidean distance:
[0140]
[0141] S43: UAV Cooperative Path Planning and UAV Swarm Traversal
[0142] After selecting the UAV pair i and j with the closest Euclidean distance, if uav i ∪uav j If ∈K, then the agent is uav i and UAV j Assign actions and remove uav from set K. i and UAV j ;like The intelligent agent is based on UAV. j The assigned action is UAV. i Assign actions and remove uav from set K. i Repeat the above steps until... Then, the first round of UAV collaborative trajectory planning at this time stamp is completed, and the process of UAV collaborative trajectory planning at the next time stamp begins;
[0143] S44: Drones achieve flight intent determination
[0144] After t seconds of collaborative trajectory planning and drone swarm traversal, the determination of whether the drones have reached the expected target formation position is made by the following formula:
[0145]
[0146] Among them, Si,t DP represents the state of drone i at time t. i S represents the target formation position of drone i. i,t and DP i Each vector consists of the UAV's longitude x, latitude y, and altitude h.
[0147] like Figure 4 The diagram shown illustrates the six-degree-of-freedom attitude flight of a UAV. Figure 5 The image shows a demonstration of a drone collaborative formation scenario.
[0148] like Figures 6-10 The image shown is a rendering of the drone collaborative formation. It can be seen that the formation trajectory planning result is a four-dimensional trajectory with timestamps and longitude, latitude, and altitude. At the same time, the trajectory granularity reaches 1 second, which can realize safe, efficient, fast, and accurate drone swarm collaborative formation trajectory planning.
[0149] The above are merely embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of the claims of the present invention pending approval.
Claims
1. A method for rapid, coordinated, and precise formation of six-degree-of-freedom quadrotor unmanned aerial vehicles (UAVs), characterized in that, The method specifically includes the following steps: Step S1: Construct a six-degree-of-freedom dynamic model of a quadcopter UAV; Step S2: Construct a collision box model of a quadcopter drone; Step S3: Agent training based on Kronecker factor trust threshold enforcer / commentator deep reinforcement learning; Step S4: Collaborative formation flight path planning based on artificial intelligence agents, specifically including: Step S41: Initialize the set of aircraft in the airspace ; Step S42: Prioritize drone pairs based on the Euclidean distance between drones; Step S43: The UAVs perform cooperative trajectory planning and swarm traversal, selecting the UAV pair with the closest Euclidean distance. , Afterwards, if Then the intelligent agent is and Assign actions and from the set Delete and ;like The intelligent agent is based on The assigned actions are Assign actions and from the set Delete Repeat the above steps until... If the drone collaborative trajectory planning for that time stamp is completed, the drone collaborative trajectory planning process for the next time stamp will begin. Step S44: The UAV determines its flight path intention; after t seconds of collaborative flight path planning and UAV swarm traversal, it is determined whether the UAV has reached the expected target formation position. The determination formula is: , in, Indicates drone exist The state at the second, Indicates drone The target formation position, and All are longitudes of drones UAV latitude drone altitude The column vector formed by these.
2. The method for rapid, coordinated, and precise formation of a six-degree-of-freedom quadcopter UAV according to claim 1, characterized in that, The steps The specific content of constructing a six-degree-of-freedom dynamic model for a quadcopter UAV is as follows: If we consider a quadcopter drone as a centrally symmetric rigid body, then its motion in space mainly consists of translation and rotation. Considering only the drone's translation, we have Newton's second law: , in, The combined force acting on the drone The lift force experienced by the drone The force of gravity acting on the drone Due to air resistance experienced by the drone, Accelerate the drone For the displacement of the drone, For drone quality; Considering the gimbal lock in Euler rotation and the motion characteristics of the UAV, it is necessary to avoid situations where the pitch angle of the UAV is ±90 degrees during operation. Therefore, the UAV rotation sequence is defined as follows: axis axis axis, corresponding to yaw angle Pitch angle and roll angle Therefore, the Euler rotation matrix of the quadcopter drone : , For a drone control system, the method to control the drone's translation and rotation is to adjust the rotational speed of the four rotors, which yields the following force balance equation for the drone: , in, for UAV acceleration in the axial direction, for UAV acceleration in the axial direction, The acceleration of the UAV in the z-axis direction; for The speed of the drone in the axial direction, for The speed of the drone in the axial direction, for The speed of the drone in the axial direction; For drones Axial drag coefficient, for Axial drag coefficient, for Axial drag coefficient; For quadcopter UAVs, the Euler rotation matrix is used. The lift coefficient of the UAV rotor. For drones propeller speed, It is the acceleration due to gravity. To generate a resultant lift force for the drone rotor; Therefore, the state transition equation for a quadcopter UAV is: , , , in, for Displacement of the UAV along the axial direction for Displacement of the UAV along the axial direction for Displacement of the UAV along the axial direction.
3. The method for rapid, coordinated, and precise formation of a quadcopter UAV with six degrees of freedom, as described in claim 1, is characterized in that... The steps The specific details of constructing the quadcopter drone collision box model are as follows: The three-dimensional coordinates of the key coordinate points of the UAV shape in the geodetic coordinate system are described as follows: , in, For the yaw angle of the drone, For the drone's pitch angle, For the roll angle of the drone; Here are the vector coordinates of the UAV in the coordinate system of the UAV's rotation center. This represents the displacement of the UAV's rotation center relative to the origin of the geodetic coordinate system. This represents the matrix cross product; in this case, , , It means that the coordinates of the key points of the drone after displacement ,at this time, It is the rotation of the UAV around the yaw axis, pitch axis, and roll axis. The geodetic coordinates after that.
4. The method for rapid, coordinated, and precise formation of a six-degree-of-freedom quadcopter UAV according to claim 1, characterized in that, The specific steps of agent training based on Kronecker factor trust threshold executor / commentator deep reinforcement learning in step S3 can be divided into: S31: Environmental design; The reward function for the drone approaching the target formation position is designed as follows: , in, For the first The first drone Current state For the first Position of the drone target formation; and All are longitudes of drones UAV latitude drone altitude The column vector formed; The reward function for maintaining a safe distance between two drones is designed as follows: , in, For drones In the A key drone location set is formed in seconds, and , in, These are the five key collision detection points for quadcopter drones: the positions of the left front rotor, right front rotor, left rear rotor, and right rear rotor. Ultimately, the two quadcopter drones received a single-step reward R at time t. t for: , S32: Agent Design; Design an internal neural network to record the backpropagation state-action pair values of the fully connected error within the UAV. neural networks ; Design an action selection strategy; the agent selects the action with the maximum value based on the e-greedy strategy. The action corresponding to the value, namely: , in, It is a relatively small positive number in the interval 0-1; S33: Environment-Agent Interaction Mode; For each environment-agent interaction, the agent is based on... Value Neural Network and e-greedy strategy for action selection Two quadcopter drones performed the maneuver. A state transition is performed, and both drones enter the next state. At the same time, the environment provides the agent with a state. Next action Feedback reward value To provide internal updates for intelligent agents Value neural network hidden parameters are used; S34: Iteration of agent action value; Initialize the agent's internal neural network as an arbitrary weight matrix. And the optimal policy estimation neural network, and the weight matrix of the optimal policy estimation neural network. Let the optimal policy estimate be... Initialize the implicit learning rate , Initialize discount factor ; intelligent agents according to Status of the two drones Make action decisions Afterwards, environmental reward points are obtained. ,remember ; Mapping , Then, based on the Kronecker factor approximation curvature expression of these two fully connected neural networks, the corresponding Fisher matrix can be derived. and : , , Then, update To reduce value; renew To reduce The value; Update drone status. Then, proceed to the next iteration.
Citation Information
Patent Citations
Four-rotor unmanned aerial vehicle formation cooperative maneuvering control method
CN114911265A
Formation control method of four-rotor unmanned aerial vehicle cluster system
CN115993846A