Unmanned aerial vehicle cluster intelligent flight path planning method based on center estimation and repulsive force mechanism
By introducing a multi-agent near-end strategy optimization algorithm based on central estimation and repulsion mechanisms into the drone cluster, the problem that drone clusters are difficult to maintain cluster formation and avoid collisions in complex environments is solved, and more efficient track planning performance is achieved.
Patent Information
- Application Number
- CN202510146913.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-13
AI Technical Summary
It is difficult for drone clusters to maintain cluster formation and avoid collisions in complex environments, and the existing technology is insufficient in flexibility and adaptability.
A multi-agent proximal strategy optimization algorithm based on central estimation and repulsion force mechanism is adopted, and reward functions are designed through deep reinforcement learning methods. Combined with the Actor-Critic network architecture, a cluster center estimation and repulsion force mechanism are introduced to optimize the track planning of the drone cluster.
It improves the track planning performance of the drone cluster in complex environments, ensures the maintenance of cluster formation and the avoidance of collisions, and improves the flexibility and adaptability of the system.
Smart Images

Figure CN119987403A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of unmanned aerial vehicles, and in particular relates to an intelligent trajectory planning method for a cluster of unmanned aerial vehicles based on center estimation and repulsive force mechanism. Background Art
[0002] UAV swarms have broad application prospects in communication support, search and rescue, and agricultural operations, and have therefore attracted widespread attention from academia and industry. In addition to communication support, the coordinated control of UAV swarms is also crucial for the successful operation of these applications. Among many technologies, swarm motion control of UAV swarms has been regarded as a basic requirement. Swarm motion is a ubiquitous collective behavior that describes group coordination and motion in nature. As early as the 1980s, Reynolds proposed the famous Boids model, which contains three basic flight motion control heuristic rules, namely cohesion, separation, and alignment. Subsequently, some algorithms for swarm motion of multi-agent systems were borrowed from the Boids model.
[0003] At the same time, there are endless studies on the problem of swarm motion control. Among them, traditional model-based methods, such as model predictive control and consensus theory, often rely on precise physical models or predefined rules, which limits their flexibility and adaptability. Unlike traditional methods, reinforcement learning algorithms are able to solve problems by directly participating in the task environment, even when the precise model is unknown. This ability to operate without a precise model is called "model-free". In particular, multi-agent deep reinforcement learning, which combines the advantages of deep learning and reinforcement learning, has shown its effectiveness in controlling swarm motion in complex dynamic environments. In the multi-agent deep reinforcement learning paradigm, the swarm motion task is called the environment. When agents interact with the environment, they collect information and improve their decision-making ability based on the reward and punishment feedback they receive. This iterative learning process enables agents to adapt to complex scenarios, optimize their cluster size, and effectively complete designated tasks. Summary of the invention
[0004] The purpose of the present invention is to propose an intelligent trajectory planning method for drone swarms based on center estimation and repulsive force mechanism for the problem of drone swarm trajectory planning, so as to effectively improve the performance of drone swarm trajectory planning in maintaining cluster formation and avoiding collision in complex environments; in order to achieve this purpose, the steps adopted by the present invention are:
[0005] Step 1: Construct the observation space and action space of the UAV cluster trajectory planning problem; the observation space consists of four parts, including target information, neighbor information, obstacle information and cluster center information; the action space is the speed and direction of the UAV; the specific construction method of the observation space and action space is:
[0006] (1) Observation space: For UAV i with a time step of , its observation value o i,t It consists of four parts, namely target information, neighbor information, obstacle information and cluster center information; further, the four parts of the observation space are specifically described as:
[0007] Target Information where g i is the absolute position of target i;
[0008] Neighbor Information Where K = || N i || is the neighbor set N i The size of; Assuming that the drone maintains a fixed number of neighbors, K is set to a fixed value;
[0009] Obstacle information When the distance between the drone and the obstacle is within the sensing range, the drone can use the sensor to accurately determine the location of the obstacle;
[0010] Cluster Center Information In order to maintain a compact cluster formation, the present invention introduces a cluster center. The cluster center of each UAV distribution estimates its local cluster center as
[0011] In summary, the environmental observation information perceived by UAV i at time step t can be expressed as
[0012] (2) Action space: A continuous action space is used to control the smooth motion of the UAV. Considering the two-dimensional plane mission area, the action a of UAV i at time step t i,t Designed for a i,t :=(ψ i,t , v i,t ), where ψ i,t ∈[0, 2π] represents the velocity direction, V i,t ∈[v min , v max ] indicates the speed; in the simulation environment, the action will be converted into speed and calculated as the next target position, so the drone will obtain the next state through the flight control system;
[0013] Step 2: Design a reward function for the deep reinforcement learning method for the UAV swarm trajectory planning problem; the reward function consists of seven parts, namely, approach target reward, neighbor avoidance reward, obstacle avoidance reward, aggregation reward, step reward, smoothness reward and boundary reward. The final reward function is a linear coupling of the above seven. The specific reward function design criteria are:
[0014] (1) Rewards for approaching the target To guide the UAV to approach the target effectively, each step should maximize the flight distance in the target direction; therefore, the reward obtained by UAV i when approaching the target at time step is defined as
[0015]
[0016] Among them, d step is the distance the drone takes each step;
[0017] (2) Neighbor Avoidance Reward Ensure that the drone maintains a safe distance from neighboring drones to avoid collisions; define the penalty for drone i to collide with its neighbors at the time step as
[0018]
[0019] Among them, ρ uav is the radius of the drone, N i is the set of neighbor nodes of the drone; the safety distance of the drone is called d su , determined by the number of drones;
[0020] (3) Obstacle Avoidance Reward Guide the UAV to keep a safe distance from surrounding obstacles; define the penalty for UAV i to collide with an obstacle at time step as
[0021]
[0022] Among them, d so is the safe distance from the drone to the obstacle, ρ ost is the radius of the obstacle, d i,m is the distance from the drone to the obstacle m∈B i ;
[0023] (4) Aggregation Rewards Guide all drones to fly towards the cluster center of the drone swarm to maintain a dense formation; the aggregate reward of drone i at time step , is defined as
[0024]
[0025] in, is the global cluster center location of the drone swarm;
[0026] (5) Step Rewards This reward function is used to promote seamless movement of the UAV, and its formula is:
[0027]
[0028] (6) Smooth Rewards The purpose of this reward function is to promote the smooth movement of the drone, and its formula is:
[0029]
[0030] (7) Boundary Rewards In order to prevent the drone from flying over the border and moving along the border, the present invention designs a circular border reward; the calculation formula of the border reward in the x-axis direction is:
[0031]
[0032] Boundary reward on the y-axis and has parity, so we only need to x in lim Replace with y lim You can calculate Therefore, the present invention defines Therefore, the total reward function r i,t The calculation formula is
[0033]
[0034] Where k represents the seven reward situations mentioned above;
[0035] Step 3: Design a multi-agent proximal strategy optimization algorithm framework based on center estimation and repulsion mechanism; the deep reinforcement learning network adopts the Actor-Critic network architecture, and the reinforcement learning algorithm is based on the multi-agent proximal strategy optimization algorithm, introducing new environmental constraints, including cluster center estimation and repulsion mechanism;
[0036] According to the joint action a t and repulsive mechanism, the current state s t Will be updated to s through the twin simulation environment t+1 ; Joint action a t The repulsive force action is superimposed to obtain the final control output; the cluster center estimated by the cluster center estimation method is the input of the decision, and the calculated global center is used to design the reward function; using the gradient and Adjust the actor and critic networks to refine the model;
[0037] Furthermore, the cluster center estimation method and repulsion mechanism proposed in the present invention are specifically defined as:
[0038] (1) Cluster center estimation
[0039] Global Cluster Center The position of can be expressed as the average position of all drones:
[0040]
[0041] Therefore, the kinematic model of the global cluster center is
[0042]
[0043] Based on the following facts:
[0044]
[0045] in is the velocity of the global cluster center;
[0046] Each drone contacts the tracking system to obtain its own position and sends it to neighboring drones; the core part of cluster center estimation is the design of the controller, where the state input includes its local position p i,t , the estimated location of the local cluster center and the central location of the neighborhood j∈N i ; The output of the controller is based on p i,t , and The velocity of the local cluster center is obtained Therefore, the consensus algorithm for the center position of the drone cluster is designed as follows:
[0047]
[0048] Among them, N i is the neighbor set of drone i, and the set size is K = || B i ||; β is the speed matching term, taking the maximum speed of the required path, that is, β=(v max , v max );
[0049] Discretizing equation (13), UAV i updates its estimated cluster center position at each time step as
[0050]
[0051] The vector form of the consensus algorithm is
[0052]
[0053] Among them, P t =(p 1,t , ..., p N,t ), L is the Laplacian matrix of the graph, and the corresponding eigenvector is 1 = (1, ..., 1) T , is the Kronecker product;
[0054] Each drone independently calculates the local cluster center And exchange cluster centers with each other, and finally reach a consensus on the location; it can be proved that the local center With the global center Asymptotic convergence;
[0055] (2) Repulsive force mechanism
[0056] First, the normal velocity of the repulsive velocity is defined according to Hooke's law; for UAV i and its neighboring UAV j, the normal velocity is set to
[0057]
[0058] Among them, p ij =p i,t -P j,t ,||p ij || is the distance between UAVs i and j; only when the distance between UAVs is less than 1.5·d su Calculate the repulsive velocity when
[0059] The present invention introduces another tangential velocity It needs to meet the following conditions:
[0060]
[0061] Then the final repulsive velocity between drones i and j It can be calculated as
[0062]
[0063] Similarly, the present invention regards obstacles as virtual drones; when the safe distance between the drone and the obstacle is within 1.5 d so Within, the same repulsive speed will be assigned to them Therefore, the total repulsive force between drone i and neighboring drones and obstacles is for
[0064]
[0065] At this time, the speed v of drone i i,t According to the repulsive speed Updated again to
[0066]
[0067] Finally, the speed of the drone is limited to a maximum value v max ,Right now
[0068]
[0069] The present invention regards cluster center estimation and repulsion mechanism as additional rules of the environment; when designing the reward function, the present invention takes cluster center into consideration; at the same time, the speed obtained from the strategy is superimposed with the influence of repulsion speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 It is the trajectory planning system model of the UAV cluster proposed by the present invention;
[0071] Figure 2 It is a multi-agent proximal strategy optimization algorithm framework based on center estimation and repulsion mechanism proposed by the present invention;
[0072] Figure 3 This is a comparison chart of the simulation results of the average reward convergence curve of the present invention and the existing algorithm;
[0073] Figure 4 It is a comparison chart of the average success rate convergence curve simulation results of the present invention and the existing algorithm. DETAILED DESCRIPTION
[0074] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.
[0075] The UAV swarm intelligent trajectory planning method based on center estimation and repulsive force mechanism proposed in the present invention first sets the following operating conditions:
[0076] 1. As attached Figure 1 As shown in the figure, the ultimate goal of this method is to navigate multiple drones from the starting area to the target area within a limited time. In this real-time path planning process, each UAV must avoid collisions with neighboring UAVs and obstacles. At the same time, all drones should be as close as possible to the center of the flight group to maintain a compact formation. Of course, the drones must also meet energy and area boundary constraints. Each drone has its own target position, and all target positions are located in the target area. The starting position and target position are set at the initialization of the simulation and remain unchanged in a single simulation.
[0077] 2. Two types of obstacles are considered, namely static obstacles and dynamic obstacles. The initial positions of these obstacles are randomly and uniformly distributed in the mission area. Dynamic obstacles are modeled as Brownian motion under the maximum speed constraint. Safety distances are set for obstacles and neighboring drones respectively. Once it is detected that the distance between a drone and an obstacle or a neighboring drone is less than the safety distance, a collision is considered to have occurred, thus terminating the motion. In addition, twice the safety distance is defined as the sensing distance. If the drone judges that the distance is less than its sensing distance, it will perform some necessary precautions.
[0078] 3. The mission area under consideration has N drones, N targets, and M obstacles. Each drone obtains its position and velocity through a combination of an inertial measurement unit (IMU) and GPS navigation. The drones exchange state information through self-organizing network communication. A discrete time scale is used, and the time interval is recorded as Δt. The position and velocity of the drone are recorded as P, t =(p 1,t , p 2,t ,…,,p N,t ) and V t =(v 1,t , v 2,t ,...,,v N,t ). Similarly, the positions of obstacles and targets are represented by b t =(b 1,t , b 2,t , ..., b M,t ) and g=(g1,g2,...,g N ), where obstacles include M1 static obstacles and M2 dynamic obstacles. The UAV starts from the initial position according to the predetermined distribution and navigates to the destination according to the real-time dynamic planning trajectory. Each UAV can be in [0, v max ] range, where v max is the maximum flight speed.
[0079] Based on the above conditions, the intelligent trajectory planning method for drone swarm based on center estimation and repulsion mechanism proposed in this invention has been implemented in Linux operating system Ubuntu and physical system, and the experimental results prove the effectiveness of this method. The specific implementation steps are as follows:
[0080] Step 1: Construct the observation space and action space of the UAV cluster trajectory planning problem. The observation space consists of four parts, including target information, neighbor information, obstacle information and cluster center information; the action space is the speed and direction of the UAV.
[0081] (1) Observation space: Observation data includes various possibilities of the agent and its surrounding environment. However, swarm motion is a partially observable Markov decision process (POMDP). Due to the limitations of sensors and communications, each UAV has a partial understanding of the environment. For UAV i with a time step of , its observation value o i,t It consists of four parts, namely target information, neighbor information, obstacle information and cluster center information. In order to improve the generalization ability of the trajectory planning decision model, relative position is used to describe the above information:
[0082] Target Information where g i is the absolute position of target i.
[0083] Neighbor Information Where K = || N i || is the neighbor set N i Assuming that the drone maintains a fixed number of neighbors, K is set to a fixed value.
[0084] Obstacle information When the distance between the drone and the obstacle is within the sensing range, the drone can use the sensor to accurately determine the location of the obstacle.
[0085] Cluster Center Information In order to maintain a compact cluster formation, the present invention introduces a cluster center. The cluster center of each UAV distribution estimates its local cluster center as
[0086] In summary, the environmental observation information perceived by UAV i at time step t can be expressed as
[0087] (2) Action space: A continuous action space is used to control the smooth motion of the UAV. Considering a two-dimensional plane mission area, the action a of UAV i at time step t i,t Designed for a i,t :=(ψ i,t , v i,t ), where ψ i,t ∈[0, 2π] represents the velocity direction, v i,t ∈[v min , V max ] indicates the speed. In the simulation environment, the action will be converted into speed and calculated as the next target position, so the drone will get the next state through the flight control system.
[0088] Step 2: Design the reward function of the deep reinforcement learning method for the UAV swarm trajectory planning problem. The reward function consists of seven parts, namely, the target approach reward, the neighbor avoidance reward, the obstacle avoidance reward, the aggregation reward, the step reward, the smoothness reward and the boundary reward. The final reward function is the linear coupling of the above seven parts.
[0089] In order to design a better training reward function, the task objectives and environmental constraints are integrated comprehensively. Specifically, the reward function is designed as follows.
[0090] (1) Rewards for approaching the target To guide the drone to approach the target effectively, each step should maximize the flight distance in the target direction. Therefore, the reward obtained by drone i when approaching the target at time step is defined as
[0091]
[0092] Among them, d step is the distance the drone travels per step.
[0093] (2) Neighbor Avoidance Reward Ensure that the drone maintains a safe distance from neighboring drones to avoid collision. Define the penalty for drone i to collide with its neighbor at time step t as
[0094]
[0095] Among them, ρ uav is the radius of the drone, N i is the set of neighbor nodes of the drone. The safety distance of the drone is called d su , which is determined by the number of drones. Typically, it contains an area slightly larger than the minimum area required to accommodate N drones without any collisions.
[0096] (3) Obstacle Avoidance Reward Guide the UAV to keep a safe distance from surrounding obstacles. The penalty for UAV i colliding with an obstacle at time step t is defined as
[0097]
[0098] Among them, d so is the safe distance from the drone to the obstacle, ρ ost is the radius of the obstacle, d i,m is the distance from the drone to the obstacle m∈B i .
[0099] (4) Aggregation Rewards Guide all drones to fly towards the center of the drone swarm to maintain a dense formation. The aggregate reward of drone i at time step t is defined as
[0100]
[0101] in, It is the global cluster center location of the drone swarm.
[0102] (5) Step Rewards This reward function is used to promote seamless movement of the UAV, and its formula is:
[0103]
[0104] (6) Smooth Rewards The purpose of this reward function is to promote the smooth movement of the drone, and its formula is:
[0105]
[0106] (7) Boundary Rewards In order to prevent the drone from flying over the border and moving along the border, the present invention designs a circular border reward. The calculation formula of the border reward in the x-axis direction is:
[0107]
[0108] Boundary reward on the y-axis and has parity, so we only need to x in lim Replace with y lim You can calculate Then, the present invention defines Therefore, the total reward function r i,t The calculation formula is
[0109]
[0110] Where k represents the seven reward situations mentioned above.
[0111] Step 4: Design a multi-agent proximal policy optimization (MAPPO) algorithm framework based on cluster center estimation and repulsion mechanism. The deep reinforcement learning network adopts the actor-critic network architecture. The reinforcement learning algorithm is based on the MAPPO algorithm and introduces new environmental constraints, including cluster center estimation and repulsion mechanism.
[0112] Attached Figure 2The framework of the multi-agent proximal strategy optimization algorithm based on center estimation and repulsion mechanism is shown. According to the joint action a t and repulsive mechanism, the current state s t Will be updated to s through the simulation environment t+1 The repulsive force mechanism is an additional constraint on the simulation training environment. t The repulsive action is superimposed to obtain the final control output. The purpose of introducing the cluster center estimation method is to enhance the clustering behavior. The estimated cluster center is the input of the decision, while the calculated global center is used to design the reward function. Therefore, the gradient and The actor and critic networks are fine-tuned to refine the model.
[0113] The cluster center estimation method and repulsion mechanism proposed by the present invention will be described in detail below.
[0114] (1) Cluster center estimation
[0115] During the training phase, the present invention adds the estimated cluster center position of the drone to the observed state, while the global cluster center position is used for the calculation of the reward function. In particular, during the distributed execution phase, the drone cannot obtain the global center information. In the drone swarm, consensus can be reached on the aggregated information through the position alignment between adjacent drones. Each drone can collect such data autonomously by interacting with nearby neighbors. With this shared information, the drone can maintain connection with other drones by moving closer to the cluster center position. This mechanism greatly enhances the cohesion of the flight cluster.
[0116] Global Cluster Center The position of can be expressed as the average position of all drones:
[0117]
[0118] Therefore, the kinematic model of the global cluster center is
[0119]
[0120] Based on the following facts:
[0121]
[0122] in is the velocity of the global cluster center.
[0123] Each drone contacts the tracking system to obtain its own position and sends it to neighboring drones. The core part of cluster center estimation is the design of a controller, where the state input includes its local position p i,t, the estimated location of the local cluster center and the central location of the neighborhood j∈N i The output of the controller is based on p i,t , and The velocity of the local cluster center is obtained Therefore, the consensus algorithm for the center position of the drone cluster is designed as follows:
[0124]
[0125] Among them, N i is the neighbor set of drone i, and the set size is K = || B i ||. β is the speed matching term, taking the maximum speed of the required path, that is, β=(v max , v max ). Calculated center speed Must be limited to maximum speed range.
[0126] Discretizing equation (34), UAV i updates its estimated cluster center position at each time step as
[0127]
[0128] The vector form of the consensus algorithm is
[0129]
[0130] Among them, P t =(p 1,t , ..., p N,t ), L is the Laplacian matrix of the graph, and the corresponding eigenvector is 1 = (1, ..., 1) T , is the Kronecker product.
[0131] Each drone independently calculates the local cluster center And exchange cluster centers with each other, and finally reach a consensus on the location. It can be proved that the local center With the global center Asymptotic convergence. Incorporating local cluster centers into local observations can enrich the final policy. While the estimated local centers may deviate from the exact value of the global center provided by the simulator, the trained policy model usually shows the ability to generalize to this error within a specified range, which is regulated by the weight hyperparameter of the cluster center reward function.
[0132] (2) Repulsive force mechanism
[0133] Although the trained strategy can prevent collisions with obstacles, adding a repulsion component to the obstacles can further enhance the safety measures of the drone. To this end, the present invention designs a repulsion function to enhance the anti-collision capability between the drone and obstacles.
[0134] First, the normal velocity of the repulsive velocity is defined according to Hooke's law. For drone i and its neighboring drone j, the normal velocity is set to
[0135]
[0136] Among them, p ij =p i,t -p j,t ,||p ij || is the distance between drones i and j. Only when the distance between drones is less than 1.5·d su The repulsion speed is calculated when , and the size of the coefficient is determined by the weight of the total reward in the reward function. After testing, it is finally selected as 0.06.
[0137] Due to the normal velocity It is completely dependent on the relative position and thus causes oscillations. To reduce these oscillations, the present invention introduces another tangential velocity It needs to meet the following conditions:
[0138]
[0139] This means that the tangential velocity and the central velocity are equal in magnitude, perpendicular in direction, and converge toward the target at the same time. Then the final repulsive velocity between drones i and j It can be calculated as
[0140]
[0141] Similarly, the present invention regards obstacles as virtual drones. When the safe distance between the drone and the obstacle is within 1.5 d so Within, the same repulsive speed will be assigned to them Therefore, the total repulsive force between drone i and neighboring drones and obstacles is for
[0142]
[0143] At this time, the speed v of drone i i,t According to the repulsive speed Updated again to
[0144]
[0145] Finally, the speed of the drone is limited to a maximum value v max ,Right now
[0146]
[0147] The present invention considers cluster center estimation and repulsion mechanism as additional rules of the environment. When designing the reward function, the present invention takes cluster center into consideration. At the same time, the speed obtained from the strategy is superimposed with the influence of repulsion speed.
[0148] The performance of the intelligent trajectory planning method for drone swarm based on center estimation and repulsive force mechanism proposed in the present invention has been simulated and verified in Linux operating system Ubuntu and physical system. In physical space, the drone is a self-developed quad-rotor drone, using Pixhawk firmware and ArduPilot software as the flight controller. The radius of the drone is 0.7 meters, the payload is 1.5 kilograms, the maximum flight speed is 15.6 meters per second, and the maximum flight time is 30 minutes. The quad-rotor drone is equipped with Raspberry Pi as an onboard computer for control and decision-making. The ground control system (GCS) is used to centrally control the drone swarm, such as takeoff and landing, and request, receive and forward data. 5G communication is used between GCS and drones, and self-organizing network communication is used between drones. In digital space, the Gazebo-ArduPilot framework is used to construct a simulation model of drone swarm and flight environment. The present invention establishes a square simulation scene with a side length of 200 meters, and the obstacles are simulated as cylinders with a radius of 6 meters, and the maximum speed of dynamic obstacles is 0.5 meters per second. The communication between drones uses a protocol stack developed based on the EXata simulator.
[0149] In the decision model, the present invention uses the PyTorch library and Python language to implement the MAPPO-based learning method. All neural networks are optimized using the Adam optimizer and the ReLU activator, where the learning rate of the actor network is 5e-4 and the learning rate of the critic network is 1e-3. The network initialization model is an orthogonal model. In general, the convergence speed and stability of the learning algorithm are mainly reflected by the learning curve. To this end, we trained three different algorithms, namely:
[0150] MAPPO: basic MAPPO algorithm;
[0151] MAPPO-C: MAPPO algorithm with cluster center estimation.
[0152] MAPPO-CR: MAPPO algorithm with cluster center estimation and repulsion mechanism.
[0153] The total number of training steps is 1e7, the round length is 200, and the time interval is 1s. The number of static and dynamic obstacles in the environment is 5. The number of drones is 12. Figure 3 and attached Figure 4 The average reward convergence curve and average success rate convergence curve of the multi-agent proximal strategy optimization algorithm based on center estimation and repulsion mechanism proposed in this invention are compared with the existing algorithm. Figure 3 and attached Figure 4 It can be seen from the simulation results that the UAV cluster intelligent trajectory planning method based on center estimation and repulsion mechanism proposed in the present invention can achieve better coordination effect than the existing UAV cluster trajectory planning method.
[0154] The contents not described in detail in the present application belong to the prior art known to the professional and technical personnel in this field.
Claims
1. An intelligent trajectory planning method for UAV swarm based on center estimation and repulsive force mechanism, the steps adopted are: Step 1: Construct the observation space and action space of the UAV swarm trajectory planning problem; The observation space consists of four parts, including target information, neighbor information, obstacle information and cluster center information; The action space is the speed and direction of the drone; the specific construction method of the observation space and action space is: (1) Observation space: For UAV i with a time step of t, its observation value o i,t It consists of four parts, namely target information, neighbor information, obstacle information and cluster center information; further, the four parts of the observation space are specifically described as: Target Information where g i is the absolute position of target i; Neighbor Information Where K = || N i || is the neighbor set N i The size of; Assuming that the drone maintains a fixed number of neighbors, K is set to a fixed value; Obstacle information When the distance between the drone and the obstacle is within the sensing range, the drone can use the sensor to accurately determine the location of the obstacle; Cluster Center Information In order to maintain a compact cluster formation, the present invention introduces a cluster center. The cluster center of each UAV distribution estimates its local cluster center as In summary, the environmental observation information perceived by UAV i at time step t can be expressed as (2) Action space: A continuous action space is used to control the smooth motion of the UAV. Considering the two-dimensional plane mission area, the action a of UAV i at time step t i,t Designed for a i,t :=(ψ i,t , v i,t ), where ψ i,t ∈[0, 2π] represents the velocity direction, v i,t ∈[v min , v max ] indicates the speed; in the simulation environment, the action will be converted into speed and calculated as the next target position, so the drone will obtain the next state through the flight control system; Step 2: Design a reward function for the deep reinforcement learning method for the UAV swarm trajectory planning problem; the reward function consists of seven parts, namely, approach target reward, neighbor avoidance reward, obstacle avoidance reward, aggregation reward, step reward, smoothness reward and boundary reward. The final reward function is a linear coupling of the above seven. The specific reward function design criteria are: (1) Rewards for approaching the target To guide the UAV to approach the target efficiently, each step should maximize the flight distance in the target direction; therefore, the reward obtained by UAV i when it approaches the target at time step t is defined as Among them, d step is the distance the drone takes each step; (2) Neighbor Avoidance Reward Ensure that the drone maintains a safe distance from neighboring drones to avoid collisions; define the penalty for drone i when it collides with its neighbors at time step t as Among them, ρ uav is the radius of the drone, N i is the set of neighbor nodes of the drone; the safety distance of the drone is called d su , determined by the number of drones; (3) Obstacle Avoidance Reward Guide the UAV to keep a safe distance from surrounding obstacles; define the penalty for UAV i to collide with an obstacle at time step t as Among them, d so is the safe distance from the drone to the obstacle, ρ ost is the radius of the obstacle, d i,m is the distance from the drone to the obstacle m∈B i ; (4) Aggregation Rewards Guide all drones to fly towards the cluster center of the drone swarm to maintain a dense formation; the aggregate reward of drone i at time step t is defined as in, is the global cluster center location of the drone swarm; (5) Step Rewards This reward function is used to promote seamless movement of the UAV, and its formula is: (6) Smooth Rewards The purpose of this reward function is to promote the smooth movement of the drone, and its formula is: (7) Boundary Rewards In order to prevent the drone from flying over the border and moving along the border, the present invention designs a circular border reward; the calculation formula of the border reward in the x-axis direction is: Boundary reward on the y-axis and has parity, so we only need to x in lim Replace with y lim You can calculate Therefore, the present invention defines Therefore, the total reward function r i,t The calculation formula is Where k represents the seven reward situations mentioned above; Step 3: Design a multi-agent proximal strategy optimization algorithm framework based on center estimation and repulsion mechanism; the deep reinforcement learning network adopts the Actor-Critic network architecture, and the reinforcement learning algorithm is based on the multi-agent proximal strategy optimization algorithm, introducing new environmental constraints, including cluster center estimation and repulsion mechanism; According to the joint action a t and exclusion mechanism, the current state s t Will be updated to s through the twin simulation environment t+1 ; Joint action a t The repulsive force action is superimposed to obtain the final control output; the cluster center estimated by the cluster center estimation method is the input of the decision, and the calculated global center is used to design the reward function; using the gradient and Adjust the actor and critic networks to refine the model; Furthermore, the cluster center estimation method and exclusion mechanism proposed in the present invention are specifically defined as: (1) Cluster center estimation Global Cluster Center The position of can be expressed as the average position of all drones: Therefore, the kinematic model of the global cluster center is Based on the following facts: in is the velocity of the global cluster center; Each drone contacts the tracking system to obtain its own position and sends it to neighboring drones; the core part of cluster center estimation is the design of the controller, where the state input includes its local position p i,t , the estimated location of the local cluster center and the central location of the neighborhood The output of the controller is based on p i,t , and The velocity of the local cluster center is obtained Therefore, the consensus algorithm for the center position of the drone cluster is designed as follows: Among them, N i is the neighbor set of drone i, and the set size is K = || B i ||; β is the speed matching term, taking the maximum speed of the required path, that is, β=(v max , v max ); Discretizing equation (13), UAV i updates its estimated cluster center position at each time step as The vector form of the consensus algorithm is Among them, P t =(p 1,t ,…,p N,t ), L is the Laplacian matrix of the graph, and the corresponding eigenvector is 1 = (1, ..., 1) T , is the Kronecker product; Each drone independently calculates the local cluster center And exchange cluster centers with each other, and finally reach a consensus on the location; it can be proved that the local center With the global center Asymptotic convergence; (2) Exclusion mechanism First, the normal velocity of the repulsive velocity is defined according to Hooke's law; for UAV i and its neighboring UAV j, the normal velocity is set to Among them, p ij =p i,t -p j,t ,||p ij || is the distance between UAVs i and j; only when the distance between UAVs is less than 1.5·d su Calculate the repulsive velocity when The present invention introduces another tangential velocity It needs to meet the following conditions: Then the final repulsive velocity between drones i and j It can be calculated as Similarly, the present invention regards obstacles as virtual drones; when the safe distance between the drone and the obstacle is within 1.5 d so Within, the same repulsive speed will be assigned to them Therefore, the total repulsive force between drone i and neighboring drones and obstacles is for At this time, the speed v of drone i i,t According to the repulsive speed Updated again to Finally, the speed of the drone is limited to a maximum value v max ,Right now The present invention regards cluster center estimation and repulsion mechanism as additional rules of the environment; when designing the reward function, the present invention takes cluster center into consideration; meanwhile, the speed obtained from the strategy is superimposed with the influence of repulsion speed.