Multi-unmanned aerial vehicle cooperative path planning method based on improved MATD3 algorithm

By improving the MATD3 algorithm and combining the PER mechanism and SA algorithm to optimize the experience replay buffer, an Actor-Critic network was constructed, which solved the computational complexity and stability problems in multi-UAV cooperative path planning and achieved efficient and stable multi-UAV cooperative navigation.

CN120909315APending Publication Date: 2025-11-07CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511182357.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing multi-UAV collaborative path planning methods suffer from high computational complexity, low learning efficiency, and poor stability in high-dimensional collaborative optimization, dynamic task allocation, and autonomous decision-making in complex environments, making it difficult to meet real-time and adaptability requirements.

Method used

An improved MATD3 algorithm is adopted, which combines the PER mechanism and SA algorithm to optimize the experience replay buffer. An Actor-Critic network is constructed, and the collaborative navigation capability of multiple UAVs is optimized by dynamically adjusting the experience sampling probability and priority. The CTDE paradigm is used to realize centralized training and distributed execution.

Benefits of technology

It significantly improves the learning speed and policy convergence efficiency of multi-UAV path planning, enhances the stability and collaborative ability of the algorithm in complex environments, and improves the mission success rate and system reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909315A_ABST
    Figure CN120909315A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of unmanned aerial vehicle path planning, and particularly relates to a multi-unmanned aerial vehicle cooperative path planning method based on an improved MATD3 algorithm. The method comprises the following steps: constructing an unmanned aerial vehicle kinematics model and a multi-unmanned aerial vehicle environment model; designing a reward function based on the two models; constructing an experience playback buffer area by adopting a PER mechanism and managing priorities by adopting a segment tree structure; an SA algorithm is adopted to optimize a PER mechanism; an MATD3 algorithm model based on an Actor-Critic network is constructed; training an MATD3 algorithm model based on the reward function and the experience playback buffer area, and completing multi-unmanned aerial vehicle cooperative path planning by using the trained MATD3 algorithm model; the method improves the performance and reliability of a multi-unmanned aerial vehicle system in complex tasks, and has a good application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of unmanned aerial vehicle path planning, and particularly relates to a multi-unmanned aerial vehicle cooperative path planning method based on an improved MATD3 algorithm. BACKGROUND

[0002] With the rapid development of unmanned aerial vehicle technology, unmanned aerial vehicle clusters have a wide range of applications in search and rescue, logistics distribution, military defense, and environmental monitoring. Compared with single unmanned aerial vehicles, multiple unmanned aerial vehicles can efficiently complete complex tasks such as path planning, obstacle avoidance, and target navigation through cooperative work. However, there are still many technical challenges in multi-unmanned aerial vehicle cooperative planning, such as high-dimensional cooperative optimization calculation complexity, real-time task allocation and priority allocation, and autonomous decision-making of unmanned aerial vehicle clusters in complex environments.

[0003] In existing technologies, unmanned aerial vehicle path planning and navigation usually adopt rule-based algorithms, traditional optimization methods, and reinforcement learning techniques. For example, path planning methods based on A* algorithm or dynamic programming generate feasible paths in static environments. However, in dynamic environments or multiple unmanned aerial vehicle cooperative work, these methods have high complexity and are difficult to meet real-time requirements. In addition, single-agent reinforcement learning such as Deep Deterministic Policy Gradient (DDPG) is difficult to solve the complex interaction between unmanned aerial vehicles when dealing with multi-unmanned aerial vehicle cooperative tasks, resulting in decreased navigation efficiency.

[0004] In recent years, Multi-Agent Reinforcement Learning (MARL) has been proposed to solve the problem of multi-unmanned aerial vehicle cooperative navigation. For example, the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm trains independent policy networks for each unmanned aerial vehicle, achieving distributed decision-making. However, existing MARL methods often face the following limitations when facing high-dimensional state space and dynamic environments: First, the standard experience replay mechanism has difficulty in efficiently utilizing high-value experience data when dealing with multi-agent tasks, resulting in slow convergence during training. Second, fixed hyperparameter settings limit the adaptability of the algorithm to different task scenarios, making it difficult to achieve optimal performance in complex dynamic environments. In addition, existing methods have poor stability in multi-unmanned aerial vehicle cooperative navigation, and are prone to collisions or task failures due to uncoordinated policy updates.

[0005] Therefore, there is an urgent need for a multi-unmanned aerial vehicle cooperative path planning method capable of improving learning efficiency, optimizing sample utilization, and enhancing dynamic environment adaptability, so as to improve the performance and reliability of the multi-unmanned aerial vehicle system in complex tasks. SUMMARY

[0006] In view of the deficiencies in the prior art, the present application provides a multi-unmanned aerial vehicle cooperative path planning method based on an improved MATD3 algorithm, which comprises the following steps:

[0007] S1: constructing an unmanned aerial vehicle kinematics model and a multi-unmanned aerial vehicle environment model;

[0008] S2: designing a reward function based on the two models;

[0009] S3: constructing an experience replay buffer using a PER mechanism and managing priorities using a segment tree structure, and optimizing the PER mechanism using an SA algorithm;

[0010] S4: constructing an MATD3 algorithm model based on an Actor-Critic network;

[0011] S5: training the MATD3 algorithm model based on the reward function and the experience replay buffer, and using the trained MATD3 algorithm model to complete multi-unmanned aerial vehicle cooperative path planning.

[0012] Preferably, the unmanned aerial vehicle kinematics model comprises:

[0013] The unmanned aerial vehicle kinematics is:

[0014]

[0015] wherein, represents the x-axis coordinate of the unmanned aerial vehicle at the next moment, represents the x-axis coordinate of the unmanned aerial vehicle at the current moment, represents the speed of the unmanned aerial vehicle at the current moment, represents the orientation angle of the unmanned aerial vehicle at the next moment, represents the time step, represents the y-axis coordinate of the unmanned aerial vehicle at the next moment, represents the y-axis coordinate of the unmanned aerial vehicle at the current moment, represents the orientation angle of the unmanned aerial vehicle at the current moment, represents the current angular velocity of the unmanned aerial vehicle;

[0016] The control input of the linear velocity and the angular velocity satisfies the physical constraint:

[0017]

[0018] wherein, represents the linear velocity of the unmanned aerial vehicle, represents a linear velocity proportionality coefficient, represents a linear velocity standard control input, represents a maximum linear velocity, represents an angular velocity of the UAV, represents an angular velocity proportionality coefficient, represents an angular velocity standard control input, represents a maximum angular velocity;

[0019] The position coordinates of the UAV satisfy the environmental constraints, and the orientation angle of the UAV is normalized to the range.

[0020] Preferably, the reward function is a sum of a distance reward, a direction reward, a time penalty, a target arrival reward, and a collision penalty, after dynamic normalization processing.

[0021] Further, the distance reward is represented as:

[0022]

[0023] wherein, represents a distance reward at time t, represents a position of the i-th UAV at time t-1, represents a position of the i-th UAV, represents a target point distance of the i-th UAV, represents a position of the i-th UAV at time t. Further, the direction reward is represented as:

[0024]

[0025]

[0026] wherein, represents a direction reward at time t, represents a velocity vector of the i-th UAV at time t, represents a small constant, represents a target point distance of the i-th UAV, represents a position of the i-th UAV at time t. Further, the collision penalty includes an obstacle collision penalty and a UAV collision penalty; the obstacle collision penalty is represented as:

[0027]

[0028]

[0029] The UAV collision penalty is represented as:

[0030] ​​​​​

[0031] where, represents the obstacle collision penalty at time t, represents the position of the i-th obstacle at time t, represents the position of the i-th drone at time t, represents the position of the i-th obstacle at time t, represents the position of the i-th obstacle at time t, represents the obstacle radius, represents the drone radius, represents the drone collision penalty at time t, represents the position of the i-th drone at time t, represents the position of the i-th drone at time t.

[0032] Preferably, the process of constructing the experience replay buffer using the PER mechanism comprises:

[0033] The experience replay buffer is denoted as:

[0034]

[0035] where, denotes the experience replay buffer, denotes the local observation state of each drone, denotes the action taken by each drone, denotes the reward obtained by each drone, denotes the local observation state of each drone at the next time step, denotes whether the drone has reached the terminal state, denotes the global state of all drones, denotes the global state of all drones at the next time step;

[0036] The PER mechanism assigns a priority to each experience, which is denoted as:

[0037]

[0038] where, denotes the priority of the i-th experience, denotes the TD error of the i-th experience, denotes a small constant; The PER mechanism controls the sampling probability through the priority, which is denoted as:

[0039]

[0040]

[0041] where, denotes the sampling probability of the i-th experience, denotes the i-th experience, denotes the i-th experience.​​ Prioritizing experience points Indicates the strength of control priority;

[0042] The PER mechanism defines a weight for the bias caused by correction priority, expressed as:

[0043]

[0044] in, Indicates the first The calculation weight of the experience points, This indicates the number of samples in the experience replay buffer. This indicates the degree of control or correction.

[0045] Preferably, the process of optimizing the PER mechanism using the SA algorithm includes:

[0046] Store the experience samples into the experience replay buffer;

[0047] The intensity of PER parameter control priority is determined based on the initial and current temperatures. With control correction degree Gaussian noise is added to the vicinity of the current solution to generate candidate parameters through random perturbation;

[0048] Calculate the performance metrics of the current solution and candidate solutions;

[0049] Based on the incremental evaluation window, the Metropolis criterion is used to determine whether to accept a candidate solution based on the performance metrics of the current solution and candidate solutions. If accepted, the current PER parameter is updated. and ;

[0050] Adjust the temperature based on the performance metrics of the current solution and candidate solutions.

[0051] Preferably, the MATD3 algorithm model based on the Actor-Critic network includes:

[0052] An independent Actor network is built for each drone, while sharing the same Criric network; the Actor takes the local state as input and generates continuous actions for each drone, while the Criric network takes the global state and the actions of all drones as input and outputs the Q-value, which is used to update the Criric network.

[0053] Maintain the target network for each Actor and Critic network.

[0054] The beneficial effects of this invention are as follows:

[0055] The application greatly improves the learning speed and strategy convergence efficiency by dynamically adjusting the experience sampling probability based on the TD error, preferentially selecting experiences with high learning value for training; the PER parameter is dynamically optimized by the simulated annealing algorithm, the priority strength and the correction degree are adaptively adjusted according to the comprehensive performance index, and the stability of the algorithm is ensured in different training stages; by adopting the CTDE paradigm, the global state and the individual state are combined, the multi-unmanned aerial vehicle cooperative navigation capability is optimized, the success rate is improved and the collision rate is reduced. The simulation results show that the multi-unmanned aerial vehicle path planning is realized by the optimized MATD3 algorithm, the efficiency, stability and cooperation capability of the task are significantly improved, the performance and reliability of the multi-unmanned aerial vehicle system in complex tasks are improved, and the application prospect is good. BRIEF DESCRIPTION OF DRAWINGS

[0056] Figure 1 Trajectory diagram of a simulation environment in the application;

[0057] Figure 2 Training framework diagram of the MATD3 algorithm model based on the Actor-Critic network in the application;

[0058] Figure 3 Comparison diagram of average reward curves of the application and comparative algorithms;

[0059] Figure 4 Comparison diagram of success rate curves of the application and comparative algorithms. DETAILED DESCRIPTION

[0060] The technical solutions in the embodiments of the application will be apparently and completely described below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the application.

[0061] The application provides a multi-unmanned aerial vehicle cooperative path planning method based on an improved MATD3 algorithm, and the method comprises the following contents.

[0062] S1: constructing an unmanned aerial vehicle kinematics model and a multi-unmanned aerial vehicle environment model.

[0063] The multi-unmanned aerial vehicle environment model is constructed, the unmanned aerial vehicle activity environment is demarcated, and the unmanned aerial vehicle position, the obstacle position and the target point position are acquired; as shown in the figure, in the training and simulation process of the application, a 500*500 two-dimensional simulation environment is designed, and three unmanned aerial vehicles, three obstacles and three target points are set. Figure 1

[0064] The unmanned aerial vehicle kinematics model is constructed: ​

[0065] The UAV kinematics is:

[0066]

[0067] wherein, represents the x-axis coordinate of the UAV at the next time, represents the x-axis coordinate of the UAV at the current time, represents the speed of the UAV at the current time, represents the orientation angle of the UAV at the next time, represents the time step, represents the y-axis coordinate of the UAV at the next time, represents the y-axis coordinate of the UAV at the current time, represents the orientation angle of the UAV at the current time, represents the current angular velocity of the UAV.

[0068] The control input of linear velocity and angular velocity satisfies the physical constraint:

[0069]

[0070] wherein, represents the linear velocity of the UAV, represents the linear velocity proportional coefficient, represents the maximum linear velocity, represents the angular velocity of the UAV, represents the angular velocity proportional coefficient, represents the linear velocity standard control input, represents the angular velocity standard control input, ; represents the maximum angular velocity;

[0071] The position coordinates of the UAV satisfy the environmental constraint: .

[0072] The orientation angle of the UAV is normalized to the range: . .

[0073] S2: Designing a reward function based on two models.

[0074] According to the environment and UAV kinematics model constructed in S1, in order to further encourage the UAV to complete the navigation task, avoid collision and optimize the path, the present application constructs a reward function, which is the sum of distance reward, direction reward, time penalty, target arrival reward and collision penalty after dynamic normalization processing; Specifically:

[0075] Distance reward :

[0076] To encourage the UAV to approach the target point, the distance reward is defined as:

[0077]

[0078] where, represents the position of the t-th UAV at time t-1, represents the position of the t-th UAV at time t, represents the distance of the t-th UAV to the target point, represents the x-coordinate of the target point of the t-th UAV, represents the y-coordinate of the target point of the t-th UAV; represents the position of the t-th UAV at time t-1, represents the position of the t-th UAV at time t, represents the Euclidean distance between the current position of the UAV and the target point; represents the distance of the t-1-th UAV to the target point. represents the distance of the t-th UAV to the target point.

[0079] represents the distance of the t-1-th UAV to the target point. represents the distance of the t-th UAV to the target point.

[0080] The direction reward is defined as:

[0081] To encourage the UAV to choose a moving direction consistent with the target direction, the direction reward is defined as:

[0082]

[0083] where, represents the velocity vector of the t-th UAV at time t, represents the Euclidean norm; dot product calculates the cosine of the angle between the target direction and the moving direction. represents a small constant, which is set to prevent the denominator from being zero. The time penalty

[0084] is defined as:

[0085] To encourage the UAV to accelerate to complete the task, the time consumption of each step is penalized:

[0086]

[0087] The target arrival reward is defined as:

[0088] When the UAV reaches the target area, a reward is given, and the target arrival reward is defined as:

[0089]

[0090] when the distance between the UAV position and the target point is less than the target radius​ reward is given.

[0091] The collision penalty includes an obstacle collision penalty and a UAV collision penalty; the obstacle collision penalty is is expressed as:

[0092]

[0093] If the distance between the UAV and any obstacle is less than a penalty is given; wherein is the radius of the obstacle; is the radius of the UAV; represents the position of the th obstacle.

[0094] The UAV collision penalty is

[0095]

[0096] wherein, represents the position of the th UAV at time t; if the distance between the UAV and other UAVs is less than twice the radius of the UAV, a penalty is given.

[0097] The comprehensive reward is:

[0098]

[0099] In order to make the training stable, the reward is processed by dynamic normalization:

[0100]

[0101] wherein prevents division by zero, is the exponential moving average of the reward mean:

[0102]

[0103] is the exponential moving average of the reward standard deviation:

[0104]

[0105] The final normalized reward is used for subsequent training.

[0106] S3: The PER mechanism is used to construct an experience replay buffer and the segment tree structure is used to manage the priority; the SA algorithm is used to optimize the PER mechanism.

[0107] As Figure 2 ​As shown, the PER mechanism is used to construct an experience replay buffer for storing experience data of the environment in the training process of multiple UAVs, including local state, action, reward, next local state, completion flag, global state and next global state; the PER mechanism assigns a priority to each piece of experience, ensuring that valuable experience can be sampled frequently in priority, accelerating the convergence of the algorithm; in particular:

[0108] The content stored in the experience replay buffer is:

[0109]

[0110] Among them, represents the experience replay buffer, represents the local observation state of each UAV, represents the action taken by each UAV, represents the reward obtained by each UAV, represents the local observation state of each UAV at the next time step, represents whether the UAV has reached the termination state, represents the global state of all UAVs, represents the global state of all UAVs at the next time step.

[0111] The PER mechanism assigns a priority to each piece of experience, and the priority is represented as:

[0112]

[0113] Among them, represents the priority of the th experience, represents a small constant, which is a very small positive number, ensuring that all experiences have the opportunity to be sampled; represents the TD error of the th experience, the TD error represents the difference between the current Q value and the target Q value, reflecting the degree of learning required by the agent, and the specific expression is as follows:

[0114]

[0115] Among them, represents the Q value predicted by the first Q function (Critic network) under the given state and action , represents the target Q value, represents the Q value predicted by the second Q function (Critic network) under the given state and action , represents the mean square error, State of the th sample, Action of the th sample.

[0116] The PER mechanism controls the sampling probability by priority, denoted as:

[0117]

[0118] where, Sampling probability of the th experience, Priority of the th experience, Intensity of controlling the priority degree; , is the total number of samples currently stored in the buffer.

[0119] The PER mechanism defines a calculation weight to correct the bias caused by priority, denoted as:

[0120]

[0121] where, Calculation weight of the th experience, Number of samples in the experience replay buffer, Intensity of controlling the correction degree.

[0122] In order to make the training more stable, the weight is normalized:

[0123]

[0124] The priority is managed by using segment tree structure (SumTree), which is an efficient data structure for storing and updating learning experience, supporting fast sampling and reducing time complexity. Its structure is a full binary tree, and the leaf node stores the priority of each experience , the non-leaf node stores the sum of the priorities of its child nodes, defined as: , the capacity is the size of the replay buffer.

[0125] The process of updating the priority of SumTree is:

[0126] 1. Update the leaf node

[0127]

[0128] where is the specified position of the buffer that needs to be updated, is the size of the buffer. The leaf node of the segment tree starts from index , then Index of the buffer is mapped to the leaf node of the segment number.

[0129] 2. Recursively update the parent node

[0130] From the leaf node upwards, update the priority and of each parent node:

[0131]

[0132] Where the value of the parent node is the sum of the priorities of its two child nodes, recursively update until the root node , which stores the sum of all priorities as .

[0133] 3. Sample query

[0134] Generate a random number , starting from the root node, recursively traverse the segment tree to find the leaf node that satisfies , and return the corresponding experience index.

[0135] Optimize the PER mechanism using the SA algorithm. SA simulates the metal annealing process, and by controlling the temperature parameter, it accepts poor solutions in the early optimization stage to avoid local optimization, and gradually converges in the later stage. The specific process is as follows:

[0136] Store the collected experience samples in the experience replay buffer:

[0137]

[0138] According to the initial temperature and the current temperature, add Gaussian noise around the current solution of the PER parameters and to generate a candidate parameter through random disturbance:

[0139]

[0140] Where: represents the strength of the candidate control priority, represents the candidate control correction degree, represents the strength of the current control priority, represents the current control correction degree; , , is the current temperature, is the initial temperature; For , the range is larger because the priority index has a more significant impact on the sampling distribution, while , the range is smaller because importance sampling requires more stable adjustments.

[0141] ​Compute the performance indicator of the current solution and the candidate solution:

[0142] PER parameter combination The performance of the combination is quantified by the performance indicator:

[0143]

[0144] wherein, represents the average reward, represents the average success rate, represents the average collision rate.

[0145] Based on the progressive evaluation window, the Metropolis criterion is used to determine whether to accept the candidate solution according to the performance indicator of the current solution and the candidate solution, and if accepted, the current PER parameter is updated and Specifically:

[0146] The SA optimizer does not update the parameters after each episode, but evaluates every SA_FREQ episodes. The specific evaluation uses a progressive evaluation window, as shown in the following Table 1 (assuming SA_FREQ=3):

[0147] Table 1 Progressive evaluation window

[0148] Episode Start_id calculation Evaluation window size Evaluation content 1 Max(0, 1, -3) = 0 1 [ep1] 2 Max(0, 2, -3) = 2 2 [ep1, ep2] 3 Max(0, 3, -3) = 3 3 [ep1, ep2, ep3] 4 Max(0, 4, -3) = 3 3 [ep2, ep3, ep4]

[0149] The Metropolis criterion is the core of the SA algorithm, which is used to balance between global search and local optimization. Its acceptance probability The formula is:

[0150]

[0151] If the performance of the candidate solution is better than the performance of the current solution , the candidate solution is unconditionally accepted to pursue higher performance; if the performance of the candidate solution is poor, the algorithm is allowed to jump out of the local optimum and explore a wider parameter space. The specific operation is to generate a random number ranging from 0 to 1, and if the calculated probability is greater than the random number, it is accepted, as shown in the following Table 2 (assuming SA_FREQ=3).

[0152] Table 2 Metropolis acceptance criterion

[0153] Episode T Current (a, b) Candidate (a, b) Performance (p) Acceptance result 30 80 (0.6,0.4) (0.62,0.38) 42.1 43.5 Accept 33 76 (0.62,0.38) (0.59,0.41) 43.5 42.8 Probability accept (prob = 0.83, random number is 0.43) 36 72.2 (0.59,0.41) (0.57,0.44) 42.8 41.2 Reject (prob = 0.31, random number is 0.57)

[0154] After updating the parameters using SA, the parameters are immediately synchronized to the PER buffer to achieve parameter updating. If is increased, it is more inclined to high TD error samples; if The importance sampling compensation is more emphasized as the increase of the temperature.

[0155] The temperature is adjusted according to the performance indicators of the current solution and the candidate solution, and specifically:

[0156] After the parameter update is completed, the cooling rate is dynamically adjusted:

[0157]

[0158] The temperature is updated as:

[0159]

[0160] Wherein, is the stop temperature, is the cooling rate.

[0161] S4: Constructing a MATD3 algorithm model based on an Actor-Critic network.

[0162] The present application adopts the CTDE paradigm, that is, centralized training and distributed execution. A MATD3 algorithm model based on an Actor-Critic network is constructed, an independent Actor network is constructed for each unmanned aerial vehicle, and a same Critic network is shared; in the training stage, the Critic network jointly evaluates the strategy Q value by using the global state and the actions of all unmanned aerial vehicles, and is used for updating the Critic network, so as to optimize the global cooperation; in the execution stage, the Actor network of each unmanned aerial vehicle only generates actions based on the local state, and does not need to access global information, thereby reducing the communication requirement.

[0163] The Actor network generates continuous actions for each unmanned aerial vehicle, inputs the local state, and follows the principle of distributed execution, and the network structure comprises:

[0164] Input layer: receiving the local state Wherein, Determined by the environment in S1, containing the normalized position, speed, orientation, relative position to the target point, obstacle and other unmanned aerial vehicle information of the unmanned aerial vehicle.

[0165] Hidden layer: containing a normalization layer, normalizing the input state, enhancing the training stability:

[0166]

[0167] Wherein, and are the mean and variance of the state respectively, Prevent division by zero

[0168] Secondly, three fully connected layers are included, with sizes of 1024, 1024, 512 respectively, and each layer is followed by a Relu activation function:

[0169]

[0170] wherein ; is the bias.

[0171] Output layer: The fully connected layer maps the 512-dimensional features to the action space, followed by activation function:

[0172]

[0173] wherein , action represents the normalized linear velocity and angular velocity.

[0174] The Critic network evaluates the joint Q value of the global state and all UAV actions, which is used for centralized training to guide the optimization of the Actor network, and its structure includes:

[0175] Input layer: receives the global state which is determined by the environment in S1, including the normalized position, velocity, orientation of the UAV itself, relative position to the target point, obstacle and other UAV information. Secondly, it receives all the actions of the UAV, which are spliced into . Finally, it receives the total dimension .

[0176] Hidden layer: contains two parallel Q networks, Q1 and Q2, each composed of four fully connected layers, wherein the Q1 network is:

[0177]

[0178] wherein

[0179] The Q2 network is:

[0180]

[0181] wherein

[0182] Output layer: outputs two Q value vectors , representing the Q value estimates of each UAV respectively. Then take the minimum value through the estimates of the two Q networks to reduce the overestimation of Q value and improve the stability of training:

[0183]

[0184] For the stability of training, the MATD3 algorithm maintains a target network for each Actor and Critic network, with parameters .

[0185] S5: Train the MATD3 algorithm model based on the reward function and the experience replay buffer, and use the trained MATD3 algorithm model to complete the cooperative path planning of multiple UAVs.

[0186] As shown in Figure 2 , the specific training of the whole model is as follows:

[0187] Before training, first construct the initial environment according to S1, initialize the experience replay pool according to S3, and construct the network structure according to S4, which contains the corresponding parameter settings.

[0188] I. Environment interaction and experience collection:

[0189] Distributed action execution:

[0190] For each UAV , use the Actor network based on the current state , add adaptive exploration noise, and clip the action to :

[0191]

[0192]

[0193] where the noise scale is dynamically adjusted according to the recent success rate

[0194] Environment interaction:

[0195] Input the action into the environment, represent the linear and angular velocities of the Nth UAV, the linear velocity includes forward and backward movement of the UAV, and the angular velocity includes counterclockwise and clockwise rotation of the UAV; At the same time, the local state , global state , reward , completion flag at the next time are obtained; Where the reward is calculated according to the reward function designed in S2.

[0196] Experience storage:

[0197] Store the experience tuple to the priority experience replay pool, and initialize the priority to the maximum value to ensure sampling:

[0198]

[0199] II. Experience collection and network update

[0200] Experience collection:

[0201] From the priority experience replay area, a small batch of experiences is collected, and at the same time, the simulated annealing optimizer is started to optimize the collected experiences, and the optimized experiences are returned to the experience tuple, index and importance sampling weight.

[0202] Network update:

[0203] Critic network update: first use the target actor network to generate the action at the next time, and add noise:

[0204]

[0205] Secondly, use the target Critic network to calculate the Q value, and take the minimum value of the Q value of the two Critic networks to reduce overestimation:

[0206]

[0207] Finally, calculate the target Q value:

[0208]

[0209] Calculate the Critic network loss: calculate the TD error of the two Q networks, and combine the calculation weight of PER to get:

[0210]

[0211] Update the Critic parameters: first apply the Adam optimizer to update the Critic parameters:

[0212]

[0213] Secondly, apply gradient clipping:

[0214]

[0215] Actor network update: take a delayed update strategy to update, and update the actor network every steps.

[0216] Actor loss calculation: use the online actor network to generate the current action , and the target actor network is used for other UAVs; calculate the loss:

[0217]

[0218] Actor parameter update: first use the Adam optimizer to update the Actor parameters:

[0219]

[0220] Secondly, gradient clipping is applied:

[0221]

[0222] Target network soft update: soft update of target network parameters:

[0223]

[0224] PER priority update: update the empirical priority according to the TD error:

[0225]

[0226] Repeat the above training steps until the training is completed, and use the trained MATD3 algorithm model to complete the multi-unmanned aerial vehicle cooperative path planning.

[0227] As shown in Figure 3 and Figure 4 Compared with other mainstream algorithms, the present application (SA_PER_MATD3) converges faster in the reward curve and has a higher success rate in the success rate curve.

[0228] In summary, by introducing the PER mechanism and using SA to dynamically optimize the parameters in PER, the present application avoids the fixed parameters from falling into a local optimal solution and optimizes the stability of the algorithm in different training stages. Meanwhile, the CTDE paradigm is adopted to combine the global state and the individual state, optimize the multi-unmanned aerial vehicle cooperative navigation capability, and improve the success rate of multi-unmanned aerial vehicle cooperative path planning.

[0229] The above examples further illustrate the purpose, technical solutions and advantages of the present application. It should be understood that the above examples are only preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made to the present application within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A multi-unmanned aerial vehicle cooperative path planning method based on an improved MATD3 algorithm, characterized in that, The application relates to a multi-unmanned aerial vehicle (UAV) cooperative path planning method based on a multi-agent temporal difference 3 (MATD3) algorithm. The method comprises the following steps: S1, constructing a UAV kinematics model and a multi-UAV environment model; S2, designing a reward function based on the two models; S3, constructing an experience replay buffer by using a PER mechanism and managing priorities by using a segment tree structure, and optimizing the PER mechanism by using an SA algorithm; S4, constructing an MATD3 algorithm model based on an Actor-Critic network; 2. The multi-UAV cooperative path planning method based on the improved MATD3 algorithm according to claim 1, characterized in that, S5, training the MATD3 algorithm model based on the reward function and the experience replay buffer, and completing multi-UAV cooperative path planning by using the trained MATD3 algorithm model. The UAV kinematics model comprises: ; wherein, represents the x-axis coordinate of the drone at the next time instant, represents the x-axis coordinate of the drone at the current time instant, represents the velocity of the drone at the current time instant, represents the heading angle of the drone at the next time instant, represents the time step, represents the y-axis coordinate of the drone at the next time instant, represents the y-axis coordinate of the drone at the current time instant, represents the heading angle of the drone at the current time instant, represents the current angular velocity of the drone; The kinematics of the UAV is: ; wherein, represents a linear velocity of the UAV, represents a linear velocity proportional coefficient, represents a linear velocity standard control input, represents a maximum linear velocity, represents an angular velocity of the UAV, represents an angular velocity proportional coefficient, represents an angular velocity standard control input, represents a maximum angular velocity; The position coordinates of the UAV satisfy the environmental constraints, and the orientation angle of the UAV is normalized to a range.

3. The multi-UAV cooperative path planning method based on the improved MATD3 algorithm according to claim 1, characterized in that, The control input of the linear velocity and the angular velocity satisfies a physical constraint:

4. The multi-UAV cooperative path planning method based on the improved MATD3 algorithm according to claim 3, characterized in that, The reward function is a sum of distance reward, direction reward, time penalty, target arrival reward and collision penalty, which is obtained after dynamic normalization processing. ; wherein, denotes the distance reward at time t, denotes the position of the t-1th UAV at time t-1, denotes the position of the t-1th UAV at time t-1, denotes the target point distance of the t-1th UAV, denotes the target point distance of the t-1th UAV, denotes the position of the t-1th UAV at time t-1, denotes the position of the t-1th UAV at time t-1.

5. The multi-UAV cooperative path planning method based on the improved MATD3 algorithm according to claim 3, characterized in that, The distance reward is expressed as: ; wherein, represents the direction reward at time t, represents the position of the ith UAV at time t, represents the velocity vector of the ith UAV, represents a small constant, represents the target point distance of the ith UAV, represents the target point distance of the ith UAV, represents the position of the ith UAV at time t, represents the position of the ith UAV at time t.

6. The multi-UAV cooperative path planning method based on the improved MATD3 algorithm according to claim 3, characterized in that, The direction reward is expressed as: The collision penalty comprises obstacle collision penalty and UAV collision penalty. ; The obstacle collision penalty is expressed as: ; wherein, represents the obstacle collision penalty at time t, represents the position of the i-th rack drone at time t, represents the position of the i-th obstacle, represents the obstacle radius, represents the drone radius, represents the drone collision penalty at time t, represents the position of the i-th rack drone at time t.

7. The multi-UAV cooperative path planning method based on the improved MATD3 algorithm according to claim 1, characterized in that, The UAV collision penalty is expressed as: The process of constructing the experience replay buffer by using the PER mechanism comprises: ; wherein, represents the experience replay buffer, represents the local observation state of each drone, represents the action taken by each drone, represents the reward obtained by each drone, represents the local observation state of each drone at the next time step, represents whether the drone reached the terminal state, represents the global state of all drones, represents the global state of all drones at the next time step; The experience replay buffer is expressed as: ; wherein, represents the TD error of the empirical priority, represents the TD error of the empirical priority, represents a small constant; The PER mechanism allocates a priority to each experience, and the priority is expressed as: ; wherein, represents the first empirical sampling probability, represents the first empirical priority, represents the strength of the control priority degree; The PER mechanism controls the sampling probability through the priority, and is expressed as: ; wherein, represents the number of samples in the experience replay buffer, represents the number of samples in the experience replay buffer, represents the number of samples in the experience replay buffer, represents the number of samples in the experience replay buffer.

8. The multi-UAV cooperative path planning method based on the improved MATD3 algorithm according to claim 1, characterized in that, The PER mechanism defines a calculation weight for correcting the bias caused by the priority, and is expressed as: The process of optimizing the PER mechanism by using the SA algorithm comprises: The strength of the PER parameter control priority is determined according to the initial temperature and the current temperature and the current solution is perturbed by adding Gaussian noise to generate candidate parameters and the current solution is perturbed by adding Gaussian noise to generate candidate parameters storing the experience sample into the experience replay buffer; Based on the progressive evaluation window, according to the performance index of the current solution and the candidate solution, the Metropolis criterion is used to judge whether to accept the candidate solution, and if accepted, the current PER parameter is updated With ; calculating the performance index of the current solution and the candidate solution; 9. The multi-UAV cooperative path planning method based on the improved MATD3 algorithm according to claim 1, characterized in that, adjusting the temperature according to the performance index of the current solution and the candidate solution. The process of constructing the MATD3 algorithm model based on the Actor-Critic network comprises: an independent Actor network is constructed for each UAV, and a same Critic network is shared; the Actor inputs a local state and generates a continuous action for each UAV, and the Critic network inputs a global state and the actions of all the UAVs to output a Q value, which is used for updating the Critic network; a target network is maintained for each Actor and Critic network.