Heterogeneous multi-unmanned aerial vehicle cooperative path planning method based on multi-agent deep reinforcement learning
By building an airspace reinforcement learning environment in the deep reinforcement learning framework of multi-agents, modeling the POMDP model and designing the MASAC-SEPR algorithm, the problems of insufficient feature extraction capabilities and low learning efficiency in the collaborative path planning of multiple drones are solved, and efficient and accurate path planning strategies are realized to adapt to dynamic and uncertain environments.
Patent Information
- Application Number
- CN202510254919.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-05
AI Technical Summary
The existing technology has problems such as insufficient feature extraction capabilities, low experience learning efficiency, and poor flexibility of collaboration strategies in terms of dynamic planning problems and agent collaboration. Especially in the coordinated path planning of multiple drones, it is difficult to meet the real-time needs in dynamic uncertain environments.
The heterogeneous multi-UAV collaborative path planning method based on multi-agent deep reinforcement learning is adopted. By building an airspace reinforcement learning environment, the UAV dynamics equation is introduced, the multi-UAV decision model is modeled as a POMDP model, and the MASAC-SEPR algorithm model is designed, combining the SEAttention module, priority experience playback memory bank and entropy network to optimize the path planning strategy.
It significantly improves the efficiency and accuracy of multi-UAV collaborative path planning, enhances feature extraction capabilities and learning efficiency, improves the flexibility and adaptability of collaborative strategies, and can more effectively respond to task requirements in dynamic and uncertain environments.
Smart Images

Figure CN120103855A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicle path planning, and specifically to a heterogeneous multi-unmanned aerial vehicle collaborative path planning method based on multi-agent deep reinforcement learning. Background Art
[0002] Path planning is the basis and core link of UAV mission execution. How to ensure that UAVs reach the target location quickly, safely and efficiently has become the focus and hot spot of current research. In the field of UAV path planning, the research process at home and abroad has experienced a continuous evolution from traditional methods to modern intelligent algorithms.
[0003] Early classical path planning algorithms showed certain practicality in relatively simple and static environments. Internationally, the Dijkstra algorithm uses its forward traversal of all nodes to find the optimal path; the A* algorithm improved by Hart et al. achieves targeted expansion of nodes by introducing an evaluation function; Stent further optimizes the A* algorithm to obtain the D* algorithm, which enables it to dynamically update nodes and replan paths based on the current state. In China, many scholars have conducted in-depth theoretical research and practical application exploration on these classical algorithms, promoting the development of the basic theory of UAV path planning. However, path planning is essentially an NP optimization problem. In a complex and changing environment, the computational complexity of classical algorithms will grow exponentially, and they often fall into the dilemma of the "curse of dimensionality", resulting in low solution efficiency or even failure to obtain effective results.
[0004] In order to cope with the limitations of classical algorithms in complex environments, heuristic algorithms have gradually attracted widespread attention from researchers at home and abroad. In foreign research, some people combined the greedy allocation strategy with the improved ant colony optimization algorithm based on variable pheromone for target point allocation and path planning of UAVs; others solved the path planning problem in static environments by implementing genetic algorithms in parallel on graphics processing units. Domestic scholars have also actively innovated and proposed methods such as hybrid particle swarm optimization algorithm fused with simulated annealing algorithm to cope with the challenges of path planning in complex environments. But in general, whether it is classical algorithms or heuristic algorithms, most of them are more suitable for static or relatively stable environments, and it is difficult to meet the needs of real-time path planning in dynamic environments.
[0005] The emergence of deep reinforcement learning (DRL) provides a new solution and method for dynamic path planning problems. Many studies have successfully solved the path planning problems in different scenarios using the DRL algorithm, such as using the improved deep deterministic policy gradient (DDPG) algorithm to plan the path of a glider in a time-varying ocean current environment, using the competitive deep Q network algorithm to achieve collision avoidance path planning at a fixed altitude, and using the double-delay DDPG (TD3) algorithm to achieve adaptive motion planning and obstacle avoidance of autonomous underwater robots. Domestic scholars have also kept up with the international frontier, using algorithms such as TD3 to enable drones to perform path planning tasks in a multi-obstacle environment with randomness and dynamics. However, most of the existing research focuses on single-agent DRL algorithms. When dealing with the problem of collaborative path planning for multiple drones, they cannot fully consider the mutual cooperation and influence between drones, resulting in low collaborative efficiency and difficulty in adapting to complex practical application needs.
[0006] In order to solve the problem of multi-agent collaborative path planning, the Multi-Agent Flexible Execution Evaluation (MASAC) algorithm came into being. Domestic and foreign scholars have conducted a series of studies on this algorithm. Based on the deep reinforcement learning framework, it introduces a flexible execution mechanism to improve the collaborative path planning ability of multi-agents in complex environments to a certain extent. It can comprehensively consider the interactions between multiple agents and make more reasonable path planning decisions by learning environmental information and interaction strategies between agents. However, the MASAC algorithm still has obvious defects: on the one hand, its ability to extract complex environmental features needs to be improved. When facing complex scenes such as complex obstacle distribution and dynamic changes in target points, it is difficult to accurately capture key information, resulting in limited path planning efficiency and accuracy. On the other hand, in the process of experience learning, the MASAC algorithm uses ordinary experience playback, which is not efficient in utilizing important experience, resulting in a slow learning speed and a long algorithm convergence time, which affects the application of multiple UAVs in scenarios with high real-time requirements. Summary of the invention
[0007] The purpose of the present invention is to solve the problems that traditional path planning algorithms in the prior art are difficult to cope with dynamic planning problems and intelligent agent collaboration problems, as well as the defects of existing multi-UAV collaborative path planning algorithms (such as MASAC) in dynamic uncertain environments, such as insufficient feature extraction capabilities, low experience learning efficiency, and poor flexibility of collaborative strategies. A heterogeneous multi-UAV collaborative path planning method based on multi-agent deep reinforcement learning is provided to solve the above problems.
[0008] In order to achieve the above object, the technical solution of the present invention is as follows:
[0009] A heterogeneous multi-UAV collaborative path planning method based on multi-agent deep reinforcement learning includes the following steps:
[0010] Analyze multi-UAV collaborative path planning tasks;
[0011] Build an airspace reinforcement learning environment;
[0012] Introduce the UAV dynamics equation;
[0013] The multi-UAV decision model in multi-UAV collaborative path planning is modeled as a POMDP model;
[0014] Design the MASAC-SEPR algorithm model to plan paths for multiple UAVs and generate a multi-agent collaborative path optimization network model;
[0015] Training of the multi-agent collaborative path optimization network model: The multi-agent collaborative path optimization network model is trained in the current airspace reinforcement learning environment to generate multi-UAV collaborative paths.
[0016] The analysis of the multi-UAV collaborative path planning task is: obtaining the UAV group path planning task, multiple UAVs avoid obstacles in a bounded airspace environment, and arrive at the target location in the form of a single master controller and multiple followers, where the master controller is responsible for avoiding obstacles and reaching the target, and the followers follow the single master controller to maintain the formation.
[0017] The construction of the airspace reinforcement learning environment is to construct a simulated airspace plane and define the speed, acceleration, and angular velocity ranges of the master UAV and the follower UAV; including the following steps:
[0018] The multi-UAV collaborative path planning simulation environment is defined as a two-dimensional plane of 700 pixels × 600 pixels, corresponding to the airspace of 7000m × 6000m in reality;
[0019] Set the animation refresh rate FPS to 60. In this environment, drones are divided into two types: master and follower. The master UAV speed range is 10-20 pixels per frame, and the acceleration control input range is -3-3m / s 2 , angular velocity range is -0.6-0.6rad / s, follower UAV velocity range is 100-300m / s, acceleration range is -6-6m / s 2 , the angular velocity range is -1.2-1.2rad / s;
[0020] It is assumed that when the environment is reset at the end of each round, the initial states of all UAVs, i.e., the x position, y position, velocity v and heading angle ψ, the target location and the positions of obstacles are randomly generated.
[0021] The introduction of the UAV dynamics equation is: taking the dynamics equation as the basic rule of UAV motion, the discrete dynamics equation of multiple UAVs from time t to t+1 contains the update formula of position, speed, and angle variables, and satisfies the constraints of flight state, control quantity, and flight path; including the following steps:
[0022] Taking the dynamic equation as the basic rule of UAV motion, the discrete dynamic equation of multiple UAVs from time t to time t+1 is:
[0023]
[0024] All variables in the above equations are between their respective minimum and maximum values;
[0025] Where i is the index of the UAV, (x, y) is the position of the UAV, v is the UAV velocity, ψ is the UAV heading angle, u is the acceleration control input of the UAV, ω is the angular velocity control input, and Δt is the time step from time t to time t+1;
[0026] Assume the obstacle is a standard circular area.
[0027] (x 0 ,y 0 ) is the center position of the obstacle, R 0 is the radius of the obstacle;
[0028] In the environment, the drone moves according to discrete dynamic equations and constraints, and its position, velocity, and heading angle states are continuously updated, and the values are always within their respective value ranges;
[0029] The dynamic equations and obstacle settings are added to the airspace reinforcement learning environment, and the drone uses the dynamic equations as the basic rules for drone movement.
[0030] Modeling the multi-UAV decision model in the multi-UAV collaborative path planning as a POMDP model includes the following steps:
[0031] Define the multi-UAV basic decision model as the POMDP model:
[0032] Defined as a seven-tuple <S,A,R,P,Z,O,γ>,
[0033] Among them, the state set S contains the state information of all UAVs, but it is unknown to the agent; the action set A is the action set of all UAVs, and the actions are known to the agent; the reward function R(s,a) represents the reward obtained by the UAV when taking action a in state s; the state transition probability P(s′|s,a) represents the probability that the UAV takes action a in state s and transfers to state s′, which is unknown to the agent; the observation set Z is the set of observation information of the UAV, which is known to the agent; the observation function O(z|s,a) represents the probability that state s produces observation z after executing action a, which is unknown to the agent; the decay factor γ∈[0,1] represents the preference and proportion of current rewards and future rewards;
[0034] For the observation space, the variables in the UAV dynamics equation are introduced, and the observation information of the i-th UAV is set as vector z i =[x i ,y i ,v i ,ψ i ,m i ] means, where [x i ,y i ,v i ,ψ i ] is the flight state quantity of the i-th UAV, mi is the observation state information related to the UAV type;
[0035] For the master UAV, m i =[x G ,y G ,O flag ], where (x G ,y G ) is the position of the target point, O flag It is the obstacle mark;
[0036] For the follower UAV, m i =[x L ,y L ,v L ], where (x L ,y L ) is the position of the master UAV, v L is the speed of the master UAV. Since there is only one master UAV, the rest of the UAVs are follower UAVs. Each follower UAV has m i =[x L ,y L ,v L ]; where O flag The definition is as follows:
[0037]
[0038] in, is the distance between the UAV and the obstacle, (x, y) represents the position of the UAV, (x 0 ,y 0 ) is the center position of the obstacle; when the distance between the UAV and the obstacle is less than twice the obstacle radius, it is considered that there is an obstacle near the UAV, and the obstacle flag position is 1, otherwise it is set to 0;
[0039] For the state space, the state information S of each UAV i =[z 1 ,z 2 ,...,z N ], including the observation status information of itself and other UAVs, a total of N UAVs;
[0040] For the action space, the action of the i-th UAV is a continuous two-dimensional vector a i =[u i ,w i ] Τ , by choosing the acceleration u i and angular velocity ω i Execute time Δt to achieve the desired speed and heading angle while satisfying the constraints;
[0041] Design reward mechanisms and reward and punishment mechanisms for different types of UAVs.
[0042] For the master UAV, including the boundary reward r e , obstacle avoidance reward r o 、Target reward r g , Formation distance reward r f and speed synergy reward r s , whose reward function R L for
[0043] R L =ω L,1 r e +ω L,2 r o +ω L,3 r g +ω L,4 r f +ω L,5 r s ,
[0044] Among them, L represents that the UAV is the master controller, ω L,1 ,ω L,2 ,ω L,3 ,ω L,4 ,ω L,5 They represent the boundary reward r of the master UAV respectively. e , obstacle avoidance reward ro 、Target reward r g , Formation distance reward r f , speed synergy reward r s The weight of
[0045] For N-1 follower UAVs, the reward function consists of the formation distance reward and the speed coordination reward. The reward function R of each follower UAV j is F,j It consists of formation distance reward and speed coordination reward.
[0046] That is R F,j =ω F,1 r f,j +ω F,2 r s,j ,
[0047] Where j (j = 1, 2, ..., N-1) is the index of the follower UAV, F represents the UAV is a follower, r f,j represents the formation distance reward of the jth follower drone, r s,j represents the speed coordination reward of the j-th follower UAV, ω F,1 ,ω F,2 Represent the weights of formation distance reward and speed coordination reward respectively,
[0048] Total reward set R = [R L ,R F,1 ,R F,2 ,...,R F,N-1 ], that is, the reward R of a master UAV L and the rewards R of the N-1 follower UAVs F,j where j (j = 1, 2, ..., N-1) is the index of the follower UAV, indicating the jth follower UAV. Then R F,1 ,R F,2 ,...,R F,N-1 Represents j=1,2,...,N-1, i.e., the rewards from the 1st, 2nd... to the N-1th follower UAV. By setting the reward weight ω, the UAV is guided to learn the optimal path strategy.
[0049] The design of the MASAC-SEPR algorithm for path planning of multiple UAVs includes the following steps:
[0050] Setting up the SEAttention module:
[0051] The SEAttention module first performs a compression operation, using adaptive average pooling to compress the input feature map into a 1×1 vector to obtain global information. Each channel generates a channel descriptor, condenses the global spatial information into a channel vector, and captures the global distribution of channel feature responses. Then, it performs an excitation operation through a sequence of fully connected layers, first using a linear layer to reduce the number of channels, then performing a nonlinear transformation through a ReLU activation function, and then restoring the number of channels through another linear layer. The channel attention weight is generated with the help of a Sigmoid activation function. Finally, the generated channel attention weight is multiplied with the original input feature map to achieve feature recalibration.
[0052] Build an Actor network and learn:
[0053] Building Critic network and learning;
[0054] Design of priority experience replay memory bank mechanism:
[0055] Prioritized Experience Replay Memory Mechanism The mechanism inputs three key parameters at runtime: experience replay memory capacity, sampling batch size, and priority index;
[0056] In the initialization phase, two double-ended queues with the same maximum length, i.e., the capacity of the experience replay memory bank, are created. One of the queues is used to store state transition experiences, which are recorded in the form of four-tuples (state, action, reward, next state); the other queue is used to store the priority corresponding to each experience.
[0057] When a new state transition experience is generated, the storage process is as follows: first determine whether the queue for storing priorities is empty. If not, obtain the maximum priority in the queue; if it is empty, set the maximum priority to 1.0. Then, add the new state transition experience to the queue for storing state transition experience, and add the obtained or set maximum priority to the queue for storing priorities.
[0058] When performing experience sampling, the specific steps are as follows: according to the length of the queue storing state transfer experience, the queue storing priority is processed accordingly and converted into a numpy array. Then, according to the priority data in the array, the sampling probability of each experience is calculated and normalized. Then, a specified number of experience indexes are randomly sampled from the queue storing state transfer experience according to the probability, and the corresponding experience samples are obtained according to the index. Finally, the state, action, reward and next state information are extracted from these samples respectively, and they are converted into torch.FloatTensor type and returned.
[0059] For the priority update operation, the mechanism receives the index list and the new priority list as input, and updates the new priority to the queue storing the priority according to the index by traversing the two lists, while the operation of obtaining the length of the memory bank directly returns the length of the queue storing the state transfer experience;
[0060] Constructing network update module and entropy network design:
[0061] Design Actor network update mechanism: Under the MASAC-SEPR algorithm framework, one of its core goals is to explore and find a new strategy π by minimizing the KL divergence. new , the Q value of the new strategy under the new strategy is greater than or equal to the Q value under the old strategy, the Q value is the action value, and the strategy π loss function of the i-th agent is expressed as:
[0062]
[0063] This loss function is used to measure the difference between the current strategy and the target strategy distribution. Optimizing this loss function allows the Actor network to generate a better strategy.
[0064] Among them, J π,i (θ) represents the strategy π loss of the i-th agent with θ as the optimization variable, D KL represents KL divergence, which is used to measure the difference between two probability distributions. θ is the parameter of the Actor network, which represents the learnable parameters of the weights of each layer in the network. By optimizing θ, the output strategy of the Actor network is adjusted. α is used as a temperature coefficient to adjust the entropy of the strategy. π i (·|s t ) indicates that agent i is in state s t The strategy under the given state is the probability distribution of taking different actions, Q i (s t ,·) is the agent i in state s t Under this condition, the Q value function of taking different actions measures the value of the action in this state, Z i (s t ) is used to normalize the distribution function to ensure the rationality of the probability distribution, which is consistent with the agent i in state s t The situation below is closely related, but has no effect on the parameter gradient of the policy network;
[0065] Design the Critic network update mechanism:
[0066] Critic Network’s Q i (s t ,a t ) represents: the i-th agent is in state s t When taking action at The current Q value obtained later is used to evaluate the value of the action in the current state.
[0067] In order to To make an accurate estimate, a neural network represented by the parameter ω is used to Approximate, that is, let Q i (s t ,a t )near In the optimization process, the mean square error is selected as the loss function to measure the degree of deviation between the estimated value and the true target value. Based on the above principle, the loss function of the Critic network of the i-th agent is expressed as:
[0068]
[0069] Among them, J Q,i (ω) represents the Q-value loss of the i-th agent with ω as the optimization variable. ω is the parameter of the Critic network. By optimizing ω, the Critic network can estimate the Q-value more accurately. Ε represents the expectation. (s t ,a t ,s t+1 ): D represents the state s sampled from the priority experience replay memory library D t 、Action a t and the next state s t+1 The sample of the priority experience replay memory D stores the historical data of the interaction between the agent and the environment;
[0070] The parameters of the target Q network (target Critic network) are optimized and updated with the help of the adaptive moment estimation algorithm. The target Q network is the target Critic network, and the parameters of the target Q network are expressed as To indicate that the update is a soft update,
[0071]
[0072] Among them, τ is the soft update rate, ω is the parameter of the Critic network;
[0073] The soft update mechanism allows the target Q network to slowly track the changes of the current network, avoiding the target value The dramatic fluctuations in the training data can improve the stability of the training;
[0074] Design the entropy network module: For agent i, the loss function of its entropy network at time t is expressed as:
[0075]
[0076] Among them, Ji (α) represents the entropy network loss of agent i with α as the optimization variable. α is the temperature coefficient, which plays the role of adjusting the entropy of the strategy. 0 is the predefined minimum policy entropy threshold, Indicates that under strategy π, for action a t Find the expectation, π(a t |s t ) is the agent in state s t Take action a t probability;
[0077] The MASAC-SEPR algorithm model is constructed based on the multi-UAV decision model, airspace reinforcement learning environment, SEAttention module, Actor network, Critic network, entropy network, priority experience replay memory library and network update module; the training method of the MASAC-SEPR algorithm model is designed:
[0078] During the training process of the MASAC-SEPR algorithm model, agent i uses local observation information z i As the input of the Actor policy network integrated with the SEAttention module, the policy network outputs action a after processing through its own internal network layer based on the state space and action space defined by the POMDP model i ,
[0079] The actions of all agents together constitute the action set A. The environment executes the action set A and obtains the state information set S′ at the next moment according to the state transition probability P(s′|s,a) in the POMDP model. At the same time, the reward set R at the current moment is calculated according to the reward function R(s,a). Subsequently, the current state information set S, action set A, reward set R and state information set S′ at the next moment are stored in the priority experience replay memory library in the form of a four-tuple (S,A,R,S′). The priority of each experience sample in the library is determined by the TD error δ. The TD error reflects the difference between the current estimated Q value and the target Q value. Its calculation formula is:
[0080] Where r represents the reward obtained by the agent after taking action at the current moment, γ is the discount factor, θ is the Actor network parameter, and π θ (S′) is the action output by the policy network (Actor network) based on the next state S′ when the parameter is θ. Indicates that under the target Critic network, the agent with index i takes the action π output by the Actor network in the next state S′ θ (S′), the action value output by the target Critic network; Q i(S, A) represents the action value output by the Critic network when the agent with index i takes action set A in the current state S, ω and are the parameters of the Critic network and its target network respectively. By calculating this error, the agent can adjust its behavior strategy to maximize the long-term reward;
[0081] After the introduction of the priority experience replay memory library mechanism, data is extracted according to priority during training to optimize the efficiency of training data utilization.
[0082] During training, a batch of data (s, a, r, s′) is extracted from the priority experience replay memory bank according to the priority, and the agent i combines the state and action information (s i ,a i ) as the input of the Critic network. The Critic network outputs two Q values Q according to the reward mechanism and observation information defined in the POMDP model. i (s t ,a t )and To evaluate the action value, for the Actor network, the state information s extracted from the priority experience replay memory is used as input to participate in the calculation of the strategy loss J π,i , in preparation for updating the parameters of the Actor network and optimizing the strategy output of the Actor network through the back-propagation algorithm;
[0083] The entropy network automatically adjusts the temperature coefficient α, balancing the relationship between the algorithm in exploring new strategies and using existing experience, giving priority to experience replay. The reward value r output by the memory bank is used to calculate Q i (s t ,a t ), Strategy loss J π,i and entropy loss J i (α) Key parameter, entropy loss J i (α) is used to update the entropy network and adjust the temperature coefficient α. As the training progresses, the value of α gradually decreases;
[0084] Finally, the Actor, Critic and entropy networks perform the operation of the network parameter update module, and update the Actor network parameters θ, Critic network parameters ω and entropy coefficient α through the back propagation algorithm. The SE mechanism enables the Actor network to output more reasonable actions, and the Critic network to more accurately evaluate the action value, making the entire update process more efficient and accurate. Finally, with the help of the SE mechanism, we gradually find the approximate optimal path planning strategy and realize the collaborative path planning of multiple UAVs in complex environments. The MASAC-SEPR algorithm model using this training method is named the multi-agent collaborative path optimization network model.
[0085] The training of the multi-agent collaborative path optimization network model in the current airspace reinforcement learning environment includes the following steps:
[0086] In the constructed airspace reinforcement learning environment, each agent in the multi-agent collaborative path optimization network model is trained. Each agent corresponds to a UAV participating in collaborative path planning, and each agent has an independent actor, critic, and entropy network structure.
[0087] Initialize the Actor network, Critic network, and entropy network corresponding to each agent in the multi-agent system, and create a priority experience replay memory library to store the agent's interaction experience, as well as a noise generator, which introduces noise into the agent's action selection in the early stage of training to enhance exploration.
[0088] During the training cycle, each agent selects an action through its own Actor network based on the current perceived environmental state. After executing the action, the agent stores the generated experience in the priority experience playback memory bank. The experience includes the current state, executed action, reward, and next state information.
[0089] When the amount of data in the priority experience replay memory reaches the preset standard, data is sampled from it to calculate the target Q value and current Q value of each agent, and then the strategy loss and entropy loss are obtained; using these calculation results, the parameters of the Actor network, Critic network and entropy network of each agent are updated through the back propagation algorithm to optimize the model and regularly save the parameters of the Actor network;
[0090] Generate a multi-agent collaborative path optimization network model after training optimization in the current airspace reinforcement learning environment, and generate a multi-UAV collaborative path based on the multi-agent collaborative path optimization network model trained in the current airspace reinforcement learning environment.
[0091] The design of the reward mechanism includes the following steps:
[0092] For the boundary reward r eWhen the UAV of the master controller approaches the boundary of the flight area, it will be punished by -1. If it does not approach the boundary, no reward or punishment measures will be implemented.
[0093]
[0094] In the obstacle avoidance reward r o By design, if the distance between the master drone and the obstacle is less than twice the radius of the obstacle, it indicates that the drone has a potential collision risk, and a penalty of -2 is given at this time; when the distance between the two is less than the radius of the obstacle, it means that the drone collides with the obstacle and the drone is damaged, in which case a penalty of -500 is given; in other cases, that is, when the distance is not less than the threshold, no reward or punishment is given to the drone.
[0095]
[0096] Among them, d O Indicates the distance between the master drone and the obstacle, R O Indicates the obstacle radius;
[0097] Construct target reward rule r g ,
[0098] When the distance d between the master UAV and the target point G Shortened to less than a certain threshold D 1 When the drone reaches the target point, it is judged to have reached the target point and a positive reward of 1000 is given. If the distance standard is not reached, a corresponding penalty mechanism is set according to the actual distance between the drone and the target point. The farther the distance, the greater the penalty, so as to guide the drone to approach the target point.
[0099]
[0100] At the same time, design the formation distance reward mechanism r f ,
[0101] When the distance between the master drone and other follower drones is maintained within the preset range, it indicates that the formation is successfully formed. In this case, no reward or punishment operation is performed; on the contrary, once the distance exceeds the range, a penalty is designed according to the distance deviation between the master drone and each follower drone. The greater the distance deviation, the heavier the penalty, so as to encourage the drones to maintain the formation state.
[0102]
[0103] Where, j (j = 1, 2, ..., N-1) is the index of the follower UAV, d jL is the distance between the jth follower UAV and the master UAV, d jL,min ,djL,max are the minimum and maximum distances between the jth follower and the master, r f,j represents the formation distance reward of the jth follower drone, then r f is j=1,2,...,N-1, i.e. the sum of the formation distance rewards from the first agent, the second agent,..., to the N-1th agent;
[0104] Build a speed collaborative reward mechanism s ,
[0105] When the speed error between the master UAV and other follower UAVs is less than the threshold value 1, it indicates that the UAVs have successfully achieved speed coordination. At this time, a reward of 1 is given to encourage maintaining this coordination state. If the speed error is greater than or equal to the threshold value 1, no reward or penalty is given to the UAV.
[0106]
[0107] Where, j (j = 1, 2, ..., N-1) is the index of the follower UAV, v L Indicates the speed of the master UAV, v j represents the speed of the jth follower, r s,j represents the speed coordination reward between the master UAV and the jth follower UAV, then r s It is j=1,2,...,N-1, that is, the sum of the speed coordination rewards of the master UAV and the first agent, the second agent,..., to the N-1th agent.
[0108] The construction of the Actor network and learning includes the following steps:
[0109] Set up the input layer and preliminary feature map:
[0110] When the state information is input to the input layer of the Actor network, it is processed by a linear transformation operation. This linear transformation operation is regarded as a mapping function, which maps the input state information to a 256-dimensional feature space. The weight of the linear mapping adopts a specific initialization strategy, that is, it is initialized according to a normal distribution with a mean of 0 and a standard deviation of 0.1.
[0111] Improve feature dimension and embed SE mechanism:
[0112] The 256-dimensional feature vector is upgraded to a 512-dimensional feature space through a linear transformation operation. In the high-dimensional feature space, the SE attention mechanism is introduced. The SE attention mechanism first performs a global average pooling operation on the 512-dimensional feature space, that is, the feature map of each channel is averaged in the spatial dimension to obtain the global feature information of each channel; then, the attention weight of each channel is generated through a sequence operation consisting of two fully connected layers and corresponding activation functions; the attention weight quantifies the importance of each channel feature for the current task, and by multiplying it element by element with the original 512-dimensional feature vector, the features of the key channels are highlighted and the information of the non-key channels is suppressed;
[0113] Perform feature dimension reduction and action probability distribution determination:
[0114] The 512-dimensional feature vector processed by the SE attention mechanism is subjected to dimensionality reduction operation, and is reduced to a 256-dimensional feature space through a linear transformation. Two linear transformations are used to output the mean and standard deviation of the action respectively. These two statistics are used to define a normal distribution. The agent samples according to the normal distribution to determine the specific action. To ensure the effectiveness and rationality of the action, the sampled action value is constrained and limited to a legal range, thereby ensuring that the agent can make decisions that meet actual needs in the multi-UAV collaborative path planning task.
[0115] The construction of the Critic network and the learning include the following steps:
[0116] Input concatenation and preliminary feature mapping: When processing information, the critic network first concatenates the state information and action information to form an integrated input vector. Then, it processes the concatenated vector using a linear transformation operation and maps it to a 256-dimensional feature space. The weight of the linear transformation is initialized according to a normal distribution with a mean of 0 and a standard deviation of 0.1.
[0117] Feature dimension enhancement and integration with SE mechanism: The Critic network enhances the 256-dimensional feature vector to a 512-dimensional feature space through linear transformation. In this high-dimensional space, the SE attention mechanism is introduced. The SE attention mechanism first performs global average pooling on the 512-dimensional features, integrates the information of each channel in the spatial dimension, and obtains the global feature representation of each channel. Then, after a sequence operation consisting of two fully connected layers and corresponding activation functions, the attention weight of each channel is generated. These weights reflect the importance of each channel feature in the action value assessment task. They are multiplied element by element with the original 512-dimensional feature vector to guide the network to focus on the key features closely related to the action value assessment.
[0118] Perform feature dimensionality reduction and action value assessment:
[0119] The 512-dimensional feature vector processed by the SE attention mechanism is reduced to 256 dimensions through linear transformation. On this basis, the Critic network uses two linear transformations to output two Q values respectively. These two Q values are used to evaluate the value of the agent taking a specific action in a given state.
[0120] Beneficial Effects
[0121] The heterogeneous multi-UAV collaborative path planning method based on multi-agent deep reinforcement learning of the present invention has an airspace reinforcement learning environment that highly simulates real scenes compared to the prior art, providing a reliable basis for algorithm training and improving the practicality and adaptability of the algorithm; the proposed MASAC-SEPR algorithm innovatively incorporates SE mechanism and priority experience playback memory mechanism, enhances feature extraction capabilities, and optimizes the learning process and multi-agent collaborative strategy. The present invention significantly improves the efficiency and accuracy of multi-UAV collaborative path planning, and provides a reliable solution for heterogeneous multi-UAV collaborative operations in dynamic uncertain environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0122] Figure 1 is a method sequence diagram of the present invention;
[0123] Figure 2 This is the core mechanism flow chart of the SEAttention module in the present invention;
[0124] Figure 3 A structural relationship diagram of the Actor network, Critic network and entropy network in the network update and entropy network of the present invention;
[0125] Figure 4 The overall block diagram of the MASAC-SEPR algorithm model training of the present invention;
[0126] Figure 5 , 6 , 7 are respectively the follower UAV sub-reward curve graph, the master controller UAV sub-reward curve graph and the total reward curve graph after algorithm training. DETAILED DESCRIPTION
[0127] In order to have a further understanding and recognition of the structural features and the effects achieved by the present invention, a preferred embodiment and accompanying drawings are used for detailed description as follows:
[0128] like Figure 1 As shown, the heterogeneous multi-UAV collaborative path planning method based on multi-agent deep reinforcement learning of the present invention comprises the following steps:
[0129] 1. Analyze the collaborative path planning task of multiple UAVs.
[0130] In actual application scenarios, after obtaining the UAV group path planning task, it is clear that multiple UAVs need to perform tasks in a bounded airspace environment. In this process, the UAVs must not only avoid various obstacles, but also reach the target location in the form of a single master controller and multiple followers. Among them, the master controller shoulders the key responsibility of avoiding obstacles and leading the formation to the target, while the followers must always closely follow the master controller to maintain a stable formation state to ensure the efficient execution of the task. This task planning method gives full play to the role of different UAVs. By clarifying the division of labor, the entire UAV group can be more orderly when performing tasks. The master controller focuses on avoiding obstacles and leading the direction. Its decisions and actions directly affect the direction and safety of the entire formation; by closely following the master controller, the followers can maintain the integrity of the formation, reduce air resistance during flight, and improve flight efficiency. At the same time, it is also convenient for unified command and coordination, so as to ensure that the task can be completed efficiently and safely.
[0131] 2. Build an airspace reinforcement learning environment.
[0132] (A) Construct a simulated airspace plane.
[0133] The multi-UAV collaborative path planning simulation environment is set as a two-dimensional plane of 700 pixels × 600 pixels, which corresponds to a 7000m × 6000m airspace in reality. The animation refresh rate FPS is set to 60 to simulate a real dynamic environment. In this environment, UAVs are divided into two types: master and follower, and the speed, acceleration, and angular velocity ranges are defined for them respectively. The master UAV speed range is 10-20 pixels per frame, and the acceleration control input range is -3-3m / s 2 , angular velocity range is -0.6-0.6rad / s; follower UAV velocity range is 100-300m / s, acceleration range is -6-6m / s 2 , the angular velocity range is -1.2-1.2rad / s. This setting highly simulates the real UAV flight environment and provides a scenario close to the actual situation for algorithm training. By setting the motion parameter range consistent with the actual situation, the algorithm can fully consider the physical performance limitations of the UAV during the training process, thereby training a path planning strategy that better meets the actual application needs, improving the practicality and reliability of the algorithm.
[0134] (ii) Randomly initialize the environment.
[0135] When the environment is reset at the end of each round, the initial states of all UAVs, including position, speed, heading angle, and the positions of target locations and obstacles, are randomly generated. This random setting greatly increases the uncertainty of the environment and is more in line with actual mission scenarios. In actual applications, the initial conditions and environmental conditions faced by drones each time they perform a mission may be different. By randomly initializing the training environment, the algorithm can learn path planning methods in various complex situations, enhancing the adaptability and generalization ability of the trained model, enabling it to better cope with various changes in reality.
[0136] 3. Introduce the UAV dynamics equation.
[0137] (i) Determine the discrete dynamic equations.
[0138] Taking the dynamic equation as the basic rule of UAV movement, the discrete dynamic equation of multiple UAVs from time t to time t+1 contains the update formula of position, velocity and angle variables. The specific equation is:
[0139]
[0140] All variables in the above equations are between their respective minimum and maximum values;
[0141] Where i is the index of the UAV, (x, y) is the position of the UAV, v is the UAV velocity, ψ is the UAV heading angle, u is the acceleration control input of the UAV, ω is the angular velocity control input, Δt is the time step from time t to time t+1, (x 0 ,y 0 ) is the center position of the obstacle, R 0 is the radius of the obstacle; these equations accurately describe the changes in the motion state of the drone at different times, providing a solid physical basis for path planning. By planning the path based on the dynamic equations, the algorithm can ensure that the movement of the drone complies with the laws of physics and avoid flight paths that are not in line with reality, thereby improving the accuracy and feasibility of path planning.
[0142] (2) Setting up obstacles.
[0143] Assume the obstacle is a standard circular area, with (x 0 ,y 0 ) is the center position of the obstacle, R 0is the radius of the obstacle. In the environment, the drone moves according to the discrete dynamic equations and constraints, and its position, speed and heading angle are constantly updated and always kept within their respective value ranges. Setting the obstacle as a standard circular area simplifies its representation and processing in the path planning environment. Describing obstacle information by the center position and radius of the circle enables the algorithm to more efficiently and accurately determine the positional relationship between the drone and the obstacle, quickly determine whether there is a collision risk and plan an obstacle avoidance path. This helps to improve the drone's obstacle avoidance capability in complex environments and ensure its flight safety.
[0144] (3) Integrate into the environment.
[0145] Integrating dynamic equations and obstacle settings into the airspace reinforcement learning environment makes the movement of drones more consistent with actual physical laws. In this environment, when planning a path, the algorithm will fully consider the physical performance limitations of the drone, calculate the appropriate acceleration and angular velocity, and ensure that the drone can avoid obstacles safely and efficiently. For example, when a drone approaches an obstacle, the algorithm will plan a safe detour route in advance based on the dynamic equations, allowing the drone to continue its mission. This fusion method enhances the authenticity of the environment and the practicality of the algorithm. When planning a path, the algorithm can comprehensively consider the dynamic characteristics of the drone and obstacle information, plan a path that is both safe and efficient, and improve the application value of the algorithm in actual scenarios.
[0146] 4. Modeling the multi-UAV decision model in multi-UAV collaborative path planning as a POMDP model
[0147] 1. Define the basic model.
[0148] The basic decision model of multiple UAVs is defined as the POMDP model: it is defined as a seven-tuple <S,A,R,P,Z,O,γ>, where the state set S contains the state information of all UAVs, but is unknown to the agent; the action set A is the action set of all UAVs, and the actions are known to the agent; the reward function R(s,a) represents the reward obtained by the UAV when taking action a in state s; the state transition probability P(s′|s,a) represents the probability that the UAV takes action a in state s and transfers to state s′, which is unknown to the agent; the observation set Z is the set of observation information of the UAV, which is known to the agent; the observation function O(z|s,a) represents the probability that state s produces observation z after executing action a, which is unknown to the agent; the decay factor γ∈[0,1] represents the preference and proportion of current rewards and future rewards; through such a definition, a unified and standardized theoretical framework is provided for subsequent algorithm learning and decision-making. In this framework, the intelligent agent can make decisions based on partially observable information and continuously learn the optimal strategy through interaction with the environment, thereby achieving collaborative path planning of multiple UAVs in complex environments.
[0149] (ii) Determine the observation, state, and action space.
[0150] For the observation space, the variables in the UAV dynamics equation are introduced, and the observation information of the i-th UAV is set as vector z i =[x i ,y i ,v i ,ψ i ,m i ] means, where [x i ,y i ,v i ,ψ i ] is the flight state of the i-th UAV, m i It is the observation status information related to the UAV type;
[0151] For the master UAV, m i =[x G ,y G ,O flag ], where (x G ,y G ) is the position of the target point, O flag It is the obstacle mark;
[0152] For the follower UAV, m i =[x L ,y L ,v L ], where (x L ,y L ) is the position of the master UAV, v L is the speed of the master UAV. Since there is only one master UAV, the rest of the UAVs are follower UAVs. Each follower UAV has m i =[x L ,y L ,v L ]; where O flag The definition is as follows:
[0153]
[0154] in, is the distance between the UAV and the obstacle, (x, y) represents the position of the UAV, (x 0 ,y 0 ) is the center position of the obstacle; when the distance between the UAV and the obstacle is less than twice the obstacle radius, it is considered that there is an obstacle near the UAV, and the obstacle flag position is 1, otherwise it is set to 0;
[0155] For the state space, the state information S of each UAVi =[z 1 ,z 2 ,...,z N ], including the observation status information of itself and other UAVs, a total of N UAVs;
[0156] For the action space, the action of the i-th UAV is a continuous two-dimensional vector a i =[u i ,w i ] Τ , by choosing the acceleration u i and angular velocity ω i Execute time Δt to achieve the desired speed and heading angle while satisfying the constraints.
[0157] Based on the above design, the observation information obtained by the intelligent agent is more targeted and practical, and can reflect the actual flight status and environmental information of the UAV. Reasonable definition of state space and action space provides a clear scope and method for the decision-making of the intelligent agent, which helps the intelligent agent to choose appropriate actions according to the current state and achieve more effective path planning. For example, the master drone can adjust the flight direction in time by observing the position of the target point and the obstacle mark, leading the formation to avoid obstacles and move towards the target point; the follower drone maintains the formation state with the master drone by observing the position and speed of the master drone.
[0158] 3. Designing a reward mechanism
[0159] 1. Master UAV Reward: For the master UAV, a boundary reward r is designed e , obstacle avoidance reward r o 、Target reward r g , Formation distance reward r f and speed synergy reward r s , whose reward function R L for
[0160] R L =ω L,1 r e +ω L,2 r o +ω L,3 r g +ω L,4 r f +ω L,5 r s ,
[0161] Among them, L represents that the UAV is the master controller, ω L,1 ,ω L,2 ,ω L,3 ,ω L,4 ,ω L,5They represent the boundary reward r of the master UAV respectively. e , obstacle avoidance reward r o 、Target reward r g , Formation distance reward r f , speed synergy reward r s Specifically, for the boundary reward r e When the UAV of the master controller approaches the boundary of the flight area, it will be punished by -1. If it does not approach the boundary, no reward or punishment measures will be implemented.
[0162]
[0163] In the obstacle avoidance reward r o By design, if the distance between the master drone and the obstacle is less than twice the radius of the obstacle, it indicates that the drone has a potential collision risk, and a penalty of -2 is given at this time; when the distance between the two is less than the radius of the obstacle, it means that the drone collides with the obstacle and the drone is damaged, in which case a penalty of -500 is given; in other cases, that is, when the distance is not less than the threshold, no reward or punishment is given to the drone.
[0164]
[0165] Among them, d O Indicates the distance between the master drone and the obstacle, R O Indicates the obstacle radius;
[0166] For the target reward rule r g , when the distance d between the master UAV and the target point G Shortened to less than a certain threshold D 1 When the drone reaches the target point, it is judged to have reached the target point and a positive reward of 1000 is given. If the distance standard is not reached, a corresponding penalty mechanism is set according to the actual distance between the drone and the target point. The farther the distance, the greater the penalty, so as to guide the drone to approach the target point.
[0167]
[0168] At the same time, design the formation distance reward mechanism r f When the distance between the master drone and other follower drones is maintained within the preset range, it indicates that the formation is successfully formed. In this case, no reward or punishment operation is performed; on the contrary, once the distance exceeds the range, a penalty is designed according to the distance deviation between the master drone and each follower drone. The greater the distance deviation, the heavier the penalty, which encourages the drones to maintain the formation state.
[0169]
[0170] Where, j (j = 1, 2, ..., N-1) is the index of the follower UAV, d jL is the distance between the jth follower UAV and the master UAV, d jL,min ,d jL,max are the minimum and maximum distances between the jth follower and the master, r f,j represents the formation distance reward of the jth follower drone, then r f is j=1,2,...,N-1, i.e. the sum of the formation distance rewards from the first agent, the second agent,..., to the N-1th agent;
[0171] For the speed synergy reward mechanism s When the speed error between the master UAV and other follower UAVs is less than the threshold value 1, it indicates that each UAV has successfully achieved speed coordination. At this time, a reward of 1 is given to encourage maintaining this coordination state. If the speed error is greater than or equal to the threshold value 1, no reward or penalty is given to the UAV.
[0172]
[0173] Where, j (j = 1, 2, ..., N-1) is the index of the follower UAV, v L Indicates the speed of the master UAV, v j represents the speed of the jth follower, r s,j represents the speed coordination reward between the master UAV and the jth follower UAV, then r s It is j=1,2,...,N-1, that is, the sum of the speed coordination rewards of the master UAV and the first agent, the second agent,..., to the N-1th agent.
[0174] By setting different rewards and penalties, the behavior of the master drone is guided. Boundary rewards encourage drones to move within the flight area to ensure flight safety; obstacle avoidance rewards encourage drones to avoid obstacles in time and avoid collisions; target rewards guide drones to quickly approach the target point to improve the efficiency of task completion; formation distance rewards and speed coordination rewards help maintain the stability of the formation and ensure the smooth progress of multi-drone collaborative operations. For example, when the master drone approaches the boundary of the flight area, the boundary reward gives a penalty to adjust its flight direction back to the safe area; when the distance to the target point is shortened, the target reward gives a positive incentive to speed up its flight to the target point.
[0175] 2. Follower UAV Reward: For N-1 follower UAVs, the reward function consists of the formation distance reward and the speed coordination reward. The reward function R of each follower UAV j is F,jIt consists of formation distance reward and speed coordination reward.
[0176] That is R F,j =ω F,1 r f,j +ω F,2 r s,j ,
[0177] Where j (j = 1, 2, ..., N-1) is the index of the follower UAV, F represents the UAV is a follower, r f,j represents the formation distance reward of the jth follower drone, r s,j represents the speed coordination reward of the j-th follower UAV, ω F,1 ,ω F,2 They represent the weights of the formation distance reward and speed coordination reward respectively. The design of the formation distance reward and speed coordination reward of the follower drone is the same as that of the master drone, and will not be repeated here.
[0178] Through the above design, the follower drones are encouraged to closely follow the master controller and maintain a stable formation. In actual flight, when encountering interference factors such as airflow that cause the formation to deviate, the reward mechanism will guide the follower drones to adjust their positions and speeds in time to restore the formation, ensuring the orderly operation of multiple drones and improving the coordination and mission execution capabilities of the entire fleet.
[0179] 3. Total reward set: The total reward set R formula is as follows:
[0180] R=[R L ,R F,1 ,R F,2 ,...,R F,N-1 ]
[0181] The reward R of a master UAV L and the rewards R of the N-1 follower UAVs F,j where j (j = 1, 2, ..., N-1) is the index of the follower UAV, indicating the jth follower UAV. Then R F,1 ,R F,2 ,...,R F,N-1 Denote j = 1, 2, ..., N-1, that is, the rewards from the 1st, 2nd ... to the N-1th follower UAV. By setting the reward weight ω, the UAV is guided to learn the optimal path strategy. This enables the entire multi-UAV system to achieve collaborative optimization under a unified reward system. By adjusting the reward weight, the behavior of the UAV can be flexibly guided according to different task requirements, balancing multiple goals such as task completion and formation maintenance, and improving the overall performance of the system.
[0182] 5. Design the MASAC-SEPR algorithm model to perform path planning for multiple UAVs.
[0183] (a) Setting up the SEAttention module.
[0184] In the development process of convolutional neural network (CNN), SENet has significantly improved the representation ability of the network with its innovative design ideas. The "Squeeze-and-Excitation" (SE) block introduced by it greatly enhances the network performance by explicitly modeling the dependencies between convolutional feature channels without increasing the computational cost. This concept is applied to the SEAttention module of this invention, which plays a key role in the collaborative path planning of multiple drones.
[0185] The core mechanism of the SEAttention module is as follows Figure 2 As shown in the figure, it mainly includes three key operations: compression, excitation, and feature recalibration. First, the compression operation is performed, and the input feature map is compressed into a 1×1 vector using adaptive average pooling. This process is similar to the compression operation in SENet. By integrating the spatial dimensions (height H, width W, and number of channels C) of the input feature map, each channel will generate a channel descriptor, and then condense the global spatial information into a channel vector, accurately capturing the global distribution of channel feature responses. This enables the SEAttention module to obtain global information, laying a solid foundation for subsequent feature processing.
[0186] After the compression is completed, the excitation operation phase begins. This operation is implemented through a sequence of fully connected layers, which is similar to the excitation mechanism of SENet. First, the linear layer is used to reduce the number of channels, remove some redundant information, and highlight the key feature factors; then the ReLU activation function is used for nonlinear transformation to explore the complex relationship between features; then another linear layer is used to restore the number of channels, and the channel attention weight is generated with the help of the Sigmoid activation function. In SENet, such an excitation mechanism can deeply explore the nonlinear interaction between channels. The same is true in this module. Different importance coefficients are assigned to each channel through the generated channel attention weights.
[0187] Finally, the generated channel attention weights are multiplied with the original input feature map to achieve feature recalibration and enhance the ability to extract complex environment features. This process is consistent with the principle of using the excitation output to recalibrate the original input feature map in SENet. Each channel of the input feature map is scaled according to the corresponding scalar in the excitation output, highlighting the information-rich features in a targeted manner and suppressing the features that contribute less to the task.
[0188] In a complex drone flight environment, there is a large amount of information, but not all information is equally important for path planning. Through the above operations, the SEAttention module can automatically focus on key environmental features and suppress the interference of irrelevant information. This reflects the advantages of SENet in explicitly modeling channel dependencies and dynamically adjusting feature importance. In the scenario of multi-drone collaborative path planning, drones need to process rich and complex environmental information including obstacle positions, shapes, dynamic changes, and target locations. The SEAttention module can focus on environmental features that play a key role in path planning, just like focusing on key image features in image recognition tasks. For example, in an urban environment, it can quickly identify the features of obstacles such as high-rise buildings, as well as the location information of the target point, ignore some irrelevant background information, avoid wasting computing resources and energy on processing a large amount of irrelevant information, and improve the accuracy and efficiency of decision-making.
[0189] At the same time, the SEAttention module inherits the advantages of SENet's lightweight design, significantly improving performance without increasing the model's burden. Its modular characteristics allow it to be flexibly inserted into a variety of architectures. In the present invention, it can be well integrated into the MASAC-SEPR algorithm model and work in conjunction with other modules. Its powerful generalization ability is also reflected. In different drone flight scenarios, such as mountainous areas, cities, forests, etc., it can adaptively adjust the attention weight according to the input feature map, accurately extract key information, and provide strong support for path planning, ensuring that multiple drones can efficiently plan paths in different scenarios and complete collaborative tasks. In addition, this module strengthens feature representation, allowing the model to understand environmental information more finely, just as SENet enables the model to understand image content more finely, improving the accuracy and robustness of drones in path planning decisions.
[0190] (2) Build an Actor network and learn it.
[0191] 1. Set the input layer and preliminary feature mapping: When the state information is input to the input layer of the Actor network, it is processed with the help of a linear transformation operation. This linear transformation operation maps the input state information to a 256-dimensional feature space. The weights of the linear mapping are initialized according to a normal distribution with a mean of 0 and a standard deviation of 0.1, laying the foundation for subsequent feature processing. Through this initialization method, the network can more effectively learn the characteristics of the state information in the early stages of training, avoiding learning difficulties or slow convergence caused by improper weight initialization. Reasonable weight initialization helps the network converge to the optimal solution faster and improves the learning efficiency and performance of the Actor network.
[0192] 2. Feature dimension enhancement and SE mechanism embedding: A linear transformation operation is used to enhance the 256-dimensional feature vector to a 512-dimensional feature space. In this high-dimensional feature space, the SE attention mechanism is introduced. The SE attention mechanism first performs a global average pooling operation on the 512-dimensional feature space to obtain the global feature information of each channel. Then, the attention weight of each channel is generated through a sequence operation consisting of two fully connected layers and corresponding activation functions. These attention weights quantify the importance of each channel feature for the current task. By multiplying the original 512-dimensional feature vector element by element, the features of the key channels are highlighted and the information of non-key channels is suppressed. The high-dimensional feature space can accommodate richer information, and the embedding of the SE mechanism further enhances the network's ability to mine complex environmental features. By highlighting key features, the Actor network can understand the environmental state more accurately, thereby providing stronger support for outputting reasonable actions, enabling the drone to make action choices that are more in line with actual needs in path planning, and improving the quality of path planning.
[0193] 3. Perform feature dimensionality reduction and determine the probability distribution of actions: The 512-dimensional feature vector processed by the SE attention mechanism is subjected to dimensionality reduction, and is reduced to a 256-dimensional feature space through a linear transformation. Then, two linear transformations are used to output the mean and standard deviation of the action respectively. These two statistics are used to define a normal distribution. The agent samples according to the normal distribution to determine the specific action. To ensure the effectiveness and rationality of the action, the sampled action values are constrained and limited to a legal range to ensure that the agent can make decisions that meet actual needs in the multi-UAV collaborative path planning task. Dimensionality reduction can remove redundant information, improve computational efficiency, and retain key information that is useful for action decision-making. By defining a normal distribution for action sampling, the randomness and exploratory nature of action selection are increased, which helps the agent discover more potential effective actions during the training process. Constraining the action value avoids the agent from outputting unreasonable actions, ensuring the safety and feasibility of the UAV in actual flight.
[0194] (3) Constructing Critic network and learning.
[0195] 1. Input concatenation and preliminary feature mapping: When processing information, the Critic network first concatenates the state information and action information to form an integrated input vector. Subsequently, the concatenated vector is processed using a linear transformation operation and mapped to a 256-dimensional feature space. The weight of the linear transformation is initialized based on a normal distribution with a mean of 0 and a standard deviation of 0.1. Concatenating state and action information as input allows the Critic network to comprehensively consider the relationship between the two and more comprehensively evaluate the value of the action in the current state. Reasonable weight initialization helps the Critic network learn the value characteristics of the state-action pair more quickly, improve the accuracy of action value assessment, and provide more reliable feedback for the Actor network.
[0196] 2. Feature dimension enhancement and SE mechanism fusion: The Critic network enhances the 256-dimensional feature vector to a 512-dimensional feature space through linear transformation. In this high-dimensional space, the SE attention mechanism is introduced. The SE attention mechanism performs global average pooling on the 512-dimensional features, integrates the information of each channel in the spatial dimension, and obtains the global feature representation of each channel. Then, after a sequence operation consisting of two fully connected layers and corresponding activation functions, the attention weight of each channel is generated. These weights reflect the importance of each channel feature in the action value assessment task. They are multiplied element by element with the original 512-dimensional feature vector to guide the network to focus on the key features closely related to the action value assessment. Increasing the feature dimension and integrating the SE mechanism enable the Critic network to mine key information related to the action value assessment in a richer feature space. By focusing on key features, the Critic network can more accurately evaluate the value of the action, reduce the evaluation error, and improve the guiding role of the Actor network strategy optimization, thereby improving the accuracy and efficiency of multi-UAV path planning.
[0197] 3. Perform feature dimensionality reduction and action value assessment: The 512-dimensional feature vector processed by the SE attention mechanism is reduced to 256 dimensions through linear transformation. On this basis, the Critic network uses two linear transformations to output two Q values respectively. These two Q values are used to evaluate the value of the agent taking a specific action in a given state. Outputting two Q values can evaluate the value of the action from different angles, complement and verify each other, and further improve the accuracy and reliability of the action value assessment. This helps the agent to more accurately judge the pros and cons of the current action, thereby selecting a better action, optimizing the path planning strategy of multiple drones, and improving the efficiency of task execution.
[0198] (iv) Design a priority experience replay memory bank mechanism.
[0199] 1. Parameters and initialization: Three key parameters need to be input when the priority experience replay memory mechanism is running: experience replay memory capacity, sampling batch size, and priority index. In the initialization phase, two double-ended queues with the same maximum length (i.e., experience replay memory capacity) are created. One queue is specifically used to store state transition experience, which is recorded in the form of a four-tuple (state, action, reward, next state); the other queue is used to store the priority corresponding to each experience. Reasonable setting of these parameters can balance the storage and use efficiency of experience. The capacity of the experience replay memory determines the number of experiences that can be stored, the sampling batch size affects the number of experiences used in each learning, and the priority index controls the priority of experience sampling. By initializing the double-ended queue, an effective data structure is provided for storing and managing the experience of the agent's interaction with the environment, which facilitates subsequent experience sampling and learning.
[0200] 2. Experience storage: When new state transition experience is generated, first determine whether the queue used to store priority is empty. If it is not empty, obtain the maximum priority in the queue; if it is empty, set the maximum priority to 1.0. After that, add the new state transition experience to the queue for storing state transition experience, and add the obtained or set maximum priority to the queue for storing priority. This storage method can ensure that the newly generated experience is recorded in a timely manner and assign an initial priority to each experience. The setting method of the maximum priority is simple and effective. When there is no historical experience in the initial stage of the experience playback memory library, it ensures that the new experience has a reasonable initial priority, which is convenient for subsequent experience sampling and learning.
[0201] 3. Experience sampling: When sampling experience, the queue storing the priority is processed accordingly according to the length of the queue storing the state transition experience, and converted into a numpy array. Next, the sampling probability of each experience is calculated based on the priority data in the array, and normalized. Subsequently, a specified number of experience indexes are randomly sampled from the queue storing the state transition experience according to the probability, and the corresponding experience samples are obtained according to the index. Finally, the state, action, reward, and next state information are extracted from these samples respectively, and they are converted to torch.FloatTensor type and returned. The priority experience playback mechanism samples according to the priority of the experience, gives priority to important experience for learning, avoids the algorithm wasting time on a large number of unimportant experiences, and improves the utilization efficiency of training data. Through sampling and data processing, valuable information is provided for network updates, which accelerates the learning and optimization process of the network, and enables the algorithm to converge to the optimal strategy faster.
[0202] 4. Priority update: For the priority update operation, the mechanism receives an index list and a new priority list as input. By traversing these two lists, the new priority is updated to the queue storing the priority according to the index, and the operation of obtaining the length of the memory bank directly returns the length of the queue storing the state transfer experience. Timely updating of the priority of the experience can keep the experience priority in the experience playback memory bank consistent with the actual situation. As the algorithm learns and the environment changes, the importance of experience will also change. By updating the priority, it is ensured that important experience can always be sampled and learned first, further improving the learning effect and convergence speed of the algorithm.
[0203] (V) Construct network update module and entropy network design.
[0204] 1. Actor network update mechanism:
[0205] Under the MASAC-SEPR algorithm framework, one of its core goals is to explore and find a new strategy π by minimizing the KL divergence. new , the Q value of the new strategy under the new strategy is greater than or equal to the Q value under the old strategy, the Q value is the action value, and the strategy π loss function of the i-th agent is expressed as:
[0206]
[0207] This loss function is used to measure the difference between the current strategy and the target strategy distribution. By optimizing this loss function, the Actor network can generate a better strategy. π,i (θ) represents the strategy π loss of the i-th agent with θ as the optimization variable, D KL represents KL divergence, which is used to measure the difference between two probability distributions. θ is the parameter of the Actor network, which represents the learnable parameters of the weights of each layer in the network. By optimizing θ, the output strategy of the Actor network is adjusted. α is used as a temperature coefficient to adjust the entropy of the strategy. π i (·|s t ) indicates that agent i is in state s t The strategy under the given state is the probability distribution of taking different actions, Q i (s t ,·) is the agent i in state s t Under this condition, the Q value function of taking different actions measures the value of the action in this state, Z i (s t ) is used to normalize the distribution function to ensure the rationality of the probability distribution, which is consistent with the agent i in state s t The situation below is closely related, but has no effect on the parameter gradient of the policy network.
[0208] Minimizing the KL divergence can prompt the Actor network to continuously adjust its strategy to make it closer to the target strategy distribution. By optimizing the strategy loss function, the Actor network can dynamically adjust the action output according to environmental changes, improving the path planning ability of the drone in complex environments. In a complex and changing flight environment, the Actor network can learn a more reasonable action selection strategy, enabling multiple drones to complete collaborative tasks and path planning tasks more efficiently, and improving the success rate and efficiency of task execution.
[0209] 2. Critic network update mechanism:
[0210] Critic Network’s Q i (s t ,a t ) represents: the i-th agent is in state s t When taking action a t The current Q value obtained later is used to evaluate the value of the action in the current state.
[0211] In order to To make an accurate estimate, a neural network represented by the parameter ω is used to Approximate, that is, let Q i (s t ,a t )near In the optimization process, the mean square error is selected as the loss function to measure the degree of deviation between the estimated value and the true target value. Based on the above principle, the loss function of the Critic network of the i-th agent is expressed as:
[0212]
[0213] Among them, J Q,i (ω) represents the Q-value loss of the i-th agent with ω as the optimization variable. ω is the parameter of the Critic network. By optimizing ω, the Critic network can estimate the Q-value more accurately. Ε represents the expectation. (s t ,a t ,s t+1 ): D represents the state s sampled from the priority experience replay memory library D t 、Action a t and the next state s t+1 The priority experience replay memory D stores the historical data of the interaction between the agent and the environment.
[0214] The mean square error loss function can effectively measure the degree of deviation between the Critic network's estimated value and the true target value. By continuously optimizing the parameters of the Critic network, it can estimate the Q value more accurately, thereby providing more reliable action value evaluation feedback to the Actor network. Accurate Q value evaluation helps the Actor network optimize its strategy, avoid choosing low-value actions, further improve the accuracy and efficiency of multi-UAV path planning, and ensure that UAVs can choose the best path to reach the target location.
[0215] 3. Target Q network (target Critic network) parameter optimization update:
[0216] The parameters of the target Q network (target Critic network) are optimized and updated with the help of the adaptive moment estimation algorithm. To indicate that the update is a soft update,
[0217]
[0218] Among them, τ is the soft update rate, ω is the parameter of the Critic network;
[0219] The soft update mechanism allows the target Q network to slowly track the changes of the current network, avoiding the target value
[0220] If soft update is not used, the target value may change drastically with the rapid update of the Critic network, resulting in unstable training and difficulty in converging to the optimal strategy.
[0221] 4. Entropy network module design:
[0222] For agent i, the loss function of its entropy network at time t is expressed as:
[0223]
[0224] Among them, J i (α) represents the entropy network loss of agent i with α as the optimization variable. α is the temperature coefficient, which plays the role of adjusting the entropy of the strategy. 0 is the predefined minimum policy entropy threshold, Indicates that under strategy π, for action a t Find the expectation, π(a t |s t ) is the agent in state s t Take action a tThe function of the entropy network is to automatically adjust the temperature coefficient α to balance the relationship between the algorithm's exploration of new strategies and the use of existing experience. When the entropy of the strategy is low, it means that the agent's action selection is more concentrated and may be trapped in a local optimum. At this time, the entropy network will increase α to encourage the agent to explore more different actions; when the entropy of the strategy is high, it means that the agent's action selection is too random. At this time, the entropy network will reduce α, making the agent more inclined to use existing experience and choose actions with higher value. As training progresses, the α value gradually decreases, allowing the algorithm to fully explore the environment in the early stages of training, and pay more attention to using the learned knowledge in the later stages.
[0225] Figure 3 The architecture and relationship between the entropy network, actor network and critic network are shown. The main function of the entropy network is to adjust the temperature coefficient α to balance the relationship between the algorithm exploring new strategies and using existing experience. In the early stage of training, a larger temperature coefficient encourages the agent to actively explore new strategies and expand the action space; as the training progresses, the temperature coefficient gradually decreases, and the agent is more inclined to use existing experience to select high-value actions. J(α) in the figure represents the loss function of the entropy network. The temperature coefficient is adjusted by optimizing the loss function to make it better adapt to the training process. The input layer of the actor network is used to receive state information, which is processed by the hidden layer, and the SEAttention module is embedded in it. This module enhances the ability to extract complex environmental features through compression, excitation and feature recalibration operations, helping the actor network to output more reasonable actions. J(θ) in the figure is the loss function of the actor network. By minimizing the loss function, the network parameter θ is updated using the back propagation algorithm to optimize the policy output of the actor network, so that the agent can make better action choices in the environment. The input layer of the critic network is used to receive the information after the state and action are spliced. The hidden layer also contains the SEAttention module to better evaluate the value of the action. The Critic network outputs two Q values to evaluate the value of the agent taking a specific action in a given state. J(ω) in the figure is the loss function of the Critic network. The Critic network is allowed to approach the target value through means such as mean square error, and the parameter ω is optimized to estimate the Q value more accurately, providing reliable action value evaluation feedback for the Actor network. The temperature coefficient α output by the entropy network will affect the training process of the Actor network and the Critic network, helping to balance exploration and utilization. The Actor network outputs actions based on the state of the environment, and the Critic network evaluates the value of these actions. The evaluation results are fed back to the Actor network for updating the strategy. The three work together to optimize the algorithm model to achieve efficient collaborative path planning for multiple drones in complex environments.
[0226] (6) Design the MASAC-SEPR algorithm model and its training method to generate a multi-agent collaborative path optimization network model.
[0227] The MASAC-SEPR algorithm model is constructed based on the multi-UAV decision model, airspace reinforcement learning environment, SEAttention module, Actor network, Critic network, entropy network, priority experience replay memory and network update module; this multi-module fusion design gives full play to the advantages of each module. The multi-UAV decision model provides the construction and theoretical framework of the UAV model, the airspace reinforcement learning environment simulates the real scene, the SEAttention module enhances the ability to extract environmental features, the Actor network is responsible for generating actions, the Critic network evaluates the value of actions, the entropy network balances exploration and utilization, the priority experience replay memory improves learning efficiency, and the network update module optimizes network parameters. Through the interaction between the agent and the environment, experience replay, network update and other processes, the algorithm can continuously learn and optimize the path planning strategy, and realize efficient collaborative path planning of multiple UAVs in complex environments.
[0228] At the same time, the training method of the MASAC-SEPR algorithm model is designed: Figure 4 As shown in the figure, during the model training process, agent i uses local observation information z i As the input of the Actor policy network that integrates the SEAttention module. The policy network outputs action a after processing through its own internal network layer based on the state space and action space defined by the POMDP model. i .
[0229] The actions of all agents together constitute the action set A. The environment executes the action set A and obtains the state information set S′ at the next moment according to the state transition probability P(s′|s,a) in the POMDP model. At the same time, the reward set R at the current moment is calculated according to the reward function R(s,a). Subsequently, the current state information set S, action set A, reward set R and state information set S′ at the next moment are stored in the priority experience replay memory library in the form of a four-tuple (S,A,R,S′). The priority of each experience sample in the library is determined by the TD error δ. The TD error reflects the difference between the current estimated Q value and the target Q value. Its calculation formula is:
[0230] Where r represents the reward obtained by the agent after taking action at the current moment, γ is the discount factor, θ is the Actor network parameter, and π θ (S′) is the action output by the policy network (Actor network) based on the next state S′ when the parameter is θ. Indicates that under the target Critic network, the agent with index i takes the action π output by the Actor network in the next state S′ θ (S′), the action value output by the target Critic network; Q i (S, A) represents the action value output by the Critic network when the agent with index i takes action set A in the current state S, ω and are the parameters of the Critic network and its target network respectively. By calculating this error, the agent can adjust its behavior strategy to maximize the long-term reward;
[0231] This training method closely combines the multi-agent system, POMDP model and priority experience replay memory. The agent continuously accumulates experience through interaction with the environment, TD error is used to measure the pros and cons of the current strategy, and the priority experience replay memory prioritizes the experience according to the TD error, so that important experience can be learned first, which accelerates the convergence speed of the algorithm and improves the training efficiency.
[0232] During training, a batch of data (s, a, r, s′) is extracted from the priority experience replay memory bank according to the priority, and the agent i combines the state and action information (s i ,a i ) as the input of the Critic network. The Critic network outputs two Q values Q according to the reward mechanism and observation information defined in the POMDP model. i (s t ,a t )and To evaluate the action value, for the Actor network, the state information s extracted from the priority experience replay memory is used as input to participate in the calculation of the strategy loss J π,i , in preparation for updating the parameters of the Actor network and optimizing the strategy output of the Actor network through the back-propagation algorithm;
[0233] Finally, the Actor, Critic and Entropy Networks perform the network parameter update module, and update the Actor network parameters θ, Critic network parameters ω and entropy coefficient α through the back propagation algorithm. The SE mechanism enables the Actor network to output more reasonable actions, and the Critic network to more accurately evaluate the action value, making the entire update process more efficient and accurate. In the continuous training iteration process, each network cooperates with each other for optimization, gradually finds the approximate optimal path planning strategy, and realizes the collaborative path planning of multiple UAVs in complex environments. The MASAC-SEPR algorithm model based on the above training method is named the multi-agent collaborative path optimization network model. The introduction of the priority experience replay memory library mechanism improves the utilization efficiency of training data, and the SE mechanism enhances the network's ability to extract and process environmental features, so that the Actor network can output more reasonable actions, and the Critic network can more accurately evaluate the action value, making the entire update process more efficient and accurate. Through continuous training iterations, each network cooperates with each other for optimization, gradually finds the approximate optimal path planning strategy, and realizes the efficient collaborative path planning of multiple UAVs in complex environments.
[0234] 6. Training of multi-agent collaborative path optimization network model.
[0235] (a) Initialization
[0236] In the established airspace reinforcement learning environment, each agent in the multi-agent collaborative path optimization network model is trained. Each agent corresponds to a drone participating in collaborative path planning, and each has an independent Actor, Critic, and entropy network structure. Initialize the Actor network, Critic network, and entropy network corresponding to each agent in the multi-agent system, and create a priority experience replay memory library to store the interactive experience of the agent, as well as a noise generator. The noise generator introduces noise to the action selection of the agent in the early stage of training to enhance exploration and prevent the algorithm from falling into a local optimal solution. For example, in the early stage of training, the noise generator will add a certain amount of random perturbation to the action selected by the agent, allowing the agent to try different action combinations and explore more possible paths.
[0237] (ii) Training cycle
[0238] During the training cycle, each agent selects an action through its own Actor network based on the current perceived state of the environment. After executing the action, the agent stores the generated experience (including the current state, executed action, reward obtained, and next state information) into the priority experience replay memory. As the training progresses, the experience replay memory continues to accumulate experience data. This experience data contains the action selection and environmental feedback of the agent in different states, and is the basis for algorithm learning and optimization. By continuously accumulating experience, the agent can better understand the environment, learn more effective path planning strategies, and gradually improve its path planning capabilities.
[0239] (iii) Network updates.
[0240] When the amount of data in the priority experience replay memory reaches the preset standard, data is sampled from it to calculate the target Q value and current Q value of each agent, and then the strategy loss and entropy loss are obtained. Using these calculation results, the parameters of the Actor network, Critic network, and entropy network of each agent are updated through the back propagation algorithm to optimize the model. In order to save the effective results of the training process, the parameters of the Actor network are saved regularly. In this way, when the training is interrupted or the effects of different training stages need to be compared, it is convenient to restore to the previous training state.
[0241] (IV) Model generation and path planning.
[0242] After multiple rounds of training, a multi-agent collaborative path optimization network model is generated after training optimization in the current airspace reinforcement learning environment. Based on the trained model, the actual mission scenario information, such as the starting position, target position, obstacle distribution, etc., is input, and the model can generate a multi-UAV collaborative path, enabling efficient and safe collaborative flight of multiple UAVs in complex environments to complete the mission objectives. The trained and optimized model can generate a reasonable collaborative path based on the actual mission scenario information, which is an important manifestation of the practicality of the algorithm. In actual applications, UAVs can fly efficiently and safely in complex environments according to the path planning generated by the model to complete various tasks, providing reliable technical support for practical applications.
[0243] (V) Comparative experiments and superiority verification.
[0244] In order to verify the performance of the MASAC-SEPR algorithm in the collaborative path planning task of heterogeneous multi-UAVs, a series of experiments were conducted and compared with the random strategy, MADDPG, and MASAC algorithms. The experimental results are shown in Table 1. The experimental results show that under the same parameter conditions, the MASAC-SEPR algorithm performs well in many aspects and can effectively cope with the challenges of collaborative path planning of heterogeneous multi-UAVs in dynamic uncertain environments.
[0245] Table 1. Performance comparison of MASAC-SEPR algorithm and other strategies
[0246]
[0247] Task completion rate: From the algorithm comparison data (see Table 1), the task completion rate of the random strategy is only 2.2%, which is almost impossible to complete the task; the task completion rate of the MADDPG algorithm is 81.0%, which has achieved certain results; the task completion rate of the MASAC algorithm has increased to 90.0%; and although the task completion rates of the MASAC-SEPR algorithm and the MASAC algorithm are both 90.0%, combined with the actual scene analysis, the MASAC-SEPR algorithm has more advantages in adaptability to complex environments. For example, in a scene with multiple obstacles and dynamically changing target positions, the MASAC-SEPR algorithm can more accurately capture environmental features with the SE mechanism, guide the drone to quickly avoid obstacles and plan a reasonable path to reach the target point. In contrast, the MASAC algorithm will have poor planning paths in some complex scenes, resulting in the inability to complete the task. This shows that the MASAC-SEPR algorithm is more reliable in ensuring task completion.
[0248] Formation retention rate: As shown in Table 1, the algorithms differ significantly in the formation retention rate indicator. The random strategy formation retention rate is 0.00, and formation cannot be achieved; the MADDPG algorithm formation retention rate is 0.67; the MASAC algorithm formation retention rate is 35.56%; the MASAC-SEPR algorithm formation retention rate reaches 41.27%, the highest among all algorithms. This is due to the formation distance reward and speed coordination reward mechanism designed by the MASAC-SEPR algorithm, as well as the collaborative learning and experience sharing method among agents. In the actual flight process, when encountering interference factors such as sudden airflow, the agent of the MASAC-SEPR algorithm can quickly adjust its own actions according to the shared experience and reward mechanism, maintain a reasonable distance and speed coordination with other drones, effectively maintain the stability of the formation, and ensure the orderly progress of multi-drone collaborative operations.
[0249] Flight time: The flight time reflects the efficiency of the algorithm in planning the path. As shown in Table 1, the flight time of the random strategy is as high as 934.67, the flight time of the MADDPG algorithm is 330.52, the flight time of the MASAC algorithm is 249.91, and the flight time of the MASAC-SEPR algorithm is 187.01, which is the shortest among all algorithms. This is mainly because the SE mechanism of the MASAC-SEPR algorithm enhances the ability to extract environmental features, enabling the drone to plan the optimal path to the target point more quickly and reduce unnecessary flight path detours. At the same time, the priority experience replay memory library mechanism speeds up the learning convergence speed of the algorithm, allowing the intelligent agent to master efficient path planning strategies more quickly, thereby significantly shortening the flight time of the drone and improving operational efficiency.
[0250] Flight trajectory and energy consumption: As shown in Table 1, in terms of flight trajectory, the flight trajectory length of the MASAC-SEPR algorithm is 100.34, which is shorter than other algorithms. A shorter flight trajectory means that the drone is more efficient during flight and reduces unnecessary flight paths, which echoes the results of flight time. In terms of energy consumption, the MASAC-SEPR algorithm is 232.95, which is also the lowest among several algorithms. This is because its optimized path planning strategy makes the changes in acceleration and angular velocity of the drone during flight more reasonable, avoiding frequent and large adjustments, thereby reducing energy consumption, reflecting the advantages of the algorithm in path planning quality and energy utilization efficiency.
[0251] Algorithm convergence and stability: By observing the follower reward, leader reward and total reward convergence curves of the MASAC-SEPR algorithm in the experiment (corresponding to the experimental results Figure 5 , Figure 6 , Figure 7 ), it can be clearly seen that the algorithm has significant advantages.
[0252] In the early stage of training, for the follower reward ( Figure 5 ), the curve fluctuates greatly and the reward values are mostly negative, because the agent is still exploring the environment and the strategy has not been optimized. However, with the priority experience replay memory mechanism, the algorithm can prioritize important experiences from a large number of experiences for learning, so that the reward value fluctuates but shows an overall upward trend, which shows that the follower agent is quickly adjusting its strategy to adapt to the environment. Compared with other algorithms that lack this mechanism, the follower may be in an inefficient exploration for a long time, and the reward value is difficult to improve.
[0253] Leader Reward Curve( Figure 6) also fluctuated violently in the early stages of training, with the reward value changing drastically between positive and negative. However, with the help of priority experience replay, the MASAC-SEPR algorithm enables the leader agent to quickly learn effective leadership strategies and reduce ineffective actions. As training progresses, the reward value gradually stabilizes at a high level, demonstrating that the algorithm allows the leader agent to efficiently guide collaborative tasks. In contrast, in some traditional algorithms, the leader agent may fail frequently during the guidance process due to the inability to quickly acquire key experience, and the reward value is difficult to stabilize.
[0254] Total Reward Curve( Figure 7 ) also fluctuated significantly in the early stage, reflecting the instability of the entire multi-agent system in the early stage of training. However, due to the priority experience replay memory mechanism, the system can quickly focus on important experiences, so that the total reward value gradually increases. This shows that the mechanism effectively promotes the coordination between agents and accelerates the optimization of the overall strategy of the system.
[0255] As training progresses, the SE mechanism begins to play a key role. Figure 5 ), after about 200 trainings, the curve becomes stable and the reward value stabilizes in a range close to 0. This is because the SE mechanism helps the follower agent to better understand the environmental information, accurately judge the action value, reduce unnecessary mistakes, and thus stably output a better strategy. For other algorithms that are not capable of extracting complex environmental features, the follower agent may continue to be affected by interference factors in the environment, and the reward value is difficult to stabilize.
[0256] Leader Reward Curve( Figure 6 ) In the later stages of training, it stabilizes at a relatively high level of around 1000. This is due to the SE mechanism, which enables the leader agent to accurately capture key information in the environment and make more reasonable decisions to guide followers, thus enhancing the stability and efficiency of the entire collaborative system. Without this mechanism, the leader agent may not be able to effectively respond to complex environmental changes, resulting in large fluctuations in reward values and difficulty in reaching a high level.
[0257] Total Reward Curve( Figure 7 ) stabilized at around 400 in the later stage, indicating that under the SE mechanism, the entire multi-agent system has a more tacit collaboration between agents, can make full use of environmental information, reduce resource waste, and achieve efficient collaborative work. However, the comparison algorithm cannot effectively handle environmental characteristics, the coordination of agents in the system is chaotic, and the total reward value is difficult to improve and is unstable.
[0258] In summary, judging from the convergence curves of the three dimensions of follower reward, leader reward and total reward, the synergy between the priority experience replay memory mechanism and the SE mechanism of the MASAC-SEPR algorithm makes it far superior to similar algorithms in terms of convergence speed and stability. It can more efficiently realize multi-agent collaborative strategy optimization and adapt to complex environments, showing significant advantages.
[0259] Adaptability to complex scenarios: The MASAC-SEPR algorithm has shown good adaptability in a variety of complex test scenarios designed, including obstacles of different numbers and distributions, dynamically changing target positions, different initial states of drones, and heterogeneous characteristics. When faced with dense obstacle scenarios, the algorithm can adjust the path of the drone in a timely manner to avoid collisions; when the target position changes dynamically, it can quickly replan the path to ensure that the drone always moves towards the target. In contrast, random strategies can hardly complete the task in complex scenarios, and the performance of the MADDPG and MASAC algorithms will show a significant decline in some complex scenarios, further verifying the effectiveness and superiority of the MASAC-SEPR algorithm in complex environments.
[0260] In summary, the MASAC-SEPR algorithm performs well in performance indicators such as mission completion rate, formation retention rate, flight time, flight trajectory and energy consumption. It also has obvious advantages in algorithm convergence speed and adaptability to complex scenarios. It effectively realizes efficient and accurate collaborative path planning of heterogeneous multi-UAVs, achieves the expected design goals, and provides reliable algorithm support for practical applications.
[0261] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions only describe the principles of the present invention. The present invention may be subject to various changes and improvements without departing from the spirit and scope of the present invention. These changes and improvements fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the attached claims and their equivalents.
Claims
1. A heterogeneous multi-UAV collaborative path planning method based on multi-agent deep reinforcement learning, characterized in that: The following steps are involved: 11) Analyze the collaborative path planning task of multiple UAVs; 12) Build an airspace reinforcement learning environment; 13) Introduce the UAV dynamics equation; 14) Model the multi-UAV decision model in multi-UAV collaborative path planning as a POMDP model; 15) Design the MASAC-SEPR algorithm model to perform path planning for multiple UAVs and generate a multi-agent collaborative path optimization network model; 16) Training of multi-agent collaborative path optimization network model: The multi-agent collaborative path optimization network model is trained in the current airspace reinforcement learning environment to generate multi-UAV collaborative paths.
2. The method for heterogeneous multi-UAV collaborative path planning based on multi-agent deep reinforcement learning according to claim 1 is characterized in that: The analysis of the multi-UAV collaborative path planning task is: obtaining the UAV group path planning task, multiple UAVs avoid obstacles in a bounded airspace environment, and arrive at the target location in the form of a single master controller and multiple followers, where the master controller is responsible for avoiding obstacles and reaching the target, and the followers follow the single master controller to maintain the formation.
3. The method for heterogeneous multi-UAV collaborative path planning based on multi-agent deep reinforcement learning according to claim 1 is characterized in that: The construction of the airspace reinforcement learning environment is to construct a simulated airspace plane and define the speed, acceleration, and angular velocity ranges of the master UAV and the follower UAV; including the following steps: 31) The multi-UAV collaborative path planning simulation environment is defined as a two-dimensional plane of 700 pixels × 600 pixels, corresponding to the airspace of 7000m × 6000m in reality; 32) Set the animation refresh rate FPS to 60. In this environment, drones are divided into two types: master and follower. The master UAV speed range is 10-20 pixels per frame, and the acceleration control input range is -3-3m / s 2 , angular velocity range is -0.6-0.6rad / s, follower UAV velocity range is 100-300m / s, acceleration range is -6-6m / s 2 , the angular velocity range is -1.2-1.2rad / s; 33) When the environment is reset at the end of each round, the initial states of all UAVs, i.e., the x position, y position, velocity v and heading angle ψ, the target location and the position of obstacles are randomly generated.
4. The method for heterogeneous multi-UAV collaborative path planning based on multi-agent deep reinforcement learning according to claim 1 is characterized in that: The introduction of the UAV dynamics equation is: taking the dynamics equation as the basic rule of UAV motion, the discrete dynamics equation of multiple UAVs from time t to t+1 contains the update formula of position, speed, and angle variables, and satisfies the constraints of flight state, control quantity, and flight path; including the following steps: 41) Taking the dynamic equation as the basic rule of UAV motion, the discrete dynamic equation of multiple UAVs from time t to time t+1 is: All variables in the above equations are between their respective minimum and maximum values; Where i is the index of the UAV, (x, y) is the position of the UAV, v is the UAV velocity, ψ is the UAV heading angle, u is the acceleration control input of the UAV, ω is the angular velocity control input, and Δt is the time step from time t to time t+1; 42) Set the obstacle as a standard circular area. (x0, y0) is the center position of the obstacle, and R0 is the radius of the obstacle; In the environment, the drone moves according to discrete dynamic equations and constraints, and its position, velocity, and heading angle states are continuously updated, and the values are always within their respective value ranges; 43) The dynamic equations and obstacle settings are added to the airspace reinforcement learning environment, and the drone uses the dynamic equations as the basic rules for drone movement.
5. The method for heterogeneous multi-UAV collaborative path planning based on multi-agent deep reinforcement learning according to claim 1 is characterized in that: Modeling the multi-UAV decision model in the multi-UAV collaborative path planning as a POMDP model includes the following steps: 51) Define the multi-UAV basic decision model as the POMDP model: Defined as a seven-tuple <S,A,R,P,Z,O,γ>, Among them, the state set S contains the state information of all UAVs, but it is unknown to the agent; the action set A is the action set of all UAVs, and the actions are known to the agent; the reward function R(s,a) represents the reward obtained by the UAV when taking action a in state s; the state transition probability P(s′|s,a) represents the probability that the UAV takes action a in state s and transfers to state s′, which is unknown to the agent; the observation set Z is the set of observation information of the UAV, which is known to the agent; the observation function O(zs,a) represents the probability that state s produces observation z after executing action a, which is unknown to the agent; the decay factor γ∈[0,1] represents the preference and proportion of current rewards and future rewards; 52) For the observation space, introduce the variables in the UAV dynamics equation and set the observation information of the i-th UAV as vector z i =[x i ,y i ,v i ,ψ i ,m i ] means, where [x i ,y i ,v i ,ψ i ] is the flight state quantity of the i-th UAV, mi is the observation state information related to the UAV type; For the master UAV, m i =[x G ,y G ,O flag ], where (x G ,y G ) is the position of the target point, O flag It is the obstacle mark; For the follower UAV, m i =[x L ,y L ,v L ], where (x L ,y L ) is the position of the master UAV, v L is the speed of the master UAV. Since there is only one master UAV, the rest of the UAVs are follower UAVs. Each follower UAV has m i =[x L ,y L ,v L ]; where O flag The definition is as follows: in, is the distance between the UAV and the obstacle, (x, y) represents the position of the UAV, and (x0, y0) is the center position of the obstacle. When the distance between the UAV and the obstacle is less than twice the obstacle radius, it is considered that there is an obstacle near the UAV. At this time, the obstacle mark position is 1, otherwise it is set to 0. For the state space, the state information S of each UAV i =[z1,z2,...,z N ], including the observation status information of itself and other UAVs, a total of N UAVs; For the action space, the action of the i-th UAV is a continuous two-dimensional vector a i =[u i ,w i ] Τ , by choosing the acceleration u i and angular velocity ω i Execute time Δt to achieve the desired speed and heading angle while satisfying the constraints; 53) Design reward mechanisms and reward and punishment mechanisms for different types of UAVs. For the master UAV, including the boundary reward r e , obstacle avoidance reward r o 、Target reward r g , Formation distance reward r f and speed synergy reward r s , whose reward function R L for R L =ω L,1 r e +oh L,2 r o +oh L,3 r g +oh L,4 r f +oh L,5 r s , Among them, L represents that the UAV is the master controller, ω L,1 ,ω L,2 ,ω L,3 ,ω L,4 ,ω L,5 They represent the boundary reward r of the master UAV respectively. e , obstacle avoidance reward r o 、Target reward r g , Formation distance reward r f , speed synergy reward r s The weight of For N-1 follower UAVs, the reward function consists of the formation distance reward and the speed coordination reward. The reward function R of each follower UAV j is F,j It consists of formation distance reward and speed coordination reward. i.e., R F,j = ω F,1 r f,j + ω F,2 r s,j , Where j (j = 1, 2, ..., N-1) is the index of the follower UAV, F represents the UAV is a follower, r f,j represents the formation distance reward of the jth follower drone, r s,j represents the speed coordination reward of the j-th follower UAV, ω F,1 ,ω F,2 Represent the weights of formation distance reward and speed coordination reward respectively, Total reward set R = [R L ,R F,1 ,R F,2 ,...,R F,N-1 ], that is, the reward R of a master UAV L and the rewards R of the N-1 follower UAVs F,j where j (j = 1, 2, ..., N-1) is the index of the follower UAV, indicating the jth follower UAV. Then R F,1 ,R F,2 ,...,R F,N-1 Represents j=1,2,...,N-1, i.e., the rewards from the 1st, 2nd... to the N-1th follower UAV. By setting the reward weight ω, the UAV is guided to learn the optimal path strategy.
6. The method for heterogeneous multi-UAV collaborative path planning based on multi-agent deep reinforcement learning according to claim 1 is characterized in that: The design of the MASAC-SEPR algorithm for path planning of multiple UAVs includes the following steps: 61) Set up the SEAttention module: The SEAttention module first performs a compression operation, using adaptive average pooling to compress the input feature map into a 1×1 vector to obtain global information. Each channel generates a channel descriptor, condenses the global spatial information into a channel vector, and captures the global distribution of channel feature responses. Then, it performs an excitation operation through a sequence of fully connected layers, first using a linear layer to reduce the number of channels, then performing a nonlinear transformation through a ReLU activation function, and then restoring the number of channels through another linear layer. The channel attention weight is generated with the help of a Sigmoid activation function. Finally, the generated channel attention weight is multiplied with the original input feature map to achieve feature recalibration. 62) Build an Actor network and learn: 63) Building Critic Network and Learning; 64) Design of a priority experience replay memory bank mechanism: Prioritized Experience Replay Memory Mechanism The mechanism inputs three key parameters at runtime: experience replay memory capacity, sampling batch size, and priority index; In the initialization phase, two double-ended queues with the same maximum length, i.e., the capacity of the experience replay memory bank, are created. One of the queues is used to store state transition experiences, which are recorded in the form of four-tuples (state, action, reward, next state); the other queue is used to store the priority corresponding to each experience. When a new state transition experience is generated, the storage process is as follows: first determine whether the queue for storing priorities is empty. If not, obtain the maximum priority in the queue; if it is empty, set the maximum priority to 1.
0. Then, add the new state transition experience to the queue for storing state transition experience, and add the obtained or set maximum priority to the queue for storing priorities. When performing experience sampling, the specific steps are as follows: according to the length of the queue storing state transfer experience, the queue storing priority is processed accordingly and converted into a numpy array. Then, according to the priority data in the array, the sampling probability of each experience is calculated and normalized. Then, a specified number of experience indexes are randomly sampled from the queue storing state transfer experience according to the probability, and the corresponding experience samples are obtained according to the index. Finally, the state, action, reward and next state information are extracted from these samples respectively, and they are converted into torch.FloatTensor type and returned. For the priority update operation, the mechanism receives the index list and the new priority list as input, and updates the new priority to the queue storing the priority according to the index by traversing the two lists, while the operation of obtaining the length of the memory bank directly returns the length of the queue storing the state transfer experience; 65) Construct network update module and entropy network design: 651) Design of Actor Network Update Mechanism: Under the framework of MASAC-SEPR algorithm, one of its core goals is to explore and find a new strategy π by minimizing KL divergence. new , the Q value of the new strategy under the new strategy is greater than or equal to the Q value under the old strategy, the Q value is the action value, and the strategy π loss function of the i-th agent is expressed as: This loss function is used to measure the difference between the current strategy and the target strategy distribution. Optimizing this loss function allows the Actor network to generate a better strategy. Among them, J π,i (θ) represents the strategy π loss of the i-th agent with θ as the optimization variable, D KL represents KL divergence, which is used to measure the difference between two probability distributions. θ is the parameter of the Actor network, which represents the learnable parameters of the weights of each layer in the network. By optimizing θ, the output strategy of the Actor network is adjusted. α is used as a temperature coefficient to adjust the entropy of the strategy. π i (·|s t ) indicates that agent i is in state s t The strategy under the given state is the probability distribution of taking different actions, Q i (s t ,·) is the agent i in state s t Under this condition, the Q value function of taking different actions measures the value of the action in this state, Z i (s t ) is used to normalize the distribution function to ensure the rationality of the probability distribution, which is consistent with the agent i in state s t The situation below is closely related, but has no effect on the parameter gradient of the policy network; 652) Design Critic network update mechanism: Critic Network’s Q i (s t ,a t ) represents: the i-th agent is in state s t When taking action a t The current Q value obtained later is used to evaluate the value of the action in the current state. In order to To make an accurate estimate, a neural network represented by the parameter ω is used to Approximate, that is, let Q i (s t ,a t )near In the optimization process, the mean square error is selected as the loss function to measure the degree of deviation between the estimated value and the true target value. Based on the above principle, the loss function of the Critic network of the i-th agent is expressed as: Among them, J Q,i (ω) represents the Q-value loss of the i-th agent with ω as the optimization variable. ω is the parameter of the Critic network. By optimizing ω, the Critic network can estimate the Q-value more accurately. Ε represents the expectation. (s t ,a t ,s t+1 ): D represents the state s sampled from the priority experience replay memory library D t 、Action a t and the next state s t+1 The sample of the priority experience replay memory D stores the historical data of the interaction between the agent and the environment; 653) The parameters of the target Q network (target Critic network) are optimized and updated by using the adaptive moment estimation algorithm. The target Q network is the target Critic network. The parameters of the target Q network are expressed as To indicate that the update is a soft update, Among them, τ is the soft update rate, ω is the parameter of the Critic network; The soft update mechanism allows the target Q network to slowly track the changes of the current network, avoiding the target value The dramatic fluctuations in the training data can improve the stability of the training; 654) Design the entropy network module: For agent i, the loss function of its entropy network at time t is expressed as: Among them, J i (α) represents the entropy network loss of agent i with α as the optimization variable. α is the temperature coefficient, which plays the role of adjusting the policy entropy. H0 is the predefined minimum policy entropy threshold. Indicates that under strategy π, for action a t Find the expectation, π(a t |s t ) is the agent in state s t Take action a t probability; 66) Based on the multi-UAV decision model, airspace reinforcement learning environment, SEAttention module, Actor network, Critic network, entropy network, priority experience replay memory library and network update module, a MASAC-SEPR algorithm model is constructed; the training method of the MASAC-SEPR algorithm model is designed: During the training process of the MASAC-SEPR algorithm model, agent i uses local observation information z i As the input of the Actor policy network integrated with the SEAttention module, the policy network outputs action a after processing through its own internal network layer based on the state space and action space defined by the POMDP model i , The actions of all agents together constitute the action set A. The environment executes the action set A and obtains the state information set S′ at the next moment according to the state transition probability P(s′|s,a) in the POMDP model. At the same time, the reward set R at the current moment is calculated according to the reward function R(s,a). Subsequently, the current state information set S, action set A, reward set R and state information set S′ at the next moment are stored in the priority experience replay memory library in the form of a four-tuple (S,A,R,S′). The priority of each experience sample in the library is determined by the TD error δ. The TD error reflects the difference between the current estimated Q value and the target Q value. Its calculation formula is: Where r represents the reward obtained by the agent after taking action at the current moment, γ is the discount factor, θ is the Actor network parameter, and π θ (S′) is the action output by the policy network (Actor network) based on the next state S′ when the parameter is θ. Indicates that under the target Critic network, the agent with index i takes the action π output by the Actor network in the next state S′ θ (S′), the action value output by the target Critic network; Q i (S, A) represents the action value output by the Critic network when the agent with index i takes action set A in the current state S, ω and are the parameters of the Critic network and its target network respectively. By calculating this error, the agent can adjust its behavior strategy to maximize the long-term reward; After the introduction of the priority experience replay memory library mechanism, data is extracted according to priority during training to optimize the efficiency of training data utilization. During training, a batch of data (s, a, r, s′) is extracted from the priority experience replay memory bank according to the priority, and the agent i combines the state and action information (s i ,a i ) as the input of the Critic network. The Critic network outputs two Q values Q according to the reward mechanism and observation information defined in the POMDP model. i (s t ,a t )and To evaluate the action value, for the Actor network, the state information s extracted from the priority experience replay memory is used as input to participate in the calculation of the strategy loss J π,i , in preparation for updating the parameters of the Actor network and optimizing the strategy output of the Actor network through the back-propagation algorithm; The entropy network automatically adjusts the temperature coefficient α, balancing the relationship between the algorithm in exploring new strategies and using existing experience, giving priority to experience replay. The reward value r output by the memory bank is used to calculate Q i (s t ,a t ), Strategy loss J π,i and entropy loss J i (α) Key parameter, entropy loss J i (α) is used to update the entropy network and adjust the temperature coefficient α. As the training progresses, the value of α gradually decreases; Finally, the Actor, Critic and entropy networks perform the operation of the network parameter update module, and update the Actor network parameters θ, Critic network parameters ω and entropy coefficient α through the back propagation algorithm. The SE mechanism enables the Actor network to output more reasonable actions, and the Critic network to more accurately evaluate the action value, making the entire update process more efficient and accurate. Finally, with the help of the SE mechanism, we gradually find the approximate optimal path planning strategy and realize the collaborative path planning of multiple UAVs in complex environments. The MASAC-SEPR algorithm model using this training method is named the multi-agent collaborative path optimization network model.
7. The method for heterogeneous multi-UAV collaborative path planning based on multi-agent deep reinforcement learning according to claim 1 is characterized in that: The training of the multi-agent collaborative path optimization network model in the current airspace reinforcement learning environment includes the following steps: 71) In the constructed airspace reinforcement learning environment, each agent in the multi-agent collaborative path optimization network model is trained. Each agent corresponds to a drone participating in collaborative path planning, and each agent has an independent Actor, Critic, and entropy network structure. 72) Initialize the Actor network, Critic network, and entropy network corresponding to each agent in the multi-agent system, and create a priority experience replay memory bank to store the agent's interaction experience, as well as a noise generator, which introduces noise into the agent's action selection at the beginning of training; 73) During the training cycle, each agent selects an action through its own Actor network based on the current perceived environmental state. After executing the action, the agent stores the generated experience in the priority experience playback memory bank. The experience includes the current state, executed action, reward obtained, and next state information; 74) When the amount of data in the priority experience replay memory reaches the preset standard, data is sampled from it to calculate the target Q value and current Q value of each agent, and then the strategy loss and entropy loss are obtained; using these calculation results, the parameters of the Actor network, Critic network and entropy network of each agent are updated through the back propagation algorithm to optimize the model and regularly save the parameters of the Actor network; 75) Generate a multi-agent collaborative path optimization network model after training and optimization in the current airspace reinforcement learning environment, and generate a multi-UAV collaborative path based on the multi-agent collaborative path optimization network model after training in the current airspace reinforcement learning environment.
8. The method for heterogeneous multi-UAV collaborative path planning based on multi-agent deep reinforcement learning according to claim 5 is characterized in that: The design of the reward mechanism includes the following steps: 81) For the boundary reward r e When the UAV of the master controller approaches the boundary of the flight area, it will be punished by -1. If it does not approach the boundary, no reward or punishment measures will be implemented. In the obstacle avoidance reward r o By design, if the distance between the master drone and the obstacle is less than twice the radius of the obstacle, it indicates that the drone has a potential collision risk, and a penalty of -2 is given at this time; when the distance between the two is less than the radius of the obstacle, it means that the drone collides with the obstacle and the drone is damaged, in which case a penalty of -500 is given; in other cases, that is, when the distance is not less than the threshold, no reward or punishment is given to the drone. Among them, d O Indicates the distance between the master drone and the obstacle, R O Indicates the obstacle radius; 82) Construct target reward rules g , When the distance d between the master UAV and the target point G When the distance is shortened to less than a specific threshold D1, it is judged that it has reached the target point and a positive reward of 1000 is given. If the distance standard is not reached, a corresponding penalty mechanism is set according to the actual distance between the drone and the target point. The farther the distance, the greater the penalty, so as to guide the drone to approach the target point. At the same time, design the formation distance reward mechanism r f , When the distance between the master drone and other follower drones is maintained within the preset range, it indicates that the formation is successfully formed. In this case, no reward or punishment operation is performed; on the contrary, once the distance exceeds the range, a penalty is designed according to the distance deviation between the master drone and each follower drone. The greater the distance deviation, the heavier the penalty, so as to encourage the drones to maintain the formation state. Where, j (j = 1, 2, ..., N-1) is the index of the follower UAV, d jL is the distance between the jth follower UAV and the master UAV, d jL,min ,d jL,max are the minimum and maximum distances between the jth follower and the master, r f,j represents the formation distance reward of the jth follower drone, then r f is j=1,2,...,N-1, i.e. the sum of the formation distance rewards from the first agent, the second agent,..., to the N-1th agent; 83) Build a speed synergy reward mechanism s , When the speed error between the master UAV and other follower UAVs is less than the threshold value 1, it indicates that the UAVs have successfully achieved speed coordination. At this time, a reward of 1 is given to encourage maintaining this coordination state. If the speed error is greater than or equal to the threshold value 1, no reward or penalty is given to the UAV. Where, j (j = 1, 2, ..., N-1) is the index of the follower UAV, v L Indicates the speed of the master UAV, v j represents the speed of the jth follower, r s,j represents the speed coordination reward between the master UAV and the jth follower UAV, then r s It is j=1,2,...,N-1, that is, the sum of the speed coordination rewards of the master UAV and the first agent, the second agent,..., to the N-1th agent.
9. The method for heterogeneous multi-UAV collaborative path planning based on multi-agent deep reinforcement learning according to claim 6 is characterized in that: The construction of the Actor network and learning includes the following steps: 91) Set the input layer and preliminary feature map: When the state information is input to the input layer of the Actor network, it is processed by a linear transformation operation. This linear transformation operation is regarded as a mapping function, which maps the input state information to a 256-dimensional feature space. The weight of the linear mapping adopts a specific initialization strategy, that is, it is initialized according to a normal distribution with a mean of 0 and a standard deviation of 0.
1. 92) Enhance feature dimension and embed SE mechanism: The 256-dimensional feature vector is upgraded to a 512-dimensional feature space through a linear transformation operation. In the high-dimensional feature space, the SE attention mechanism is introduced. The SE attention mechanism first performs a global average pooling operation on the 512-dimensional feature space, that is, the feature map of each channel is averaged in the spatial dimension to obtain the global feature information of each channel; then, the attention weight of each channel is generated through a sequence operation consisting of two fully connected layers and corresponding activation functions; the attention weight quantifies the importance of each channel feature for the current task, and by multiplying it element by element with the original 512-dimensional feature vector, the features of the key channels are highlighted and the information of the non-key channels is suppressed; 93) Perform feature dimension reduction and action probability distribution determination: The 512-dimensional feature vector processed by the SE attention mechanism is subjected to dimensionality reduction operation, and is reduced to a 256-dimensional feature space through a linear transformation. Two linear transformations are used to output the mean and standard deviation of the action respectively. These two statistics are used to define a normal distribution. The agent samples according to the normal distribution to determine the specific action. To ensure the effectiveness and rationality of the action, the sampled action value is constrained and limited to a legal range, thereby ensuring that the agent can make decisions that meet actual needs in the multi-UAV collaborative path planning task.
10. The method for heterogeneous multi-UAV collaborative path planning based on multi-agent deep reinforcement learning according to claim 6 is characterized in that: The construction of the Critic network and the learning include the following steps: 101) Input concatenation and preliminary feature mapping: When processing information, the Critic network first concatenates the state information and the action information to form an integrated input vector. Then, the concatenated vector is processed by a linear transformation operation and mapped to a 256-dimensional feature space. The weight of the linear transformation is initialized according to a normal distribution with a mean of 0 and a standard deviation of 0.
1. 102) Feature dimension enhancement and integration with SE mechanism: The Critic network enhances the 256-dimensional feature vector to a 512-dimensional feature space through linear transformation. In this high-dimensional space, the SE attention mechanism is introduced. The SE attention mechanism first performs global average pooling on the 512-dimensional features, integrates the information of each channel in the spatial dimension, and obtains the global feature representation of each channel. Then, after a sequence operation consisting of two fully connected layers and corresponding activation functions, the attention weight of each channel is generated. These weights reflect the importance of each channel feature in the action value assessment task, and are multiplied element by element with the original 512-dimensional feature vector to guide the network to focus on the key features closely related to the action value assessment. 103) Perform feature dimension reduction and action value evaluation: The 512-dimensional feature vector processed by the SE attention mechanism is reduced to 256 dimensions through linear transformation. On this basis, the Critic network uses two linear transformations to output two Q values respectively. These two Q values are used to evaluate the value of the agent taking a specific action in a given state.
Citation Information
Patent Citations
Voltage reactive power optimization method based on double-layer reinforcement learning power grid-user cooperation
CN115313407A
Multi-agent collaborative navigation method based on deep reinforcement learning
CN116579372A
Unmanned aerial vehicle safety path planning method based on maximum entropy multi-agent reinforcement learning
CN117908565A
Extensible deep reinforcement learning multi-unmanned aerial vehicle path planning cooperation method
CN117930864A
Multi-service robot dynamic space-time path planning method based on reinforcement learning
CN118502418A
Cited By
Discrete manufacturing system toughness enhancing method and system based on Petri net
CN120338457A
Unmanned aerial vehicle obstacle avoidance control method and system based on deep reinforcement learning
CN120445231A
Decision method for patrol path of unmanned aerial vehicle in complex environment
CN120521607A
Three-dimensional data acquisition management system and method
CN120726257A
Autonomous inspection and return control method and system for blow-off pipeline robot
CN120742875A