Multi-unmanned aerial vehicle multi-task planning method based on deep reinforcement learning of attention mechanism

By adopting a deep reinforcement learning method based on attention mechanism in drone task planning, combining the Dragonfly algorithm and the REINFORCE algorithm, the problem of inefficiency in high-complexity and large-scale task planning is solved, and efficient collaborative planning of multi-UAV systems in complex task scenarios is realized.

CN119937597APending Publication Date: 2025-05-06CHONGQING QINGLING TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510113708.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Traditional methods are inefficient when dealing with high-complexity and large-scale drone mission planning problems, making it difficult to achieve efficient collaborative planning of multiple drone systems in complex mission scenarios.

Method used

A deep reinforcement learning method based on attention mechanism is adopted, combined with the Dragonfly algorithm and the REINFORCE algorithm, initial task allocation is performed through the fuzzy C-mean algorithm, a mathematical model for multi-drone task planning is established, and a mask strategy is designed to ensure the effectiveness of task planning.

Benefits of technology

It significantly improves the efficiency and effectiveness of collaborative task planning of multiple drones in complex mission environments, improves the efficiency of task planning, model calculation time and optimization strategy stability, and reduces task execution time and cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119937597A_ABST
    Figure CN119937597A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of unmanned aerial vehicle task planning and scheduling, discloses a multi-unmanned aerial vehicle multi-task planning method based on deep reinforcement learning of an attention mechanism, and aims to solve the problem of low efficiency of a traditional method in processing high-complexity and large-scale task planning problems. The method comprises the following steps: S1, task decomposition and initial allocation; s2, establishing a multi-unmanned aerial vehicle task planning mathematical model; s3, establishing a Markov decision process; s4, designing a deep reinforcement learning model based on an attention mechanism under a mask strategy; s5, strategy optimization and updating are carried out in combination with a dragonfly algorithm and a REINFORCE algorithm; and S6, performing task planning on the multiple unmanned aerial vehicles by using the trained model. When facing a high-complexity and multi-constraint large-scale task scene, compared with other algorithms, the method has the advantages that the calculation time is shorter, the planning result is more efficient, the larger the scale is, the more obvious the comparison effect is, the better effect can be achieved in different task number environments, and the generalization ability is higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of unmanned aerial vehicle (UAV) mission planning and scheduling, and in particular, aims at path planning and task allocation problems in large-scale mission scenarios, and relates to a multi-UAV multi-task planning method based on deep reinforcement learning of an attention mechanism. Background Art

[0002] As a rapidly developing high-tech technology, drones have been widely used in many fields due to their advantages such as high efficiency, flexibility, autonomy, low cost and low risk. However, with the increasing complexity and scale of mission environments, the demand for autonomous flight and autonomous decision-making of drones is also increasing, especially in mission planning, which involves complex relationships and couplings between multiple constraints. Traditional intelligent optimization algorithms and heuristic algorithms often find it difficult to solve problems online within an acceptable time. Therefore, efficient and intelligent algorithms are urgently needed to solve the shortcomings of traditional intelligent optimization algorithms and heuristic algorithms in large-scale drone mission planning problems.

[0003] In recent years, deep reinforcement learning has become a hot topic of research due to its powerful performance in complex decision-making problems. Deep reinforcement learning has made significant progress in the fields of robot control and game strategy optimization by simulating the interaction between intelligent agents and the environment and continuously optimizing decision-making strategies. However, reinforcement learning still faces many challenges in practical applications, such as instability during training, high dependence on data, and difficulty in ensuring convergence to the global optimal solution. Therefore, in UAV mission planning, how to effectively solve these challenges and achieve the stability and efficiency of reinforcement learning algorithms remains an important topic. The dragonfly algorithm is derived from the swarm intelligence optimization algorithm, which simulates the predation and migration behavior of dragonflies in nature, and guides individuals to find the best solution in the search space through rules such as dispersion, aggregation, migration, attraction and repulsion. The original intention of the dragonfly algorithm was to solve complex multi-objective optimization problems. By enhancing the global search capability and local development capability, the convergence and optimization accuracy of the algorithm are significantly improved. It performs well in complex application scenarios such as UAV mission planning.

[0004] This paper effectively combines deep reinforcement learning with the dragonfly algorithm, using the global search capability of the dragonfly algorithm and the decision-making optimization capability of deep reinforcement learning to explore a new path to solve the UAV mission planning problem. Summary of the invention

[0005] In view of this, the purpose of the present invention is to provide a multi-UAV multi-task planning method based on deep reinforcement learning of the attention mechanism, so as to solve the problem of low efficiency of traditional methods in dealing with high-complexity and large-scale task planning problems, realize efficient collaborative planning of multi-UAV systems in complex mission scenarios, and provide an efficient solution for multi-UAV heterogeneous task planning.

[0006] The present invention provides a multi-UAV multi-task planning method based on deep reinforcement learning of attention mechanism, comprising the following steps:

[0007] S1: Task decomposition and initial assignment;

[0008] The fuzzy C-means algorithm is used for initial task allocation, the target area is divided into multiple clusters according to the distance criterion, and each cluster is assigned to a specific UAV;

[0009] S2: Establish a mathematical model for multi-UAV mission planning;

[0010] Obtain the task set and task coordinate point information, and then establish a multi-UAV task allocation mathematical model with UAV range and task completion as constraints;

[0011] S3: Establishing Markov decision process;

[0012] S4: Design of a deep reinforcement learning model based on attention mechanism under mask strategy;

[0013] S5: Combine the Dragonfly algorithm and REINFORCE algorithm to optimize and update the strategy;

[0014] S6: Use the trained model to perform multi-UAV mission planning.

[0015] Further, the step S1 includes the following sub-steps:

[0016] S1.1 Set the number of clusters C and the membership matrix U, and initialize the values ​​in the U matrix to ensure that the sum of the membership of each data point to all clusters is 1, that is,

[0017] Among them, U ij Represents data point x i The degree of membership to cluster j;

[0018] S1.2 Calculate the center v of each cluster based on the initial membership matrix j :

[0019]

[0020] Where n is the total number of data points; x i is the coordinate of the ith data point; v j is the center of the jth cluster; m is the fuzzy index;

[0021] S1.3 Based on the cluster center calculated in step S1.2, update the membership of each data point in the membership matrix U:

[0022]

[0023] Among them, ‖x i -v j ‖ represents the data point x i With cluster center v j The Euclidean distance between k represents the center of the kth cluster;

[0024] S1.4 repeatedly calculates cluster centers and updates the membership matrix until the stopping condition is met, then divides the target area into multiple clusters based on the distance criterion, and assigns each cluster to a specific UAV;

[0025] The stopping condition is the maximum number of iterations of the algorithm or meeting a set distance standard.

[0026] Furthermore, the multi-UAV mission planning mathematical model established in step S2 is as follows:.

[0027]

[0028] Where L is the total flight distance of the UAV; T = T1, T2, ..., T k , T is the UAV’s task execution sequence, i.e., the node index order; t i represents the flight time from the current task to task i; t i,Return represents the time to return to the starting point after completing task i; t Remain Indicates the remaining flight time of the drone; x i Indicates that task i can be executed only once.

[0029] Furthermore, in step S3, a corresponding state space S, action space A, reward function r and state transition probability τ are defined for any drone.

[0030] Further, in step S4, the deep reinforcement learning model based on the attention mechanism includes an encoder and a decoder;

[0031] The encoder includes a multi-head self-attention layer, a feedforward neural network layer, and a residual connection and layer normalization, which encodes different task types into a unified feature vector input to the encoder. The encoder extracts context information and long-distance dependencies from the input sequence through a self-attention mechanism and a feedforward neural network, generates a context representation of each input, and passes the generated context representation to the decoder;

[0032] The decoder is used to decode the output of the encoder. The decoder includes a masked multi-head self-attention layer, a feedforward neural network layer, a residual connection and layer normalization. The masked multi-head self-attention layer decodes the task according to the mask strategy to ensure that the drone selects a task that meets the flight time limit and ensures that each task is executed only once.

[0033] Further, the step S5 includes the following sub-steps:

[0034] S5.1 uses the reward value as the fitness value of the dragonfly algorithm to guide the dragonfly individual to find the optimal position through five behaviors. In each strategy update step, the REINFORCE algorithm is used to calculate the policy gradient;

[0035] The five behaviors include dispersal behavior, aggregation behavior, migration behavior, attraction behavior and repulsion behavior;

[0036] S5.2 passes the gradient information calculated in step S5.1 to the Dragonfly algorithm. The Dragonfly algorithm performs an optimization search in the neighborhood of the current strategy parameters based on the gradient information, selects a set of parameters with the best performance for update, and then updates the strategy using the strategy parameters optimized by the Dragonfly algorithm.

[0037] Beneficial effects:

[0038] 1. Aiming at various constraints in the mission scenario, the present invention proposes a multi-UAV multi-task planning method based on deep reinforcement learning of attention mechanism. Firstly, the fuzzy C-means algorithm is used to perform preliminary division of the mission area, and the heterogeneous tasks are uniformly represented in the divided mission area and the mask strategy is used to ensure the effectiveness of the mission planning. By adopting a deep reinforcement learning model based on the attention mechanism and combining the hybrid optimization method of the dragonfly algorithm and the REINFORCE algorithm, the efficiency and effect of multi-UAV collaborative mission planning in complex mission environments are significantly improved.

[0039] 2. The technical solution of the present invention has broad application prospects in multi-UAV collaborative mission planning, and is particularly suitable for large-scale complex mission scenarios that require timely responses, such as disaster relief and border patrols. The present invention significantly improves the efficiency of mission planning, model calculation time, and the stability of optimization strategies. The technology can reduce mission execution time and cost, and can respond quickly when executing missions, improving the overall operational efficiency of the UAV system, and has high market application potential and significant economic benefits.

[0040] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 This is a flow chart of a multi-UAV multi-task planning method based on deep reinforcement learning of an attention mechanism of the present invention;

[0042] Figure 2 The initial task allocation diagram described in the embodiment of the present invention;

[0043] Figure 3 This is a schematic diagram of the structure of an encoder according to an embodiment of the present invention;

[0044] Figure 4 The figure is a schematic diagram of the structure of a decoder according to an embodiment of the present invention. DETAILED DESCRIPTION

[0045] In order to make the technical solutions, advantages and purposes of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the described embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the protection scope of this application.

[0046] In view of the shortcomings of low solution efficiency and serious time consumption of traditional methods in solving large-scale and highly complex task planning problems, this paper proposes a multi-UAV multi-task planning method based on deep reinforcement learning of attention mechanism. This paper adopts a deep reinforcement learning model based on attention mechanism and combines the hybrid optimization method of dragonfly algorithm and REINFORCE algorithm to obtain an efficient task execution plan.

[0047] like Figure 1 As shown, the present invention provides a multi-UAV multi-task planning method based on deep reinforcement learning of attention mechanism, comprising the following steps:

[0048] Step 1: Task decomposition and initial allocation;

[0049] The fuzzy C-means (FCM) algorithm is used for initial task allocation, the target area is divided into multiple clusters according to the distance criterion, and each cluster is assigned to a specific UAV;

[0050] The specific steps are as follows: Set the number of clusters C (usually pre-set) and the membership matrix U, where U ij Represents data point x i The membership degree of each data point to cluster j is initialized to ensure that the sum of the membership degrees of each data point to all clusters is 1, that is,

[0051] According to the initial membership matrix, calculate the center (centroid) v of each cluster j :

[0052]

[0053] Where n is the total number of data points; x i is the coordinate of the ith data point; v j is the center of the jth cluster; m is the fuzzy index;

[0054] According to the current cluster center, update the membership of each data point in the membership matrix U:

[0055]

[0056] Among them, ‖x i -v j ‖ represents the data point x i With cluster center v j The Euclidean distance between k Represents the center of the kth cluster. This formula ensures that data points closer to the cluster center have higher membership, and vice versa. Repeat the calculation of cluster centers and update of membership matrix until the stop condition is met to obtain the final partition result. Figure 2 shown.

[0057] The stopping condition is the maximum number of iterations of the algorithm or meeting a set distance standard.

[0058] Step 2: Establish a mathematical model for multi-UAV mission planning;

[0059] The mathematical model in the present invention is as follows:

[0060]

[0061] Where L is the total flight distance of the UAV; T = T1, T2, ..., T k , T is the UAV’s task execution sequence, i.e., the node index order; t i represents the flight time from the current task to task i; t i,Return represents the time to return to the starting point after completing task i; t Remain Indicates the remaining flight time of the drone; x iIt means that task i can only be executed once. It requires that the total flight distance of the drone is minimal, each task can only be executed once, and the flight time of the drone cannot exceed the maximum flight limit.

[0062] Step 3: Establish a Markov decision process;

[0063] In step 3, the corresponding state space S, action space A, reward function r and state transition probability τ are defined for any drone.

[0064] Among them, the state space S is the state of all tasks, and the action a is the agent strategy π θ The strategy is implemented by a neural network model represented by the parameter θ. The reward r is the total flight distance of the drone, and the transition probability τ = p(s′|s,a) represents the probability of reaching the next state s′ after selecting action a.

[0065] Step 4: Design a deep reinforcement learning model based on attention mechanism under the mask strategy;

[0066] In step S4, the deep reinforcement learning model based on the attention mechanism includes an encoder and a decoder;

[0067] The encoder includes a multi-head self-attention layer, a feedforward neural network layer, and a residual connection and layer normalization. Different task types are encoded into a unified feature vector input to the encoder. The encoder extracts context information and long-distance dependencies from the input sequence through the self-attention mechanism and feedforward neural network, generates a context representation for each input, and passes the generated context representation to the decoder.

[0068] The decoder is used to decode the output of the encoder. The decoder includes a masked multi-head self-attention layer, a feedforward neural network layer, a residual connection and layer normalization. The masked multi-head self-attention layer decodes the task according to the mask strategy to ensure that the drone selects the task that meets the flight time limit and ensures that each task is executed only once.

[0069] The specific steps are as follows: through a unified task feature representation method, heterogeneous tasks are encoded into feature vectors of the same size as encoder input;

[0070] X i =(x1,y1)||(x2,y2)||A||I Type

[0071] In the formula, (x1, y1) is the key position of the task; (x2, y2) is the end position of the task; A represents the radius of the monitoring task and the length of the delivery task, and the value of A for the patrol task is 0; I Type Coding indicators for each task type.

[0072] Considering the characteristics of the task and the resource constraints of the UAV, the mask strategy is introduced. The mask strategy takes into account the constraints when selecting actions at each step, thereby ensuring the effectiveness of the UAV task planning. The mask is composed of the completion mask and time limit mask composition.

[0073]

[0074] The completion mask indicates whether the task has been completed. If the task is completed, the corresponding mask element is set to 1. The time mask ensures that the drone does not select a task that cannot be completed within its remaining flight time. In the formula, ∨ represents the logical "or" operation.

[0075] The encoder computes the initial node embeddings through a linear projection layer Initial node embedding Update through N attention layers to generate new node embeddings:

[0076]

[0077] Node embedding after processing by the Multi-Head Self-Attention (MHA) layer Further processing through a feed-forward neural network layer produces updated node embeddings

[0078]

[0079] Graph Embedding is calculated as the final node embedding The mean of :

[0080]

[0081] Both node embeddings and graph embeddings are passed as input to the decoder. At each decoding step t, the context information of the decoder comes from the output of the encoder at time t, which includes the current node T t , the previous node (the last node) T t-1 and the first node T1.

[0082]

[0083] Where [.,.,.] is a horizontal connection operator. The decoder generates a new query vector Q' through the MHA layer. The query vector includes the context node h c and partial solution information T 1:k , which is initially zero and is updated after the first task is selected. Encoder output node embedding After being processed by the decoder MAH layer, it is converted into (h1′,h′2,...,h′ k ). Multi-head attention is used during the decoding process to calculate the attention score of each task node and mask the nodes that cannot be accessed at the current time t.

[0084]

[0085] a′=softmax(u′)(10)

[0086] In the formula, m i is the value of mask M. Mask M reduces the probability of completed tasks to zero in the softmax function by setting the attention score of the completed task to negative infinity, ensuring that the completed task will not be selected again. The softmax function converts the attention score u i ′ is converted into probability distribution a′, and then the next task node is selected by sampling from probability distribution a′.

[0087] After selecting the next mission, the decoder updates the remaining flight time budget t remain , and some solutions T 1:t-1 to reflect the current decoding status. The updated partial solution information includes the completed tasks and the embedding of new tasks, and the remaining flight time budget is updated to t remain Subtract the completion time of the selected task. The encoder and decoder structure of the embodiment of the present invention is as follows Figure 3 , Figure 4 shown.

[0088] Step 5: Combine the Dragonfly algorithm and REINFORCE algorithm to optimize and update the strategy;

[0089] The policy gradient estimation in the REINFORCE algorithm updates the policy parameters θ according to the return for each time step t:

[0090]

[0091]

[0092] Where N is the batch size; k is the number of tasks; π θ (a i,t |s i,t ) is state s i,t Take action a i,t strategy; is the gradient of the policy; is the cumulative reward; α is the learning rate. The algorithm is trained on 1.28 million task instances, with a batch size of 512, 100 cycles, a learning rate of 1e-4, and the Adam optimizer.

[0093] Combine the REINFORCE algorithm with the dragonfly algorithm. By simulating the predation and migration behaviors of dragonflies in nature, as well as the local interactions between individuals and the exchange of global information, efficient exploration can be achieved in a broad search space to optimize complex problems. Each dragonfly in the dragonfly population adjusts its position and speed through the following five behavioral rules. The position of the dragonfly with the best fitness value will be used as the optimal parameter.

[0094]

[0095]

[0096]

[0097] F i =X + -X i (16)

[0098] E i =X - +X i (17)

[0099] Among them, the dispersion behavior S i To avoid individuals concentrating near the local optimal solution, x i Indicates the current position of the dragonfly, x j represents the position of the jth adjacent dragonfly individual, N is the number of dragonflies; the aggregation behavior A i Improve the overall search efficiency of the population, V j represents the flight speed of the jth dragonfly; the migration behavior C i Guide dragonfly individuals to move toward the average position of neighboring individuals to maintain the cohesion of the population; attraction behavior F i Guide the population towards the potential optimal solution area, X + Indicates food source; rejection behavior E i , avoid falling into high-consumption or suboptimal solution areas, and help dragonflies explore new possibilities. - Indicates the position of the natural enemy. According to the above five behavior rules, update the speed and position of the dragonfly individual:

[0100] ΔX i+1 =(sS i +aA i +cC i +fF i +eE i )+wΔX i (18)

[0101] X i+1 =X i +ΔX i+1(19)

[0102] Where ΔX i is the position increment of the i-th dragonfly; s is the separation weight; a is the alignment weight; c is the aggregation weight; f is the food factor; e is the natural enemy factor; and w is the inertia weight.

[0103] We use the reward value as the fitness value of the dragonfly algorithm to guide the dragonfly individual to find the optimal position through five behaviors. In each policy update step, the REINFORCE algorithm is used to calculate the policy gradient, but the gradient is not directly used for update. Instead, the gradient information is passed to the dragonfly algorithm, and the dragonfly algorithm optimizes and searches in the neighborhood of the current policy parameters based on this information, selects a set of parameters with the best performance for update, and then updates the policy parameters optimized by the dragonfly algorithm.

[0104] Step 6: Use the trained model to perform multi-UAV mission planning.

[0105] It is hereby declared that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention is described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should be included in the scope of the claims of the present invention.

Claims

1. A multi-UAV multi-task planning method based on deep reinforcement learning of attention mechanism, characterized in that: The following steps are involved: S1: Task decomposition and initial assignment; The fuzzy C-means algorithm is used for initial task allocation, the target area is divided into multiple clusters according to the distance criterion, and each cluster is assigned to a specific UAV; S2: Establish a mathematical model for multi-UAV mission planning; Obtain the task set and task coordinate point information, and then establish a multi-UAV task allocation mathematical model with UAV range and task completion as constraints; S3: Establishing Markov decision process; S4: Design of a deep reinforcement learning model based on attention mechanism under mask strategy; S5: Combine the Dragonfly algorithm and REINFORCE algorithm to optimize and update the strategy; S6: Use the trained model to perform multi-UAV mission planning.

2. The multi-UAV multi-task planning method based on deep reinforcement learning of attention mechanism according to claim 1, characterized in that: The step S1 comprises the following sub-steps: S1.1 Set the number of clusters C and the membership matrix U, and initialize the values ​​in the U matrix to ensure that the sum of the membership of each data point to all clusters is 1, that is, Among them, U ij Represents data point x i The degree of membership to cluster j; S1.2 Calculate the center v of each cluster based on the initial membership matrix j : Where n is the total number of data points; x i is the coordinate of the ith data point; v j is the center of the jth cluster; m is the fuzzy index; S1.3 Based on the cluster center calculated in step S1.2, update the membership of each data point in the membership matrix U: Among them, ‖x i -v j ‖ represents the data point x i With cluster center v j The Euclidean distance between k represents the center of the kth cluster; S1.4 repeatedly calculates cluster centers and updates the membership matrix until the stopping condition is met, then divides the target area into multiple clusters based on the distance criterion, and assigns each cluster to a specific UAV; The stopping condition is the maximum number of iterations of the algorithm or meeting a set distance standard.

3. The multi-UAV multi-task planning method based on deep reinforcement learning of attention mechanism according to claim 2 is characterized in that: The multi-UAV mission planning mathematical model established in step S2 is as follows: Where L is the total flight distance of the UAV; T = T1, T2, ..., T k , T is the UAV’s task execution sequence, i.e., the node index order; t i represents the flight time from the current task to task i; t i,Return represents the time to return to the starting point after completing task i; t Remain Indicates the remaining flight time of the drone; x i Indicates that task i can be executed only once.

4. The multi-UAV multi-task planning method based on deep reinforcement learning of attention mechanism according to claim 3 is characterized by: In step S3, a corresponding state space S, action space A, reward function r and state transition probability τ are defined for any drone.

5. The multi-UAV multi-task planning method based on deep reinforcement learning of attention mechanism according to claim 4, characterized in that: In step S4, the deep reinforcement learning model based on the attention mechanism includes an encoder and a decoder; The encoder includes a multi-head self-attention layer, a feedforward neural network layer, and a residual connection and layer normalization, which encodes different task types into a unified feature vector input to the encoder. The encoder extracts context information and long-distance dependencies from the input sequence through a self-attention mechanism and a feedforward neural network, generates a context representation of each input, and passes the generated context representation to the decoder; The decoder is used to decode the output of the encoder. The decoder includes a masked multi-head self-attention layer, a feedforward neural network layer, a residual connection and layer normalization. The masked multi-head self-attention layer decodes the task according to the mask strategy to ensure that the drone selects a task that meets the flight time limit and ensures that each task is executed only once.

6. The multi-UAV multi-task planning method based on deep reinforcement learning of attention mechanism according to claim 5, characterized in that: The step S5 comprises the following sub-steps: S5.1 uses the reward value as the fitness value of the dragonfly algorithm to guide the dragonfly individual to find the optimal position through five behaviors. In each strategy update step, the REINFORCE algorithm is used to calculate the policy gradient; The five behaviors include dispersal behavior, aggregation behavior, migration behavior, attraction behavior and repulsion behavior; S5.2 passes the gradient information calculated in step S5.1 to the Dragonfly algorithm. The Dragonfly algorithm performs an optimization search in the neighborhood of the current strategy parameters based on the gradient information, selects a set of parameters with the best performance for update, and then updates the strategy using the strategy parameters optimized by the Dragonfly algorithm.

Citation Information

Cited By

  • Unmanned aerial vehicle inspection path planning method based on attention mechanism

    CN120576772A

  • A drone inspection path planning method based on attention mechanism

    CN120576772B

  • Multi-starting-point sequence decision reinforcement learning method based on dynamic mask attention

    CN121119029A