Unmanned aerial vehicle cluster dynamic task re-planning framework for real-time decision
By proposing a two-layer framework and multi-agent reinforcement learning algorithm in the task re-planning of drone clusters, the problem of task re-allocation of drone clusters in dynamic environments is solved, and efficient and stable real-time decision-making and task execution are achieved.
Patent Information
- Application Number
- CN202510137371.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-30
AI Technical Summary
The existing UAV cluster task re-planning decision-making method lacks a task re-allocation method that solves the long-term task execution of UAV clusters in distributed environments, cannot achieve robustness, adaptability, and converge to global optimal results, and is difficult to cope with sparse rewards and increased task complexity in extended state space.
A two-layer framework for re-planning of drone cluster cluster missions for real-time decision-making is proposed. By establishing a two-layer optimization problem model, using K-medoids clustering algorithm that simulates annealing enhancement and the multi-agent reinforcement learning algorithm for goal-oriented belief space, we calculate the optimal task grouping and drone formation deployment, and realize real-time task redistribution.
This framework can realize efficient real-time decision-making of drone clusters in dynamic environments, improve deployment adaptability and stability, be better than existing algorithms that lack dynamic adjustment and redundancy balance, and greatly improve training efficiency and ability to converge to optimal strategies.
Smart Images

Figure CN120066067A_ABST
Abstract
Description
(1) Technical Field
[0001] The present invention belongs to the technical field of reliability engineering, and particularly relates to a dynamic task replanning framework for an unmanned aerial vehicle (UAV) swarm facing real-time decision-making. (2) Background Art
[0002] The application of UAV swarms in the long-term urban service field is showing an increasing trend, covering aspects such as geographical hazard monitoring, data collection, and medical supply delivery. When ensuring its continuous operation in a geographically distributed environment, the task planning work must prioritize the UAV deployment and task allocation to achieve efficient cooperation between systems and task coherence. Although the integrated scheme of multi-UAV deployment and task allocation can effectively manage ground users in different spatial positions, UAVs will inevitably encounter some unforeseen situations during operation, resulting in dynamic requirements exceeding the pre-set planning scope. This requires not only re-planning the on-demand scheduling strategy but also the UAVs to have the ability of real-time decision-making to ensure an effective response to emergencies caused by changes in task requests. These challenges highlight the necessity of constructing a customized framework and require advanced methods to coordinate and resolve the coupled sub-problems.
[0003] However, there are multiple deficiencies in the current UAV swarm task replanning decision-making methods. Specifically, there is a lack of a task reallocation method that can solve the long-term task execution of UAV swarms in a distributed environment; it is unable to achieve robustness, self-adaptability, and convergence to the global optimal result for potentially changing long-term tasks; there is a lack of effective strategies to cope with sparse rewards and increasing task complexity in the extended state space to assist decision-makers in effectively exploring and improving learning efficiency. Therefore, the present invention aims to address the above deficiencies and proposes a framework for replanning to achieve efficient real-time decision-making for UAV swarms in a dynamic environment. (3) Summary of the Invention
[0004] The purpose of the present invention is to provide a two-layer framework for UAV swarm task replanning facing real-time decision-making, comprehensively considering various factors such as predefined task locations and dynamic requirements, and effectively processing these key elements and achieving the goals by developing improved algorithms for the sub-problems of each layer. To achieve the above purpose, the present invention provides the following framework:
[0005] S100: Establish a problem scenario and a hierarchical optimization problem model for UAV swarm task execution;
[0006] S101: Abstract the scenario of cluster decentralized task execution;
[0007] S102: Establish an optimization model for UAV swarm deployment problem based on dynamic subgroup division;
[0008] S103: Establish an optimization model for the UAV mission reassignment problem for task reassignment in subgroups;
[0009] S200: Calculate the optimal task grouping and UAV formation deployment based on the K-medoids clustering algorithm enhanced by simulated annealing;
[0010] S201: Initialize all task groupings and the formation deployments of all UAVs;
[0011] S202: Use the K-medoids algorithm to group the tasks to obtain several task subgroups;
[0012] S203: Assign and deploy the UAVs to specific task subgroups and form formations;
[0013] S204: If the annealing condition or the maximum number of annealing rounds is not satisfied, return to step S201;
[0014] S300: Calculate the real-time task reassignment result based on the goal-oriented belief space multi-agent reinforcement learning algorithm;
[0015] S301: For each task subgroup and the corresponding UAV formation, initialize the agent environment, the action-value function, and the objective function;
[0016] S302: Use the belief-guided experience replay mechanism to train the multi-agent through reinforcement learning;
[0017] S303: Perform real-time task reassignment within the UAV formation based on the trained multi-agent.
[0018] In step S100, establish the problem scenario and the hierarchical optimization problem model for the UAV cluster task execution, and the specific process is as follows.
[0019] Establish a mathematical model for the task execution scenario, and establish the UAV cluster task execution problem as a two-layer optimization problem model. The upper layer establishes an optimization problem model for the global optimal task grouping and UAV deployment, and the lower layer establishes a real-time decision-making problem model for task reassignment for each group of tasks and the corresponding UAV formation.
[0020] Among them, in step S101, abstract the decentralized task execution scenario of the cluster, considering that the tasks include accessing M static tasks with limited access, and the task set is denoted as Γ = {τ 1 , ω 2 , …, τ M}. The UAV cluster for task execution includes a total of N UAVs, and the UAV cluster set is denoted as V = {v 1 , v 2 , …, v N}. For each UAV vi , for \(i = 1, 2, \ldots, N\), consider modeling it as a five - tuple \(\langle l\) i , s i , \(\theta\) i , c i , f i \(\rangle\), where \(l\) i \(=(x\) i , y i , z i ), \(s\) i \(>0\), \(\theta\) i \(\in(0, \pi)\), f i \(>0\), representing the position coordinates, constant flight speed, observation angle, ability level, and maximum flight distance of the UAV respectively.
[0021] Among them, in step S102, an optimization model for the UAV swarm deployment problem based on dynamic subgroup partitioning is established. The task set is divided into \(K\) groups, denoted as \(\{A\) 1 , A 2 , \ldots, A K \}\). Taking the maximization of the balance of remaining capacity BRC as the optimization objective, the UAV swarm is also divided into \(K\) groups, denoted as \(V=\{V\) 0 , V 1 , V 2 , \ldots, V K \}, representing the UAV formations deployed to each task dynamic subgroup respectively, and taking the UAV swarm formation deployment plan as the decision variable of this optimization problem.
[0022] Among them, in step S103, taking the subgroup as the unit, an optimization model for the UAV task reassignment problem facing task reassignment is established. The task reassignment problem within each UAV formation can be established as a decentralized partially observable Markov decision process, denoted as Each UAV within the UAV formation is abstracted as an agent, denoted as where \(N\) k represents the number of UAVs in the \(k\) - th formation. \(\gamma\) is the discount factor of the reward, and the other elements are defined as follows:
[0023] State space: The environmental state at time \(t\) is defined as \(s\) t \(\in S\), which includes the state \(\langle l\) i , s i , \(\theta\) i , c i , f i \rangle\) of each UAV and the state of the task.
[0024] Action space: The action of each UAV \(v\) i at time \(t\) is defined as where The first four elements respectively represent that the drone chooses to fly east, west, south, and north, and the last element represents staying still in place.
[0025] Observation space: Due to the observation range and communication limitations of each drone, the observation information relied on by each agent when making decisions at time t only includes the task location and task requirements.
[0026] State transition probability: The transition probability P satisfies P(s t+1 |s t ,a t ):s t ×a t ×s t+1 →[0,1], indicating the probability of transitioning from state s t by choosing action a t to state s t+1 .
[0027] Reward function: For choosing t action in state s , the reward obtained will be
[0028] Belief space: b t ∈b represents the probability of a new task request occurring at time t, and its calculation method is as follows:
[0029]
[0030] Among them, the calculation method of p is as follows:
[0031]
[0032] Goal space: The space containing the final goal and intermediate goals is denoted as G, indicating an increase in the belief level. g is a mapping function that satisfies g:(b t ,s t )→G.
[0033] In step S200, the optimal task grouping and UAV formation deployment are calculated based on the simulated annealing enhanced K-medoids clustering algorithm, and the specific process is as follows.
[0034] The outer loop uses the simulated annealing algorithm. In each internal iteration of the simulated annealing, after initializing all task groupings and all UAV formation deployments, the K-medoids algorithm is used to group the tasks to obtain several task subgroups, and then the UAVs are assigned and deployed to specific task subgroups to form formations. After meeting the conditions of the simulated annealing, the UAV deployment in the upper layer of the framework is completed.
[0035] Among them, in step S201, the first Euclidean centroid μ of the cluster is randomly sampled and initialized from the task set 1 , and then based on the distances from the known centroids, the other k - 1 centroids are iteratively selected. Points closer to all the known centroids among all the targets are more likely to be selected as centroids. At the same time, all the UAV formations and task subgroups are initialized as empty sets.
[0036] Among them, in step S202, the K - medoids algorithm is used to group the tasks to obtain several task subgroups. The specific process is as follows.
[0037] Step1: Calculate the Euclidean distance d of each task τ j to all the centroids jk = ||τ j - μ k || 2 ;
[0038] Step2: Find the minimum distance r = argmin k d jk , and add the task τ j to the task subgroup with the minimum distance;
[0039] Step3: Calculate the density d of task subgroup k k , if it is greater than the threshold, reject this clustering decision, select the second - nearest task subgroup and retest the threshold; otherwise, update the Euclidean centroid μ of task subgroup k k .
[0040] Among them, in step S203, the UAVs are assigned and deployed to specific task subgroups and formed into formations. The specific process is as follows.
[0041] Step 1: Randomly deploy the UAVs in the UAV cluster to K task subgroups;
[0042] Step 2: For each UAV v t , calculate a global remaining capacity balance BRC new ;
[0043] Step 3: For the above - selected UAV v t , sequentially select all the subgroups C k , and sequentially transfer v t to the task subgroup C k ;
[0044] Step 4: Calculate the remaining capacity balance BRC now at this time and the change ΔBRC = BRC now - BRC new ;
[0045] Step 5: If ΔBRC > 0, or then accept the transfer in Step 3; otherwise reject the transfer operation in Step 3.
[0046] Among them, for step S204, if the simulated annealing round is greater than the threshold N I then exit the loop to obtain the final task grouping and UAV cluster deployment; otherwise, return to step S201 for a new round of iterative optimization.
[0047] In step S300, based on the goal-oriented belief space multi-agent reinforcement learning algorithm, the task reallocation result is calculated in real time, and the specific algorithm process is as follows.
[0048] Among them, for step S301, for each task subgroup and the corresponding UAV formation, initialize the agent environment, action-value function, and objective function, and the specific steps are as follows.
[0049] Step 1: Define the task environment where the multi-agent is located, including information such as the task area, user demand distribution, and UAV performance parameters, providing a basic environment for the agent to perceive and make decisions;
[0050] Step 2: Initialize the action-value function Q for each agent π (b t ,s t ,a t ) = E[U t |b t ,s t ,a t , where this function is used to evaluate the expected cumulative discounted reward of the agent taking action a t under different belief states b t and environmental states s t , and a reasonable initial value range is assigned to it during initialization so as to be gradually optimized during subsequent learning;
[0051] Step 3: Determine the objective function of the multi-agent collaborative task;
[0052] Among them, for step S302, use the belief-oriented off-policy experience replay mechanism to train the multi-agent through reinforcement learning, and the specific steps are as follows.
[0053] Step 1: Perform belief-oriented experience storage. At each time step t, after the agent selects action a t according to the current policy and observation information, store the experience Tr t = {b t ,s t ,at , r t , b t+1 , s t+1} stored in the dataset Tr = {Tr 1 , Tr 2 , …, Tr M}. Among them, the belief state b t reflects the probability of the existence of dynamic task requests at time step t. The environmental state contains information such as the positions of ground users and task requirements. The reward is calculated according to the reward function.
[0054] Step 2: Experience sampling and update. Randomly sample the experience Tr t from the dataset, and generate new experience according to the objective function g (the mapping function g from the belief state b t and the environmental state s t to the target space G: g:(b t , s t ) → G where is greater than r when the UAV moves in the correct direction t . Then, use these new experiences to update the agent's action-value function and belief state.
[0055] Step 3: Multi-agent deep Q-learning update. Based on the sampled experience, adopt the temporal difference training method to update the parameters θ of the action-value function through the gradient descent algorithm to minimize the difference between the predicted value and the target value, so that the agent can learn a better policy. At the same time, in a multi-agent environment, considering the interaction and cooperation between agents, introduce the value decomposition network. Assume that the agents cooperate completely, that is, share the same reward function. Through the decomposition formula satisfy decompose the global value into local values, alleviate the instability in the learning process, and ensure that the local optimal combination is the global optimal.
[0056] Among them, in step S303, based on the trained multi-agent, perform real-time task reallocation within the UAV formation. When the multi-agent completes training, during actual operation, each agent selects an action a t according to the current environmental state s t and its own belief state b t based on the learned policy π to achieve real-time task reallocation within the UAV formation. For example, the agent decides the flight path, task execution order, and resource allocation of the UAV according to factors such as task priority, UAV position, and load conditions to efficiently complete the task and adapt to changes in dynamic task requirements. In the face of sudden task requests or environmental changes, the agent can quickly adjust the decision to ensure the continuity of the task and the stability of the system.
[0057] Compared with the prior art, the present invention has the following beneficial effects: A dynamic task replanning framework for UAV swarms facing real-time decision-making is proposed, which can establish a model to describe the two-layer task replanning problem. The SADCK-Medoid algorithm is proposed in the upper layer for task grouping and UAV deployment, which can effectively cope with potential changes in requirements, environment, etc. in long-term applications, improve deployment adaptability and stability, and is superior to existing algorithms lacking dynamic adjustment and redundancy balance; the GOBS-MARL algorithm is proposed in the lower layer for task reallocation, which greatly improves the training efficiency, accelerates convergence to the optimal policy, solves the problem of multi-UAV task allocation, and surpasses existing methods restricted by sparse rewards and complex environments. This framework can effectively guide UAV swarms to efficiently complete hierarchical distributed task replanning. (4) BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 It is a flowchart of a dynamic task replanning framework for UAV swarms facing real-time decision-making (5) SPECIFIC IMPLEMENTATION MANNER
[0059] The exemplary implementation of the present invention will be described in detail below with reference to the accompanying drawings. The following description includes specific details to assist understanding, but these specific details should only be shown as exemplary. Therefore, those of ordinary skill in the art will recognize that various changes and modifications can be made to the various examples described herein without departing from the scope and spirit of the present disclosure. In addition, descriptions of well-known functions and structures are omitted for clarity and conciseness.
[0060] The terms and words used in the following description and claims are not limited to their literal meanings, but are only used by the inventor for a clear and consistent understanding of implementing the present invention. Therefore, those skilled in the art should clearly understand that the following description of the various exemplary implementations of the present invention is only provided for illustrative purposes and is not intended to limit the present invention defined by the appended claims and their equivalents.
[0061] Step S100: Establish a task execution scenario and a hierarchical optimization problem model for the UAV swarm, mainly considering that the typical task scenario is that the UAV swarm provides parcel delivery services and emergency medical supplies for the community. The parameters of tasks and UAVs are shown in Table 1.
[0062] Table 1 Parameter settings
[0063]
[0064] Establish the upper-layer UAV deployment optimization problem, and the minimization of BRC is defined as the global optimization goal. For the internal autonomous reallocation decision problem of the UAV formation in the lower layer, the average reward is used as the optimization goal.
[0065] Step S200: Calculate the optimal task grouping and UAV formation deployment based on the simulated annealing enhanced K-medoids clustering algorithm, and cluster 152 ground users into five task areas. On this basis, detailed deployment results are obtained. Initially, the number of ground users in each task area is 9, 17, 9, 17, and 12 respectively, and as a result, 4, 6, 3, 7, and 6 UAVs are initially deployed, and 14 UAVs are stationed at the base. After the first emergency, the number of ground users in each task area increases except for TR1. The deployment decision remains unchanged. After the second emergency, the increase in each task area is significant, and 6, 9, 5, 10, and 7 UAVs are redeployed respectively. At the same time, in the first stage, using the SADCK-Medoid algorithm, the BRC converges to 19.11. In the second stage, the BRC value rises from 19.11 to 27.27, without exceeding the threshold, so redeployment is not carried out. In the third stage, along with the emergence of a new event sequence, the BRC climbs and exceeds the threshold, but with the assistance of the SADCK-Medoid algorithm calculation, it is driven to converge to an acceptable level again.
[0066] Step S300: According to the formation deployment results, calculate the task reallocation results in real time based on the goal-oriented belief space multi-agent reinforcement learning algorithm. The task reallocation results are generated after the agents trained by the GOBS-MARL algorithm make decisions. In the initial stage, the UAVs are positioned around the existing targets according to the strategy to increase the probability of detecting unpredictable events. When the ground user requirements are updated, the UAVs execute tasks according to the BRC global goal and their own capabilities, and when new ground users are encountered, the redeployment mechanism is activated with the assistance of additional UAVs to ensure comprehensive coverage and achieve precise task allocation and dynamic coordination.
[0067] A core feature of the framework of the present invention is that the system has the ability to make deployment decisions offline, which reserves sufficient computing time for the algorithm to deeply optimize the solution and effectively ensure that the long-term performance reaches the optimal state. When the deployment process is completed, each UAV immediately starts an independent operation mode, autonomously identifies and selects tasks based on the built-in intelligent decision-making module, and does not require excessive external intervention throughout the process, efficiently realizing the deep integration of the autonomy and flexibility of task allocation.
[0068] Particularly prominent is that the invention framework has good processing efficiency for scriptless tasks, enabling drones to quickly respond and make precise decisions in complex and ever-changing emergency scenarios. Once new task requirements emerge, the system can immediately initiate the redeployment process, efficiently allocate resources, and comprehensively cover the new tasks to ensure the coherence and integrity of task execution. The collaborative operation between deployment and task allocation is achieved through the BRC threshold mechanism, which is flexibly regulated based on the real-time state of the system, greatly enhancing the overall adaptability and dynamic response flexibility of the framework, and laying a foundation for the stable and efficient operation of the multi-drone system in complex task environments.
[0069] The above-described examples have detailed the implementation methods of various parts of the present invention. The specific implementation forms of the present invention are not limited thereto. For those of ordinary skill in the art, all obvious changes made without departing from the spirit of the method described in the present invention and the scope of the claims are within the protection scope of the present invention.
Claims
1. A dynamic task re-planning framework for UAV swarms for real-time decision making, characterized by The steps include: S100: Establish UAV swarm mission execution scenario and hierarchical optimization problem model; S101: Abstract cluster distributed task execution scenario; S102: Establish an optimization model for drone swarm deployment problem based on dynamic subgroup division; S103: Taking subgroups as units, an optimization model for the UAV task reallocation problem oriented to task reallocation is established; S200: Calculate optimal task grouping and UAV formation deployment based on K-Medoids clustering algorithm enhanced by simulated annealing; S201: Initialize all task groups and the formation deployment of all UAVs; S202: grouping the tasks using a K-medoids algorithm to obtain a number of task subgroups; S203: Allocate and deploy the drones to specific task subgroups and form a formation; S204: If the annealing condition or the maximum annealing rounds are not met, return to step S201; S300: Real-time calculation of task redistribution results based on goal-oriented belief space multi-agent reinforcement learning algorithm; S301: For each task subgroup and corresponding UAV formation, initialize the agent environment, action-value function and objective function; S302: Using a belief-guided post-experience replay mechanism to train multiple agents through reinforcement learning; S303: Perform real-time task redistribution within the UAV formation based on the trained multi-agent.
2. The UAV swarm task re-planning framework for real-time decision-making according to claim 1 is characterized in that: The Markov decision process problem model based on belief space described in step S103 specifically includes: For each group of tasks and the corresponding UAV formation, a real-time decision-making model for task redistribution is established. The task redistribution problem within each UAV formation can be established as a decentralized partially observable Markov decision process, denoted as Z =<v,S,O,b,A,R,P,T,G,g,γ> The drones in each drone formation are abstracted as intelligent agents, denoted as v = {1, 2..., N k }, where N k represents the number of drones in the kth formation. γ is the discount factor of the reward, and the other elements are defined as follows: State space: The state of the environment at time t is defined as s t ∈S, which contains the state of each drone <l i ,s i ,θ i , c i , f i >, and the status of the task. Action space: Each drone v i The action at time t is defined as in The first four elements indicate that the drone chooses to fly east, west, south, and north, respectively, and the last element indicates staying still. Observation space: Due to the observation range and communication limitations of each drone, each agent relies on the observation information when making decisions at time t Only the mission location and mission requirements are included. State transition probability: The transition probability P satisfies P(s t+1 |s t , a t ):s t ×a t ×s t+1 →[0, 1], indicating that in s t Status selection a t Transfer to state s t+1 probability. Reward function: For In s t Select in status Action will be rewarded Belief space: b t ∈b represents the probability of a new task request at time t, which is calculated as follows: The calculation method of p is as follows: Goal space: The space containing the final goal and the intermediate goal is denoted by G, which represents the increase in the belief level. g is a mapping function that satisfies g: (b t ,s t )-→G.
3. The UAV swarm task re-planning framework for real-time decision-making according to claim 1 is characterized in that: In step S200, the optimal task grouping and UAV formation deployment are calculated based on the K-medoids clustering algorithm enhanced by simulated annealing, which specifically includes: The outer loop uses the simulated annealing algorithm. The following steps are performed inside each simulated annealing. After initializing all task groups and the formation deployment of all drones, the K-medoids algorithm is used to group the tasks to obtain several task subgroups. The drones are then assigned and deployed to specific task subgroups and formed into formations. After the conditions of simulated annealing are met, the deployment of drones in the upper layer of the framework is completed. In step S201, the first Euclidean centroid μ1 of the cluster is randomly sampled from the task set, and then the other k-1 centroids are iteratively selected based on the distance from the known centroid. Points closer to all targets and the known centroid are more likely to be selected as centroids. At the same time, all drone formations and task subgroups are initialized to empty sets. In step S202, the tasks are grouped using the K-medoids algorithm to obtain a number of task subgroups. The specific process is as follows. Step 1: Calculate τ for each task j Euclidean distance d to all centroids jk =||τ j -μ k || 2 ; Step 2: Find the minimum distance r=argmin among all task subgroups k d jk , the task τ j Join the task subgroup with the smallest distance; Step 3: Calculate the density d of task subgroup k k , if it is greater than the threshold, reject the clustering decision, select the second closest task subgroup and retest the threshold; otherwise update the Euclidean centroid μk of task subgroup k. Among them, step S203 allocates and deploys the drones to specific task subgroups and forms a formation. The specific process is as follows. Step 1: Randomly deploy drones in the drone cluster into K task subgroups; Step 2: For each drone v t , calculate the global remaining capacity balance BRC new ; Step 3: For the drone selected above t , select all subgroups C in turn k , v t Transfer to task subgroup C in sequence k ; Step 4: Calculate the remaining capacity balance BRC at this time now And the change ΔBRC = BRC now -BRC new ; Step 5: If ΔBRC>0, or If yes, the transfer in Step 3 is accepted; otherwise, the transfer in Step 3 is rejected. In step S204, if the number of simulated annealing rounds is greater than the threshold N I Then exit the loop and obtain the final task grouping and drone cluster deployment, otherwise return to step S201 for a new round of iterative optimization.
4. The method for optimizing the location of security facilities considering priority positioning and protection according to claim 1, characterized in that: In step S300, the task reallocation result is calculated in real time based on the goal-oriented belief space multi-agent reinforcement learning algorithm. include: Among them, step S301, for each task subgroup and corresponding drone formation, initialize the agent environment, action-value function and objective function, the specific steps are as follows. Step 1: Define the mission environment of multiple agents, including mission area, user demand distribution, drone performance parameters and other information, to provide the agents with a basic environment for perception and decision-making; Step 2: Initialize the action-value function Q for each agent π (b t ,s t , a t )=E[U t |b t ,s t , a t ],in This function is used to evaluate the agent in different belief states b t , environmental status t Take action a t The expected cumulative discounted reward is given a reasonable initial value range during initialization so that it can be gradually optimized in subsequent learning; Step 3: Determine the objective function of the multi-agent collaboration task; Among them, step S302 utilizes the belief-oriented post-experience replay mechanism to train multiple agents through reinforcement learning, and the specific steps are as follows. Step 1: Perform belief-guided experience storage. At each time step t, the agent selects action a based on the current strategy and observation information. t After that, the experience Tr t = {b t ,s t , a t , r t , b t+1 ,s t+1 } is stored in the data set Tr = {Tr1, Tr2, ..., Tr M }. Among them, belief state b t It reflects the probability of the existence of a dynamic task request at time step t. The environment state includes information such as the location of the ground user and the task requirements. The reward is calculated based on the reward function. Step 2: Experience sampling and updating, randomly sampling experience Tr from the data set t , and according to the objective function g (from the belief state b t and environmental state s t The mapping function g to the target space G is: (b t ,s t )→G generates new experience in When the drone is moving in the correct direction, it is greater than r t These new experiences are then used to update the agent’s action-value function and belief state. Step 3: Multi-agent deep Q learning update, based on sampled experience, using the time difference training method, The parameter θ of the action-value function is updated through the gradient descent algorithm to minimize the difference between the predicted value and the target value, so that the agent can learn a better strategy. At the same time, in a multi-agent environment, the interaction and collaboration between agents are considered and a value decomposition network is introduced. Assuming that the agents fully cooperate, that is, share the same reward function, the decomposition formula satisfy Decomposing the global value into local values can alleviate the instability in the learning process and ensure that the local optimal combination is the global optimal. In step S303, based on the trained multi-agents, real-time task redistribution is performed within the UAV formation. After the multi-agents have completed training, in actual operation, each agent will perform tasks according to the current environment state s t and one's own belief state b t , select action a based on the learned strategy π t , achieving real-time task redistribution within the UAV formation. For example, the agent determines the flight path, task execution order, and resource allocation of the UAV based on factors such as task priority, UAV location, and load conditions, so as to efficiently complete tasks and adapt to changes in dynamic task requirements. When faced with sudden task requests or environmental changes, the agent can quickly adjust its decisions to ensure task continuity and system stability.
Citation Information
Cited By
Multi-robot collaborative boarding method and system based on reinforcement learning
CN120469431A
Unmanned aerial vehicle inspection path planning method and device and storage medium
CN121115887A
Unmanned aerial vehicle agent task execution and optimal trajectory generation method based on deep reinforcement learning
CN121742501A