Multi-task-oriented unmanned aerial vehicle group convening method
Through the hierarchical information collection and diffusion mechanism and multi-agent deep reinforcement learning algorithm, the resource scheduling and task response problems of drone groups in dynamic multitasking scenarios are solved, efficient task response and resource utilization are achieved, and the system's adaptability is enhanced.
Patent Information
- Application Number
- CN202510706146.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-08-08
AI Technical Summary
It is difficult for the existing technology to achieve efficient resource scheduling and task response of drone groups in dynamic multitasking scenarios, especially in large-scale clusters, where task generation time is uncertain and resource coupling is complex, and existing algorithms face the problems of large decision space and uncontrollable time.
A hierarchical information collection and diffusion mechanism is adopted to make task decisions through fine-grained nearest neighbor perception and coarse-grained long-distance estimation, combining the task selection and decision-making mechanism of neighbor information and scarce resources, and using attention mechanisms and multi-agent deep reinforcement learning algorithms.
It improves the task response efficiency and resource scheduling accuracy of drone groups in a multi-task environment, enhances the system's adaptability to dynamic environmental changes, and is suitable for complex unmanned system collaboration scenarios with limited resources and intensive tasks.
Smart Images

Figure CN120456119A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of autonomous networking and control of unmanned aerial vehicles (UAVs), and in particular relates to a multi-task oriented UAV group convening method. Background Art
[0002] In recent years, with advances in various devices and algorithms, drone platforms have demonstrated unique advantages in a variety of fields, including emergency rescue, urban governance, logistics and transportation, and smart agriculture. However, due to the limited payload and battery capacity of a single drone, it is difficult to cope with the various challenges encountered in complex and dynamic multi-task scenarios. Therefore, it is necessary to assemble multiple conventional drones with heterogeneous resources into a task force. Through information exchange, resource sharing, and collaborative cooperation among these drones, various tasks in dynamic scenarios can be completed.
[0003] In dynamic multi-tasking scenarios, task execution relies on the coordinated scheduling of multiple resources within the drone fleet, including perception, computing, and communication capabilities. Once a task is generated, drones with the appropriate resources must be properly dispatched based on their specific needs to ensure the feasibility of overall resource allocation. In a multi-tasking environment, different tasks have varying generation times and resource requirements. Therefore, drones must be guided to select tasks that best match their resource profiles. In particular, for drones carrying scarce resources, ensuring efficient resource utilization within the mission should be prioritized to improve overall mission completion efficiency and resource efficiency.
[0004] Existing research on the recruitment problem can be broadly categorized into two main groups: centralized and distributed approaches. Centralized approaches, such as solutions like mixed-integer linear programming and the multi-traveling salesman problem, are suitable for small and medium-sized groups. When faced with complex constraints, heuristic algorithms such as genetic algorithms, particle swarm optimization, and simulated annealing are often combined to solve the problem. However, these approaches often rely on global access to task information and struggle to cope with decision delays caused by dynamic task changes. To improve the dynamic adaptability of systems, distributed task allocation schemes have attracted attention. These approaches enable drones to make autonomous decisions based on their local state, often employing game theory models, auction mechanisms, and contract network protocols to achieve dynamic coordination of heterogeneous resources. Some research has further introduced reinforcement learning and multi-agent coordination mechanisms to address uncertain task objectives and complex resource coupling.
[0005] Despite extensive progress in centralized and distributed scheduling, most research has failed to delve into the entire process of task generation and dissemination, including UAV perception and participation in execution. This research has neglected the interplay between task information dissemination, resource selection strategies, and response mechanisms. In large-scale clusters, the heterogeneity of tasks and the dynamic nature of resources present three major challenges to traditional approaches. First, the uncertain timing of task generation leads to time-varying resource requirements, making global state maintenance prohibitively expensive. Second, UAVs carry multiple resources, including communication, computation, and perception (hereinafter referred to as "resources"), which are coupled to each other. Completing a specific task often requires the simultaneous support of multiple resource types, and the distribution of these resources is uneven. Therefore, task matching and node recruitment must prioritize the scarcest resources in the scenario, prioritizing their availability to achieve coordinated matching and efficient allocation of resources. Third, for the real-time response of large-scale UAVs, existing algorithms face the challenges of a large decision space and high computational overhead, making it difficult to generate timely and reasonable decisions within a tolerable timeframe. Summary of the Invention
[0006] To solve the above technical problems, the present invention provides a multi-task oriented drone group summoning method, which improves the task response efficiency and resource scheduling accuracy of the drone group in a multi-task environment, and also enhances the system's adaptability to dynamic environmental changes.
[0007] The technical solution adopted by the present invention is: a multi-task oriented drone group summoning method, the specific steps are as follows:
[0008] S1. Establish a system model for heterogeneous drone mobilization scenarios in a multi-task environment. Rasterize the area, initialize the drone's ID, location, various resources it carries, and the probability of generating a task in each grid, and establish a task model and a scarce resource model.
[0009] S2. Each drone maintains a hierarchical neighbor resource table and periodically obtains various information about its neighbors through a hierarchical information collection mechanism, including the location and resources carried by the neighbors, while also receiving real-time mission information disseminated in the scene.
[0010] S3, update the grid status, capture the task if a new task is generated, update the existing task information, and spread the information to the drones in the scene through the directed diffusion mechanism;
[0011] S4, a task selection and decision-making mechanism based on neighbor information and scarce resources, treats each UAV as an intelligent agent, uses an attention mechanism to encode neighbor information, selects candidate tasks based on its own information and task feature information, and uses a reinforcement learning algorithm combined with the encoded neighbor features to make decisions on the candidate tasks and decide whether to join the candidate task;
[0012] Among them, the task selection and decision-making mechanism based on neighbor information and scarce resources includes: a neighbor feature encoding mechanism based on the attention mechanism, a task selection mechanism based on task and resource matching, and a task decision-making mechanism based on multi-agent deep reinforcement learning.
[0013] Furthermore, the step S1 is specifically as follows:
[0014] Assume that the heterogeneous UAV summoning scenario in the multi-task environment is (X max ,Y max ,Z max ), the task area is rasterized and divided into L×W grids with a side length of d.
[0015] Among them, X max Indicates the maximum spatial range of the scene in the X-axis direction, Y max Indicates the maximum spatial range of the scene in the Y-axis direction, Z max Indicates the maximum spatial range of the scene in the Z-axis direction, and the coordinate of each grid center is g i,j ={(x i,j ,y i,j ,0)|i=[1,2,…,L],j∈[1,2,…,W]}, and set p i,j Indicates the probability of the task occurring in this grid.
[0016] In the scenario, there are N drones with heterogeneous resources. The resource set of drone i is represented as
[0017] in, Indicates the number of sensing resources, Indicates the number of computing resources, Indicates the number of communication resources.
[0018] Then, the task model and scarce resource model are established as follows:
[0019] A1, Task model;
[0020] Set the task st at time t m Feature information F m (t) is expressed as follows:
[0021]
[0022] Among them, ID m Indicates the ID of the task, pos m =(x m ,y m ,z m ) represents the location where the task occurs, t pro Indicates the time when the feature information is generated. Indicates the task start time. Indicates the deadline for task call. represents the amount of perception, computation, and communication resources required by the task at time t It represents the amount of sensing, computing, and communication resources that the mission has at time t. The mission information changes with the addition of drones and is updated and diffused.
[0023] A2, Scarce Resource Model;
[0024] If the resources owned by a drone are greater than the sum of the resources of its surrounding drones, it can be considered as a scarce resource. Assume that the resource scarcity of drone i is expressed as
[0025] Among them, type represents one of the sensing resources s, computing resources c or communication resources b. Indicates the number of corresponding resources carried by UAV i, NH i represents the set of drones surrounding drone i.
[0026] Furthermore, the step S2 is specifically as follows:
[0027] Each drone node periodically collects resource information from its neighbors and maintains a hierarchical neighbor resource table. This table is divided into detailed information within two hops and coarse regional information beyond two hops, based on the network distance between the neighboring nodes and the local node. This allows for fine-grained local resource status awareness and coarse-grained remote resource estimation.
[0028] Assume the current drone is U i , whose communication radius is r c , then the one-hop neighbor set The expression is defined as follows:
[0029]
[0030] Two-hop neighbor set The expression is defined as follows:
[0031]
[0032] Among them, p i Indicates drone Ui location.
[0033] For neighbors within two hops, UAV U i Periodically collect its specific resource information, including: UAV location p j , the number of perceived resources Number of computing resources Number of communication resources Idle state indicator δ j ∈{0,1}, where δ j =1 means the drone is idle, δ j =0 means the drone is busy.
[0034] For distant nodes beyond two hops, a regional statistical and aggregation approach is used to construct a coarse resource view. When each drone obtains resource information from a distant region, it no longer collects specific parameters for each node, but instead aggregates key statistical indicators, including the number of reachable idle nodes, the total amount of available resources, and the average resource value.
[0035] The number of reachable idle nodes The expression is as follows:
[0036]
[0037] Total available resources The expression is as follows:
[0038]
[0039] Resource average The expression is as follows:
[0040]
[0041] Each drone is then required to maintain a multi-level, structured resource information view, which includes: local resource information Remote area statistics Neighboring detailed resource information
[0042] Among them, local resource information By the current UAV U i Real-time updates of the drone itself, including: drone location p j , the number of perceived resources Number of computing resources Number of communication resources Idle state indicator δ j ; Remote area statistics Information from neighboring nodes more than two hops away does not include specific details of the drone, but is forwarded by other nodes, including: the number of idle nodes in the area Total available resources Resource average Neighboring detailed resource information It is obtained by periodically interacting with neighboring nodes within two hops. The expression is as follows:
[0043]
[0044] Furthermore, the step S3 is specifically as follows:
[0045] S31, task generation and capture, the system performs task generation determination on the grid space at each moment;
[0046] Assume that at a certain time step t, at grid g i,j Newly generated task st m , its task parameters include: the identification number ID of the task m , the location where the task occurs pos m , the amount of sensing, computing, and communication resources required by the task Task start time Task deadline call time
[0047] The drone closest to the corresponding grid center captures the task, flies to the grid center and hovers, then changes its identity to a summoning drone. Based on the resource requirements of the task, it summons other drones in the scene to form a drone swarm to perform the task.
[0048] S32, directed diffusion propagation;
[0049] UAVU i First, based on the neighbor resource table maintained by itself, the resource matching situation in each propagation direction is estimated. All possible diffusion spatial directions are recorded as a set D = {d1, d2, ..., d P For direction d p ∈D, p∈[1, P], define its resource matching function Φ i (d p ,st m )The expression is as follows:
[0050]
[0051] Where P represents the number of directions in space, Indicates drone U j The number of various resources carried, Indicates task st mThe number of various resources required, β1,β2,β3∈[0,1] represents the weighted coefficient of resource matching, U i (d p ) represents the set of reachable drones in that direction.
[0052] Finally determine the optimal diffusion direction at the current moment And give priority to broadcasting the task st to the idle drones in this direction m Meta-information includes task parameters, response deadlines, and propagation path identifiers. This process can be recursively performed hop by hop, with each relay node re-evaluating the resource matching direction to form an efficient diffusion path that "advances in the direction of resource aggregation."
[0053] Among them, each drone only determines the direction and forwards the mission information received for the first time, and discards the repeated mission information.
[0054] Furthermore, the step S4 is specifically as follows:
[0055] S41, Neighborhood feature encoding mechanism based on attention mechanism;
[0056] Setting up the drone i The number of two-hop neighbors is N i,TH , then the neighbor feature input in the attention mechanism is represented as
[0057] First calculate the attention scores of each input neighbor i =W attention G, performs a Softmax operation on the attention scores of all neighbor features to obtain the normalized attention weights
[0058] Among them, W attention represents the weight matrix.
[0059] Then, the neighbor features are weighted and summed according to the normalized attention weights, and all weighted neighbor features are summed to obtain the aggregate feature representation
[0060] Among them, g i Represents the characteristic information of a single neighbor.
[0061] Finally, the aggregated features are input into the fully connected layer to obtain the final encoding feature y i,TH =fc out (z).
[0062] Among them, fc out represents a fully connected layer.
[0063] S42, Task selection mechanism based on task and resource matching;
[0064] Setting up the drone i The set of task feature information perceived at time t is F Ui (t)={F1(t),F2(t),…F m (t)}.
[0065] Among them, each task feature information F m (t) corresponds to a specific task.
[0066] For the feature information F m (t), calculate the distance between the UAV and the corresponding mission center position And considering the distance between the two and the flight speed of the UAV, ensure that the UAV reaches the center of the mission before the mission deadline, that is,
[0067] Among them, (x i ,y i ,z i ) indicates drone U i The three-dimensional coordinates, (x m ,y m ,z m ) indicates the task st m The center coordinates of ST i,opt (t) represents the UAV U i The set of subtasks that can be selected at time t, Indicates task st m The call deadline, V U Indicates the flight speed of the drone.
[0068] Then evaluate the adaptability of the drone to the task and set the resource matching degree
[0069] in, Indicates drone U i and mission st m Resource matching, Indicates drone U i The number of resources of the corresponding type carried, Indicates task st m The required quantity of resources of the corresponding type, Indicates task st m The number of resources of the corresponding type currently summoned.
[0070] After considering the distance and resource matching, UAV U i Calculate the comprehensive selection weight W for each character i , expressed as
[0071] Among them, ω1 and ω2 represent weight coefficients, that is, the preference of the drone for distance and resource matching.
[0072] Finally, for the task st m UAVU i The task selection strategy based on location and resource matching is:
[0073] Among them, P i,m (t) represents the UAV U i Select subtask st at time t m probability.
[0074] S43. Task decision-making mechanism based on multi-agent deep reinforcement learning;
[0075] B1. First, consider each drone as an agent and define the environment state, action, and reward of each agent as follows:
[0076] (1) Status;
[0077] The status includes: basic information of the current drone The vector y after encoding the neighbor feature information i,TH , the task feature information of the candidate task Concatenate them together as the input of the neural network.
[0078] Among them, C i (t) represents the coordinates of UAV i at time t, Represents the characteristic information of the candidate task selected after step S42. The state space S of drone i at time t i (t) is expressed as follows:
[0079]
[0080] Among them, the basic information of drone i includes: current position pos i and the sensing resources, computing resources, and communication resources carried
[0081] (2) Action;
[0082] For the drone, its decision on the candidate task is to accept or refuse to join the task, so its action space A i (t) is expressed as follows:
[0083] A i (t)={0,1}
[0084] Among them, Ai (t)=0 means that the UAV refuses to join the candidate mission and remains idle. i (t) = 1 means that the UAV accepts the candidate task and changes to the busy state.
[0085] (3) Rewards;
[0086] Supplementary definition of resource usage, for UAV i and candidate task st m The existence relation expression is as follows:
[0087]
[0088] R used =min(R give ,R demand )
[0089] R unused =R give -R used
[0090] R beyond =max(R give -R demand ,0)
[0091] Among them, R give Indicates the resources carried by the drone, R demand Indicates the resources required for the task, R used Indicates the resources used by the task, R unused Indicates the resources provided by the drone but not used by the mission, R beyond Indicates resources that exceed the task's requirements.
[0092] The reward function of the drone is divided into three parts, namely resource contribution reward, scarce resource utilization reward, and task completion reward.
[0093] Resource Contribution Rewards task , rewards the drone for successfully joining the mission and contributing resources to the mission. The expression is as follows:
[0094]
[0095] Among them, ω task represents the resource contribution reward coefficient, Indicates the number of resources of the corresponding type carried by the drone, Indicates the quantity of resources of the corresponding type required by the task.
[0096] Scarce resource utilization reward scarce , when the drone uses scarce resources to perform tasks, it will be given a certain reward. The overall expression is as follows:
[0097]
[0098] Among them, ω scarce Represents the scarcity coefficient.
[0099] Mission Completion Rewards finish , when a task is completed, a global reward is given to drive the drone to complete the task as much as possible. The expression is as follows:
[0100] r finish =ω finish ·F
[0101] Among them, ω finish represents the task completion reward coefficient, F represents the reward value for task completion, and is a constant.
[0102] In summary, the reward R obtained by the drone for taking action i (t) is expressed as follows:
[0103] R i (t) = r task +r scarce +r finish
[0104] B2. Use the dual deep Q network (DDQN) reinforcement learning algorithm to initialize the policy network and target network of each agent and initialize the shared experience replay pool.
[0105] B3: The drone constructs the environment state vector S and uses the ε-greedy strategy to select the action A corresponding to this state, that is, whether to join the candidate task;
[0106] B4. The drone performs action A to interact with the environment and obtains a new state S′ and the corresponding reward R. ′ ) is stored in the global shared experience replay pool;
[0107] B5. Each drone learns and updates the policy network and value function network parameters based on the experience replay pool;
[0108] B6. Sample a fixed number of samples in the experience replay pool and calculate the Q value of the policy network.
[0109] B7. Calculate the error and update the network parameters according to back propagation;
[0110] B8. Update the target network parameters to meet the target network update frequency;
[0111] B9. Repeat steps B3-B8 to continuously train the drone until the mission in the scene is completed.
[0112] Beneficial effects of the present invention: The method of the present invention divides information interaction into fine-grained neighbor perception and coarse-grained long-distance estimation through a hierarchical information collection and diffusion mechanism, sharing detailed resource information and regional summary information respectively, and then based on the task selection and decision-making mechanism of neighbor information and resource scarcity, after receiving multiple task information, the drone comprehensively considers the resource requirements, distance, and time constraints of the task, and combines the resource availability of the neighbors and the scarcity of the required resources in the scene to select candidate tasks and decide whether to join the task. The method of the present invention takes into account both communication efficiency and scheduling accuracy, which not only improves the task response efficiency and resource scheduling accuracy of the drone group in a multi-task environment, but also enhances the system's adaptive ability in the face of dynamic environmental changes. It has good scalability and robustness, and is suitable for complex unmanned system collaboration scenarios with limited resources and intensive tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0113] Figure 1 This is a flow chart of a multi-task oriented drone group summoning method of the present invention.
[0114] Figure 2 Schematic diagram of a heterogeneous UAV convening scenario in a multi-task environment according to an embodiment of the present invention. DETAILED DESCRIPTION
[0115] The method of the present invention is further described below with reference to the accompanying drawings and embodiments.
[0116] like Figure 1 As shown in FIG, a flowchart of a multi-task drone group summoning method of the present invention is shown, and the specific steps are as follows:
[0117] S1. Establish a system model for heterogeneous drone mobilization scenarios in a multi-task environment. Rasterize the area, initialize the drone's ID, location, various resources it carries, and the probability of generating a task in each grid, and establish a task model and a scarce resource model.
[0118] In this embodiment, the hierarchical information collection and diffusion mechanism adopted in steps S2-S3 includes: a hierarchical information collection mechanism and a directed diffusion mechanism. Step S4 adopts a task selection and decision-making mechanism based on neighbor information and scarce resources.
[0119] S2. Each drone maintains a hierarchical neighbor resource table and periodically obtains various information about its neighbors through a hierarchical information collection mechanism, including the location and resources carried by the neighbors, while also receiving real-time mission information disseminated in the scene.
[0120] S3, update the grid status, capture the task if a new task is generated, update the existing task information, and spread the information to the drones in the scene through the directed diffusion mechanism;
[0121] S4, a task selection and decision-making mechanism based on neighbor information and scarce resources, treats each UAV as an intelligent agent, uses an attention mechanism to encode neighbor information, selects candidate tasks based on its own information and task feature information, and uses a reinforcement learning algorithm combined with the encoded neighbor features to make decisions on the candidate tasks and decide whether to join the candidate task;
[0122] Among them, the task selection and decision-making mechanism based on neighbor information and scarce resources includes: a neighbor feature encoding mechanism based on the attention mechanism, a task selection mechanism based on task and resource matching, and a task decision-making mechanism based on multi-agent deep reinforcement learning.
[0123] The feature encoding mechanism weights the influence of different neighbors through the attention mechanism, which can effectively deal with the problem of inconsistent input dimensions caused by the non-fixed number of neighbors and provide a unified and discriminative state representation for reinforcement learning.
[0124] The task selection mechanism can comprehensively consider multiple factors such as distance, resources and execution time, and select one as a candidate task from multiple task information received by the drone.
[0125] The task decision mechanism can combine the UAV's own resource status, location, candidate tasks and the resource distribution of its neighbors to determine whether it should join the candidate task, thereby achieving the priority allocation of scarce resources and maximizing the efficiency of group collaboration.
[0126] like Figure 2 As shown, in this embodiment, the step S1 is specifically as follows:
[0127] Assume that the heterogeneous UAV summoning scenario in the multi-task environment is (X max ,Y max ,Z max ), the task area is rasterized and divided into L×W grids with a side length of d.
[0128] Among them, X max Indicates the maximum spatial range of the scene in the X-axis direction, Y max Indicates the maximum spatial range of the scene in the Y-axis direction, Z max Indicates the maximum spatial range of the scene in the Z-axis direction, and the coordinate of each grid center is g i,j ={(x i,j ,y i,j ,0)|i=[1,2,…,L],j∈[1,2,…,W]}, and set p i,j Indicates the probability of the task occurring in this grid.
[0129] In the scenario, there are N drones with heterogeneous resources. The resource set of drone i is represented as
[0130] in, Indicates the number of sensing resources, Indicates the number of computing resources, Indicates the number of communication resources.
[0131] When recruiting drones, it's necessary to disseminate current mission information so they can determine whether to join the mission. Therefore, key aspects of the mission are extracted as feature information. When resource requirements change as drones join the mission, the feature information can be modified synchronously to ensure consistency between the feature information and the mission itself. This allows for the creation of a mission model.
[0132] Real-world multi-task, multi-drone collaboration scenarios often involve the scheduling and allocation of multiple resources. Especially in more complex environments, mission demands and environmental uncertainties can cause certain resources to become bottlenecks that limit the swarm's mission execution, making these resources scarce. Scarce resources are limited within the current mission execution environment and critical to mission success. For example, for sensing resources in a swarm, some high-precision sensors are typically not universally available on all drones due to cost and size constraints. Therefore, sensing resources can become scarce in certain missions. Similarly, computing and communication resources can also become scarce in specific circumstances. Reasonable resource scheduling and allocation, especially the optimal utilization of scarce resources, are key factors in improving swarm efficiency and mission success. Therefore, a scarce resource model is developed.
[0133] Then, the task model and scarce resource model are established as follows:
[0134] A1, Task model;
[0135] Set the task st at time t m Feature information F m (t) is expressed as follows:
[0136]
[0137] Among them, ID m Indicates the ID of the task, pos m =(x m ,y m ,z m ) represents the location where the task occurs, t pro Indicates the time when the feature information is generated. Indicates the task start time. Indicates the deadline for task call. represents the amount of perception, computation, and communication resources required by the task at time t, It represents the amount of sensing, computing, and communication resources that the mission has at time t. The mission information changes with the addition of drones and is updated and diffused.
[0138] A2, Scarce Resource Model;
[0139] The definition of scarce resources is based on the comparison between a certain resource carried by each drone and the total amount of the resource in the surrounding environment. If a drone has more resources than the sum of the resources of its surrounding drones, it can be considered a scarce resource. Let the resource scarcity of drone i be expressed as
[0140] Among them, type represents one of the sensing resources s, computing resources c or communication resources b. Indicates the number of corresponding resources carried by UAV i, NH i Represents the set of drones surrounding drone i. Scarcity L i,type The larger the value, the more scarce the resource is in the environment around UAV i, and it is expected that this UAV can make better use of its scarce resources when participating in the mission.
[0141] In this embodiment, step S2 is specifically as follows:
[0142] To support efficient task selection and scheduling, each drone node periodically collects resource information from its neighbors and maintains a hierarchical neighbor resource table. This table is divided into detailed information within two hops and coarse regional information beyond two hops, based on the network distance between the neighboring nodes and the local node. This allows for fine-grained local resource status awareness and coarse-grained remote resource estimation.
[0143] Assume the current drone is U i , whose communication radius is r c , then the one-hop neighbor set The expression is defined as follows:
[0144]
[0145] Two-hop neighbor set The expression is defined as follows:
[0146]
[0147] Among them, p i Indicates drone U i location.
[0148] For neighbors within two hops, UAV U iPeriodically collect its specific resource information, including: UAV location p j , the number of perceived resources Number of computing resources Number of communication resources Idle state indicator δ j ∈{0,1}, where δ j =1 means the drone is idle, δ j = 0 means the drone is busy. This part of the information constitutes the local resource status of the drone, with high accuracy and real-time performance.
[0149] For distant nodes beyond two hops, due to the lack of direct communication capabilities and the need for real-time information, this embodiment uses a regional statistical and aggregation approach to construct a coarse resource view. When each drone obtains resource information from a distant region, it no longer collects specific parameters for each node, but instead aggregates key statistical indicators, including the number of reachable idle nodes, the total amount of available resources, and the average resource value.
[0150] The number of reachable idle nodes The expression is as follows:
[0151]
[0152] Total available resources The expression is as follows:
[0153]
[0154] Resource average The expression is as follows:
[0155]
[0156] The above aggregated information can be pre-processed by relay nodes or nodes within the region during the information dissemination process and disseminated along with the task information, thereby providing the node with an ability to approximately perceive the resource layout of the remote region and avoid excessive communication overhead due to information redundancy.
[0157] In order to support subsequent task diffusion, node convening and resource matching, each drone maintains a multi-level, structured resource information view, which includes: local resource information Remote area statistics Neighboring detailed resource information
[0158] Among them, local resource information By the current UAV U i Real-time updates of the drone itself, including: drone location p j , the number of perceived resources Number of computing resources Number of communication resources Idle state indicator δ j ; Remote area statistics Information from neighboring nodes more than two hops away does not include specific details of the drone, but is forwarded by other nodes, including: the number of idle nodes in the area Total available resources Resource average Neighboring detailed resource information It is obtained by periodically interacting with neighboring nodes within two hops. The expression is as follows:
[0159]
[0160] Through this hierarchical resource information collection mechanism, each drone can build a global resource view under the premise of controllable communication load. When faced with key steps such as task diffusion, direction evaluation, and candidate task selection, it has sufficient data support to achieve more efficient and reasonable group resource scheduling and task matching.
[0161] In this embodiment, step S3 is specifically as follows:
[0162] S31, task generation and capture, the system performs task generation determination on the grid space at each moment;
[0163] Assume that at a certain time step t, at grid g i,j Newly generated task st m , its task parameters include: the identification number ID of the task m , the location where the task occurs pos m , the amount of sensing, computing, and communication resources required by the task Task start time Task deadline call time
[0164] The drone closest to the corresponding grid center captures the task, flies to the grid center and hovers, then changes its identity to a summoning drone. Based on the resource requirements of the task, it summons other drones in the scene to form a drone swarm to perform the task.
[0165] S32, directed diffusion propagation;
[0166] UAVU i First, based on the neighbor resource table maintained by itself, the resource matching situation in each propagation direction is estimated. All possible diffusion spatial directions are recorded as a set D = {d1, d2, ..., d P For direction d p∈D, p∈[1, P], define its resource matching function Φ i (d p ,st m )The expression is as follows:
[0167]
[0168] Where P represents the number of directions in space, Indicates drone U j The number of various resources carried, Indicates task st m The number of various resources required, β1, β2, β3∈[0,1] represent the weighted coefficients of resource matching, which are used to express the preference for different resource types, U i (d p ) represents the set of drones that can be reached in this direction. Matching index Φ i (d p ,st m ) is larger, indicating that the resources in this direction are more sufficient and more suitable for the current task.
[0169] Finally determine the optimal diffusion direction at the current moment And give priority to broadcasting the task st to the idle drones in this direction m Meta-information includes task parameters, response deadlines, and propagation path identifiers. This process can be recursively performed hop by hop, with each relay node re-evaluating the resource matching direction to form an efficient diffusion path that "advances in the direction of resource aggregation."
[0170] Among them, each drone only determines the direction and forwards the mission information received for the first time, and discards the repeated mission information.
[0171] In this embodiment, step S4 is specifically as follows:
[0172] S41, Neighborhood feature encoding mechanism based on attention mechanism;
[0173] Setting up the drone i The number of two-hop neighbors is N i,TH , then the neighbor feature input in the attention mechanism is represented as
[0174] First calculate the attention scores of each input neighbor i =W attention G, performs a Softmax operation on the attention scores of all neighbor features to obtain the normalized attention weights
[0175] Among them, Wattention represents the weight matrix.
[0176] Then, the neighbor features are weighted and summed according to the normalized attention weights, and all weighted neighbor features are summed to obtain the aggregate feature representation
[0177] Among them, g i Represents the characteristic information of a single neighbor.
[0178] Finally, the aggregated features are input into the fully connected layer to obtain the final encoding feature y i,TH =fc out (z).
[0179] Among them, fc out represents a fully connected layer.
[0180] S42, Task selection mechanism based on task and resource matching;
[0181] Setting up the drone i The set of task feature information perceived at time t is:
[0182] Among them, each task feature information F m (t) corresponds to a specific task.
[0183] For the feature information F m (t), calculate the distance between the UAV and the corresponding mission center position And considering the distance between the two and the flight speed of the UAV, ensure that the UAV reaches the center of the mission before the mission deadline, that is,
[0184] Among them, (x i ,y i ,z i ) indicates drone U i The three-dimensional coordinates, (x m ,y m ,z m ) indicates the task st m The center coordinates of ST i,opt (t) represents the UAV U i The set of subtasks that can be selected at time t, Indicates task st m The call deadline, V U Indicates the flight speed of the drone.
[0185] Then evaluate the adaptability of the drone to the task and set the resource matching degree
[0186] in, Indicates drone U i and mission st m Resource matching, Indicates drone U i The number of resources of the corresponding type carried, Indicates task st m The required quantity of resources of the corresponding type, Indicates task st m The number of resources of the corresponding type currently summoned.
[0187] After considering the distance and resource matching, UAV U i Calculate the comprehensive selection weight W for each character i , expressed as
[0188] Among them, ω1 and ω2 represent weight coefficients, that is, the preference of the drone for distance and resource matching. These two coefficients are adjusted to adapt to different scene conditions.
[0189] Finally, for the task st m UAVU i The task selection strategy based on location and resource matching is:
[0190] Among them, P i,m (t) represents the UAV U i Select subtask st at time t m probability.
[0191] In this way, the probability of drone selection not only considers the distance of the task, but also integrates the matching degree of resources, avoiding the over-selection of tasks that are too close or have low resource matching degree, thereby reducing resource waste and optimizing task allocation.
[0192] S43. Task decision-making mechanism based on multi-agent deep reinforcement learning;
[0193] B1. First, consider each drone as an agent and define the environment state, action, and reward of each agent as follows:
[0194] (1) Status;
[0195] The status includes: basic information of the current drone The vector y after encoding the neighbor feature information i,TH , the task feature information of the candidate task Concatenate them together as the input of the neural network.
[0196] Among them, C i(t) represents the coordinates of UAV i at time t, Represents the characteristic information of the candidate task selected after step S42. The state space S of drone i at time t i (t) is expressed as follows:
[0197]
[0198] Among them, the basic information of drone i includes: current position pos i and the sensing resources, computing resources, and communication resources carried
[0199] (2) Action;
[0200] For the drone, its decision on the candidate task is to accept or refuse to join the task, so its action space A i (t) is expressed as follows:
[0201] A i (t)={0,1}
[0202] Among them, A i (t)=0 means that the UAV refuses to join the candidate mission and remains idle. i (t) = 1 means that the UAV accepts the candidate task and changes to the busy state.
[0203] (3) Rewards;
[0204] Supplementary definition of resource usage, for UAV i and candidate task st m The existence relation expression is as follows:
[0205]
[0206] R used =min(R give ,R demand )
[0207] R unused =R give -R used
[0208] R beyond =max(R give -R demand ,0)
[0209] Among them, R give Indicates the resources carried by the drone, R demand Indicates the resources required for the task, R used Indicates the resources used by the task, R unusedIndicates the resources provided by the drone but not used by the mission, R beyond Indicates resources that exceed the task's requirements.
[0210] The reward function of the drone is divided into three parts, namely resource contribution reward, scarce resource utilization reward, and task completion reward.
[0211] Resource Contribution Rewards task , rewards the drone for successfully joining the mission and contributing resources to the mission. The expression is as follows:
[0212]
[0213] Among them, ω task represents the resource contribution reward coefficient, Indicates the number of resources of the corresponding type carried by the drone, Indicates the quantity of resources of the corresponding type required by the task.
[0214] Scarce resource utilization reward scarce , when the drone uses scarce resources to perform tasks, it will be given a certain reward. The overall expression is as follows:
[0215]
[0216] Among them, ω scarce Represents the scarcity coefficient.
[0217] Mission Completion Rewards finish , when a task is completed, a global reward is given to drive the drone to complete the task as much as possible. The expression is as follows:
[0218] r finish =ω finish ·F
[0219] Among them, ω finish represents the task completion reward coefficient, F represents the reward value for task completion, and is a constant.
[0220] In summary, the reward R obtained by the drone for taking action i (t) is expressed as follows:
[0221] R i (t) = r task +r scarce +r finish
[0222] By setting this reward function, the UAV can gradually learn the appropriate decision-making method during the interaction with the environment to ensure the realization of the optimization goal.
[0223] B2. Use the dual deep Q network (DDQN) reinforcement learning algorithm to initialize the policy network and target network of each agent and initialize the shared experience replay pool.
[0224] B3: The drone constructs the environment state vector S and uses the ε-greedy strategy to select the action A corresponding to this state, that is, whether to join the candidate task;
[0225] B4. The drone performs action A to interact with the environment and obtains a new state S′ and the corresponding reward R. ′ ) is stored in the global shared experience replay pool;
[0226] B5. Each drone learns and updates the policy network and value function network parameters based on the experience replay pool;
[0227] B6. Sample a fixed number of samples in the experience replay pool and calculate the Q value of the policy network.
[0228] B7. Calculate the error and update the network parameters according to back propagation;
[0229] B8. Update the target network parameters to meet the target network update frequency;
[0230] B9. Repeat steps B3-B8 to continuously train the drone until the mission in the scene is completed.
[0231] In summary, by constructing a hierarchical information collection and dissemination mechanism, combined with a task selection and decision-making mechanism based on neighbor resource awareness and scarce resource prioritization, efficient and adaptive scheduling and collaboration of drone groups in multi-task environments can be achieved. By precisely understanding the resource status of neighboring nodes locally and maintaining a coarse-grained regional resource view over long distances, drones can effectively perceive the overall resource distribution, dynamically select diffusion paths, and guide the dissemination of task information toward more resource-compatible directions. Furthermore, by leveraging feature encoding through an attention mechanism and strategy optimization through reinforcement learning, drones can intelligently select appropriate tasks and decide whether to participate based on their own status and neighbor information, thereby improving the responsiveness of group collaboration, resource utilization, and task success rate.
[0232] Those skilled in the art will appreciate that the embodiments described herein are intended to aid the reader in understanding the principles of the present invention, and it should be understood that the scope of the present invention is not limited to such specific descriptions and embodiments. Various modifications and variations are readily apparent to those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention are intended to be included within the scope of the claims.
Claims
1. A multi-task oriented drone group summoning method, the specific steps are as follows: S1. Establish a system model for heterogeneous drone mobilization scenarios in a multi-task environment. Rasterize the area, initialize the drone's ID, location, various resources it carries, and the probability of generating a task in each grid, and establish a task model and a scarce resource model. S2. Each drone maintains a hierarchical neighbor resource table and periodically obtains various information about its neighbors through a hierarchical information collection mechanism, including the location and resources carried by the neighbors, while also receiving real-time mission information disseminated in the scene. S3, update the grid status, capture the task if a new task is generated, update the existing task information, and spread the information to the drones in the scene through the directed diffusion mechanism; S4, a task selection and decision-making mechanism based on neighbor information and scarce resources, treats each UAV as an intelligent agent, uses an attention mechanism to encode neighbor information, selects candidate tasks based on its own information and task feature information, and uses a reinforcement learning algorithm combined with the encoded neighbor features to make decisions on the candidate tasks and decide whether to join the candidate task; in, The task selection and decision-making mechanism based on neighbor information and scarce resources includes: a neighbor feature encoding mechanism based on the attention mechanism, a task selection mechanism based on task and resource matching, and a task decision-making mechanism based on multi-agent deep reinforcement learning.
2. A multi-task oriented drone group summoning method according to claim 1, characterized in that: The step S1 is specifically as follows: Assume that the heterogeneous UAV summoning scenario in the multi-task environment is (X max ,Y max ,Z max ), the task area is rasterized and divided into L×W grids with a side length of d; Among them, X max Indicates the maximum spatial range of the scene in the X-axis direction, Y max Indicates the maximum spatial range of the scene in the Y-axis direction, Z max Indicates the maximum spatial range of the scene in the Z-axis direction, and the coordinate of each grid center is g i,j ={(x i,j ,y i,j ,0)|i=[1,2,…,L],j∈[1,2,…,W]}, and set p i,j Indicates the probability of the task occurring in this grid; In the scenario, there are N drones with heterogeneous resources. The resource set of drone i is represented as in, Indicates the number of sensing resources, Indicates the number of computing resources, Indicates the number of communication resources; Then, the task model and scarce resource model are established as follows: A1, Task model; Set the task st at time t m Feature information F m (t) is expressed as follows: Among them, ID m Indicates the ID of the task, pos m =(x m ,y m ,z m ) represents the location where the task occurs, t pro Indicates the time when the feature information is generated. Indicates the task start time. Indicates the deadline for task call. represents the amount of perception, computation, and communication resources required by the task at time t, It represents the amount of sensing, computing, and communication resources that the mission has at time t; and the mission information changes with the addition of drones, and is updated and diffused; A2, Scarce Resource Model; If a drone has more resources than the sum of the resources of its surrounding drones, it can be considered a scarce resource. The resource scarcity of drone i is expressed as Among them, type represents one of the sensing resources s, computing resources c or communication resources b. Indicates the number of corresponding resources carried by UAV i, NH i represents the set of drones surrounding drone i.
3. The multi-task oriented drone group summoning method according to claim 1, characterized in that: The step S2 is specifically as follows: Each drone node periodically collects resource information from its neighbors and maintains a hierarchical neighbor resource table. This table is divided into detailed information within two hops and rough regional information beyond two hops, based on the network distance between the neighboring node and the local node. This allows for fine-grained local resource status perception and coarse-grained remote resource estimation. Assume the current drone is U i , whose communication radius is r c , then the one-hop neighbor set The expression is defined as follows: Two-hop neighbor set The expression is defined as follows: Among them, p i Indicates drone U i location; For neighbors within two hops, UAV U i Periodically collect its specific resource information, including: UAV location p j , the number of perceived resources Number of computing resources Number of communication resources Idle state indicator δ j ∈{0,1}, where δ j =1 means the drone is idle, δ j =0 means the drone is busy; For distant nodes beyond two hops, a regional statistical and aggregation method is used to construct a rough resource view. When each drone obtains resource information from a distant area, it no longer collects specific parameters for each node, but aggregates key statistical indicators, including the number of reachable idle nodes, the total amount of available resources, and the average resource value. The number of reachable idle nodes The expression is as follows: Total available resources The expression is as follows: Resource average The expression is as follows: Each drone is then required to maintain a multi-level, structured resource information view, which includes: local resource information Remote area statistics Neighboring detailed resource information Among them, local resource information By the current UAV U i Real-time updates of the drone itself, including: drone location p j , the number of perceived resources Number of computing resources Number of communication resources Idle state indicator δ j ; Remote area statistics Information from neighboring nodes more than two hops away does not include specific details of the drone, but is forwarded by other nodes, including: the number of idle nodes in the area Total available resources Resource average Neighboring detailed resource information It is obtained by periodically interacting with neighboring nodes within two hops. The expression is as follows:
4. The multi-task oriented drone group summoning method according to claim 1, characterized in that: The step S3 is specifically as follows: S31, task generation and capture, the system performs task generation determination on the grid space at each moment; Assume that at a certain time step t, at grid g i,j Newly generated task st m , its task parameters include: the identification number ID of the task m , the location where the task occurs pos m , the amount of sensing, computing, and communication resources required by the task Task start time Task deadline call time The drone closest to the corresponding grid center captures the task, flies to the grid center and hovers, then switches its identity to a summoning drone. Based on the resource requirements of the task, it summons other drones in the scene to form a drone swarm to perform the task. S32, directed diffusion propagation; UAVU i First, based on the neighbor resource table maintained by itself, the resource matching situation in each propagation direction is estimated; all possible diffusion spatial directions are recorded as a set D = {d1, d2, ..., d P }; For direction d p ∈D, p∈[1, P], define its resource matching function Φ i (d p ,st m )The expression is as follows: Where P represents the number of directions in space, Indicates drone U j The number of various resources carried, Indicates task st m The number of various resources required, β1,β2,β3∈[0,1] represents the weighted coefficient of resource matching, U i (d p ) represents the set of reachable drones in that direction; Finally determine the optimal diffusion direction at the current moment And give priority to broadcasting the task st to the idle drones in this direction m Meta-information includes task parameters, response deadlines, and propagation path identifiers. This process can be performed recursively hop by hop, with each relay node re-evaluating the resource matching direction to form an efficient diffusion path that "advances in the direction of resource aggregation." Among them, each drone only determines the direction and forwards the mission information received for the first time, and discards the repeated mission information.
5. The multi-task oriented drone group summoning method according to claim 1, characterized in that: The step S4 is specifically as follows: S41, Neighborhood feature encoding mechanism based on attention mechanism; Setting up the drone i The number of two-hop neighbors is N i,TH , then the neighbor feature input in the attention mechanism is represented as First calculate the attention scores of each input neighbor i =W attention G, performs a Softmax operation on the attention scores of all neighbor features to obtain the normalized attention weights Among them, W attention represents the weight matrix; Then, the neighbor features are weighted and summed according to the normalized attention weights, and all weighted neighbor features are summed to obtain the aggregate feature representation Among them, g i Represents the characteristic information of a single neighbor; Finally, the aggregated features are input into the fully connected layer to obtain the final encoding feature y i,TH =fc out (z); Among them, fc out represents a fully connected layer; S42, Task selection mechanism based on task and resource matching; Setting up the drone i The set of task feature information perceived at time t is F Ui (t)={F1(t),F2(t),…F m (t)}; Among them, each task feature information F m (t) all correspond to a specific task; For the feature information F m (t), calculate the distance between the UAV and the corresponding mission center position And considering the distance between the two and the flight speed of the UAV, ensure that the UAV reaches the center of the mission before the mission deadline, that is, Among them, (x i ,y i ,z i ) indicates drone U i The three-dimensional coordinates, (x m ,y m ,z m ) indicates the task st m The center coordinates of ST i,opt (t) represents the UAV U i The set of subtasks that can be selected at time t, Indicates task st m The call deadline, V U Indicates the flight speed of the drone; Then evaluate the adaptability of the drone to the task and set the resource matching degree in, Indicates drone U i and mission st m Resource matching, Indicates drone U i The number of resources of the corresponding type carried, Indicates task st m The required quantity of resources of the corresponding type, Indicates task st m The number of resources of the corresponding type currently summoned; After considering the distance and resource matching, UAV U i Calculate the comprehensive selection weight W for each character i , expressed as Among them, ω1 and ω2 represent weight coefficients, i.e., the preference of the UAV for distance and resource matching; Finally, for the task st m UAVU i The task selection strategy based on location and resource matching is: Among them, P i,m (t) represents the UAV U i Select subtask st at time t m probability; S43. Task decision-making mechanism based on multi-agent deep reinforcement learning; B1. First, consider each drone as an agent and define the environment state, action, and reward of each agent as follows: (1) Status; The status includes: basic information of the current drone The vector y after encoding the neighbor feature information i,TH , the task feature information of the candidate task Concatenate them together as input to the neural network; Among them, C i (t) represents the coordinates of UAV i at time t, Represents the characteristic information of the candidate task selected after step S42; the state space S of drone i at time t i (t) is expressed as follows: Among them, the basic information of drone i includes: current position pos i and the sensing resources, computing resources, and communication resources carried (2) Action; For the drone, its decision on the candidate task is to accept or refuse to join the task, so its action space A i (t) is expressed as follows: A i (t)={0,1} Among them, A i (t)=0 means that the UAV refuses to join the candidate mission and remains idle. i (t) = 1 means that the UAV accepts the candidate task and changes to the busy state; (3) Rewards; Supplementary definition of resource usage, for UAV i and candidate task st m The existence relation expression is as follows: R used =min(R give ,R demand ) R unused =R give -R used R beyond =max(R give -R demand ,0) Among them, R give Indicates the resources carried by the drone, R demand Indicates the resources required for the task, R used Indicates the resources used by the task, R unused Indicates the resources provided by the drone but not used by the mission, R beyond Indicates resources that exceed task requirements; The reward function of the drone is divided into three parts, namely resource contribution reward, scarce resource utilization reward, and task completion reward; Resource Contribution Rewards task , rewards the drone for successfully joining the mission and contributing resources to the mission. The expression is as follows: Among them, ω task represents the resource contribution reward coefficient, Indicates the number of resources of the corresponding type carried by the drone, Indicates the quantity of resources of the corresponding type required by the task; Scarce resource utilization reward scarce , when the drone uses scarce resources to perform tasks, it will be given a certain reward. The overall expression is as follows: Among them, ω scarce represents the scarcity coefficient; Mission Completion Rewards finish , when a task is completed, a global reward is given to drive the drone to complete the task as much as possible. The expression is as follows: r finish =ω finish ·F Among them, ω finish represents the task completion reward coefficient, F represents the reward value for task completion, which is a constant; In summary, the reward R obtained by the drone for taking action i (t) is expressed as follows: R i (t)=r task +r scarce +r finish B2. Use the dual deep Q network (DDQN) reinforcement learning algorithm to initialize the policy network and target network of each agent and initialize the shared experience replay pool. B3: The drone constructs the environment state vector S and uses the ε-greedy strategy to select the action A corresponding to this state, that is, whether to join the candidate task; B4. The drone performs action A to interact with the environment and obtains a new state S′ and the corresponding reward R. ′ ) is stored in the global shared experience replay pool; B5. Each drone learns and updates the policy network and value function network parameters based on the experience replay pool; B6. Sample a fixed number of samples in the experience replay pool and calculate the Q value of the policy network. B7. Calculate the error and update the network parameters according to back propagation; B8. Update the target network parameters to meet the target network update frequency; B9. Repeat steps B3-B8 to continuously train the drone until the mission in the scene is completed.
Citation Information
Cited By
Cluster control method and system for underwater unmanned equipment
CN121433263A
Multi-target-based unmanned aerial vehicle group collaborative confrontation method, system and application
CN121806982A