Multi-uav cooperative resource allocation method based on hierarchical graph attention network
By proposing a multi-UAV collaborative resource allocation method based on hierarchical graph attention networks, the efficiency and fairness issues of collaborative offloading and resource allocation in multi-UAV assisted mobile edge computing are solved, achieving efficient task offloading and resource allocation, and improving the overall performance and long-term service balance of the system.
Patent Information
- Application Number
- CN202610736320.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-26
- Publication Date
- 2026-07-10
AI Technical Summary
In multi-drone-assisted mobile edge computing scenarios, existing technologies struggle to effectively coordinate task offloading and resource allocation among drones, leading to localized overload, resource waste, and service imbalance. Furthermore, it is difficult to balance user service fairness with drone load fairness.
A multi-UAV collaborative resource allocation method based on hierarchical graph attention network is adopted. By constructing a multi-agent partially observable Markov decision process, combining local observation vectors, mixed action space, joint reward function and optimization objective function, the hierarchical graph attention network is used to encode collaborative observation features, and a multi-agent reinforcement learning approach with centralized training and decentralized execution is adopted to update policy parameters.
It achieves intelligent balancing of performance indicators such as task offloading, trajectory control, system latency, energy consumption, and service fairness in dynamic mobile user scenarios, improving the efficiency of multi-UAV collaborative task offloading and resource allocation and long-term service balance.
Smart Images

Figure CN122373058A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of wireless communication and edge intelligence technology, specifically relating to a multi-UAV collaborative resource allocation method based on a hierarchical graph attention network. Background Technology
[0002] In multi-UAV-assisted mobile edge computing scenarios, the high mobility of user terminals, continuous task arrival, and uneven spatial distribution pose significant challenges to resource-constrained edge service systems. Specifically, frequent user movement leads to changes in network topology, continuous task arrival places high demands on the real-time performance of edge service systems, and uneven spatial distribution can easily cause local overload or resource idleness. Simultaneously, the available computing resources, wireless communication bandwidth, and onboard energy of UAVs are limited due to their size, weight, and endurance. Therefore, how to efficiently complete task access and offloading between users and UAVs, rationally allocate UAV computing and communication resources, and achieve collaborative computing among multiple UAVs under these constraints has become a key technical problem that urgently needs to be solved in this field.
[0003] To address the aforementioned issues, existing research has proposed various technical solutions, such as: maximizing system computational efficiency by employing a multi-agent deep deterministic policy gradient algorithm to jointly optimize UAV trajectories, computational resource allocation, communication resource allocation, and task offloading ratios; minimizing system latency by dividing priority queues according to task urgency and using a multi-agent deep deterministic policy gradient algorithm to implement task migration decisions; and constructing and encoding a spatially heterogeneous graph using a hierarchical graph attention network to extract local environmental structure and dynamic state features, and then employing a multi-agent dual-delay deep deterministic policy gradient algorithm to achieve collaborative coverage of UAV swarms.
[0004] However, existing technologies still have certain limitations. First, current research focuses primarily on efficiency metrics such as latency and energy consumption. While some methods consider trajectory planning or offloading ratio optimization, they often neglect the collaborative relationships between drones, which can easily lead to tasks being concentrated on a few drones, causing local overload and resource waste. Second, when introducing fairness constraints, they typically only focus on one aspect of user-side service balancing or resource-side load balancing, rarely simultaneously characterizing user service fairness and drone load fairness, making it difficult to balance long-term service balancing and system efficiency. Finally, while multi-agent reinforcement learning can improve distributed decision-making capabilities, in multi-drone scenarios, it is still difficult to simultaneously model the collaborative propagation relationships between drones and the service interaction relationships between drones and users, resulting in limited effectiveness in collaborative offloading and resource allocation.
[0005] Based on this, the present invention proposes a multi-UAV collaborative resource allocation method based on a hierarchical graph attention network, which can jointly model UAV collaborative relationships, user task states, and dual fairness constraints, thereby achieving more effective collaborative perception, load transfer, and long-term service balancing in dynamic and complex scenarios. Summary of the Invention
[0006] The purpose of this invention is to provide a multi-UAV collaborative resource allocation method based on hierarchical graph attention networks, which enables multiple UAVs to achieve efficient collaboration under local observation conditions. In dynamic mobile user scenarios, it intelligently balances multiple coupled performance indicators such as task offloading, trajectory control, system latency, energy consumption, and service fairness, thereby improving the overall performance and long-term service balance of multi-UAV assisted mobile edge computing systems.
[0007] To achieve the above objectives, the technical solution adopted by this invention is: a multi-UAV collaborative resource allocation method based on a hierarchical graph attention network, comprising the following steps: A multi-UAV cooperative resource allocation method based on hierarchical graph attention networks includes the following steps: S1. Scenario Construction: Construct a multi-drone assisted mobile edge computing scenario, which includes multiple drones equipped with edge computing servers and multiple mobile user devices; S2. Process Modeling: The resource allocation problem in the scenario is modeled as a multi-agent partially observable Markov decision process, and the local observation vector, mixed action space, joint reward function, optimization objective function and their constraints are defined. S3. Collaborative Feature Representation: Based on the user equipment movement and task generation, and combined with the UAV coverage relationship, task access is completed, and a collaborative observation feature is constructed for each UAV, which includes its own collaborative observation feature, neighbor collaborative observation feature, and covered user collaborative observation feature. S4. Joint Action Decision: The collaborative observation features are encoded using a hierarchical graph attention network to output a joint action that includes movement control, cooperative target selection, and unloading ratio control. S5. Strategy Training and Update: Each UAV completes task unloading and collaborative execution according to the joint action, and updates the strategy parameters using a multi-agent reinforcement learning approach with centralized training and decentralized execution to obtain the final collaborative task unloading and resource allocation strategy.
[0008] Furthermore, the specific steps for scene construction in step S1 include: Let the set of drones be The user equipment set is ;No. A drone in a time slot The position is represented as , No. Individual users in time slots The position is represented as ; Drones establish neighborly cooperation relationships based on their perception range for task forwarding and collaborative processing; The drone establishes a service relationship with the user based on the coverage radius, and accesses the drone to perform task unloading through the task unloading link.
[0009] Furthermore, the local observation vector is composed of its own state, the state of neighboring UAVs, and the state of covered users, and is represented as follows: In the formula, Indicates the first Observation of the drone's own status. This indicates the status observation of neighboring drones of this drone. Indicates the first The drones cover user status observation; among them, the status of the drones themselves includes location information, load information and fair pressure information, the status of neighboring drones includes relative location information, load information and fair pressure information, and the status of the covered users includes relative location information, task size information, computational complexity information, waiting time information and underservice flag information. The hybrid action space can be represented as the first... A drone in a time slot Actions: In the formula, Indicates the first A drone in a time slot The movement action, Indicates the first The collaborative target drone of the drone, Indicates the first The percentage of unloaded tasks for drones; The joint reward function is weighted by user service fairness, drone load fairness, normalized total system latency, and normalized total system energy consumption, and is expressed as follows: In the formula, These are non-negative weighting coefficients. and These represent the normalized total system delay and total system energy consumption, respectively.
[0010] Furthermore, the expression for the optimization objective function is: In the formula, For the joint strategy to be designed, For the expected operation, This represents the total number of time slots. As a discount factor, For time slots Fairness of user services For time slots Fairness of drone payload; The expression for the constraint condition is: In the formula, and They represent the first A drone in a time slot plane coordinates, and These represent the length and width of the task area, respectively. Indicates drone With drones In the time slot Euclidean distance, This represents the minimum safe distance allowed between any two drones. Indicates the first A drone in a time slot Task uninstallation ratio Indicates the first A drone in a time slot Selected cooperative target drone, Indicates the first A drone in a time slot A collection of neighboring drones, Indicates the drone's sensing range. Indicates the first A drone in a time slot The user set within the coverage area This indicates the ground coverage radius of the drone. Indicates user With drones The horizontal distance between them.
[0011] Furthermore, the fairness of user service is jointly determined by the user's cumulative service frequency, long-term service coverage, and underservice status, and its expression is: In the formula, Indicates user Deadline slot The cumulative number of services, This indicates its long-term service coverage. This indicates that it lacks a service tag. and Let represent the mean and standard deviation of the cumulative number of services, respectively. This represents the fairness threshold coefficient. This represents the Jain Fairness Index.
[0012] Furthermore, the drone load fairness is calculated based on the payload of each drone, and its expression is: In the formula, Indicates the first The original payload of the drone Indicates the first The percentage of underserved users within the coverage area of the drone. Indicates the first The effective payload of the drone This represents the Jain Fairness Index.
[0013] Furthermore, the expression for the collaborative observation feature vector is: In the formula, Indicates the first Observation of the drone's own status. This indicates the status observation of neighboring drones of this drone. Indicates the first The system monitors the status of all covered drones. The drones' status includes location information, load information, and fair pressure information. The status of neighboring drones includes relative location information, load information, and fair pressure information. The status of the covered users includes relative location information, task size information, computational complexity information, waiting time information, and underservice flag information.
[0014] Furthermore, step S3 specifically includes: The local observation vectors of each UAV are mapped to low-dimensional core state feature vectors through its own state encoder, as expressed by: In the formula, Encoding function for its own state. These are the learnable parameters of the encoder; A neighbor topology graph is dynamically constructed based on the real-time physical location of drones, and a weighted aggregation of neighbor drone information is obtained through a two-layer graph attention network to obtain a neighbor cooperation feature vector. In the formula, For the first A drone in a time slot A collection of neighboring drones, For graph attention aggregation function; Regarding the first For user devices within the coverage area of the drone, the relative location of the covered users, task input size, computational complexity, waiting time, and underservice flags are encoded to obtain user-side feature vectors: In the formula, For user state encoding function, These are the learnable parameters of the encoder; Using neighbor collaboration features as query features and overriding user features as key and value features, a cross-attention module adaptively filters key user-side information to obtain a user context feature vector: In the formula, Represents the cross-attention fusion function; The system concatenates its own state features, neighbor cooperation features, and user context features, and outputs the final collaborative representation feature vector through a feature fusion network. In the formula, Indicates the first The final collaborative feature vector of the drones.
[0015] Furthermore, in step S5, a multi-agent twin-delay deep deterministic policy gradient algorithm is used to update the policy parameters. The algorithm process includes: updating the dual critic network based on the experience replay pool, calculating the temporal difference target value according to the target critic network, updating the Actor network at a lower update frequency than the critic network, and separating the actions of other drones from the computation graph when updating the current drone policy, so as to reduce gradient coupling in the multi-agent joint training process.
[0016] The beneficial effects of the above technical solution are as follows: 1. This invention achieves global performance balance in multi-objective optimization scenarios by simultaneously introducing user service fairness and drone load fairness as optimization objectives in multi-drone assisted mobile edge computing scenarios. Specifically, it designs a user fairness measurement mechanism that includes cumulative service counts, long-term service coverage, and underservice flags, as well as a drone load fairness measurement mechanism based on effective load. This integrates drone trajectory control, collaborative target selection, and task offloading ratio into a unified decision-making framework, and achieves dynamic trade-offs through a multi-objective reward function. This ensures that the system maintains task collaborative processing efficiency while balancing long-term service balance and resource utilization efficiency, avoiding local overload and service imbalance caused by long-term task concentration on a few drones.
[0017] 2. This invention improves the efficiency and intelligence of distributed collaboration through a hierarchical graph attention mechanism. Specifically, the UAV agent dynamically filters key collaborative information based on its own state, neighbor states, and covered user states through graph attention branches, assigning different importance weights to different neighbors, thereby effectively overcoming the "information silo" limitation under local observation. At the same time, user features are extracted through user-side cross-attention branches, enhancing the UAV's ability to perceive the task pressure of covered users. While strengthening collaborative perception among UAVs, it also takes into account the differences in user-side tasks, significantly improving the decision-making efficiency of multi-UAV collaborative task offloading and resource allocation.
[0018] 3. This application employs a Siamese-delayed deep deterministic policy gradient algorithm for policy optimization, ensuring the stability and robustness of the algorithm training. Specifically, the multi-agent Siamese-delayed deep deterministic policy gradient algorithm constructs a temporal difference objective by taking the smaller value through a dual-critic network to suppress overestimation of the value function. It uses a delayed policy update mechanism to avoid training oscillations and separates other UAV actions from the computational graph to reduce gradient coupling, thus ensuring the convergence and stability of the algorithm training. Simultaneously, separating other UAV actions from the computational graph during the shared policy update process reduces gradient coupling in multi-agent joint training, allowing each UAV to retain its adaptability to local environmental differences while sharing a basic collaborative policy, thereby learning a more robust and better-performing joint policy. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the multi-UAV assisted mobile edge computing scenario model architecture in an embodiment of the present invention; Figure 2 This is a framework diagram of the HGA-MATD3 algorithm in an embodiment of the present invention; Figure 3 This is a comparison chart of the average reward convergence curves of different algorithms in the embodiments of the present invention; Figure 4This is a comparison chart of the average delay curves of different algorithms in the embodiments of the present invention; Figure 5 This is a comparison chart of the average energy consumption curves of different algorithms in the embodiments of the present invention; Figure 6 This is a comparison chart of the average fairness curves of different algorithms in the embodiments of the present invention; Figure 7 This is a comparison chart of the latency of different algorithms as the number of drones changes in an embodiment of the present invention; Figure 8 This is a comparison chart of energy consumption of different algorithms with varying numbers of drones in embodiments of the present invention; Figure 9 This is a fairness comparison chart of different algorithms in this invention as the number of drones changes; Figure 10 This is a comparison chart of latency for different algorithms as the number of users changes in an embodiment of the present invention; Figure 11 This is a comparison chart of energy consumption of different algorithms as the number of users changes in this embodiment of the invention; Figure 12 This is a fairness comparison chart of different algorithms as the number of users changes in an embodiment of the present invention. Detailed Implementation
[0020] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0021] It should be noted that, unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.
[0022] Before providing a detailed description of the embodiments of the present invention, the following explanations are given for some of the terms used in the embodiments: Hierarchical Graph Attention Network (HGA): A multi-layer feature extraction structure that combines graph attention mechanisms with user-side cross-attention mechanisms. Its core features are: firstly, it utilizes graph attention mechanisms to adaptively weight and aggregate neighbor nodes in the UAV's neighborhood topology to strengthen the modeling of collaborative relationships between UAVs; secondly, it uses cross-attention mechanisms to filter coverage of user task states to extract user context information that is more critical to the current unloading decision.
[0023] Local observation refers to the subset of information about the state of a single UAV, its neighboring UAVs, and the overlying users that it can directly perceive at a specific decision-making moment. In partially observable environments, local observation is insufficient to fully describe the global state of the entire system.
[0024] Cooperative Feature Vector: This refers to a low-dimensional feature representation generated by jointly encoding and fusing the local state information of a single UAV, the state information of neighboring UAVs, and the state information of the overlying user tasks through the hierarchical graph attention network described in this invention. It is used to characterize the cooperative environment and service pressure of the UAV in the current time slot.
[0025] Centralized Training, Decentralized Execution (CTDE): A multi-agent reinforcement learning training paradigm. During the training phase, a centralized server is used to acquire joint observation and action information to improve training stability; during the execution phase, each UAV independently outputs actions based solely on local observations to meet the needs of decentralized deployment and real-time decision-making.
[0026] User-Service Fairness: A metric used to measure whether different users have an equal opportunity to receive services during long-term task services. It takes into account the cumulative number of service visits, long-term service coverage, and underservice status.
[0027] UAV Load Fairness: A metric used to measure whether the task load distribution is balanced among different UAVs, taking into account both the original load of the UAVs and the proportion of underserved users within the coverage area.
[0028] To address the issues of local overload, service imbalance, and limited effectiveness of collaborative offloading and resource allocation in existing multi-UAV assisted mobile edge computing systems, this invention proposes a multi-UAV collaborative resource allocation method based on a hierarchical graph attention network, comprising the following steps: S1. Scenario Construction: Construct a multi-drone assisted mobile edge computing scenario, which includes multiple drones equipped with edge computing servers and multiple mobile user devices; S2. Process Modeling: The resource allocation problem in the scenario is modeled as a multi-agent partially observable Markov decision process, and the local observation vector, mixed action space, joint reward function, optimization objective function and their constraints are defined. S3. Collaborative Feature Representation: Based on the user equipment movement and task generation, and combined with the UAV coverage relationship, task access is completed, and a collaborative observation feature is constructed for each UAV, which includes its own collaborative observation feature, neighbor collaborative observation feature, and covered user collaborative observation feature. S4. Joint Action Decision: The collaborative observation features are encoded using a hierarchical graph attention network to output a joint action that includes movement control, cooperative target selection, and unloading ratio control. S5. Strategy Training and Update: Each UAV completes task unloading and collaborative execution according to the joint action, and updates the strategy parameters using a multi-agent reinforcement learning approach with centralized training and decentralized execution to obtain the final collaborative task unloading and resource allocation strategy.
[0029] The steps of this invention will be described in detail below: S1, Scene Construction like Figure 1 As shown, the multi-UAV assisted mobile edge computing scenario in this embodiment includes multiple unmanned aerial vehicles (UAVs) and multiple user equipment (UEs). Each UAV carries an edge computing server, providing wireless access, task offloading, and collaborative computing services to the UEs at a fixed flight altitude. The multiple UEs are randomly distributed within a two-dimensional task area and continuously move according to a predetermined mobility model, while periodically generating computing tasks in each discrete time slot.
[0030] Let the set of drones be The user equipment set is . No. A drone in a time slot The position is represented as , No. Individual users in time slots The position is represented as .
[0031] Unmanned aerial vehicles (UAVs) establish cooperative links based on their sensing range, forming neighborly cooperative relationships. This is possible if the distance between two UAVs does not exceed their sensing radius. If the two are considered to form a neighbor connection within that time slot, they can perform task forwarding and collaborative processing.
[0032] Based on coverage radius, the relationship between drones and users Establish a service relationship. If the distance between the user equipment and the drone does not exceed the coverage radius. If the user equipment is located within the coverage area of the UAV, it is assumed that the user equipment is connected to the UAV via the task offloading link to perform task offloading.
[0033] S2, Process Modeling To transform the multi-UAV collaborative task offloading problem into a form suitable for reinforcement learning, this invention models the resource allocation problem in a multi-UAV assisted mobile edge computing system as a multi-agent partially observable Markov decision process, and defines the local observation vector, hybrid action space, joint reward function, optimization objective function and its constraints.
[0034] Each drone acts as an independent intelligent agent, making action decisions based solely on its own local observations. The combined actions of all drones, however, collectively determine the evolution of the environmental state and the system's reward.
[0035] The local observation vector is composed of its own state, the states of neighboring UAVs, and the states of covered users, and is represented as follows: In the formula, Indicates the first Observation of the drone's own status. This indicates the status observation of neighboring drones of this drone. Indicates the first The system monitors the status of all covered drones. The drones' status includes location information, load information, and fair pressure information. The status of neighboring drones includes relative location information, load information, and fair pressure information. The status of the covered users includes relative location information, task size information, computational complexity information, waiting time information, and underservice flag information.
[0036] The hybrid action space can be represented as the first... A drone in a time slot Actions: In the formula, Indicates the first A drone in a time slot The movement action, Indicates the first The collaborative target drone of the drone, Indicates the first The percentage of unloaded tasks for drones.
[0037] In this embodiment, the joint optimization objective of the multi-UAV assisted mobile edge computing system simultaneously considers user service fairness, UAV load fairness, total system latency, and total system energy consumption. Its joint reward function can be expressed as: In the formula, These are non-negative weighting coefficients. and These represent the normalized total system delay and total system energy consumption, respectively.
[0038] The expression for its optimization objective function is: In the formula, For the joint strategy to be designed, For the expected operation, This represents the total number of time slots. As a discount factor, For time slots Fairness of user services For time slots Fairness of drone payload.
[0039] The expression for the constraint is: In the formula, and They represent the first A drone in a time slot plane coordinates, and These represent the length and width of the task area, respectively. Indicates drone With drones In the time slot Euclidean distance, This represents the minimum safe distance allowed between any two drones. Indicates the first A drone in a time slot Task uninstallation ratio Indicates the first A drone in a time slot Selected cooperative target drone, Indicates the first A drone in a time slot A collection of neighboring drones, Indicates the drone's sensing range. Indicates the first A drone in a time slot The user set within the coverage area This indicates the ground coverage radius of the drone. Indicates user With drones The horizontal distance between them.
[0040] Constraints C1 and C2 restrict each drone to always be within the mission area; constraint C3 restricts the distance between any two drones to no less than the minimum safe distance to avoid collisions; constraint C4 restricts the mission offload ratio to between 0 and 1; constraint C5 restricts the cooperative target to only the drone itself or its neighboring drones; constraint C6 defines the drone neighbor set, i.e. drones whose distance does not exceed the perception range; constraint C7 defines the user set within the drone's coverage area, i.e. user devices whose distance does not exceed the coverage radius.
[0041] S3, Collaborative Feature Representation To enhance the collaborative perception capabilities of UAVs under local observation conditions, this invention employs a hierarchical graph attention network to encode the local observations. Specifically, based on user device movement and task generation, and combined with UAV coverage relationships, task access is completed, and collaborative observation features of each UAV, its neighbors, and covered users are constructed. The overall feature extraction process is as follows: Figure 2 As shown on the left.
[0042] First, the UAV's own observations are mapped into core state feature vectors using its own state encoder: ,in, This represents the state encoding function. These are its learnable parameters.
[0043] Subsequently, a neighbor topology graph is dynamically constructed based on the real-time physical location of the drones, and a weighted aggregation of neighbor drone information is obtained through a two-layer graph attention network to obtain a neighbor cooperation feature vector: In the formula, For the first A drone in a time slot A collection of neighboring drones, This is the graph attention aggregation function.
[0044] Furthermore, the user task information is encoded to obtain the user-side feature vector. ,in, For user state encoding function, These are the learnable parameters of the encoder.
[0045] Then, using the neighbor collaboration feature vector as the query feature and the user-side feature vector as the key and value features, the user context feature vector is extracted through a cross-attention module. ,in, This represents the cross-attention fusion function.
[0046] Finally, the user's own state feature vector, neighbor cooperation feature vector, and user context feature vector are fused to generate the final cooperative representation feature vector. ,in, Indicates the first The final collaborative feature vector of the drones.
[0047] S4, Joint Action Decision The collaborative observation features are encoded using a hierarchical graph attention network, outputting a joint action that includes movement control, cooperative target selection, and unloading ratio control; the action decoding process is as follows: Figure 2 As shown in the middle.
[0048] Specifically, after obtaining the collaborative representation feature vector, the decoder outputs the joint actions of the UAV. These joint actions include UAV motion control variables, cooperative target selection variables, and task offloading ratios. Specifically, the UAV motion control variables determine the UAV's direction and intensity of movement; the cooperative target selection variables select cooperative execution nodes from among the current UAV and its neighboring UAVs; and the task offloading ratios determine the proportion of tasks currently undertaken by the UAV that require further collaborative forwarding.
[0049] During task execution, the user equipment first associates with the corresponding drone based on the coverage relationship between the user equipment and the drone. When a user equipment is covered by multiple drones, a unique service drone is determined according to the principle of minimizing the distance between the user equipment and the drone. Subsequently, the service drone distributes the assigned task between local execution and execution by cooperating drones based on the offload ratio.
[0050] When the target of cooperation is itself, the task is processed locally by the current drone; when the target is a neighboring drone, the task is forwarded by the current drone to the neighboring drone for collaborative processing. The total task latency can be expressed as: In the formula, This indicates the latency when the task is processed locally on the current drone. This indicates the latency when the task is processed by the collaborative drone.
[0051] The action decision-making process satisfies the following constraints: the UAV's position is always within the mission area; the distance between any two UAVs is not less than the minimum safe distance. Task unloading ratio Always limited to the range Within; cooperative targets can only be selected from the current drone itself or its neighboring drones.
[0052] S5, Strategy Training Update This invention employs a multi-agent twin-delay deep deterministic policy gradient algorithm for training, with centralized training and distributed execution. For example... Figure 2 As shown, the HGA-MATD3 algorithm proposed in this invention includes a hierarchical graph attention collaborative feature extraction module on the left, a decentralized Actor decision module in the middle, and a centralized dual critic training module on the right.
[0053] Figure 2 The left side shows the hierarchical graph attention collaborative feature extraction module. This module corresponds to the technical solution in the claim of "constructing collaborative observation features for each UAV, its neighbors, and the covered users." Specifically, firstly, the UAV's own state, the states of neighboring UAVs, and the states of covered users are input into the feature encoding unit to obtain core state features. Subsequently, in the graph attention branch, a dynamic graph structure is constructed based on the spatial adjacency relationship between UAVs, and the graph attention layer is used to adaptively weight and aggregate the features of neighboring UAVs, thereby extracting neighborhood collaborative features that reflect the cooperative relationship between UAVs. Simultaneously, in the user attention branch, key, query, and value mapping is performed on the task state of the covered users, and the user context features most relevant to the current UAV decision are extracted through a cross-attention aggregation mechanism. Finally, the neighborhood collaborative features and user context features are fused and projected to generate a collaborative feature vector for subsequent decision-making. Through this structure, the present invention can simultaneously model the cooperative dependency relationship between UAVs and the differences in user task requirements, improving the state representation capability under local observation conditions.
[0054] Figure 2 The decentralized Actor decision-making module is shown in the middle. This module corresponds to the technical solution in the claim that "encodes the collaborative observation features using a hierarchical graph attention network to output a joint action including movement control, cooperative target selection, and task unloading ratio control." Specifically, the Actor network first receives the local observation vector of the current UAV and extracts local basic features through an encoder; then, the local basic features are input into the left-hand hierarchical graph attention collaborative feature extraction module to obtain enhanced features containing neighbor cooperation relationships and task-related user information; next, the local basic features and enhanced features are concatenated and fused, and input into the decoder to output the original action vector; the original action vector is further mapped to UAV movement control quantity, cooperative target selection result, and task unloading ratio. Among them, the movement control quantity is used to adjust the position change of the UAV within the task area, the cooperative target selection result is used to determine whether the current task is handled by the local UAV or forwarded to a neighboring UAV for collaborative handling, and the task unloading ratio is used to determine the allocation relationship between the local execution part and the cooperative forwarding part. Thus, each UAV does not need to obtain the global state during the execution phase and can complete decentralized real-time decision-making based only on local observations.
[0055] Figure 2 The right side shows the centralized dual-commentator training module. This module corresponds to the technical solution in the claim that "policy parameters are updated using a multi-agent reinforcement learning approach with centralized training and decentralized execution." During the training phase, the states, actions, rewards, and next-moment states generated by environmental interactions constitute state transition samples and are stored in the experience replay pool. The centralized server samples a small batch of samples from the experience replay pool, inputs the next-moment observation into the target Actor network to obtain the next-moment action, and then inputs this action into the two target commentator networks. The smaller value between their outputs is used to construct a temporal difference target to suppress overestimation of the value function and improve training stability.
[0056] To verify the effectiveness of the HGA-MATD3 algorithm of this invention, a simulation environment for multi-UAV assisted mobile edge computing was constructed, and the method described in this invention was compared with various benchmark algorithms. The results are as follows: Figure 3-12 As shown. The comparison algorithms used include: MATD3, GAT-TD3, Centralized SAC, and randomized strategy methods.
[0057] Figure 3 The convergence curves of the average reward for different algorithms are compared. It can be seen that the HGA-MATD3 algorithm proposed in this invention can quickly improve the average reward in the early stages of training and maintain a high and stable convergence level in the later stages, outperforming Attention-MATD3, MATD3, MADDPG, Centralized SAC, GAT-TD3, and randomized policy methods overall. This is because this invention simultaneously models the cooperative relationships between drones and covers the user's task context information through a hierarchical graph attention network, enabling Actors to extract more discriminative cooperative features under local observation conditions, thereby generating joint actions more effectively.
[0058] Figure 4 The comparison results of average latency for different algorithms are presented. It can be seen that the HGA-MATD3 algorithm proposed in this invention maintains a low average latency throughout the entire training process and achieves optimal performance after convergence, indicating that this invention can more effectively reduce task completion time in multi-UAV assisted mobile edge computing scenarios. This is because this invention jointly introduces motion control, cooperative target selection, and task offloading ratio control in the action space, enabling UAVs to flexibly select better service nodes and forwarding paths based on local cooperative characteristics, thereby reducing latency overhead caused by queuing, repeated forwarding, and invalid offloading.
[0059] Figure 5The comparison results of average energy consumption of different algorithms are presented. It can be seen that the HGA-MATD3 algorithm proposed in this invention can continuously reduce system energy consumption and ultimately achieve better energy consumption performance than other algorithms. The reason is that this invention, through collaborative feature extraction and joint action optimization, enables the UAV to meet mission service requirements while minimizing unnecessary long-distance migration, redundant calculations, and inefficient forwarding, thereby reducing overall communication and computing energy consumption.
[0060] Figure 6 The average fairness curves of different algorithms are compared. It can be seen that the method of this invention maintains a high level of long-term service fairness, indicating that dual fairness modeling can effectively improve the problems of user service opportunity imbalance and drone load imbalance.
[0061] Figure 7 The latency comparison results of different algorithms under varying drone numbers are presented. It can be seen that as the number of drones increases from 3 to 7, the overall latency of all algorithms shows a decreasing trend, indicating that the participation of more drones in collaboration can improve task sharing capabilities and reduce system queuing pressure. Further observation reveals that the HGA-MATD3 algorithm proposed in this invention does not show a particularly large gap with some benchmark algorithms in small-scale scenarios, but its advantages become increasingly apparent as the number of drones increases, especially in scenarios with 5 or more drones, where it can consistently maintain a lower latency level. This indicates that this invention can more fully utilize the collaborative gains brought by scale expansion and maintain stronger scheduling efficiency in medium-to-large-scale networks through hierarchical collaborative feature extraction and joint action decision-making. In contrast, while MATD3, MADDPG, and Centralized SAC also show some improvement with increasing scale, the improvement is relatively limited, indicating that they do not fully utilize the newly added collaborative resources.
[0062] Figure 8 The energy difference results of different algorithms under varying drone numbers are presented. It can be seen that as the number of drones increases, the overall energy difference of each algorithm decreases, indicating that the system can alleviate the load pressure on single nodes and thus reduce energy consumption after the scale of cooperation expands. From the curve trends, the HGA-MATD3 algorithm proposed in this invention remains optimal overall, and its gap with other algorithms is more stable in medium-scale scenarios, indicating that this invention can more rationally coordinate the cooperative relationship between drones during task allocation, reducing energy consumption caused by invalid forwarding and redundant computation. In contrast, although MATD3 and MADDPG can also reduce some energy consumption after the scale increases, their reduction process fluctuates more, indicating that they still have instability in the utilization of cooperative resources; the overall energy-saving effect of Centralized SAC is limited, reflecting the insufficient optimization efficiency of centralized high-dimensional decision-making in complex environments.
[0063] Figure 9 The fairness comparison results of different algorithms under varying drone numbers are presented. It can be seen that as the number of drones increases, the overall fairness of each algorithm fluctuates, and the fairness of some benchmark algorithms decreases with increasing scale. This indicates that as the system scales up, without effective coordination and fairness constraints, it is more prone to uneven task and load distribution. In contrast, the HGA-MATD3 algorithm proposed in this invention maintains a high level of fairness across different scales, and its advantage becomes more stable as the number of drones increases. This demonstrates that by introducing user service fairness and drone load fairness into the reward function and combining it with collaborative feature modeling, this invention can effectively suppress excessive task concentration and local node overload. In comparison, while MATD3 and Attention-MATD3 possess some learning ability, their characterization of fairness objectives is insufficient, making them more prone to imbalance degradation as the scale increases. The fairness improvements of MADDPG and Centralized SAC are relatively limited.
[0064] Figure 10 The results compare the latency differences of different algorithms under varying user numbers. It can be seen that as the number of users increases from 30 to 80, the absolute value of the latency difference for each algorithm gradually increases, indicating that task scheduling and resource contention become more pronounced with increased user load, and the system latency optimization space also increases accordingly. Among them, the HGA-MATD3 algorithm proposed in this invention maintains optimal performance under various user scales, and its advantages over Centralized SAC, MADDPG, and Random methods become increasingly apparent as the number of users increases, especially when the number of users is high, it can maintain a greater reduction in latency. Compared with MATD3, this invention has a stable advantage in low-to-medium load scenarios and can still maintain good latency control capabilities in high-load scenarios, indicating that this invention can more fully utilize cooperative features and joint action decision-making capabilities when the user scale expands, thereby improving the system's adaptability to large-scale user access scenarios. In contrast, although Centralized SAC can reduce latency to some extent, its reduction is significantly limited; MADDPG's improvement trend is weak, indicating its insufficient optimization capability in complex joint scheduling problems.
[0065] Figure 11The energy consumption comparison results of different algorithms under varying user numbers are presented. It can be seen that as the number of users increases, the overall energy consumption of each algorithm rises, indicating that a larger user scale brings heavier task processing pressure and communication overhead. During this process, the HGA-MATD3 algorithm proposed in this invention maintains a consistently low energy consumption level and exhibits relatively stable performance, demonstrating that the proposed method can effectively suppress energy consumption caused by invalid forwarding and additional computation when the user scale changes. From the curve changes, the energy-saving advantage of this invention compared to the Centralized SAC and Random methods is significant, and it maintains a stable low energy consumption state even after the number of users increases, indicating its strong energy robustness. Compared to MATD3, this invention is generally in the same excellent range, with small differences between the two under certain user scales, but the energy consumption fluctuation of this invention is smaller, indicating that hierarchical graph attention collaborative feature extraction can help the strategy to distribute load more stably. The overall energy consumption level of MADDPG is higher than that of this invention, and it fluctuates to some extent with changes in the number of users, reflecting its relatively insufficient energy efficiency control capability in complex scenarios. Due to the lack of an effective resource coordination mechanism, the energy consumption of the Random method increases most significantly with the number of users, demonstrating the obvious disadvantage of non-cooperative strategies in large-scale scenarios.
[0066] Figure 12 The fairness comparison results of different algorithms under varying user numbers are presented. It can be seen that as the number of users increases from 30 to 80, the overall fairness of each algorithm improves, indicating that with a larger user scale, the system serves a richer range of users, and the optimization effect of the fairness objective is more easily reflected. Among them, the HGA-MATD3 algorithm proposed in this invention maintains the highest fairness across all user scales and continues to show a stable upward trend even with an increasing number of users. This demonstrates that by explicitly introducing user service fairness and drone load fairness constraints into the reward function, this invention can continuously promote service opportunity balance and load balance. Compared to MATD3, the fairness advantage of this invention already exists in small-scale user scenarios and remains stable in medium-scale scenarios, without showing an unbounded expansion trend. Instead, it exhibits a "consistently leading and smoother improvement," indicating that the fairness improvement of this invention is more inclined towards stable enhancement rather than drastic fluctuations. Compared to MADDPG and Centralized SAC, this invention has a more significant fairness advantage across all user scales, indicating that collaborative perception and joint decision-making can effectively alleviate the problems of task concentration and local congestion. The Random method consistently exhibits the lowest fairness, indicating that without learning and coordination mechanisms, the system struggles to achieve long-term balanced service.
[0067] In summary, this invention achieves joint optimization of collaborative task offloading and resource allocation in multi-UAV-assisted mobile edge computing scenarios through a hierarchical graph attention mechanism, a dual-fairness reward design, and a multi-agent twin-delay deep deterministic policy gradient training framework. This can effectively improve system efficiency, long-term service balance, and algorithm robustness in dynamic and complex environments.
[0068] Finally, it should be noted that any parts of this invention not described in detail are prior art. Those skilled in the art will understand that the above descriptions are merely preferred embodiments of the invention and are not intended to limit the invention. Although the invention has been described in detail with reference to the foregoing examples, those skilled in the art can still modify the technical solutions described in the foregoing examples or make equivalent substitutions for some of the technical features. All modifications and equivalent substitutions made within the spirit and principles of the invention should be included within the scope of protection of the invention.
Claims
1. A multi-UAV cooperative resource allocation method based on hierarchical graph attention networks, characterized in that, Includes the following steps: S1. Scenario Construction: Construct a multi-drone assisted mobile edge computing scenario, which includes multiple drones equipped with edge computing servers and multiple mobile user devices; S2. Process Modeling: The resource allocation problem in the scenario is modeled as a multi-agent partially observable Markov decision process, and the local observation vector, mixed action space, joint reward function, optimization objective function and their constraints are defined. S3. Collaborative Feature Representation: Based on the user equipment movement and task generation, and combined with the UAV coverage relationship, task access is completed, and a collaborative observation feature is constructed for each UAV, which includes its own collaborative observation feature, neighbor collaborative observation feature, and covered user collaborative observation feature. S4. Joint Action Decision: The collaborative observation features are encoded using a hierarchical graph attention network to output a joint action that includes movement control, cooperative target selection, and unloading ratio control. S5. Strategy Training and Update: Each UAV completes task unloading and collaborative execution according to the joint action, and updates the strategy parameters using a multi-agent reinforcement learning approach with centralized training and decentralized execution to obtain the final collaborative task unloading and resource allocation strategy.
2. The multi-UAV cooperative resource allocation method based on a hierarchical graph attention network according to claim 1, characterized in that, The specific steps for scene construction in step S1 include: Let the set of drones be The user equipment set is ;No. A drone in a time slot The position is represented as , No. Individual users in time slots The position is represented as ; Drones establish neighborly cooperation relationships based on their perception range for task forwarding and collaborative processing; The drone establishes a service relationship with the user based on the coverage radius, and accesses the drone to perform task unloading through the task unloading link.
3. The multi-UAV cooperative resource allocation method based on a hierarchical graph attention network according to claim 1, characterized in that, The local observation vector is composed of its own state, the states of neighboring UAVs, and the states of covered users, and is represented as follows: In the formula, Indicates the first Observation of the drone's own status. This indicates the status observation of neighboring drones of this drone. Indicates the first The drones cover user status observation; among them, the status of the drones themselves includes location information, load information and fair pressure information, the status of neighboring drones includes relative location information, load information and fair pressure information, and the status of the covered users includes relative location information, task size information, computational complexity information, waiting time information and underservice flag information. The hybrid action space can be represented as the first... A drone in a time slot Actions: In the formula, Indicates the first A drone in a time slot The movement action, Indicates the first The collaborative target drone of the drone, Indicates the first The percentage of unloaded tasks for drones; The joint reward function is weighted by user service fairness, drone load fairness, normalized total system latency, and normalized total system energy consumption, and is expressed as follows: In the formula, These are non-negative weighting coefficients. and These represent the normalized total system delay and total system energy consumption, respectively.
4. The multi-UAV cooperative resource allocation method based on a hierarchical graph attention network according to claim 1, characterized in that, The expression for the optimization objective function is: In the formula, For the joint strategy to be designed, For the expected operation, This represents the total number of time slots. As a discount factor, For time slots Fairness of user services For time slots Fairness of drone payload; The expression for the constraint condition is: In the formula, and They represent the first A drone in a time slot plane coordinates, and These represent the length and width of the task area, respectively. Indicates drone With drones In the time slot Euclidean distance, This represents the minimum safe distance allowed between any two drones. Indicates the first A drone in a time slot Task uninstallation ratio Indicates the first A drone in a time slot Selected cooperative target drone, Indicates the first A drone in a time slot A collection of neighboring drones, Indicates the drone's sensing range. Indicates the first A drone in a time slot The user set within the coverage area This indicates the ground coverage radius of the drone. Indicates user With drones The horizontal distance between them.
5. A multi-UAV cooperative resource allocation method based on a hierarchical graph attention network according to claim 3, characterized in that, The fairness of user services is determined by the user's cumulative service frequency, long-term service coverage, and underservice status, and its expression is: In the formula, Indicates user Deadline slot The cumulative number of services, This indicates its long-term service coverage. This indicates that it lacks a service tag. and Let represent the mean and standard deviation of the cumulative number of services, respectively. This represents the fairness threshold coefficient. This represents the Jain Fairness Index.
6. A multi-UAV cooperative resource allocation method based on a hierarchical graph attention network according to claim 3, characterized in that, The drone load fairness is calculated based on the payload of each drone, and its expression is as follows: In the formula, Indicates the first The original payload of the drone Indicates the first The percentage of underserved users within the coverage area of the drone. Indicates the first The effective payload of the drone This represents the Jain Fairness Index.
7. A multi-UAV cooperative resource allocation method based on a hierarchical graph attention network according to claim 1, characterized in that, The expression for the collaborative observation feature vector is: In the formula, Indicates the first Observation of the drone's own status. This indicates the status observation of neighboring drones of this drone. Indicates the first The system monitors the status of all covered drones. The drones' status includes location information, load information, and fair pressure information. The status of neighboring drones includes relative location information, load information, and fair pressure information. The status of the covered users includes relative location information, task size information, computational complexity information, waiting time information, and underservice flag information.
8. A multi-UAV cooperative resource allocation method based on a hierarchical graph attention network according to claim 1, characterized in that, Step S3 is as follows: The local observation vectors of each UAV are mapped to low-dimensional core state feature vectors through its own state encoder, as expressed by: In the formula, Encoding function for its own state. These are the learnable parameters of the encoder; A neighbor topology graph is dynamically constructed based on the real-time physical location of drones, and a weighted aggregation of neighbor drone information is obtained through a two-layer graph attention network to obtain a neighbor cooperation feature vector. In the formula, For the first A drone in a time slot A collection of neighboring drones, For graph attention aggregation function; Regarding the first For user devices within the coverage area of the drone, the relative location of the covered users, task input size, computational complexity, waiting time, and underservice flags are encoded to obtain user-side feature vectors: In the formula, For user state encoding function, These are the learnable parameters of the encoder; Using neighbor collaboration features as query features and overriding user features as key and value features, a cross-attention module adaptively filters key user-side information to obtain a user context feature vector: In the formula, Represents the cross-attention fusion function; The system concatenates its own state features, neighbor collaboration features, and user context features, and outputs the final collaborative representation feature vector through a feature fusion network. In the formula, Indicates the first The final collaborative feature vector of the drones.
9. A multi-UAV cooperative resource allocation method based on a hierarchical graph attention network according to claim 1, characterized in that, In step S5, the multi-agent twin-delay deep deterministic policy gradient algorithm is used to update the policy parameters. The algorithm process includes: updating the dual critic network based on the experience replay pool, calculating the temporal difference target value according to the target critic network, updating the Actor network at a lower update frequency than the critic network, and separating the actions of other drones from the computation graph when updating the current drone policy, so as to reduce gradient coupling in the multi-agent joint training process.