Edge computing resource allocation method based on unmanned aerial vehicle assistance and related equipment
Through drone-assisted edge computing, node status data is acquired and aggregated, discrete encoding and multi-agent reinforcement learning are carried out, which solves the problems of high delay and low efficiency of computing tasks in remote areas, and achieves the improvement of throughput and fairness.
Patent Information
- Application Number
- CN202510155444.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-07-04
AI Technical Summary
In remote areas or after natural disasters, there are few fixed base stations near user terminal equipment, which makes it difficult to process computing-intensive tasks in a timely manner, resulting in problems of large delays and low efficiency.
Through drone-assisted edge computing, the status data of drone nodes and adjacent nodes are obtained, feature extraction and aggregation are performed, and discrete coding and multi-agent reinforcement learning are combined to optimize resource allocation and path selection, adapt to dynamically change complex topology, and improve the throughput and fairness of user terminal equipment.
Effectively integrate multi-source data, optimize resource allocation and path selection, reduce task processing delays, improve the throughput and fairness of user terminal equipment, and adapt to dynamic changes in complex environments.
Smart Images

Figure CN120256084A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of resource allocation, and in particular, to a method for allocating edge computing resources assisted by an unmanned aerial vehicle and related devices. Background Art
[0002] Generally, the computing tasks generated by user equipment can be calculated by offloading them to a fixed edge server. However, in some special areas, some user terminal devices are set in areas where there are no people or the environment is relatively harsh, and the user terminal devices still generate a lot of compute-intensive tasks. In this scenario, since the computing servers on the ground are usually deployed in fixed base stations (BSs) and are far from the user equipment that generates computing tasks, offloading the computing tasks to the fixed base station (BS) will cause a large delay and cannot meet the requirements of the delay-sensitive tasks generated by the user terminal equipment. At the same time, in some remote areas or after natural disasters, etc., there are few fixed base stations (BSs) near the user terminal equipment, and even no fixed base stations are deployed. With the construction and deployment of the network, the compute-intensive tasks generated by the user equipment in these areas are difficult to be processed in a timely manner, resulting in problems such as large delay and low efficiency. Summary of the Invention
[0003] The present disclosure proposes a method for allocating edge computing resources assisted by an unmanned aerial vehicle and related devices to solve technical problems such as large task processing delay and low efficiency to a certain extent.
[0004] In a first aspect of the present disclosure, there is provided a method for allocating edge computing resources assisted by an unmanned aerial vehicle, including:
[0005] Obtaining first observation data of a first unmanned aerial vehicle node, where the first observation data includes first status data of the first unmanned aerial vehicle node and second status data of adjacent nodes;
[0006] Performing feature extraction and feature aggregation based on the first status data and the second status data to obtain an aggregated feature;
[0007] Performing discretization encoding on the aggregated feature to obtain a first discretized feature, and receiving a second discretized feature from a second unmanned aerial vehicle node;
[0008] Combining the first discretized feature and the second discretized feature to obtain an updated feature;
[0009] Determining a resource allocation strategy for the first unmanned aerial vehicle node based on the updated feature.
[0010] In a second aspect of the present disclosure, there is provided a device for allocating edge computing resources assisted by an unmanned aerial vehicle, including:
[0011] An acquisition module, configured to acquire first observation data of a first UAV node, where the first observation data includes first status data of the first UAV node and second status data of adjacent nodes;
[0012] A feature aggregation module, configured to perform feature extraction and feature aggregation based on the first status data and the second status data to obtain an aggregated feature;
[0013] An encoding module, configured to perform discretized encoding on the aggregated feature to obtain a first discretized feature;
[0014] The acquisition module is further configured to receive a second discretized feature from a second UAV node;
[0015] A feature update module, configured to combine the first discretized feature and the second discretized feature to obtain an updated feature;
[0016] A resource allocation module, configured to determine a resource allocation strategy for the first UAV node based on the updated feature.
[0017] In a third aspect of the present disclosure, an electronic device is provided, including one or more processors, a memory; and one or more programs, where the one or more programs are stored in the memory and executed by the one or more processors, and the programs include instructions for executing the method according to the first aspect.
[0018] In a fourth aspect of the present disclosure, a non-volatile computer-readable storage medium containing a computer program is provided. When the computer program is executed by one or more processors, the processors are caused to execute the method according to the first aspect.
[0019] In a fifth aspect of the present disclosure, a computer program product is provided, including computer program instructions. When the computer program instructions are executed on a computer, the computer is caused to execute the method according to the first aspect.
[0020] As can be seen from the above, a method and related devices for UAV-assisted edge computing resource allocation provided by the present disclosure effectively fuse multi-source data by extracting features of UAVs, ground base stations, and user terminal devices, and aggregating discrete communications between multiple UAVs, so as to jointly optimize resource allocation and path selection in a complex environment, adapt to dynamic and complex topologies, make appropriate task scheduling decisions, and improve the throughput and fairness of user terminal devices. Description of the Drawings
[0021] To more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings described below are only the embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0022] Figure 1 It is a schematic diagram of the architecture of the edge computing resource allocation method based on UAV assistance according to an embodiment of the present disclosure.
[0023] Figure 2 It is a schematic diagram of the hardware structure of an exemplary electronic device according to an embodiment of the present disclosure.
[0024] Figure 3 It is a schematic flowchart of the edge computing resource allocation method based on UAV assistance according to an embodiment of the present disclosure.
[0025] Figure 4 It is a schematic diagram of the principle of the discrete communication model according to an embodiment of the present disclosure.
[0026] Figure 5 It is a schematic diagram of the hybrid network according to an embodiment of the present disclosure.
[0027] Figure 6 It is a schematic diagram of the edge computing resource allocation device based on UAV assistance according to an embodiment of the present disclosure. Detailed implementation manners
[0028] To make the purpose, technical solutions, and advantages of the present disclosure clearer and more understandable, the following further elaborates on the present disclosure in detail with reference to specific embodiments and the accompanying drawings.
[0029] It should be noted that unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the ordinary meaning understood by those of ordinary skill in the field to which the present disclosure belongs. The "first", "second", and similar terms used in the embodiments of the present disclosure do not indicate any order, quantity, or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before the term cover the elements or objects listed after the term and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0030] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to users and the authorization of users should be obtained through appropriate means in accordance with relevant laws and regulations.
[0031] For example, when receiving an active request from a user, a prompt message is sent to the user to clearly prompt that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operations of the technical solutions of the present disclosure according to the prompt message.
[0032] It is understandable that the above process of notifying and obtaining user authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0033] Figure 1 FIG. shows a schematic diagram of the architecture of the drone-assisted edge computing resource allocation method according to an embodiment of the present disclosure. Refer to Figure 1 , the architecture 100 of the drone-assisted edge computing resource allocation method may include a server 110, a terminal 120, and a network 130 that provides a communication link. The server 110 and the terminal 120 can be connected through the wired or wireless network 130. Among them, the server 110 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, security services, and CDN.
[0034] The terminal 120 can be implemented by hardware or software. For example, when the terminal 120 is implemented by hardware, it can be various electronic devices with a display screen and supporting page display, including but not limited to smart phones, tablet computers, e-book readers, laptop portable computers, and desktop computers, etc. When the terminal 120 device is implemented by software, it can be installed in the above-listed electronic devices; it can be implemented as multiple software or software modules (for example, software or software modules for providing distributed services), or it can be implemented as a single software or software module, which is not specifically limited herein.
[0035] It should be noted that the drone-assisted edge computing resource allocation method provided by the embodiments of the present application can be executed by the terminal 120 or by the server 110. It should be understood that Figure 1 the numbers of terminals, networks, and servers in are only illustrative and are not intended to limit them. According to the implementation requirements, there can be any number of terminals, networks, and servers.
[0036] Figure 2 shows a schematic diagram of the hardware structure of the exemplary electronic device 200 provided by the embodiments of the present disclosure. As Figure 2 shown, the electronic device 200 may include: a processor 202, a memory 204, a network module 206, a peripheral interface 208, and a bus 210. Among them, the processor 202, the memory 204, the network module 206, and the peripheral interface 208 are communicatively connected to each other inside the electronic device 200 through the bus 210.
[0037] The processor 202 may be a central processing unit (CPU), a neural network processor (NPU), a microcontroller (MCU), a programmable logic device, a digital signal processor (DSP), an application specific integrated circuit (ASIC), or one or more integrated circuits. The processor 202 may be used to execute functions related to the technologies described in the present disclosure. In some embodiments, the processor 202 may further include multiple processors integrated as a single logic component. For example, as Figure 2 shown, the processor 202 may include multiple processors 202a, 202b, and 202c.
[0038] The memory 204 may be configured to store data (e.g., instructions, computer code, etc.). As Figure 2 shown, the data stored in the memory 204 may include program instructions (e.g., program instructions for implementing the drone-assisted edge computing resource allocation method according to the embodiments of the present disclosure) and data to be processed (e.g., the memory may store configuration files of other modules, etc.). The processor 202 may also access the program instructions and data stored in the memory 204, and execute the program instructions to operate on the data to be processed. The memory 204 may include a volatile storage device or a non-volatile storage device. In some embodiments, the memory 204 may include a random access memory (RAM), a read-only memory (ROM), an optical disc, a magnetic disk, a hard disk, a solid state drive (SSD), a flash memory, a memory stick, etc.
[0039] The network module 206 can be configured to provide communication with other external devices to the electronic device 200 via a network. The network can be any wired or wireless network capable of transmitting and receiving data. For example, the network can be a wired network, a local wireless network (such as Bluetooth, WiFi, Near Field Communication (NFC), etc.), a cellular network, the Internet, or a combination of the above. It can be understood that the type of the network is not limited to the above specific examples. In some embodiments, the network module 206 can include any combination of any number of network interface controllers (NICs), radio frequency modules, transceivers, modems, routers, gateways, adapters, cellular network chips, etc.
[0040] The peripheral interface 208 can be configured to connect the electronic device 200 to one or more peripheral devices to achieve information input and output. For example, the peripheral devices can include input devices such as a keyboard, a mouse, a touchpad, a touch screen, a microphone, various sensors, etc. and output devices such as a display, a speaker, a vibrator, an indicator light, etc.
[0041] The bus 210 can be configured to transmit information between various components of the electronic device 200 (such as the processor 202, the memory 204, the network module 206, and the peripheral interface 208), such as an internal bus (such as a processor - memory bus), an external bus (USB port, PCI - E bus), etc.
[0042] It should be noted that although the architecture of the above - mentioned electronic device 200 only shows the processor 202, the memory 204, the network module 206, the peripheral interface 208, and the bus 210, in the specific implementation process, the architecture of the electronic device 200 can also include other components necessary for normal execution. In addition, those skilled in the art can understand that the architecture of the above - mentioned electronic device 200 can also only include the components necessary to implement the solution of the embodiments of the present disclosure, and does not necessarily include all the components shown in the figure.
[0043] Edge computing aims to solve the problems of data processing, storage, and transmission encountered in cloud computing. It can be deployed distributively using edge gateways to collect, process, perform protocol conversion, and analyze data locally, greatly alleviating the pressure on network transmission and data centers. At the same time, on the edge side, it ensures low latency for business terminals, faster processing speed, strong service response performance, and good real-time performance. Edge computing significantly improves the network service quality of edge user devices in the region, bringing a better network experience to nearby users. Generally, the computing tasks generated by user devices can be offloaded to fixed edge servers for computing. However, in some special areas, some user terminal devices are set up in areas with no people or harsh environments, and these user terminal devices still generate many compute-intensive tasks. In this scenario, since ground computing servers are usually deployed in fixed base stations (BS), which are far from the user devices generating the computing tasks, offloading the computing tasks to the fixed base stations (BS) will result in a large latency and cannot meet the requirements of latency-sensitive tasks generated by user terminal devices. At the same time, in some remote areas or after natural disasters, there are few or even no fixed base stations (BS) near the user terminal devices. With the construction and deployment of the network, the compute-intensive tasks generated by user devices in these areas are difficult to be processed in a timely manner, resulting in problems such as large latency and low efficiency.
[0044] In recent years, unmanned aerial vehicles (UAVs) have been widely used in edge computing services. At the same time, UAV-assisted computing offloading applications have gradually emerged. Considering the currently well-developed UAV-assisted edge computing services, in remote areas or after natural disasters, edge servers can be deployed on UAVs. During the flight of the UAVs, they can provide auxiliary edge computing offloading services for ground user terminal devices within the coverage area. For computing tasks, the decisions of task scheduling and the latency and cost brought by transmission cannot be ignored. In the UAV-assisted computing task scheduling in this scenario, each computing node carried by the UAV moves in a specific patrol area. The computing tasks can be executed locally or processed by the UAV or ground base stations. In edge computing scenarios, most are limited to using UAVs for centralized control to assist user terminal devices in computing. However, how to cope with the dynamically changing complex topology, make appropriate task scheduling decisions, improve the throughput and fairness of user terminal devices, reduce task processing latency, improve task processing efficiency, reasonably allocate resources, and provide timely and effective computing power has become a technical problem that urgently needs to be solved.
[0045] In view of this, the embodiments of the present disclosure provide a method for allocating edge computing resources assisted by drones and related devices. By extracting the characteristics of drones, ground base stations, and user terminal devices, and aggregating the discrete communications between multiple drones, multi-source data is effectively fused to jointly optimize resource allocation and path selection in complex environments, while adapting to the dynamically changing complex topology and making appropriate task scheduling decisions, and improving the throughput and fairness of user terminal devices.
[0046] The embodiments of the present disclosure provide a computing task scheduling architecture assisted by inspection drones in an edge computing scenario. N user terminal devices (PNT) are distributed in the network, and M drones (UAV) carry computing servers to conduct inspections in this area. At the same time, they can provide computing capabilities for PNT or act as repeaters to hand over computing tasks to the ground base station (BS). Some of the PNT are far from the ground network and cannot be connected to the ground network. Usually, M ≤ N. The present disclosure uses a triple to represent the initial task generated on the i-th PNT, where L i represents the size of the computing task, cpu i represents the number of computing cycles required for the task, represents the maximum tolerable processing delay of the task. Define the total service time in the scenario as T, which is divided into k timestamps, and the length of each timestamp
[0047] represents the location of user terminal device i, represents the location of drone i' at time t. For each drone i', there is an observation range OR and a transmission range TR. The observation range OR is a circular area with a radius of OR. In this observation range, the drone can observe other nearby drones and user terminal devices, etc., and obtain their location status. OR i (t) represents the observation range of drone UAV i’ at time t. The transmission range is a circular area with a radius of TR. If within timestamp k, user terminal device PNT i is located within the transmission range of drone UAV i’ , that is, it satisfies:
[0048]
[0049] then this UAV i’ can communicate with user terminal device PNT i . Within each timestamp, corresponding to different UAV position distributions and computing task situations, but in fact each timestamp makes a decision based on the information at the start of this timestamp, assuming that all tasks can be completed within one timestamp.
[0050] As can be seen from the above system modeling, the total power consumption is not only determined by the energy consumption of computing and transmission, but also the offloading decision of computing tasks and the resource allocation of the server are important factors affecting the overall energy consumption of the system. At the same time, for user equipment, throughput is one of the important evaluation indicators to judge the quality of computing services obtained. Throughput is usually expressed as the amount of data transferred from one communication entity (such as a UAV) to another entity (such as a user) per unit time. In this disclosure, the definition of throughput revolves around the amount of data collected and transmitted by the UAV within a given time, which specifically depends on the transmission range, location of the UAV, and the data requirements of the user. Each UAV has a fixed transmission range TR i (t), and data transmission can only be carried out when the user is within this range. The definition of the transmission range can be expressed as:
[0051]
[0052] where represents the location of user terminal device i, represents the location of UAV i’ at time t. PNT i being within the transmission range means that data transmission can be carried out.
[0053] The optimization objective is to minimize the total energy consumption of the uplink transmission energy consumption and the task computing energy consumption within the entire service time T, and improve the throughput of the PNT device. The derivation is as follows:
[0054]
[0055] Among them, μ and ν respectively represent the weights of the two optimization objectives. Throughput and overhead usually restrict each other: to increase throughput, more computing resources, bandwidth, or energy consumption may be required; while reducing overhead may sacrifice throughput. Therefore, through the multi-objective optimization method, a balance point can be found to make both throughput and overhead reach a satisfactory level. When optimizing the above objectives, in many works, it is usually necessary to maintain the fairness of all PNT devices at the same time, and most of them are evaluated by Qos. However, within the entire service time T, the UAV conducts inspections in a specified area, and the globally optimal resource allocation and path planning may include temporarily unfair steps, so we need to consider long-term fairness.
[0056] The Resource Uniformity Index can include:
[0057]
[0058] Among them, N represents the number of entities participating in resource allocation (such as users, services, etc.). x l represents the amount of resources obtained by each entity l (such as bandwidth, throughput, etc.). The value of the resource uniformity index varies between 0 and 1, indicating the uniformity of resource allocation in the system. When the index is 1, it means that the resource allocation is completely uniform, that is, all entities obtain equal amounts of resources. When the index is close to 0, it means that the resource allocation is extremely uneven, that is, some entities obtain the vast majority of the resources, while other entities obtain very little or no resources. Based on the above resource uniformity index, the fairness of the PNT throughput is modeled as follows:
[0059]
[0060] The larger the value of, the more similar the values of all PNT devices are, that is, the better the fairness represents. At the same time, The larger the value of, the higher the system throughput. Therefore, one of the optimization objectives can be to maximize the product of the above two metrics:
[0061]
[0062] In summary, the optimization problem of this model is associated with the resource allocation strategy and can be:
[0063] See Figure 3 , Figure 3 shows a schematic flowchart of an unmanned aerial vehicle (UAV)-assisted edge computing resource allocation method according to an embodiment of the present disclosure. The UAV-assisted edge computing resource allocation method according to an embodiment of the present disclosure can be deployed on a terminal or a server side. Figure 3 In, the UAV-assisted edge computing resource allocation method 300 can further include the following steps.
[0064] In step S310, obtain first observation data of a first UAV node, where the first observation data includes first status data of the first UAV node and second status data of adjacent nodes.
[0065] Among them, the observation data can refer to the information about the environment, state, or other variables collected from a certain observation point or sensor. In a drone system, the observation data may include the position, speed, direction, sensor readings, etc. of the drone. The first observation data can refer to a specific set of observation data collected from the first drone node, which can be collected at a certain specific time point or time period. Adjacent nodes can refer to other nodes that have direct connections or interactions with a specific node. In a drone network, adjacent nodes can refer to the nodes that are directly connected or directly communicate with the first drone node. The first state data can refer to the state information about the first drone node itself, such as position, speed, etc. The second state data can refer to the state information about the adjacent nodes themselves, such as position, speed, etc.
[0066] In step S320, feature extraction and feature aggregation are performed based on the first state data and the second state data to obtain aggregated features.
[0067] The Feature Extraction Network (FEN) is used to process the local information within the observation range OR i (k) of the drone, mainly responsible for the partial observation Observation in MARL. The purpose of this GNN layer is to use the graph convolutional network to efficiently integrate the local feature information extracted by the UAV, so as to help the agents (UAVs) in the system make better decisions. The input of the entire FEN layer is the initial feature vector of the drone node (including position, computing resource amount, etc.), the feature vectors of neighbor nodes (including the position states of other drones, the position states of PNT devices), and the relationship types of edges.
[0068] Currently, in the research on applying graph neural networks to computing task scheduling, most of the work uses MPNN (Message Passing Neural Network) to update the features of nodes through a message passing mechanism. In MPNN, each node receives information from its neighbor nodes, and this information is called "messages". Then, these messages are combined with its own features through an aggregation operation (such as summation or averaging) to update the node state. This mechanism is usually unweighted, that is, the influence of each neighbor node on the target node is the same. Therefore, MPNN cannot distinguish the importance of neighbor nodes. So, a two-layer GAT (Graph Attention Network) graph attention network can be used to process the observations of a single drone.
[0069] Compared with MPNN, GAT introduces an Attention Mechanism to enhance the weight assignment in the message passing process. In GAT, the contributions of different neighbor nodes to the target node are different. GAT calculates attention coefficients to determine which neighbors have a greater impact on the target node. In this disclosure, the position of the drone changes at each different timestamp, which is suitable for scenarios that require dynamically adjusting the importance of neighbor nodes.
[0070] GAT selectively transmits and aggregates information between nodes through an adaptive attention mechanism, enabling each node to focus on more relevant neighbors, thereby optimizing the node representation. Therefore, for the graph feature extraction in this disclosure, a two-layer GAT graph attention network is used to process the observations of a single drone. The structure of GAT consists of two stages: the attention calculation stage and the feature aggregation stage. In the attention calculation stage, each node calculates an attention score using learnable weights based on its own and neighbor node features to determine the influence of neighbor nodes. Specifically, for each target node, GAT performs a linear transformation and concatenation on its features and neighbor node features, calculates the attention score through a learnable attention weight vector and a non-linear activation (such as LeakyReLU), and normalizes it using Softmax to obtain the attention weights of neighbor nodes, thereby representing the contribution of neighbors to the information aggregation of the target node. In the feature aggregation stage, the node uses these attention weights to perform a weighted sum of the features of neighbors, thereby generating an updated node feature representation. In addition, GAT usually adopts a multi-head attention mechanism, calculating multiple weight sets through multiple independent attention heads. Each attention head learns different attention patterns, and the final output is the concatenation or averaging of the results of all heads, enhancing the robustness and expressive power of the model. GAT is suitable for tasks with graph-structured data such as social network analysis, knowledge graphs, and recommendation systems. Especially when the graph structure changes dynamically, GAT can better adapt to the complex relationships between nodes.
[0071] In some embodiments, performing feature extraction and feature aggregation based on the first state data and the second state data to obtain aggregated features includes:
[0072] Performing feature extraction on the first state data to obtain first features, and performing feature extraction on the second state data to obtain second features;
[0073] Determining the terminal weight of the terminal features corresponding to the adjacent terminal nodes in the adjacent nodes relative to the first features, and the drone weight of the drone features corresponding to the adjacent drone nodes in the adjacent nodes relative to the first features;
[0074] Aggregate features based on the terminal weights, the terminal features, the UAV weights, the UAV features, and the first feature to obtain the aggregated features.
[0075] Among them, to effectively process multiple neighbor nodes observed by the UAV, FEN uses a multi-head self-attention mechanism to assign different weights to each neighbor node, and these weights represent the importance of the neighbor node to the current decision of the UAV. For each neighbor node, FEN calculates its attention weight with the UAV. This process is divided into two parts, respectively dealing with the relationship between the UAV and the power grid terminal device terminal (PNT) and the relationship between the UAV and other UAVs, and records the feature value of the UAV within the current timestamp k i′ is PNT i is
[0076] For each UAV i′ , the feature value of PNT i obtained through the feature extraction network FEN first performs a linear transformation on the features j′ of the UAV and the feature value of PNT i , and then calculates the attention weight of PNT For the UAV j The attention weight j′ of the UAV represents the importance of PNT j to the decision of the UAV i′ :
[0077]
[0078] where σ1 is the activation function LeakyReLU, is the weight vector used to calculate the attention, and Xavier initialization is adopted to ensure that the input and output variances are consistent. The Concat function is used to link two matrices, and W UAV and W PNT are both trainable weight matrices (such as the initial weight matrix), which are respectively used to transform the features of the UAV and PNT.
[0079] For each UAV i′ , the feature value of UAV j′ obtained through the feature extraction network FEN also calculates the attention weight of UAV j′ For the UAV i′ The attention weight represents the importance of UAV j′ to the decision of the UAV i′ :
[0080]
[0081] FEN can adopt the multi - head attention mechanism. Each attention head independently calculates a set of attention weights and performs a weighted sum of the features of neighboring nodes. The final feature is the concatenation of the outputs of multiple attention heads, as follows:
[0082] For PNT j For UAV i′ The relationship of:
[0083]
[0084] Among them, the Concat function is used to concatenate matrices, Head is the number of multi - head attention, and σ2 is the non - linear activation function ReLU.
[0085] For UAV j′ For UAV i′ The relationship of:
[0086]
[0087] Among them, the Concat function is used to concatenate matrices, and Head is the number of multi - head attention.
[0088] After passing through the multi - head attention mechanism, FEN aggregates the features from other devices through a layer of MLP and generates updated features. The output of FEN is the updated feature vector of UAV - BS at the current time step This feature vector combines the key information from the observed PNT and UAV. This information will be used for policy decisions in reinforcement learning (the moving direction and resource allocation strategy of the UAV).
[0089] In step S330, the aggregated feature is discretely encoded to obtain the first discretized feature, and the second discretized feature from the second UAV node is received.
[0090] Among them, in MARL, in order to make the actions of the entire system approach the global optimum as much as possible, each agent needs to exchange information in cooperative tasks to obtain more information. From the eigenvalue observed by each of the above drones and the observation state in the POMDP model, it can be seen that if the agent directly transmits the feature information, the required communication overhead is extremely large. Therefore, in the discrete communication model, the communication between agents is abstracted into a discrete communication network (DCN, Discrete Communication Networks), and the policy information is transmitted between UAVs through discrete communication edges, thereby promoting the cooperation between UAVs. Through communication, each UAV can not only make decisions based on its own observations, but also adjust its policy according to the information received from other UAV-BSs. The communication between UAVs is different from simple observations. Communication usually requires the exchange of policy information, such as how to allocate resources and how to plan flight paths. Since in the scenario of power grid UAV inspection, the communication link is limited by bandwidth, transmitting information through this discrete communication network can greatly reduce the communication overhead.
[0091] The FEN layer encodes the features of the UAV into discrete symbols, then transmits the symbols through discrete communication edges, and decodes these symbols at the receiver for information aggregation. The input of the encoder is the feature vector of the current UAV The output is the probability distribution of each discrete symbol, that is, the probability that each possible symbol is selected:
[0092]
[0093] To reduce the bandwidth overhead of communication, the FEN layer compresses the information transmitted between UAVs into discrete symbols. During the training process, Perturb-and-MAP is used to approximate the sampling process of discrete symbols while maintaining end-to-end differentiability. Specifically, each UAV will have a chance to select its discrete symbol i Add Gumbel noise:
[0094] Then select the maximum value after adding the noise as the final choice of the discrete variable:
[0095]
[0096] In this way, the discrete symbol transmission can retain the gradient information during the training process, thereby achieving end-to-end optimization. During the actual execution process, directly select the discrete symbol for transmission.
[0097] After the UAV receives discrete symbols from other UAVs, it will decode these symbols into useful information through a decoder and aggregate this information to update its own features. The input to the decoder is the aggregated symbol vector, which represents the discrete information received from other UAVs; the output is the updated UAV-BS feature, which is part of the subsequent policy decision. Finally, the FEN is updated through graph convolution:
[0098]
[0099] where decoder and encoder represent the decoder and encoder functions respectively, which are used to encode the features of the UAV into discrete symbols and then decode the aggregated discrete symbols into effective feature information. update is the vertex update function, which is used to combine the current UAV and the feature information from other UAVs, and the gated recurrent unit GRU is used to parameterize the update function.
[0100] In summary, the entire discrete communication model (Discrete Communication Model) is as Figure 4 shown. The local features sensed by the UAV, the features of other user devices within the sensing range of the UAV, and the adjacency matrix are input to the FEN layer. The FEN layer includes a two-layer GAT network. The first-layer GAT-shallow network is used to observe the adjacent features of the current UAV, and then a deeper outer-layer feature is obtained through a layer of GAT-deep network. The output is the aggregated feature of the domain information of the current UAV. Then this feature is used as the input to the discrete communication layer. The discrete communication layer includes an encoder and a decoder. The message sender encoder is an MLP that discretizes the message through Perturb-and-MAP. The receiver uses GRU as the decoder to extract information from the message and combines the current feature and historical information to update the feature of the current UAV. The final new feature is transformed into the target data through a layer of MLP, which is the Q-Value of the current agent.
[0101] In step S340, an updated feature is obtained by combining the first discretized feature and the second discretized feature.
[0102] The Encoder-Decoder architecture is a deep learning model widely used in sequence-to-sequence (seq2seq) tasks. Its core idea is to gradually compress the input sequence into a hidden state representation of a fixed dimension through the Encoder, and then the Decoder generates the target sequence based on this hidden state. The task of the Encoder is to encode the information in the input sequence into a "context vector" that represents the input context. This process usually consists of recurrent neural networks (such as RNN, LSTM, GRU) or Transformer encoder layers, which are particularly suitable for capturing context information in the sequence. The Encoder processes each element of the input sequence one by one, and under the gradual compression of the multi-layer network, finally generates a hidden state that synthesizes the input features. In this way, the global information of the sequence can be extracted and passed to the Decoder in a compact form, providing strong support for the subsequent decoding steps.
[0103] The Decoder generates each element in the target sequence step by step by using the hidden state information output by the Encoder. The Decoder usually consists of multiple RNN, LSTM, GRU or Transformer decoding layers, and predicts the current element based on the previously generated output and the Encoder hidden state at each decoding step. Especially in machine translation or text generation tasks, the output of the Decoder is an autoregressive process, that is, each generated output will be used as the input for the next step to generate the complete target sequence. This iterative generation method allows the Decoder to flexibly adjust the subsequent generation path according to the generated part of the content. To improve the performance in processing long sequences, the attention mechanism is introduced in many Encoder-Decoder architectures. The attention mechanism allows the Decoder to dynamically select the most relevant part of the Encoder at each generation step, rather than relying only on the global information in the context vector. In this way, the model can more accurately associate the corresponding relationships between various parts of the input and output, especially significantly improving the effect in long sequence tasks.
[0104] The Encoder-Decoder architecture has strong flexibility and adaptability, so it is widely used in various sequence-to-sequence tasks. In different tasks, different types of neural networks can be selected for the encoder and decoder to optimize specific requirements. For example, in the task of image caption generation, a convolutional neural network (CNN) is usually used in the encoder part to extract image features, and then the decoder generates the descriptive text; while in text tasks such as machine translation, the encoder and decoder mostly adopt a Transformer-based structure to improve the parallel processing ability. With the continuous expansion and optimization of applications, many variants of the Encoder-Decoder architecture have emerged in practice, such as the encoding-decoding structure based entirely on the Attention mechanism in Transformer. Generally speaking, the Encoder-Decoder architecture meets diverse needs with its stable performance and flexible design, and has become one of the most classic and fundamental models in various sequence-to-sequence modeling tasks.
[0105] In step S350, a resource allocation strategy for the first drone node is determined based on the updated features.
[0106] Multi-Agent Reinforcement Learning (MARL) is a method that extends traditional Reinforcement Learning (RL) and is used to learn effective strategies in complex environments containing multiple agents. Each agent has independent goals, observations, and behavioral decisions in such an environment, and there may be competitive, cooperative, or mixed relationships among them. The core of multi-agent reinforcement learning is that agents learn optimal strategies through continuous interaction with the environment and other agents to achieve optimal behavior at the individual or overall level. In multi-agent reinforcement learning, each agent not only needs to consider its own state but also the cooperation and mutual influence with other agents. Therefore, compared with the MDP model of reinforcement learning, MARL needs to be further extended to a partially observable Markov decision process (POMDP): where represents the set of all possible states in the system, represents the set of actions that each agent can execute. Actions are selected based on the local information observed by the agent, represents the set of observed states, represents the state transition function, is the reward function. At time step t, the environment is in state s t . Each agent i receives its observation and makes a decision based on this observation. Each agent i selects an action based on its policy according to the current observation The action combinations of all agents form a joint action, which acts on the environment. After the joint action is executed, the system updates to the next state s according to the state transition function . Each agent i receives the corresponding immediate reward according to the reward function t+1 . The goal of each agent is to maximize its long-term cumulative reward, i.e., the expected return by the policy π, where γ is the discount factor, which is used to balance the influence of immediate rewards and future rewards. t
[0107] In some embodiments, determining the resource allocation policy of the first drone node based on the updated feature includes:
[0108] Determining the resource allocation policy based on the updated feature and the trained decision network; wherein, the trained decision network is trained based on training samples, and the training samples include local observation samples of drone samples, communication coding samples from adjacent drones, and actual global values.
[0109] In some embodiments, the trained decision network is trained based on training samples, including:
[0110] Performing feature extraction and feature aggregation on the local observation samples to obtain aggregated feature samples;
[0111] Performing discretization encoding on the aggregated feature samples to obtain discretized feature samples;
[0112] Updating the discretized samples based on the communication coding samples to obtain updated feature samples;
[0113] Determining the sample local value of the drone sample based on the updated feature samples;
[0114] Combining the local value functions of all the drone samples to obtain the sample global value;
[0115] Adjusting the model parameters of the decision network to minimize the difference between the sample global value and the actual global value, to obtain the trained decision network.
[0116] Among them, in the edge computing scenario of the disclosed embodiments, each UAV base station is regarded as an agent, and the overall environment includes UAVs, ground base stations, user terminal devices, and relevant discrete communication rules. In the previous part, the MARL problem has been modeled as a partially observable Markov decision process (POMDP). To represent the complex relationships of different devices in the system, the research models the entire system as a heterogeneous graph. This heterogeneous graph contains three types of nodes and four types of edges. Each UAV processes these relationships through graph convolution to make decisions in a partially observable environment. Among them, the three types of nodes and four types of edges can include: UAV nodes: Each UAV is regarded as a node. UAV vertices are mainly responsible for being "service providers" in the network. They optimize the coverage of ground terminal devices by adjusting their flight paths and communicate and cooperate with other UAVs / BSs. UAV nodes update their states through their observation information and communication with other UAVs and ground terminal devices. These are also the nodes for making decisions. N UAV ={1, 2, ……, M UAV}, M UAV is the set of the number of UAVs. In this node attribute, it includes the location information of the UAV, the UAV computing information, and the information of other BSs or UAVs observed. PNT nodes: Each PNT device is also represented as a node. This node does not make decisions. PNT vertices are "service requesters". They need to obtain wireless coverage services from UAVs and hope to obtain efficient and stable computing services through UAVs. N PNT ={1, 2, ……, N PNT}, N PNT is the set of PNT devices. BS nodes: Each ground base station BS is also a node that does not make decisions. It also serves as a service provider. It can directly serve PNT devices or accept the computing tasks transmitted by UAV nodes as relay devices. N BS ={1, 2, ……, N BS}, N BS is the set of ground base stations. The definition of edges can include: Feature extraction edges: Feature extraction edges are mainly responsible for part of the observation Observation in MARL. There will be a feature extraction edge between two nodes if and only if the distance between the UAV and other devices is within the observation range OR i (k) of the current UAV. In each time stamp k, the system dynamically updates these edges according to the position of the current UAV. Feature extraction edges are represented as an adjacency matrix. U-P: Represents the PNT devices that the UAV can observe in the environment. The adjacency matrix represents the PNT devices that each UAV can observe at the current time stamp k, where M UAV, N PNT respectively represent the number of UAVs and PNTs. Matrix element represents that at timestamp k, UAV i can observe PNT j , otherwise U-U: Represents other UAVs in the environment that a UAV can observe. Adjacency matrix represents other UAVs that each UAV can observe at the current timestamp k, where M UAV represents the number of UAVs. U-B: Represents the ground base stations in the environment that a UAV can observe. Adjacency matrix represents other BSs that each UAV can observe at the current timestamp k, where M UAV , N BS respectively represent the number of UAVs and BSs. Discrete communication edges: These edges are mainly used to represent the discrete communication relationships between UAVs and are responsible for the communication between agents in MARL. If there is a discrete communication edge between two UAVs, it means that these two UAVs can share information and cooperate through a wireless channel. Two UAVs can communicate if and only if the distance between them is less than the discrete communication range TR i (k) of the current UAV. Different from the U-U feature extraction edges, discrete communication edges not only represent observing each other, but also mean that U-U can actively exchange policy information to jointly optimize the system performance. Similarly, discrete communication edges are also represented as an adjacency matrix, C-UU: Represents that efficient discrete communication can occur between two UAVs.
[0117] The feature extraction FEN layer is specifically used to process the observation information of UAVs, that is, the information of other user devices and UAVs that a UAV can directly observe. Using the Graph Attention Mechanism, different weights are assigned to the neighbors of each UAV (i.e., the detected PNTs and other UAVs). These weights are calculated based on the relevance between the neighbors and the current UAV, so as to filter out the most important information. The discrete communication DCN layer is used for information sharing between different UAVs. To reduce communication overhead, this layer adopts a discrete message passing mechanism. Specifically, a UAV generates discrete symbols through an encoder, performs approximate sampling through Perturb-and-MAP, and then sends these symbols to other UAVs. The receiving end then decodes these discrete information through a decoder and uses it to update its own state.
[0118] Based on the idea of the QMIX algorithm, the model training and execution use the strategy of centralized training and distributed execution (CTDE). The entire network is similar to the "centralized network" and "distributed executors" of the AC framework. Centralized training: During the training phase, the observations, actions, and rewards of all UAVs share a centralized training framework. This framework can access the global observation information of all UAVs and build a heterogeneous graph, as Figure 4 shown. It includes the local visual information of each UAV and the communication information with other UAVs, enabling end-to-end model optimization. The model updates its parameters through experience replay and minimizing the TD error. See Figure 5 , Figure 5 which shows a schematic diagram of the hybrid network according to an embodiment of the present disclosure. During the execution phase, each UAV makes independent decisions based on its own local observations and the shared information of the DCN layer. The execution of the model does not rely on global information. Specifically, based on the current state and the experience of other agents, each agent independently selects actions to schedule computing tasks and plan the UAV path. At the beginning of each timestamp, the action space of a certain agent can be defined as: the selection of the moving speed direction of the UAV and the selection of whether the UAV is used as a computing device or a relay node. The specific process of DGDCMM is as follows:
[0119] Local Q-network: Each UAV acts as an independent agent and first calculates its action-value function Q value through its own local Q-network. The information input into the local Q-network is the feature information abstracted by the local observation layer and the states of other power grid terminal devices. The output local Q value Q i represents the expected return of the current agent.
[0120] Hybrid network: The hybrid network is similar to QMIX and is a non-linear network used to integrate the local Q values of all UAVs into a global Q mix . The hybrid network needs to satisfy the monotonicity constraint. When any local Q value increases, the entire global Q mix value will not decrease. This property ensures that even when each agent executes in a decentralized manner, the locally optimal choice can be closer to the globally optimal.
[0121] Calculation of the global Q value: The hybrid network is constructed using a multi-layer perceptron (MLP). Each layer uses the state vector as the input, and the Q mix calculation formula is as follows:
[0122]
[0123] Among them, ELU is an activation function used to increase the non - linear relationship. w1 and b1 are the weight matrix and bias term of the local Q - value, and w final and b final are the weight matrix and bias term of the global Q - value, which are used to make the network satisfy the monotonicity constraint.
[0124] The global Q - value is used for policy optimization: during the training process, the TD - error between Q mix and the actual reward is calculated to optimize the parameters of the local Q - network. In this way, by continuously adjusting the local Q - network and the hybrid network, the system gradually learns how to make the locally optimal choice improve the global Q - value, thus achieving effective cooperation among multiple agents.
[0125] As can be seen, the method according to the embodiments of the present disclosure proposes a UAV-assisted computing task scheduling framework in an edge computing scenario. Under this model framework, multiple unmanned aerial vehicles (UAVs) carry computing servers and act as auxiliary devices to help user equipment perform computing task scheduling. Among them, the UAV can be used as a computing power node to process computing tasks or as a repeater to offload computing tasks to a fixed base station (BS) on the ground. Multiple UAVs fly in a fixed area and provide computing services for user equipment along the line. A resource allocation evaluation index, the Resource Uniformity Index, is also proposed to ensure the maximization of the throughput of ground user terminal equipment and the long-term fairness of resource allocation, and a distributed resource allocation and path selection problem for UAV-ground equipment (UAV-PNT) is defined. Considering the dynamic variability of the scenario, we use Heterogeneous Graph Neural Networks to represent the relationship between ground user equipment (PNT), UAVs, and ground base stations (BS). A two-layer GAT discrete communication model is constructed through a graph attention network and an encoder-decoder architecture for the dynamic variability of the scenario and the communication loss between UAVs. The outer layer features of the second-order neighbors are obtained through two layers of FEN, and the output is the aggregated feature of the current UAV. Then this feature is used as the input to the discrete communication layer GCN, and the communication loss is reduced by transmitting discretized feature information between UAVs. A multi-agent reinforcement learning method based on a two-layer graph attention network and a discrete communication model (DGDCMM, Double-layer GAT and Discrete Communication Model based MARL) is also proposed. The upper layer of the method uses a two-layer GAT discrete communication model, and the lower layer regards each UAV as an agent, and constructs the resource allocation and path selection of UAVs as an observable Markov decision process (POMDP). Using FEN and DCN to replace the independent agent training process in MARL is more suitable for collaborative tasks such as discrete action spaces and resource allocation, and significantly improves the overall throughput of user terminal equipment.
[0126] It should be noted that the method according to the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only execute one or more steps of the method according to the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.
[0127] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0128] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure further provides a method and apparatus for allocating edge computing resources assisted by an unmanned aerial vehicle (UAV). Refer to Figure 6 The method and apparatus for allocating edge computing resources assisted by an unmanned aerial vehicle, the apparatus includes:
[0129] An acquisition module, configured to acquire first observation data of a first UAV node, where the first observation data includes first status data of the first UAV node and second status data of adjacent nodes;
[0130] A feature aggregation module, configured to perform feature extraction and feature aggregation based on the first status data and the second status data to obtain an aggregated feature;
[0131] An encoding module, configured to perform discretization encoding on the aggregated feature to obtain a first discretized feature;
[0132] The acquisition module is further configured to receive a second discretized feature from a second UAV node;
[0133] A feature update module, configured to combine the first discretized feature and the second discretized feature to obtain an updated feature;
[0134] A resource allocation module, configured to determine a resource allocation strategy for the first UAV node based on the updated feature.
[0135] For convenience of description, when describing the above apparatus, it is divided into various modules according to functions for separate description. Of course, when implementing the present disclosure, the functions of each module can be implemented in one or more software and / or hardware.
[0136] The apparatus of the above embodiment is used to implement the corresponding method for allocating edge computing resources assisted by an unmanned aerial vehicle in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated herein.
[0137] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the method for drone-assisted edge computing resource allocation described in any of the above embodiments.
[0138] The computer-readable media of this embodiment include both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0139] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the method for drone-assisted edge computing resource allocation described in any of the above embodiments and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0140] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; under the concept of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present disclosure as described above. For the sake of brevity, they are not provided in detail.
[0141] In addition, for simplicity of explanation and discussion, and so as not to make the embodiments of the present disclosure difficult to understand, well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the devices may be shown in block diagram form in order to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that details regarding the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (i.e., these details should be entirely within the understanding of those skilled in the art). In cases where specific details (such as circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure may be practiced without these specific details or with variations of these specific details. Accordingly, these descriptions should be regarded as illustrative rather than restrictive.
[0142] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0143] Embodiments of the present disclosure are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Accordingly, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the embodiments of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A method for allocating edge computing resources assisted by drones, comprising: Obtaining first observation data of a first drone node, where the first observation data includes first status data of the first drone node and second status data of adjacent nodes; Performing feature extraction and feature aggregation based on the first status data and the second status data to obtain an aggregated feature; Performing discretization encoding on the aggregated feature to obtain a first discretized feature, and receiving a second discretized feature from a second drone node; Combining the first discretized feature and the second discretized feature to obtain an updated feature; Determining a resource allocation strategy for the first drone node based on the updated feature.
2. The method according to claim 1, wherein, Performing feature extraction and feature aggregation based on the first status data and the second status data to obtain an aggregated feature, including: Performing feature extraction on the first status data to obtain a first feature, and performing feature extraction on the second status data to obtain a second feature; Determining a terminal weight of a terminal feature corresponding to an adjacent terminal node in the adjacent nodes relative to the first feature, and a drone weight of a drone feature corresponding to an adjacent drone node in the adjacent nodes relative to the first feature based on the second feature; Performing feature aggregation based on the terminal weight and the terminal feature, the drone weight and the drone feature, and the first feature to obtain the aggregated feature.
3. The method according to claim 2, wherein Determining a terminal weight of a terminal feature corresponding to an adjacent terminal node in the adjacent nodes relative to the first feature, and a drone weight of a drone feature corresponding to an adjacent drone node in the adjacent nodes relative to the first feature based on the second feature, including: Adjacent terminal node PNT j For the first unmanned aerial vehicle node UAV i′ Terminal weight of: Adjacent UAV nodes j′ For the first UAV node i′ UAV weight: Among them, σ1 is the activation function LeakyReLU, is the attention weight, W UAV and W PNT are both weight matrices.
4. The method according to claim 3, wherein Performing feature aggregation based on the terminal weight and the terminal feature, the drone weight and the drone feature, and the first feature to obtain the aggregated feature, including: Adjacent terminal node PNT j For the first unmanned aerial vehicle node UAV i′ The first attention feature of: Adjacent UAV nodes j′ For the first UAV node i′ The second attention feature of: Performing aggregation based on the first attention feature, the second attention feature, and the first feature to obtain the aggregated feature.
5. The method according to claim 4, wherein Combining the first discretized feature and the second discretized feature to obtain an updated feature, including: Where decoder and encoder respectively represent decoder and encoder functions, and update is a vertex update function.
6. The method according to claim 1, wherein, Determining a resource allocation strategy for the first drone node based on the updated feature, including: Determining the resource allocation strategy based on the updated feature and a trained decision network; where the trained decision network is obtained by training based on training samples, and the training samples include local observation samples of drone samples, communication coding samples from adjacent drones, and actual global values.
7. The method according to claim 6, wherein The trained decision network is obtained by training based on training samples, including: Performing feature extraction and feature aggregation on the local observation samples to obtain aggregated feature samples; Performing discretization encoding on the aggregated feature samples to obtain discretized feature samples; Updating the discretized samples based on the communication coding samples to obtain updated feature samples; Determining a sample local value of the drone samples based on the updated feature samples; Combining the local value functions of all the drone samples to obtain a sample global value; Adjust the model parameters of the decision network to minimize the difference between the sample global value and the actual global value, and obtain the trained decision network.
8. An edge computing resource allocation device assisted by a drone, comprising: An acquisition module, configured to acquire first observation data of a first drone node, where the first observation data includes first state data of the first drone node and second state data of adjacent nodes; A feature aggregation module, configured to perform feature extraction and feature aggregation based on the first state data and the second state data to obtain an aggregated feature; An encoding module, configured to perform discretization encoding on the aggregated feature to obtain a first discretized feature; The acquisition module is further configured to receive a second discretized feature from a second drone node; A feature update module, configured to combine the first discretized feature and the second discretized feature to obtain an updated feature; A resource allocation module, configured to determine a resource allocation strategy for the first drone node based on the updated feature.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium, where the non-transitory computer-readable storage medium stores computer instructions for causing a computer to execute the method according to any one of claims 1 to 7.