Joint optimization method for task unloading and resource allocation of space-air-ground edge computing network
Through the joint optimization method of GAT and PPO, the complex dependencies and dynamic resource allocation problems of user tasks in the integrated air-space-ground network are solved, efficient computing services are achieved at the disaster site, and task latency and energy consumption are optimized.
Patent Information
- Application Number
- CN202510901684.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-10-03
AI Technical Summary
Traditional ground mobile edge computing networks suffer from degraded service quality in extreme environments such as disaster sites due to damaged communication infrastructure or limited resources. Existing technologies are also unable to effectively handle the efficient offloading of user tasks and the rational allocation of computing resources in integrated air-space-ground networks, especially in optimizing latency and energy consumption under complex task dependencies and dynamic resource states.
A joint optimization method of graph attention network (GAT) and proximal policy optimization (PPO) is adopted. By constructing an air-space-ground edge computing network model, GAT is used to extract the topology and features of the directed acyclic graph (DAG) of the task, and combined with the PPO deep reinforcement learning network, the task offloading location and resource allocation are dynamically adjusted to minimize latency and energy consumption.
It significantly improves the execution efficiency of user tasks and network resource utilization in the integrated air-space-ground network, and optimizes the user task completion delay and system energy consumption.
Smart Images

Figure CN120751445A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of wireless communications, and more specifically, it relates to a method, system, and terminal device for jointly optimizing task offloading and resource allocation in an air-space-ground-ground edge computing network. Background Art
[0002] With the rapid development of technologies such as mobile communications, drone communications, and satellite communications, user tasks are increasingly characterized by large data volumes, intensive computations, and complex inter-task dependencies. While traditional terrestrial mobile edge computing (MEC) networks can provide certain computing services to users, in extreme environments such as disaster sites, damaged communication infrastructure or resource constraints often lead to service degradation or even interruption. To address these issues, integrated air-ground-space edge computing networks—those that deeply integrate space, air, and ground networks—have gained widespread attention in recent years. Edge computing servers carried by drones and low-orbit satellites effectively supplement the terrestrial network, significantly improving the computing service capabilities of user tasks and network coverage. However, due to the heterogeneous and dynamic nature of integrated air-ground-space networks, efficient offloading of user tasks and the rational allocation of computing resources remain pressing technical challenges. On the one hand, user tasks are typically composed of multiple interdependent subtasks, each with a clear execution order. Traditional offloading strategies struggle to accurately capture these complex task dependencies. On the other hand, existing technologies struggle to balance the optimization objectives of dynamically adjusting resource allocation based on the real-time computing resource status of drones and low-orbit satellites in the network to minimize latency and energy consumption. To address these challenges, a joint optimization approach based on a graph attention network (GAT) and proximal policy optimization (PPO) has emerged as an effective solution. GAT efficiently extracts and characterizes the topological structure and subtask characteristics of a task's directed acyclic graph (DAG), while PPO accurately determines the task's offloading location (local ground, drone, or low-orbit satellite) and the proportion of computing resources allocated based on the current state. This joint optimization mechanism dynamically adjusts the decision-making policy through a deep reinforcement learning framework, significantly improving the overall efficiency of user task execution and network resource utilization in integrated air-ground-space networks. Summary of the Invention
[0003] The present invention provides a method, system and terminal device for joint optimization of task offloading and resource allocation in an air-space-ground-ground edge computing network, which is beneficial to optimizing user task completion delay and system energy consumption.
[0004] In a first aspect, the present invention proposes a method for joint optimization of task offloading and resource allocation in an air-ground-space edge computing network, the method comprising:
[0005] S1: In the air-ground-space edge computing network, the network models of air-based network, space-based network and ground network, user-associated task model, communication model and computing model are constructed respectively;
[0006] S2: Calculate the total delay and total energy consumption required to complete the task through the model; construct a normalized cost function, and use the normalized cost function to measure the joint indicator of user task delay and system energy consumption;
[0007] S3: Under the joint constraints of user task latency and overall system energy consumption, the problem is formulated as a joint non-convex optimization problem, and the normalized cost function is used as the objective function of the optimization problem;
[0008] S4: The optimal solution is obtained by using the joint decision-making method of GAT and PPO. GAT encodes the user task DAG and extracts node feature vectors that can characterize the task characteristics through the multi-head attention mechanism.
[0009] S5: Input the feature vector as the state into the PPO deep reinforcement learning network, build an Actor network with dual output heads, and use a deep neural network to enhance PPO's learning of states and actions, so that the agent can gradually minimize the total task delay and total system energy consumption during the interaction with the environment.
[0010] In one possible implementation, the network models, user-associated task models, communication models, and computing models of the air-based network, space-based network, and ground-based network are specifically as follows:
[0011] There are U ground users GU in the ground network g ,in Each ground user GU g The device will generate 1 task; there are A drones U in the space-based network a ,in There is a low-orbit satellite L in the space-based network s , L s To interact with all drones on the air-based platform, communications between ground users and low-orbit satellites need to be relayed by drones.
[0012] In actual scenarios, user tasks are composed of subtasks, which are interrelated and interdependent. g The resulting task is denoted as ψ g , assuming Where W g Indicates the size of user task input data, F gIndicates the number of CPU cycles required to process the task; g Modeled as DAG, DAG can provide dependencies between tasks and use structured representation to improve the efficiency of task scheduling and resource allocation, which is expressed as where the vertex set v represents the subtask, and N is the total number of subtasks in DAG, v g,n Indicates ground user GU g The nth subtask in the task DAG, assuming
[0013] In the air-ground mobile edge computing network, Orthogonal Frequency Division Multiple Access (OFDMA) technology is used between users and drones, between drones, and between drones and low-orbit satellites; drones use the L band to communicate with ground users, use the C band to communicate with adjacent drones, and low-orbit satellites use the Ka band to communicate with drones; in the ground-ground link, due to the Doppler effect caused by the rapid movement of drones in the space-based network, the small-scale fading of the ground-ground link occurs. Rayleigh fading is usually used to model the communication link between the ground and the air, and the Shannon formula of Rayleigh fading is used to calculate the ground user GU g With UAV a The uplink and downlink data transmission rates between them are: In the space link, due to the long distance and complex environment of space communication, the free space path loss is greater, and the signal propagation is also affected by rain fading. The communication channel model of the space link is modeled as a channel with small-scale fading and Ricean fading. The Shannon formula of Ricean fading is used to calculate the UAV U a With low-orbit satellite L s The uplink and downlink transmission rates of the two UAVs are affected by free space path loss and small-scale Rayleigh fading. The Shannon formula of Rayleigh fading is used to calculate the UAV U a With UAV a′ The link transmission rate between them;
[0014] In the described air-space-ground mobile edge computing network, users, drones, and low-orbit satellites will perform computational processing on subtasks, use the ratio of task data size to transmission rate to construct transmission delay, and use the product of transmission power and transmission delay to construct transmission energy consumption; use the ratio of the number of CPU cycles required for task execution to the computing resources allocated by the mobile edge service to construct computing delay, and use the product of the number of CPU cycles required for task execution and the square of the computing resources allocated by the mobile edge service to construct computing energy consumption.
[0015] In one possible implementation, the total latency, total energy consumption, and normalized cost function required to complete the calculation task are specifically:
[0016] To determine the total delay of the task, it is necessary to calculate the subtask v g,n Preparation delay at each node; the subtask v g,n The preparation delay of the predecessor task Completion delay and calculation result return delay If the predecessor task and the successor task are on the same device, the calculation result return delay is 0; therefore, the ground user execution delay is expressed as the sum of the preparation delay and the calculation delay, and the execution energy consumption is expressed as the sum of the calculation energy consumption and the transmission energy consumption. The UAV execution delay is expressed as the sum of the preparation delay, the calculation delay and the transmission delay, and the execution energy consumption is expressed as the sum of the calculation energy consumption and the transmission energy consumption. The low-orbit satellite execution delay is expressed as the sum of the preparation delay, the calculation delay and the transmission delay, and the execution energy consumption is expressed as the sum of the calculation energy consumption and the transmission energy consumption. The total task delay is the longest task completion delay of all exit tasks, and the total energy consumption is the sum of the execution energy consumption of the ground user, the UAV, and the low-orbit satellite. The total delay and total energy consumption are normalized and multiplied by their respective weight factors, and the sum is added to obtain the normalized cost function.
[0017] In one possible implementation, under the joint constraints of user task latency and overall system energy consumption, the non-convex optimization problem is:
[0018] The normalized cost function is used as the objective function of the non-convex optimization problem, and the non-convex optimization problem is constructed to minimize the normalized cost function by optimizing the offloading strategy, drone position, and resource allocation of each computing node; and the user task offloading decision constraints, computing resource allocation constraints, maximum tolerable delay constraints, maximum tolerable energy consumption constraints, and drone flight range constraints are constructed.
[0019] In one possible implementation, GAT encodes the user task DAG and uses a multi-head attention mechanism to extract node feature vectors that can characterize task characteristics. The specific steps are:
[0020] GAT is used to encode the user task DAG. First, the feature vector of each subtask in the DAG is defined. Then, the attention coefficient of the node is calculated through the weight matrix and the shared attention mechanism. The attention coefficient is then normalized using the Softmax function. Finally, the final output feature vector of the DAG is obtained through the multi-head attention mechanism.
[0021] In one possible implementation, the feature vector is input as the state into the PPO deep reinforcement learning network, and a dual-output-head Actor network is constructed. The method of using a deep neural network to enhance PPO is as follows:
[0022] The user task offloading process is modeled as a Markov process. The input sequence of tasks is prioritized and recursively sorted starting from the exit task. The prioritized subtasks are sequentially passed through the PPO Actor network according to their priority. The PPO method state space, action space, and reward function are constructed. The feature vector obtained after the subtask is processed by the GAT module, the offloading status of the current user task DAG, and the status information of the MEC server in the environment are input into the Actor network as state. The Actor network with a multi-head perceptron structure with two heads and two layers is used to output discrete and continuous mixed actions. A Critic network is constructed to train and optimize the Actor network. The clipping objective function is designed to limit the amplitude of each policy update to prevent the policy update from being too large, thereby maintaining the stability of training.
[0023] In the second aspect, the present invention proposes a joint optimization device for task offloading and resource allocation in an air-space-ground edge computing network, comprising a processing unit, an acquisition unit, a computing unit and an output unit: the processing unit is used to process user tasks, generate a DAG with mutually related subtasks, and perform GAT encoding on the DAG; the computing unit is used to calculate the preparation delay, computing delay, transmission delay, computing energy consumption and transmission energy consumption of the DAG; the acquisition unit is used to collect the computing resources of each computing node in the current environment and the feature vector of the user task after GAT encoding; the output unit is used to process the environmental information and the user task feature vector to obtain a user task offloading strategy and a computing node resource allocation strategy.
[0024] In a third aspect, the present invention proposes a terminal device, which includes a processor and a memory, and the processor and the memory are connected to each other; wherein the memory is used to store a computer program, and the computer program includes program instructions; the processor is specifically used to call the computer program from the memory to execute the method proposed in the first aspect above.
[0025] In a fourth aspect, the present invention proposes a computer-readable storage medium, which stores computer-readable instructions. When the computer-readable instructions are executed on a communication device, the communication device executes the method proposed in the first aspect and any possible implementation thereof.
[0026] In a fifth aspect, the present invention proposes a computer program product, which, when running on a computer, enables the computer to execute the method proposed in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings described herein are used to provide a further understanding of the embodiments of the present invention, constitute a part of the present invention, and do not constitute a limitation of the embodiments of the present invention, wherein:
[0028] Figure 1 A diagram of the air-space-ground mobile edge computing network scenario provided for an example of the present invention;
[0029] Figure 2 A flow chart of the air-ground mobile edge computing network provided for an example of the present invention;
[0030] Figure 3 A schematic diagram of a directed acyclic graph provided by an example of the present invention;
[0031] Figure 4 GAT encoding flow chart provided for the example of the present invention;
[0032] Figure 5 Schematic diagram of the multi-head attention mechanism calculation of feature vectors provided by the example of the present invention;
[0033] Figure 6 Flow chart of the GAT-PPO method provided by the example of the present invention;
[0034] Figure 7 A schematic diagram of the structure of a terminal device provided by an example of the present invention;
[0035] Figure 8 It is a structural schematic diagram of a device provided by an example of the present invention. DETAILED DESCRIPTION
[0036] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following examples are only schematic illustrations of the basic concept of the present invention. In the absence of conflict, the following examples and the features in the examples can be combined with each other.
[0037] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the examples of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.
[0038] The same or similar numbers in the drawings of the examples of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0039] See also Figure 1 and Figure 2 , which is a network scenario diagram and flow chart of air-space-ground mobile edge computing provided by an example of the present invention, in which a three-layer network of air-space-ground mobile edge computing is constructed in a disaster scenario;
[0040] Specifically, there are U ground users GU in the ground network. g ,in Each ground user GU g The device will generate 1 task, so there are G=U tasks in total. Assume There are A drones U in the space-based network a ,in These drones will cover all users in the area; there is a low-orbit satellite L in the air-based network. s , L s Interact with all drones on the air-based platform. Since the transmission power of ground users is relatively low, the communication between ground users and low-orbit satellites needs to be relayed by drones.
[0041] Use binary uninstall strategy P g,n To express the offloading strategy of subtasks in user tasks:
[0042]
[0043] in Indicates the binary uninstall decision, Represents subtask v g,n Processing is done locally, Represents subtask v g,n Unload to UAV a To process, Represents subtask v g,n Unload to low-orbit satellite L s For processing, the subtask can only be processed on one device in the local or mobile edge computing server, so the task offloading decision needs to meet the following requirements:
[0044]
[0045] In the ground-to-air link, the Doppler effect caused by the rapid movement of drones in the space-based network leads to small-scale fading of the ground-to-air link. Rayleigh fading is usually used to model the communication link between the ground and the air. g With UAV a The uplink data transmission rate between is expressed as:
[0046]
[0047] Among them, B ua is the bandwidth of the space-ground link; P u is the transmission power of the ground user; G ua is the antenna gain from the ground user to the UAV, assuming the antenna is omnidirectional, G ua is a constant; is the free space path loss, which is calculated as:
[0048]
[0049] where d ua is the distance between the user and the drone; f ua is the carrier frequency between the user and the drone; is the Rayleigh fading coefficient, represents the non-line-of-sight component of the complex Gaussian distribution; B0 represents the noise power of the system, N0 = κ B T S B ua ; where κ B is the Boltzmann constant; T S is the total noise temperature of the system;
[0050] Similarly, ground user GU g With UAV a The downlink data transmission rate between is expressed as:
[0051]
[0052] Among them B ua =B au ;P a is the transmitting power of the UAV; G au is the antenna gain from the UAV to the ground user;
[0053] In the space link, due to the long distance and complex environment of space communication, the free space path loss is greater, and the signal propagation is also affected by rain fading. The communication channel model of the space link is modeled as a channel with small-scale fading and Rice fading. aWith low-orbit satellite L s The uplink transmission rate is expressed as:
[0054]
[0055] Among them, B as is the bandwidth of the space link; P a is the transmitting power of the UAV; G as is the antenna gain from the UAV to the LEO satellite; is the free space path loss; Losses due to rain attenuation; is the channel fading coefficient, and its expression is:
[0056]
[0057] Where F is the Rice fading factor; represents the non-line-of-sight component of the complex Gaussian distribution;
[0058] Similarly, you can get drone U a With low-orbit satellite L s The downlink transmission rate is expressed as:
[0059]
[0060] Among them, B as =B sa ;P s is the transmission power of the low-orbit satellite; G sa is the antenna gain from the LEO satellite to the UAV;
[0061] During the mission execution, the UAV a The subsequent tasks of the subtasks executed on the UAV may be offloaded to other UAVs UAVs. a′ Therefore, it is necessary to define the link transmission rate between UAVs; assuming that the link between two UAVs is affected by free space path loss and small-scale Rayleigh fading, the UAV U a With UAV a′ The link transmission rate between them is:
[0062]
[0063] Among them, B aa′ is the bandwidth for communication between UAVs; P a is the transmitting power of the UAV; G aa′ is the antenna gain from drone to another drone; is the free space path loss; is the Rayleigh fading coefficient,
[0064] See also Figure 3 , which is a schematic diagram of a directed acyclic graph provided by an example of the present invention, considering that a user task is composed of many subtasks, and the subtasks are interrelated and interdependent;
[0065] Ground User GU g The resulting task is denoted as ψ g , assuming Where W g Indicates the size of user task input data, F g Indicates the number of CPU cycles required to process the task; g Modeled as DAG, DAG can provide dependencies between tasks and use structured representation to improve the efficiency of task scheduling and resource allocation, which is expressed as The vertex set represents a subtask, and is the total number of subtasks in DAG, v g,n Indicates ground user GU g The nth subtask in the task DAG, assuming Subtask v g,n The calculation result after completion is Assumptions The directed edge set ε represents the dependency relationship between subtasks, which is expressed as ε=(v g,i ,v g,j ), where v g,i is the predecessor task, v g,j A subsequent task can only be executed after the predecessor task is completed.
[0066] After the user task is generated, the subtask will be assigned to local execution or offloaded to the corresponding mobile edge computing server for processing according to the decision of the intelligent agent; by constructing the user task ψ g The calculation model is used to analyze the total delay and total energy consumption during task completion. Due to the dependency between subtasks, subsequent tasks must wait for the completion of predecessor tasks before they can be executed. The completion time of subtasks can be divided into three parts: transmission delay, preparation delay, and calculation delay. User energy consumption mainly includes transmission energy consumption and calculation energy consumption.
[0067] Subtasks may be offloaded to three different types of devices. The task computation model covers three cases:
[0068] When the user subtask v g,n When performing calculations locally, subtasks do not need to be transmitted, so the transmission delay of the task is 0; if the subtask executed locally is a predecessor task of a task, the calculation results need to be transmitted to the UAV for further calculation or relay transmission; define the ground user GUg UAV a The transmission task delay is:
[0069]
[0070] Where b>1 indicates transmission loss; Indicates the current user device GU g The data size of the transfer task;
[0071] Ground User GU g UAV a The energy consumption of the transmission task is expressed as:
[0072]
[0073] Subtask v g,n On the ground user GU g When performing the calculation, the resulting delay is expressed as:
[0074]
[0075] in, Indicates ground user GU g The subtask v g,n allocated computing resources; Indicates user equipment GU g The number of CPU cycles required to execute the current subtask in. g,n On the ground user GU g The energy consumption generated by the calculation is expressed as:
[0076]
[0077] where ι is the energy factor, which depends on the chip structure of the device;
[0078] As a relay node in the air, the drone can not only offload the user's subtasks to other drones or low-orbit satellites, but also act as a mobile edge server to calculate the user's subtasks. a Upon receiving the ground user GU g Uninstalled subtask v g,n After that, it will be further offloaded to other edge computing servers or completed locally according to the task requirements, and the completed subtasks may transmit the results to other devices; UAV U a May provide GU to ground users g 、Other UAVs a′ , low-orbit satellite L s Transmission subtask; UAV U a Towards low-orbit satellite L sThe latency of the transmission task is expressed as:
[0079]
[0080] in, Indicates the current drone U a The data size of the transmission task. a Towards low-orbit satellite L s The energy consumption of the transmission task is expressed as:
[0081]
[0082] UAVU a To another drone U a′ The latency of the transmission task is expressed as:
[0083]
[0084] UAVU a To another drone U a′ The energy consumption of the transmission task is expressed as:
[0085]
[0086] UAVU a To ground users GU g The latency of the transmission task is expressed as:
[0087]
[0088] UAVU a To ground users GU g The energy consumption of the transmission task is expressed as:
[0089]
[0090] Subtask v g,n In UAV U a When performing the calculation, the resulting delay is expressed as:
[0091]
[0092] in, Indicates drone U a For subtask v g,n allocated computing resources; Indicates drone U a The number of CPU cycles required to execute the current subtask in. g,n In UAV U a When performing the calculation, the resulting energy consumption is expressed as:
[0093]
[0094] As a key node in space, low-orbit satellites have extremely high computing power, but also have higher energy consumption; the subtasks of ground users v g,n The drone is used to relay and upload to the low-orbit satellite, so the low-orbit satellite L s Only with UAV a For communication. Low Earth Orbit Satellite L s UAV a The latency of the transmission task is expressed as:
[0095]
[0096] in, Indicates the current low-orbit satellite L s The data size of the transmission mission. s UAV a The energy consumption of the transmission task is expressed as:
[0097]
[0098] Subtask v g,n In low-orbit satellite L s When performing the calculation, the resulting delay is expressed as:
[0099]
[0100] in, Indicates low-orbit satellite L s For subtask v g,n allocated computing resources; Indicates low-orbit satellite L s The number of CPU cycles required for the execution of the current subtask;
[0101] Subtask v g,n In low-orbit satellite L s When performing the calculation, the resulting energy consumption is expressed as:
[0102]
[0103] Next, calculate the subtask v g,n Preparation delay on each node; subtask v g,n The preparation delay of the predecessor task Completion delay and calculation result return delay If the predecessor task and the successor task are on the same device, the calculation result return delay is 0; to better represent the preparation delay, define the current subtask v g,n Prerequisites The collection is Then the current subtask v g,n The preparation delay is expressed as:
[0104]
[0105] Therefore, subtask v g,n On the ground user local GU g The completion time during execution is:
[0106]
[0107] Subtask v g,n On the ground user local GU g The energy consumption during execution is:
[0108]
[0109] Subtask v g,n In UAV U a The completion time during execution is:
[0110]
[0111] Subtask v g,n In UAV U a The energy consumption during execution is:
[0112]
[0113] Subtask v g,n In low-orbit satellite L s The completion time during execution is:
[0114]
[0115] Subtask v g,n In low-orbit satellite L s The energy consumption during execution is:
[0116]
[0117] Therefore, the task ψ g The total delay is expressed as:
[0118]
[0119] Among them, Q represents the set of exit tasks;
[0120] Task ψ g The total energy consumption is expressed as:
[0121]
[0122] In order to meet the user mission requirements in the space-ground integrated network while minimizing the system cost as much as possible, a cost function of the normalized total delay and total energy consumption of the joint mission is constructed, which is expressed as:
[0123]
[0124] Among them, T max , E max are the user's maximum tolerable delay and the system's maximum energy consumption, respectively; μ1 and μ2 are the weights of delay and energy consumption, respectively; by finding a joint strategy consisting of task offloading decision and computing resource allocation, the cost function is optimized to minimize the task's delay and the system's energy consumption; the joint optimization problem is expressed as:
[0125]
[0126] Among them, A n represents the uninstallation strategy, J = {J1, J2, J3, ..., J a} indicates drone U a location, Indicates that ground users, drones, and low-orbit satellites are subtasks v g,n allocated computing resources;
[0127] C1 represents the user subtask v g,n It is a binary uninstall;
[0128] C2 represents the user subtask υ n It can only be executed at one location, either locally on the user, on a drone, or in a low-orbit satellite;
[0129] C3, C4, and C5 represent ground users, drones, and low-orbit satellites as subtasks v g,n The total amount of allocated computing resources cannot exceed the amount of computing resources allocated to the system.
[0130] C6 means that the total delay of task completion must be lower than the user's maximum tolerance delay;
[0131] C7 indicates that the energy consumption of the task must be lower than the maximum energy consumption of the system;
[0132] C8 represents the flight range limit of the drone, where x a 、y a Indicates the horizontal and vertical coordinates of the drone. The flight altitude of the drone is fixed at h a ;
[0133] C9 ensures that multiple drones maintain a certain distance between each other to avoid collisions.
[0134] See also Figures 4-5, which is a GAT encoding flow chart provided by an example of the present invention and a multi-head attention mechanism calculation diagram of a feature vector provided by an example of the present invention, using GAT to encode the DAG of the user task;
[0135] Ground User GU g Task ψ g The DAG is represented as In the DAG, for each subtask v g,n Define a feature vector To represent its properties and status, where c i Indicates the current node type, which is divided into entry task, normal task, and exit task; q i Indicates the number of predecessor tasks of the current task, q i It will be updated with each network update; d represents the dimension of the feature vector;
[0136] The feature vectors of all subtasks are stacked by nodes to form the ground user task ψ g The characteristic matrix The nth column of matrix X corresponds to node v g,n The eigenvector x g,n ;Through the feature matrix, the node attributes of the entire task DAG can be expressed;
[0137] In order to express the task DAG more accurately, it is necessary to perform a learnable linear transformation on the feature matrix, apply a shared linear transformation to each subtask, and define a weight matrix controlled by the parameter θ To parameterize the linear transformation; node v g,n The transformed eigenvector is:
[0138] e g,n =W θ x g,n , (38)
[0139] in,
[0140] At node v g,n The shared attention mechanism is performed on To calculate the node v g,n For Node Attention coefficient:
[0141]
[0142] in, is a learnable attention weight vector controlled by parameter θ; Represents the node v g,m For node v g,nThe importance of ; the symbol || represents the vector splicing operation; if the node is an entry node, it will be spliced with itself;
[0143] The masked attention mechanism is introduced to restrict nodes to only pay attention to their first-order neighbor nodes, and assumes Represents node v g,n The first-order neighbor nodes are combined;
[0144] Through the masked attention mechanism, the model can more accurately capture and utilize the structural features of DAG;
[0145] In order to make the attention coefficients comparable between different nodes, the Softmax function is used to normalize the attention coefficients to obtain the normalized attention weights, which are expressed as:
[0146]
[0147] Through the attention mechanism, different importance weights can be implicitly assigned to different nodes in the neighborhood of each node;
[0148] Then, node v g,n The output feature vector of It can be obtained by weighted summation of the attention weights of all first-order neighbor nodes:
[0149]
[0150] Among them, σ is a nonlinear activation function, and this paper uses the exponential linear unit;
[0151] To further improve the stability of the model and its ability to express DAG, a multi-head attention mechanism is used to independently execute K attention heads at each node. Each attention head calculates its own node feature vector and then concatenates them into node v g,n The final output feature vector;
[0152] The final output feature vector is expressed as:
[0153]
[0154] After the multi-head attention mechanism, the feature dimension contained in the feature vector is converted from the original N′ to KN′; when the multi-head attention mechanism is applied to the final output layer, the output feature vector will be the average of the feature vectors of each node instead of splicing, thus ensuring that the feature dimension of the feature vector output by the output layer is N′. The feature vector output by the output layer is expressed as:
[0155]
[0156] in, Contains subtask v g,n All important features of the DAG are input into the decision module of PPO; the encoding method design of GAT for DAG is shown in Table 1;
[0157] Table 1
[0158]
[0159]
[0160] See also Figure 6 This is a flowchart of the GAT-PPO method provided by an example of the present invention. This method inputs the feature vector as the state into the PPO deep reinforcement learning network, constructs an Actor network with dual output heads, with the discrete head outputting subtask offloading decisions and the continuous head outputting the computing resource ratio. A deep neural network is used to enhance PPO's learning of states and actions.
[0161] Specifically, the user task offloading process is modeled as an MDP. First, the input sequence of tasks needs to be prioritized to ensure that the predecessor task of each offloaded subtask has been offloaded. The priority expression formula is defined as:
[0162]
[0163] Among them, the definition of priority adopts a recursive method to sort from the exit task, succ(v g,n ) represents subtask v g,n The post-task set; the prioritized subtasks will pass through the PPO Actor network in order of priority;
[0164] The state space, action space and reward function of PPO are defined as follows:
[0165] State Space: The input state of the algorithm consists of three parts, represented as:
[0166] s t =[x′ g,n ||c t ||o t ], (45)
[0167] in, It is the feature vector obtained by the GAT module of the user subtask, which includes all the features of the current subtask and its related first-order neighbor tasks; It is the current remaining computing resources of the MEC server, which is used to measure the maximum computing resources that can be obtained by offloading the subtask; It is a collection of subtasks that have been offloaded, including the offloading strategy and the allocated computing resources;
[0168] Action space: To achieve joint optimization of task offloading and computing resources, the action space is designed as a hybrid action space, consisting of a discrete action and a continuous action. The action space of the PPO algorithm is expressed as:
[0169]
[0170] Among them, discrete actions The offloading target selected for the current subtask is one-hot encoded, and the agent executes the task locally or offloads it to a device such as a drone or a low-orbit satellite.
[0171] Continuous Action Indicates the proportion of computing resources allocated to the device providing computing services for this subtask;
[0172] Reward function: The goal of the optimization problem is to find the optimal task offloading and resource allocation strategy to minimize the normalized cost function. The reward function is defined as:
[0173] R=-C; (47)
[0174] Next, PPO assigns the user subtask v g,n The feature vector obtained after processing by the GAT module, the unloading status of the current user task DAG and the status information of the MEC server in the environment are used as the state s t Input into the Actor network;
[0175] First, after being processed by the Actor network shared feature layer, the features in the state are extracted and expressed as:
[0176]
[0177] in, is the parameter of the shared hidden layer controlled by the parameter θ. This paper adopts the Actor network with a two-layer MLP structure;
[0178] Next, two action heads are output in parallel, and the user subtask offloading strategy output by the discrete head is expressed as:
[0179]
[0180] Among them, W d , b d are the parameters of the discrete action head controlled by the parameter θ;
[0181] The ratio of computing resources allocated by MEC to this subtask output by the continuous head is expressed as:
[0182]
[0183] Among them, W c , b c are the parameters of the continuous action head controlled by the parameter θ;
[0184] In order to obtain the optimal policy output, the Critic network is used to train and optimize the Actor network. The number of hidden layers and input of the Critic network are consistent with those of the Actor network. The output is the estimated value of the current state, which is expressed as:
[0185] V(s t ;θ)=f θ (s t ), (51)
[0186] where f θ () is a function of the network parameter θ;
[0187] Let π θ (a|s t ) represents the action probability output by the current Actor network, Represents the output probability of the old Actor network, and defines the strategy ratio as:
[0188]
[0189] And by designing a clipping objective function to limit the magnitude of each policy update, prevent the policy update from being too large, and thus maintain the stability of training; the clipping objective function is expressed as:
[0190]
[0191] Among them, ∈ is the clipping threshold hyperparameter, which limits r(θ) to [1-∈, 1+∈]; A t is the advantage function, which is expressed using the generalized advantage estimate:
[0192]
[0193] Where γ is the discount factor and λ is the attenuation coefficient of the advantage function, which can adjust the bias-variance trade-off. The advantage function can measure how good the current action is compared to the average action.
[0194] Finally, the optimization objective function of PPO can be expressed as:
[0195] L PPO (θ)=L clip (θ)-c υ L value (θ)+c s H(πθ ), (55)
[0196] Among them, c v , c s Represents the hyperparameters of each loss function; H(π θ ) represents the entropy of the strategy, which can increase the exploratory nature of the algorithm and avoid premature convergence; L value (θ) is the mean square error loss of the Critic network, expressed as:
[0197]
[0198] The pseudo code of the proposed algorithm is shown in Table 2.
[0199] Table 2
[0200]
[0201]
[0202] See also Figure 7 A structural diagram of a terminal device provided for an example of the present invention, the terminal includes a processor 701 and a memory 702, the memory 702 is coupled to the processor 701, the memory is used to store computer program code, the computer program code includes computer instructions, when the processor reads the computer instructions from the memory, so that the electronic terminal executes a joint optimization method for air-space-ground edge computing network task offloading and resource allocation.
[0203] See also Figure 8 1 is a schematic structural diagram of a device provided by an embodiment of the present invention, comprising:
[0204] The first processing unit 801 is configured to construct a user task into a DAG with interrelated subtasks;
[0205] A second calculation unit 802 is used to calculate the preparation delay, calculation delay and transmission delay of the user task DAG;
[0206] The third calculation unit 803 is used to calculate the computing energy consumption and transmission energy consumption of the user task DAG;
[0207] The fourth processing unit 804 is configured to encode the user task DAG to generate a feature vector containing first-order neighbor information;
[0208] A fifth acquisition unit 805 is configured to collect computing resources of each computing node in the current environment and a feature vector of the user task after GAT encoding;
[0209] A sixth output unit 806 is configured to process the environment information and the user task feature vector to obtain a user task offloading strategy and a computing node resource allocation strategy;
[0210] An example of the present invention also provides a computer-readable storage medium, which may include a computer program or instructions. When the computer program or instructions are run on a computer, the computer is enabled to execute the graph-based joint optimization method of task offloading and resource allocation in the space-ground edge computing network described in the above example.
[0211] An example of the present invention also provides a computer program product, including a computer program or instructions. When the computer program or instructions are run on a computer, the computer is enabled to execute the graph-based joint optimization method of task offloading and resource allocation in the space-ground edge computing network described in the above example.
[0212] The above examples may be implemented in whole or in part using software, hardware, firmware, or any combination thereof. When implemented using a software program, the above examples may appear in whole or in part in the form of a computer program product, which includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the examples of this application are generated in whole or in part.
[0213] It should be understood that the proposed systems, devices and methods can be implemented in other ways. The device examples described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0214] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of this example method.
[0215] Finally, it should be noted that the above examples are only used to illustrate the technical method of the present invention and are not limiting. Although the present invention has been described in detail with reference to preferred examples, ordinary technicians in this field should understand that the technical method of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical method, which should be included in the scope of the claims of the present invention.
Claims
1. A joint optimization method for task offloading and resource allocation in air-space-ground edge computing networks, characterized in that: The method comprises: S1: In the air-ground-space edge computing network, the network models of air-based network, space-based network and ground network, user-associated task model, communication model and computing model are constructed respectively; S2: Calculate the total delay and total energy consumption required to complete the task through the model; construct a normalized cost function, and use the normalized cost function to measure the joint indicator of user task delay and system energy consumption; S3: Under the joint constraints of user task latency and overall system energy consumption, the problem is formulated as a joint non-convex optimization problem, and the normalized cost function is used as the objective function of the optimization problem; S4: Optimal solutions are obtained using a joint decision-making approach using a Graph Attention Network (GAT) and Proximal Policy Optimization (PPO). GAT encodes the user task directed acyclic graph (DAG) and extracts node feature vectors containing task characteristics through a multi-head attention mechanism. S5: Input the feature vector as the state into the PPO deep reinforcement learning network, build an Actor network with dual output heads, and use a deep neural network to enhance PPO's learning of states and actions, so that the agent can gradually minimize the total task delay and total system energy consumption during the interaction with the environment.
2. The method for joint optimization of task offloading and resource allocation in air-space-ground edge computing network according to claim 1 is characterized in that: The step S1 is specifically as follows: Consider the air-ground mobile edge computing network scenario under disaster scenarios, including ground network, space-based network, and air-based network. Specifically, there are U ground users GU in the ground network. g ,in Each ground user GU g The device will generate 1 task, so there are G=U tasks in total. Assume There are A drones U in the space-based network a ,in These drones will cover all users in the area; there is a low-orbit satellite L in the air-based network. s , L s To interact with all drones on the air-based platform, communications between ground users and low-orbit satellites need to be relayed by drones; In actual scenarios, user tasks are composed of subtasks, which are interrelated and interdependent. Some subtasks can only be executed after other subtasks are completed. The execution of tasks must follow a certain order. g The resulting task is denoted as ψ g , assuming Where W g Indicates the size of user task input data, F g Indicates the number of CPU cycles required to process the task; g Modeled as DAG, DAG can provide dependencies between tasks and use structured representation to improve the efficiency of task scheduling and resource allocation, which is expressed as The vertex set represents a subtask, and N is the total number of subtasks in DAG, v g,n Indicates ground user GU g The bth subtask in the task DAG, and assuming In the air-ground mobile edge computing network, Orthogonal Frequency Division Multiple Access (OFDMA) technology is used between users and drones, between drones, and between drones and low-orbit satellites; drones use the L band to communicate with ground users, use the C band to communicate with adjacent drones, and low-orbit satellites use the Ka band to communicate with drones; in the ground-ground link, due to the Doppler effect caused by the rapid movement of drones in the space-based network, the small-scale fading of the ground-ground link occurs. Rayleigh fading is usually used to model the communication link between the ground and the air, and the Shannon formula of Rayleigh fading is used to calculate the ground user GU g With UAV a The uplink and downlink data transmission rates between them are: In the space link, due to the long distance and complex environment of space communication, the free space path loss is greater, and the signal propagation is also affected by rain fading. The communication channel model of the space link is modeled as a channel with small-scale fading and Ricean fading. The Shannon formula of Ricean fading is used to calculate the UAV U a With low-orbit satellite L s The uplink and downlink transmission rates of the two UAVs are affected by free space path loss and small-scale Rayleigh fading. The Shannon formula of Rayleigh fading is used to calculate the UAV U a With UAV a′ The link transmission rate between them; In the described air-space-ground mobile edge computing network, users, drones, and low-orbit satellites will perform computational processing on subtasks, use the ratio of task data size to transmission rate to construct transmission delay, and use the product of transmission power and transmission delay to construct transmission energy consumption; use the ratio of the number of CPU cycles required for task execution to the computing resources allocated by the mobile edge service to construct computing delay, and use the product of the number of CPU cycles required for task execution and the square of the computing resources allocated by the mobile edge service to construct computing energy consumption.
3. The method for joint optimization of task offloading and resource allocation in air-space-ground edge computing networks according to claim 1, characterized in that: The step S2 is specifically as follows: To determine the total delay of the task, it is necessary to calculate the subtask v g,n Preparation delay at each node; the subtask v g,n The preparation delay of the predecessor task Completion delay and calculation result return delay If the predecessor task and successor task are on the same device, the calculation result return delay is 0; Therefore, the execution delay of ground users is expressed as the sum of preparation delay and calculation delay, and the execution energy consumption is expressed as the sum of calculation energy consumption and transmission energy consumption. The execution delay of UAVs is expressed as the sum of preparation delay, calculation delay and transmission delay, and the execution energy consumption is expressed as the sum of calculation energy consumption and transmission energy consumption. The execution delay of low-orbit satellites is expressed as the sum of preparation delay, calculation delay and transmission delay, and the execution energy consumption is expressed as the sum of calculation energy consumption and transmission energy consumption. The total mission delay is the longest mission completion delay of all exit missions, and the total energy consumption is the sum of the execution energy consumption of ground users, drones, and low-orbit satellites; The total delay and total energy consumption are normalized and multiplied by their respective weight factors, and the sum is added to obtain a normalized cost function.
4. The method for joint optimization of task offloading and resource allocation in air-space-ground edge computing networks according to claim 1, characterized in that: The step S3 is specifically as follows: The normalized cost function is used as the objective function of the non-convex optimization problem, and the non-convex optimization problem is constructed to minimize the normalized cost function by optimizing the offloading strategy, drone position, and resource allocation of each computing node; and the user task offloading decision constraints, computing resource allocation constraints, maximum tolerable delay constraints, maximum tolerable energy consumption constraints, and drone flight range constraints are constructed.
5. The method for joint optimization of space-ground edge computing network task offloading and resource allocation according to claim 1, characterized in that: The step S4 is specifically as follows: GAT is used to encode the user task DAG. First, a feature vector is defined for each subtask of the DAG. The attention coefficient of the node is then calculated using the weight matrix and shared attention mechanism. The Softmax function is then used to normalize the attention coefficient. Finally, the multi-head attention mechanism is used to obtain the final output feature vector of the DAG.
6. The method for joint optimization of task offloading and resource allocation in air-space-ground edge computing network according to claim 1, characterized in that: The step S5 is specifically as follows: The user task offloading process is modeled as a Markov decision process. The input sequence of tasks is prioritized and recursively sorted starting from the exit task. The prioritized subtasks are sequentially passed through the PPO Actor network according to their priority. The PPO algorithm state space, action space, and reward function are constructed. The feature vector obtained after the subtask is processed by the GAT module, the offloading status of the current user task DAG, and the status information of the mobile edge computing server in the environment are input into the Actor network as state. The Actor network outputs a mixed discrete and continuous action using a multi-layer perceptron structure with a double head and two layers. A Critic network is constructed to train and optimize the Actor network. A clipping objective function is designed to limit the magnitude of each policy update, preventing excessive policy updates and thus maintaining training stability.
7. A graph-based joint optimization device for task offloading and resource allocation in an air-space-ground edge computing network, characterized in that: include: A first processing unit is used to construct a user task into a DAG with interconnected subtasks; A second calculation unit is used to calculate the preparation delay, calculation delay and transmission delay of the user task DAG; A third calculation unit, configured to calculate the computing energy consumption and transmission energy consumption of the user task DAG; a fourth processing unit, configured to encode the user task DAG to generate a feature vector containing first-order neighbor information; A fifth acquisition unit is used to collect the computing resources of each computing node in the current environment and the feature vector of the user task after GAT encoding; The sixth output unit is used to process the environmental information and the user task feature vector to obtain a user task offloading strategy and a computing node resource allocation strategy.
8. A terminal device, characterized in that: The terminal device includes a processor and a memory, and the processor and the memory are connected to each other, wherein the memory is used to store a computer program, and the computer program includes program instructions. The processor is specifically used to call the computer program from the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that The computer storage medium stores computer-readable instructions, and when the computer-readable instructions are executed on the communication device, the communication device is caused to execute the method according to any one of claims 1 to 6.
10. A computer program product, characterized in that When the computer program product is run on a computer, the computer is caused to execute the method according to any one of claims 1 to 6.
Citation Information
Cited By
TD3-based cross-domain ocean network task unloading and computing resource allocation method
CN121985036A