UAV-Assisted Vehicular Network Task Offloading Strategy Based on GAT-DDPG
The GAT-DDPG algorithm optimizes the unloading strategy of unmanned aerial vehicle networking tasks, which solves the complexity of task unloading in multiple intelligent scenarios, reduces the total delay and edge computing delay, and improves the computing power and reliability of the system.
Patent Information
- Application Number
- CN202410788719.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-19
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2044-06-19
AI Technical Summary
The existing task offloading strategies fail to effectively deal with the complex problems of multiple agents and fail to reduce the latency of edge computing tasks. Deep reinforcement learning networks lack appropriate attention and weighting in multi-agent environments.
The unloading strategy of unmanned aerial vehicle-assisted vehicle networking task based on GAT-DDPG is adopted. By setting the objective function of minimizing total delay, combined with the graph attention network and deep deterministic strategy gradient algorithm, the task unloading strategy in multi-vehicle and multi-intelligent scenarios is optimized.
It has realized the optimization of vehicle tasks offloading in complex scenarios of multiple vehicles and multiple agents, reducing the total delay and reducing the delay of edge computing tasks, and improving the reliability and computing capabilities of the system.
Smart Images

Figure CN118574161B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of vehicle networking, and particularly relates to a drone-assisted vehicle networking task offloading strategy based on GAT-DDPG. Background Art
[0002] The application of artificial intelligence technology in vehicles has high requirements for computing latency and reliability. At the same time, on-vehicle computing resources are relatively limited. Therefore, intelligent vehicles need to transfer some or all of the computing tasks to other devices with sufficient computing power and high performance for processing. This process of transferring computing tasks is called task offloading.
[0003] The current task offloading strategies do not consider the handling of complex problems involving multiple agents. On the other hand, the current task offloading strategies do not consider reducing the latency of edge computing tasks. Deep reinforcement learning networks often lack proper attention and weighting to important elements in the environment and tend to focus on the interaction between a single agent and the environment, which leads to challenges in dealing with complex problems involving multiple agents. Therefore, the existing task offloading strategies need to be improved. Summary of the Invention
[0004] To solve the above problems in the prior art, that is, the problems of handling multiple agents and the latency of edge computing tasks, the present invention provides a drone-assisted vehicle networking task offloading strategy based on GAT-DDPG, which can optimize the vehicle task offloading in a complex scenario of multiple vehicles and multiple agents, reduce the total latency, and can also reduce the latency of edge computing tasks.
[0005] In the first aspect of the present invention, a drone-assisted vehicle networking task offloading strategy based on GAT-DDPG is proposed. Determining the task offloading strategy includes: setting an objective function for minimizing the total latency according to the estimated time for vehicle local task processing, the estimated total execution time of drone tasks, the estimated total execution time of edge server tasks, and the task offloading strategy combination; the constraint conditions of the objective function include: restricting the time for the drone to process tasks not to exceed the time the vehicle stays within the drone service coverage area, at the same time slot, the vehicle can select at most one edge server or drone for task offloading, and the offloading ratio is from 0 to 1; using the GAT-DDPG algorithm to determine the optimal solution of the objective function and obtain an optimized task offloading strategy.
[0006] In the second aspect of the present invention, a UAV-assisted vehicle network task offloading strategy is proposed. Determining the task offloading strategy includes: setting an objective function according to the estimated local task processing time of the vehicle, the estimated total execution time of the UAV tasks, the estimated total execution time of the edge server tasks, and the task offloading strategy combination; the constraint conditions of the objective function include: restricting the time for the UAV to process tasks so that it cannot exceed the time the vehicle stays within the UAV service coverage area, at the same time slot, the vehicle can select at most one edge server or UAV for task offloading, and the offloading ratio is from 0 to 1; determining the optimal solution of the objective function to obtain an optimized task offloading strategy.
[0007] Specifically, the objective function is specifically:
[0008]
[0009] Among them, , ; represents the set of task offloading strategy sets A and the set of task offloading ratio sets P ; represents the total computing resources of the UAV and the edge server; , represents the set of vehicles, n represents the vehicle number, N represents the total number of vehicles; , represents the set of UAVs, m represents the UAV number, M represents the total number of UAVs; , represents the vehicle n offloading the task to the UAV m task offloading strategy, represents the vehicle n offloading the task to the edge server task offloading strategy, the task offloading strategy being 1 means offloading, and the task offloading strategy being 0 means not offloading; , represents the vehicle n offloading the task to the UAV m ratio, represents the vehicle n offloading the task to the edge server ratio; represents the UAV m within the service cycle total computing resources; represents the total computing resources of the edge server within the service cycle ; , represents the service cycle,t Indicates the time slot number, T Indicates the total number of time slots, and each time slot is a service cycle of ; Indicates the total estimated delay within the time slot t , , Indicates the estimated time for the vehicle to process local tasks within the time slot t ; n Indicates the total estimated execution time of the drone task for the vehicle to unload to the drone within the time slot ; t Indicates the total estimated execution time of the edge server task for the vehicle to unload to the edge server within the time slot; n ; m Indicates the residence time of the vehicle within the service coverage of the drone ; t Indicates the maximum tolerable delay of the vehicle task n ; Indicates the computing resources allocated to the drone within the time slot n ; m Indicates the computing resources allocated to the edge server within the time slot ; n Indicates the computing resources allocated by the drone to the vehicle within the time slot ; m Indicates the computing resources allocated by the edge server to the vehicle within the time slot t ; Indicates the amount of data for the vehicle to unload tasks within the time slot t ; Indicates the total amount of data for the vehicle to unload tasks within the service cycle t ; m ; n ; ; t ; n ; ; t ; n ; ; ; n ;
[0010] The beneficial effects of the above-mentioned drone-assisted vehicle networking task offloading strategy based on GAT-DDPG include:
[0011] (1) It can optimize the vehicle task offloading delay problem in scenarios involving multiple vehicles and multiple agents, obtain the vehicle task offloading optimization strategy in complex scenarios of multiple vehicles and multiple agents, reduce the total delay, and can also reduce the delay of edge computing tasks.
[0012] (2) A framework combining the Graph Attention Network (GAT) and the Deep Deterministic Policy Gradient (DDPG) is adopted. GAT can better understand the communication and computing resource allocation among agents. By utilizing the superior feature extraction performance of GAT on multi-vehicle multi-agent topological data and inputting the extracted features into DDPG, the feature extraction efficiency and fitting efficiency of GAT-DDPG can be effectively improved. Compared with a pure DDPG network, it is more suitable for determining task offloading strategies in complex multi-vehicle multi-agent scenarios, effectively reducing the latency of edge computing tasks, and providing a feasible and effective guarantee for realizing a multi-UAV-assisted vehicle-to-everything (V2X) system.
[0013] The beneficial effects of the above-mentioned UAV-assisted V2X task offloading strategy include: being able to optimize the vehicle task offloading latency problem in scenarios involving multiple vehicles and multiple agents, obtaining an optimized vehicle task offloading strategy in complex multi-vehicle multi-agent scenarios, reducing the total latency, and being able to reduce the latency of edge computing tasks.
[0014] Other features and advantages of the present invention will be described in the following specification. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings. Brief Description of the Drawings
[0015] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objectives, and advantages of the present application will become more apparent:
[0016] Figure 1 is a flowchart for determining the UAV-assisted V2X task offloading strategy based on GAT-DDPG provided by an embodiment of the present invention;
[0017] Figure 2 is a schematic diagram of the application framework of the GAT-DDPG algorithm in complex multi-vehicle multi-agent scenarios provided by an embodiment of the present invention;
[0018] Figure 3 is a training convergence comparison diagram of the GAT-DDPG algorithm when applied in complex multi-vehicle multi-agent scenarios provided by an embodiment of the present invention;
[0019] Figure 4 is the average system latency curve when the number of UAVs is different;
[0020] Figure 5 is a flowchart for determining a UAV-assisted V2X task offloading strategy provided by an embodiment of the present invention. Detailed Description of the Embodiment
[0021] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for ease of description, only the parts related to the relevant invention are shown in the drawings.
[0022] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.
[0023] An embodiment of the present invention provides a UAV-assisted vehicle network task offloading strategy based on GAT-DDPG. As Figure 1 shown, determining the aforementioned task offloading strategy includes steps S1 and S2:
[0024] S1: Set an objective function for minimizing the total delay according to the estimated time for local task processing of the vehicle, the estimated total execution time of the UAV tasks, the estimated total execution time of the edge server tasks, and the task offloading strategy combination. The constraint conditions of the objective function include: restricting the time for the UAV to process tasks so that it cannot exceed the time the vehicle stays within the UAV service coverage area; at the same time slot, the vehicle can select at most one edge server or UAV for task offloading, and the offloading ratio is from 0 to 1.
[0025] S2: Use the GAT-DDPG algorithm to determine the optimal solution of the objective function and obtain an optimized task offloading strategy.
[0026] In this embodiment, by establishing a system model, the problem of optimizing the UAV-assisted vehicle network edge computing offloading strategy is transformed into a multi-agent optimization problem, which can optimize the vehicle task offloading delay problem in scenarios involving multiple vehicles and multiple agents, obtain the vehicle task offloading optimization strategy in a complex scenario of multiple vehicles and multiple agents, reduce the total delay, and can also reduce the delay of edge computing tasks.
[0027] DDPG is a strategy optimization algorithm in reinforcement learning. Based on DQN (Deep Q-Network), DDPG adds support for continuous action spaces and is suitable for more complex control tasks. However, in complex topology scenarios, the feature extraction efficiency and fitting efficiency of DDPG are relatively low, so DDPG is not suitable for direct application in complex scenarios of multiple vehicles and multiple agents. In the application scenario of the present invention, tasks will be offloaded from vehicles to UAVs or edge servers, and there is no situation where UAVs or edge servers offload tasks to vehicles. Therefore, the vehicle is the task sender, and the UAV or edge server is the task receiver.
[0028] The complex scenario of multiple vehicles and multiple agents is as Figure 2As shown in the multi-UAV-assisted vehicle networking environment, multiple UAVs provide computing offloading services for freely moving vehicles at a unified height above the road. The position of each UAV can be dynamically adjusted according to requirements, and each is equipped with a server with certain computing resources. In addition to UAVs, edge servers deployed on the roadside can also provide computing offloading services for vehicles. The coverage range of each UAV for vehicles is the same, and the coverage radius is R , using to represent the position of UAV m , and to represent the ground vertical projection position of UAV m , H represents the height of UAV m , and the service period of the UAV-assisted vehicle networking system is discretized into T time slots with a length of . Use to represent the position of the vehicle in time slot t , , represents the service period, t represents the time slot number, T represents the total number of time slots. Freely flowing vehicles conform to a Poisson random process. The speed of vehicle n within the service coverage of UAV m is , n represents the vehicle number, m represents the UAV number. The residence time n of vehicle m within the service coverage of UAV is , that is, . The distance n between vehicle m and UAV t in time slot is represented as . UAV services include wireless transmission and computing of offloading tasks.
[0029] In this system, the communication between UAVs and vehicles is line-of-sight communication. In some specific embodiments, determining the transmission rate at which a vehicle offloads a task to a UAV includes:
[0030] By , determine the transmission rate t at which vehicle n offloads a task to UAV m in time slot , where represents the distance between vehicle t and UAV n in time slot mThe channel bandwidth between represents a time slot t the wireless signal transmission power of the vehicle n within represents the shadow fading component following a gamma distribution represents the wireless signal transmission path loss exponent represents a time slot t the vehicle n and the drone m the distance between represents the natural noise power.
[0031] The location of the edge server deployed in the drone-assisted vehicle networking system is represented by In the time slot t the vehicle n and the drone m the distance between is expressed as . The communication between the vehicle and the edge server is line-of-sight communication. The data transmission rate for offloading tasks from the vehicle n to the edge server through the wireless link is as follows: represents the time slot t the vehicle n and the channel bandwidth between the edge server represents the small-scale fading parameter.
[0032] To overcome the defect that DDPG has low feature extraction efficiency and fitting efficiency in the complex scenario of multi-vehicle and multi-agent, this embodiment combines GAT and DDPG and applies them to the vehicle task offloading scenario of multi-vehicle and multi-agent. GAT is a neural network architecture based on graph-structured data. By stacking layers, nodes can participate in the features of neighbors, and different weights can be implicitly assigned to different nodes in the neighborhood without any costly matrix operations (such as inversion), nor the need to know the structure of the graph in advance. The deep deterministic policy gradient algorithm with a graph attention network can effectively capture the cooperation mode between multi-agents, thereby obtaining a computational offloading strategy that minimizes the system delay.
[0033] Therefore, by adopting a framework combining GAT and DDPG, GAT can better understand the communication and computing resource allocation among agents. Using GAT to learn the multi-vehicle multi-agent node topology data can improve the feature extraction ability. Inputting the extracted features into the DDPG network to learn the optimal task offloading strategy can effectively improve the feature extraction efficiency and fitting efficiency of GAT-DDPG. Compared with a pure DDPG network, it is more suitable for determining the task offloading strategy in a complex multi-vehicle multi-agent scenario, effectively reducing the latency of edge computing tasks, and providing a feasible and effective guarantee for the realization of a multi-UAV-assisted vehicle networking system.
[0034] In some specific embodiments, the objective function in the above embodiment is:
[0035]
[0036] where , ; represents the set of task offloading strategies A and the set of task offloading ratios P ; represents the total computing resources of the UAVs and the edge server; , represents the set of vehicles, n represents the vehicle number, N represents the total number of vehicles; , represents the set of UAVs, m represents the UAV number, M represents the total number of UAVs; , represents the task offloading strategy for vehicle n to offload tasks to UAV m , represents the task offloading strategy for vehicle n to offload tasks to the edge server, being 1 means vehicle n offloads tasks to UAV m , being 1 means vehicle n offloads tasks to the edge server, and the task offloading strategy being 0 means not offloading; , represents the ratio of vehicle n to offload tasks to UAV m , represents the ratio of vehicle nThe proportion of offloading tasks to the edge server, where the offloading proportion is between 0 and 1. An offloading proportion of 1 means that the task is completely offloaded to the drone or edge server for execution. An offloading proportion greater than 0 and less than 1 means that part of the task is offloaded to the drone or edge server for execution, and the other part of the task is executed locally. An offloading proportion of 0 means that the task is completely executed locally on the vehicle.
[0037] In the objective function, represents the total computing resources of the drone m within the service period ; represents the total computing resources of the edge server within the service period ; When the vehicle is not within the service coverage of the drone, it can establish a wireless signal link with the edge server and offload the computing task to the edge server for processing.
[0038] In the objective function, , represents the service period, t represents the time slot number, T represents the total number of time slots, and each time slot is of the service period represents the total estimated delay within the time slot t , , represents the estimated time for local task processing of the vehicle t within the time slot n , represents the total estimated execution time of the drone tasks offloaded from the vehicle t to the drone n within the time slot m , represents the total estimated execution time of the edge server tasks offloaded from the vehicle t to the edge server within the time slot n ; It considers the delay situation of one vehicle to multiple drones, and the delay situation of multiple vehicles to multiple drones, and also considers the delay situation of multiple vehicles to the edge server; represents the residence time of the vehicle within the service coverage of the drone n ; m represents the maximum tolerable delay of the vehicle task; n represents the computing resources allocated to the drone within the time slot m , t involving the delay situation of one drone to multiple vehicles; represents the time slot t within which the drone m allocates to the vehiclen Computing resources . Indicates the computing resources allocated to the edge server within the time slot t involving the latency situation of the edge server for multiple vehicles; Indicates the time slot t The computing resources allocated by the edge server to the vehicle within n . .
[0039] In the above objective function, Indicates the amount of data of the offloading task of the vehicle within the time slot t n . Indicates the total amount of data of the offloading task of the vehicle within the service period n .
[0040] During the entire service period of the UAV , each vehicle only offloads tasks to one edge node, that is, a UAV or an edge server. Within the time slot t , the vehicle n will generate a task, represented by the set . Indicates the central processing unit (CPU) cycles required for the vehicle t task within the time slot n .
[0041] The constraint conditions of the objective function include C1 to C7. C1 restricts that the time for the UAV to process tasks cannot exceed the time the vehicle stays within the service range of the UAV. C2 indicates that each task can be processed within its maximum allowable processing latency. C3 indicates that the vehicle can only select one edge node for task offloading and the value of the offloading decision is 0 or 1. C4 indicates that the offloading ratio is from 0 to 1. C5 and C6 limit that the computing resources allocated by the edge server and the UAV to the vehicle cannot exceed their own computing resources. C7 indicates that all computing tasks can be processed within the service period of the UAV.
[0042] The above objective function considers different offloading methods when the vehicle is within the service coverage of the UAV and when it is not within the service coverage of the UAV, considers the latency situation of one vehicle for multiple UAVs in a multi-vehicle multi-agent scenario, and the latency situation of multiple vehicles for multiple UAVs, also considers the latency situation of multiple vehicles for the edge server, and the latency situation of one UAV for multiple vehicles, and considers the overall latency situation of multiple vehicles for multiple agents, and can more accurately determine the total latency.
[0043] In some specific implementation manners, determining the total estimated execution time of the UAV task includes:
[0044] Based on the transmission delay of the vehicle unloading the task to the UAV, the computing delay of the UAV executing the vehicle unloading task, and the queue delay of the UAV executing the vehicle unloading task, the total estimated execution time of the UAV task is obtained.
[0045] In some specific embodiments, through , the time slot t for the vehicle n to unload the task to the UAV m is obtained, where the total estimated execution time of the UAV task , and among them, represents the transmission delay of the vehicle t unloading the task to the UAV n within the time slot m , represents the computing delay of the UAV t executing the vehicle m unloading task within the time slot n , represents the queue delay of the UAV t executing the vehicle m unloading task within the time slot n . The amount of task data after the UAV processes is very small compared to the amount of data of the input unloading task. Therefore, when determining , the backhaul delay does not need to be considered.
[0046] Through , the transmission delay of the vehicle t unloading the task to the UAV n within the time slot m is obtained, where represents the proportion of the vehicle n unloading the task to the UAV m , represents the amount of data of the vehicle t unloading the task within the time slot n , represents the transmission rate of the vehicle t unloading the task to the UAV n within the time slot m .
[0047] Within the time slot t , the computing delay m of the UAV n executing the vehicle unloading task is expressed as .
[0048] The queue delay t of the UAV m executing the vehicle n unloading task within the time slot is expressed as , is the time slot t in which the computing resources m allocated to the vehicle j by the UAV are located, is the time slot t in which the computing resources j required for the task unloaded from the vehicle to the UAV are located, and the vehicle j is the vehicle m ahead of the vehicle n in the task processing queue of the UAV.
[0049] In some specific embodiments, the total estimated execution time of the edge server task is obtained according to the transmission delay of the task unloaded from the vehicle to the edge server and the execution delay of the task by the edge server. For example, when the vehicle is not within the UAV service coverage area, the vehicle unloads the task to the edge server. Since the computing power of the edge server is very high, the queue delay caused by task accumulation and the feedback delay of the processing result are not considered. The total estimated execution time t of the edge server task unloaded from the vehicle n in the time slot is expressed as . represents the transmission delay t in the time slot n when the vehicle unloads the task to the edge server, represents the computing delay t in the time slot n when the edge server executes the task unloaded by the vehicle, , , represents the proportion n when the vehicle unloads the task to the edge server, represents the data volume t in the time slot n when the vehicle unloads the task, represents the data transmission rate n at which the wireless link unloads the task from the vehicle to the edge server, t represents the central processing unit (CPU) cycles n required for the vehicle task in the time slot t and n represents the computing resources
[0050] allocated to the vehicle by the edge server in the time slot k Indicates the agent number, K Indicates the total number of agents, M Indicates the total number of drones. Since the agents include k several drones and an edge server, so .
[0051] In some specific embodiments of step S2, the GAT-DDPG algorithm is adopted to determine the optimal solution of the objective function and obtain an optimized task offloading strategy. As Figure 2 shown, it includes: using a graph attention network to learn the relationship feature vectors between each agent; according to the relationship feature vectors, determining the input of the deep deterministic policy gradient network, and training to obtain an optimized task offloading strategy.
[0052] Furthermore, in some specific embodiments of step S2, in the graph attention network, the observation state of the agent is determined based on the agent position, vehicle position, computing resources, task quantity, and resources required to execute the task. For example, in time slot t the state space of agent k is expressed as:
[0053] , using and to represent the position of vehicle t in time slot n , represents the computing resources of agent k in time slot t ; the action space of agent t in time slot k is expressed as: , where is the set of offloading decisions of agent k , is the set of offloading ratios of agent k , is the set of computing resources allocated by agent k to each vehicle user. Since the optimization goal is to minimize the system delay, a negative correlation relationship is established between the reward function and the system delay.
[0054] Specifically, the reward function is:
[0055] where is the reward value of agent t in time slot k , n represents the vehicle number, N represents the total number of vehicles, t represents the time slot number, represents the time slott Intra-vehicle n Estimated time for local task processing, Indicating time slot t Intra-vehicle n Unloading to the drone m Total estimated execution time of the drone task, m Indicating the drone number, M Indicating the total number of drones, Indicating time slot t Intra-vehicle n Total estimated execution time of the edge server task for unloading to the edge server, Indicating the vehicle n Unloading task to the drone m Task unloading strategy, Indicating the vehicle n Task unloading strategy for unloading task to the edge server.
[0056] Regarding a single agent as a node to form a topological network. Obtaining the observation state t Observed by the agent node k in the time slot , , Indicating the observation space formed by the observation state, , Indicating the position information of the agent and the vehicle, , Indicating the computing resources of the agent, , Indicating the agent k Computing resources. Indicating the size of the computing task, Indicating time slot t Intra-vehicle n Data volume of the unloading task, . Indicating the computing resources required for each computing task, Indicating time slot t Intra-vehicle n CPU cycles required for the task, .
[0057] Using a multi-layer perceptron to encode the observation state of the agent k in the time slot t into a feature vector ; , where is the weight factor corresponding to the agent k in the weight matrix, is the bias vector, and ReLU represents the activation function.
[0058] Combine the feature vectors of each node to obtain the time slot t Inner observation space Feature vector representation of , .
[0059] The graph attention layer introduces attention-related coefficients for adaptive aggregation according to the importance of neighbor nodes. For nodes k and node m* , the attention weights can be calculated by the following formula: , represents the adjacent agent k and m* the attention weight between , is the set of all one-hop neighbor agent nodes of agent k , g is the adjacent agent number, is the adjacent agent k and m* attention-related coefficient, is the adjacent agent k and g attention-related coefficient, and LeakyReLU is the activation function.
[0060] The output of the graph attention layer is a new feature obtained by weighted summation of the features corresponding to the nodes according to the attention coefficients. According to , determine the feature vector output by the first-layer graph attention network layer, is the activation function, W is the weight matrix, h m*,t is the time slot t the feature vector of the observation state encoding of agent m* within; according to the feature vector output by the first-layer graph attention network layer, output the feature vector of the second-layer graph attention network layer.
[0061] According to , determine the input of the deep deterministic policy gradient network. As Figure 2 shown, the feature vector finally output by the graph attention network layer is output to the DDPG network after passing through the fully connected layer to obtain the input state of the DDPG network: , where is the time slot t the observation state of agent k within At the input state corresponding to the deep deterministic policy gradient network, is the weight factor corresponding to the agent in the weight matrix k and is the bias vector. The DDPG network includes a Critic estimation network, a Critic target network, an Actor estimation network, and an Actor target network.
[0062] When training the DDPG network, the agent obtains rewards from the environment and stores the experience tuple in the experience replay buffer . represents the input state of DDPG within the time slot t , represents the action space of the time slot t , represents the input state of DDPG in the next time slot. In the deep deterministic policy gradient network, the reward value in the experience tuple is obtained through the reward function.
[0063] Initialize the parameters of the Actor estimation network and the Critic estimation network to and . The Critic estimation network outputs , and the Critic target network outputs . Input the target action along with the experience tuple into the Critic target network to obtain the target value , , where is the parameter of the Critic target network, is the discount factor, , is the parameter of the Actor target network, is the action value approximately estimated by the Actor target network. The data output by the Critic target network is used to guide the Actor network to output a benchmark for policy improvement, that is, the Actor network will update its parameters according to the evaluation made by the Critic network, thereby improving the policy.
[0064] The Critic estimation network is updated by minimizing the squared error between the current value and the target value:
[0065] , represents the loss function with respect to the Critic estimation network, represents the mean function, represents the Q-value function of the current state-action output by the Critic estimation network, Indicates the current state, Indicates the current action.
[0066] The Actor estimation network is updated through the policy gradient algorithm:
[0067]
[0068] Indicates the gradient function of Indicates the policy, is the action function of the current state under the Indicates the gradient function of
[0069] Both the Critic target network and the Actor target network adopt soft update:
[0070]
[0071] Among them, is the soft update coefficient, generally a constant much smaller than 1, such as 0.001, etc.
[0072] To evaluate the effectiveness of the proposed task offloading strategy based on the GAT-DDPG algorithm, simulation comparison experiments are carried out. The other three task offloading strategies selected are: (1) Local Computing (LC): A task offloading strategy that completely offloads to local computing; (2) A task offloading strategy based on the Random algorithm; (3) A task offloading strategy based on the MADDPG algorithm. For the LC task offloading strategy, the task is completely executed locally and does not require training to converge. Figure 3 Shows the training convergence curves of the GAT-DDPG algorithm, the MADDPG algorithm, and the Random algorithm. The horizontal axis represents the number of training rounds, and the vertical axis represents the system reward. From Figure 3It can be seen that in the initial stage of training, the rewards of the three algorithms fluctuate at a low level because the algorithms have just started training and have not accumulated experience yet. When training for 20 to 25 rounds, the MADDPG algorithm converges rapidly. When training for 30 to 55 rounds, the GAT-DDPG algorithm converges rapidly. It can be seen that the convergence speed of the MADDPG algorithm is slightly faster than that of the GAT-DDPG algorithm. The reason for this phenomenon is that compared with the MADDPG algorithm, the GAT-DDPG algorithm is applicable to multi-agent problems with higher complexity, so the convergence speed will be slightly slower. The RANDOM algorithm always fluctuates at a very low level, indicating that its performance is the worst. From the perspective of the fluctuation amplitude after convergence, the GAT-DDPG algorithm has the smallest fluctuation amplitude, the MADDPG algorithm ranks second, and the RANDOM algorithm has the largest fluctuation amplitude, which also proves that the GAT-DDPG algorithm has the best performance compared with other baseline algorithms.
[0073] Figure 4 It records the changing trend of the average system delay per unit time slot with the increase in the number of UAVs. Figure 4 In [graph (a) above], the average system delays under different numbers of UAVs are compared. The horizontal axis represents the number of UAVs, and the vertical axis represents the average system delay per unit time slot. The computational task data volumes of different strategies are fixed at 120 megabytes for comparison. From Figure 4 graph (a) of [reference], it can be seen that the system delay of the LC offloading strategy remains unchanged all the time because the LC offloading strategy does not introduce UAVs and the tasks are all executed locally on the vehicle, so the system delay has nothing to do with the number of UAVs. The other three curves all show a downward trend. After the number of UAVs increases to 6, the system delays of the GAT-DDPG offloading strategy and the MADDPG offloading strategy both remain in a stable state, indicating that at the beginning, as the number of UAVs increases, the task computing ability is enhanced and the system delay can decrease, but when the number of UAVs increases beyond a certain value, it has no contribution to the reduction of the system delay. From Figure 4 graph (a) of [reference], it can be seen that the UAV-assisted vehicle network task offloading strategy based on GAT-DDPG provided by the embodiments of the present invention has a lower average system delay than other strategies when there are different numbers of UAVs in the system, can better meet the strict requirements of vehicles for delay, and make the system more reliable.
[0074] Figure 4 In [reference], in graph (b) below, it is the average system delay of the GAT-DDPG algorithm under different task data volumes and different numbers of UAVs. The horizontal axis represents the number of UAVs, and the vertical axis represents the average system delay per unit time slot. From Figure 4As can be seen from the curve graph (b), as the number of drones increases, the average system delay under different amounts of computing task data decreases, indicating that the computing resources provided by the drones can improve the system's task processing ability. After the number of drones is set to 6, the curve tends to be stable, which is consistent with Figure 4 the characteristics of the system delay reflected in the curve graph (a), and the larger the amount of computing task data, the higher the average system delay.
[0075] In the above embodiments, although the steps are described in the above sequential order, those skilled in the art can understand that in order to achieve the effects of this embodiment, different steps do not have to be executed in such an order, and they can be executed simultaneously (in parallel) or in a reversed order, and these simple changes are all within the protection scope of the present invention.
[0076] Another embodiment of the present invention provides a drone-assisted vehicle network task offloading strategy. As Figure 5 shown, determining the aforementioned task offloading strategy includes steps P1 and P2:
[0077] P1: Set an objective function for minimizing the total delay according to the estimated time for local task processing of the vehicle, the estimated total execution time of the drone tasks, the estimated total execution time of the edge server tasks, and the task offloading strategy combination; the constraint conditions of the objective function include: restricting the time for the drone to process tasks to not exceed the time the vehicle stays within the service coverage of the drone, and at the same time slot, the vehicle can select at most one edge server or drone for task offloading, and the offloading ratio is from 0 to 1.
[0078] P2: Determine the optimal solution of the objective function to obtain an optimized task offloading strategy.
[0079] In this embodiment, by establishing a system model, the problem of optimizing the drone-assisted vehicle network edge computing offloading strategy is transformed into a multi-agent optimization problem, which can optimize the vehicle task offloading delay problem in scenarios involving multiple vehicles and multiple agents, obtain a vehicle task offloading optimization strategy in a complex scenario of multiple vehicles and multiple agents, reduce the total delay, and can also reduce the delay of edge computing tasks.
[0080] In some specific embodiments, the objective function is specifically:
[0081]
[0082] where , ; represents the set of task offloading strategy sets A and task offloading ratio sets P ; Represents the total computing resources of the UAV and the edge server; , represents the vehicle set, n represents the vehicle number, N represents the total number of vehicles; , represents the UAV set, m represents the UAV number, M represents the total number of UAVs; , represents vehicle n offloading tasks to the UAV m task offloading strategy, represents vehicle n offloading tasks to the edge server. The task offloading strategy of 1 means offloading, and the task offloading strategy of 0 means not offloading; , represents vehicle n offloading tasks to the UAV m ratio, represents vehicle n offloading tasks to the edge server ratio; represents the UAV m total computing resources within the service period ; represents the total computing resources of the edge server within the service period ; , represents the service period, t represents the time slot number, T represents the total number of time slots. Each time slot is of ; represents the total estimated delay within the time slot t , , represents the estimated time for local task processing of vehicle t within the time slot n , represents the estimated total execution time of the UAV tasks when vehicle t within the time slot n offloads to the UAV m , represents the estimated total execution time of the edge server tasks when vehicle t within the time slot n offloads to the edge server; represents the residence time of vehicle n within the service coverage of the UAV m ; represents the maximum tolerable delay of vehicle n tasks; Denote the drone m Allocate to the time slot t The computing resources within Denote the computing resources allocated by the edge server to the time slot t The computing resources within Denote the time slot t The computing resources of the drone within m Allocated to the vehicle n The computing resources of Denote the time slot t The computing resources allocated by the edge server to the vehicle within n The computing resources of Denote the time slot t The vehicle within n The amount of data for the offloading task Denote within the service cycle The vehicle within n The total amount of data for the offloading task.
[0083] The above objective function takes into account the different offloading methods of the vehicle when it is within the service coverage of the drone and when it is not within the service coverage of the drone, considers the latency situation of one vehicle to multiple drones in a multi-vehicle multi-agent scenario, and the latency situation of multiple vehicles to multiple drones, also considers the latency situation of multiple vehicles to the edge server, and the latency situation of one drone to multiple vehicles, and considers the overall latency situation of multiple vehicles to multiple agents, and can more accurately determine the total latency.
[0084] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.
Claims
1. A method for determining a task offloading strategy of an unmanned aerial vehicle (UAV)-assisted vehicle-to-everything (V2X) network based on Graph Attention Network (GAT)-Deep Deterministic Policy Gradient (DDPG), characterized in that, The method includes: Setting an objective function for minimizing the total delay according to the estimated time for local task processing of the vehicle, the total estimated execution time of the UAV tasks, the total estimated execution time of the edge server tasks, and the combination of task offloading strategies; The constraint conditions of the objective function include: restricting the time for the UAV to process tasks to not exceed the time the vehicle stays within the UAV service coverage area, the vehicle can select at most one edge server or UAV for task offloading in the same time slot, and the offloading ratio is from 0 to 1; Using the GAT-DDPG algorithm to determine the optimal solution of the objective function and obtain an optimized task offloading strategy: Regarding each UAV and edge server as an independent agent; Using the Graph Attention Network (GAT) to learn the relationship feature vectors between each agent; where the relationship feature vector refers to the observation state of the agent encoded by the GAT, and the observation state in the GAT is obtained based on the agent's position, the vehicle's position, computing resources, the number of tasks, and the resources required to execute the tasks; Determining the input of the Deep Deterministic Policy Gradient Network (DDPG) according to the relationship feature vectors and training to obtain an optimized task offloading strategy; The objective function is specifically: Among them, represents the set of task offloading policy set A and task offloading ratio set P; represents the total computing resources of the drone and the edge server; represents the vehicle set, n represents the vehicle number, and N represents the total number of vehicles; represents the drone set, m represents the drone number, and M represents the total number of drones; represents the task offloading policy for vehicle n to offload tasks to drone m, represents the task offloading policy for vehicle n to offload tasks to the edge server. The task offloading policy of 1 means offloading, and the task offloading policy of 0 means not offloading; represents the ratio of vehicle n to offload tasks to drone m, represents the ratio of vehicle n to offload tasks to the edge server; represents the total computing resources of drone m during the service period ; represents the total computing resources of the edge server during the service period ; represents the service period, t represents the time slot number, T represents the total number of time slots, and each time slot is of T total (t) represents the total estimated delay within time slot t, represents the estimated time for vehicle n to process local tasks within time slot t, represents the total estimated execution time of the drone tasks for vehicle n to offload to drone m within time slot t, represents the total estimated execution time of the edge server tasks for vehicle n to offload to the edge server within time slot t; represents the residence time of vehicle n within the service coverage of drone m; represents the maximum tolerable delay of vehicle n's tasks; represents the computing resources allocated to drone m within time slot t, represents the computing resources allocated to the edge server within time slot t, represents the computing resources allocated by drone m to vehicle n within time slot t, represents the computing resources allocated by the edge server to vehicle n within time slot t, D n (t) represents the data volume of vehicle n to offload tasks within time slot t, D n represents within the service period the total data volume of vehicle n to offload tasks; The constraint conditions of the objective function include C1 to C7. C1 restricts the time for the UAV to process tasks to not exceed the time the vehicle stays within the UAV service range. C2 indicates that each task can be processed within its maximum allowed processing delay. C3 indicates that the vehicle can only select one edge node for task offloading and the offloading decision takes values of 0 or 1. C4 indicates that the offloading ratio is from 0 to 1. C5 and C6 limit that the computing resources allocated by the edge server and UAV to the vehicle cannot exceed their own computing resources. C7 indicates that all computing tasks can be processed within the service cycle of the UAV.
2. The method for determining the UAV-assisted vehicle network task offloading strategy based on GAT-DDPG according to claim 1, characterized in that The total estimated execution time of the UAV tasks is obtained according to the transmission delay for the vehicle to offload tasks to the UAV, the computing delay for the UAV to execute the vehicle offloaded tasks, and the queue delay for the UAV to execute the vehicle offloaded tasks.
3. The method for determining the UAV-assisted vehicle network task offloading strategy based on GAT-DDPG according to claim 2, characterized in that By obtain the total estimated execution time of the UAV task for vehicle n unloaded to UAV m within time slot t wherein represents the transmission delay for vehicle n to unload the task to UAV m within time slot t, represents the computing delay for UAV m to execute the unloading task of vehicle n, represents the queue delay for UAV m to execute the unloading task of vehicle n; By obtain the transmission delay for vehicle n to offload tasks to drone m within time slot t, where represents the proportion of tasks offloaded by vehicle n to drone m, D n (t) represents the amount of data of tasks offloaded by vehicle n within time slot t, R n,m (t) represents the transmission rate for vehicle n to offload tasks to drone m within time slot t.
4. The method for determining the UAV-assisted vehicle network task offloading strategy based on GAT-DDPG according to claim 3, characterized in that By determining R n,m (t), where B n,m (t) represents the channel bandwidth between vehicle n and drone m within time slot t, P n (t) represents the wireless signal transmission power of vehicle n within time slot t, ρ1 represents the shadow fading component following a gamma distribution, α represents the wireless signal transmission path loss exponent, d n,m (t) represents the distance between vehicle n and drone m within time slot t, η 2 represents the natural noise power.
5. The method for determining the UAV-assisted vehicle network task offloading strategy based on GAT-DDPG according to claim 1, characterized in that The total estimated execution time of the edge server tasks is obtained according to the transmission delay for the vehicle to offload tasks to the edge server and the task execution delay of the edge server.
6. The method for determining the UAV-assisted vehicle network task offloading strategy based on GAT-DDPG according to claim 1, characterized in that In the graph attention network, the observation state of the agent is determined based on the agent's position, vehicle position, computing resources, number of tasks, and resources required to execute the tasks; Encode the observation state of agent k at time slot t into a feature vector h using a multi-layer perceptron k,t ; By determining the attention weight α between adjacent agents k and m* k,m* , is the set of all one-hop neighbor agent nodes of agent k, g is the adjacent agent number, e k,m* is the attention correlation coefficient between adjacent agents k and m*, e k,g is the attention correlation coefficient between adjacent agents k and g, and LeakyReLU is the activation function; According to determine the feature vector h′ output by the first-layer graph attention network layer k,t , where σ is the activation function, W is the weight matrix, and h m*,t is the feature vector encoded by the observation state of agent m* within time slot t; According to the feature vector h' output by the first-layer graph attention network layer k,t , output the feature vector h'' output by the second-layer graph attention network layer k,t ; According to s t (o k,t ) = w k h″ k,t + b k , determine the input of the deep deterministic policy gradient network, where s t (o k,t ) is the observation state O of agent k within time slot t k,t at the input state corresponding to the deep deterministic policy gradient network, w k is the weight factor corresponding to agent k in the weight matrix, b k is the bias vector; In the deep deterministic policy gradient network, the reward value in the experience tuple is obtained through a reward function, and the reward function is: In it, r t,k is the reward value of agent k in time slot t, n represents the vehicle number, N represents the total number of vehicles, t represents the time slot number, represents the estimated local task processing time of vehicle n in time slot t, represents the total estimated execution time of the UAV task when vehicle n unloads to UAV m in time slot t, m represents the UAV number, M represents the total number of UAVs, T n,s (t) represents the total estimated execution time of the edge server task when vehicle n unloads to the edge server in time slot t, represents the task offloading strategy for vehicle n to offload tasks to UAV m, represents the task offloading strategy for vehicle n to offload tasks to the edge server.
Citation Information
Patent Citations
Calculation unloading and resource allocation method based on GAT mixed action multi-agent reinforcement learning
CN117098189A