Edge computing task unloading method and system based on unmanned aerial vehicle group

By introducing node state transition model and multi-agent MADDPG algorithm in the drone cluster edge computing system, the problem of task offloading strategies in the drone system affecting energy consumption and processing delay is solved, and efficient resource allocation and task processing are achieved.

CN120091365APending Publication Date: 2025-06-03SHENYANG AEROSPACE UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510241772.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In the existing drone-driven edge computing systems, the task offloading strategy affects UAV's energy consumption and calculation task processing delays, and the single-drone system has a slow processing speed, and the lack of collaboration between multiple drone systems leads to waste of resources.

Method used

A method of offloading edge computing task based on drone population is proposed. By introducing a node state transition model and multi-agent MADDPG algorithm, negotiation and resource sharing between drone nodes are realized, and task offloading strategy is optimized.

Benefits of technology

It improves resource utilization, shortens task processing delay, enhances dynamic environment adaptability, optimizes multi-task concurrent processing, and reduces the computing load and energy consumption of the drone.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120091365A_ABST
    Figure CN120091365A_ABST
Patent Text Reader

Abstract

The invention discloses an edge computing task unloading method and system based on an unmanned aerial vehicle group, and relates to the technical field of unmanned aerial vehicles. The invention aims to solve the problems of load balancing and task unloading in a distributed wireless communication network assisted by an unmanned aerial vehicle, and defines the unmanned aerial vehicle to be in different states by introducing a GAF node conversion model so as to save energy consumption and prolong the service time of the unmanned aerial vehicle. Then, a DDPG algorithm is introduced to train a single unmanned aerial vehicle to obtain an optimal action, the optimal action is expanded to the multi-agent field through an MADDPG algorithm, and the whole unmanned aerial vehicle group is regarded as a whole to provide services for user equipment, so that an efficient resource allocation strategy is realized. The computing resources of the unmanned aerial vehicle can be effectively distributed in a complex scene, and the purpose of shortening task processing time delay is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicles, and particularly to an edge computing task offloading method and system based on a group of unmanned aerial vehicles. Background Art

[0002] With the rapid progress of communication technologies, the number of Internet of Things (IoT)-based applications has increased sharply. However, mobile devices equipped with IoT components lack sufficient communication and computing capabilities. Mobile edge computing (MEC) can reduce task processing latency, save bandwidth resources, and enhance computing capabilities by pushing computing tasks to network edge nodes. Due to the excellent characteristics of unmanned aerial vehicles (UAVs) such as convenient deployment, high flexibility, and low cost, UAV-driven MEC has received extensive attention. In scenarios with dense crowds and a large number of computing tasks, UAVs are usually deployed with certain computing and caching resources as edge computing nodes. User equipment can delegate some of its computing tasks to nearby UAV computing nodes. This method of being able to distribute computing tasks between UAVs and mobile devices is called partial task offloading.

[0003] In a UAV-driven edge computing system, the task offloading strategy not only affects the energy consumption and computing task processing latency of UAVs, but also has an impact on the user's quality of service experience. Therefore, designing an efficient and applicable task offloading strategy has become the key to solving the problems of limited computing capabilities, resource shortages of UAVs, and diverse requirements of user equipment. Currently, there are already UAVs used as edge computing nodes to solve such problems, but most are single-UAV or multi-UAV systems lacking cooperation, and there is no negotiation and cooperation between UAV nodes. In a single-UAV system, there are often problems such as an overly long task queue and slow processing of task requirements of other user equipment. In a multi-UAV system, due to the lack of cooperation and communication between each other, some UAV nodes are idle, resulting in unnecessary energy consumption and resource waste. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention proposes an edge computing task offloading method and system based on a group of unmanned aerial vehicles. By introducing a node state transition model, different states are given to UAV nodes, and they negotiate with each other to save energy consumption. A MADDPG algorithm suitable for multi-agent systems is proposed, which can effectively handle systems in dynamic environments and solve the problem of task allocation for user equipment. This offloading strategy has excellent applicability in various scenarios that require rapid response, real-time data processing, and improved resource utilization.

[0005] The technical solution of the present invention is as follows:

[0006] In a first aspect of the present invention, an edge computing task offloading method based on a group of unmanned aerial vehicles is provided, including:

[0007] Build a UAV edge computing system model;

[0008] Define the states of UAV edge computing nodes in the UAV edge computing system model and the conditions for state transitions;

[0009] Regard each UAV edge computing node as an agent, construct a multi-agent environment, and define the global state space of the multi-agent, the state of a single agent, the joint actions of the multi-agent, the actions of a single agent, the reward function, and the Q-value function;

[0010] According to the defined global state space of the multi-agent, the state space of a single agent, the joint actions of the multi-agent, the actions of a single agent, the reward function, and the Q-value function, apply the multi-agent deep deterministic policy gradient algorithm MADDPG to centrally train several agents in the multi-agent environment to obtain trained agents;

[0011] Each trained agent uses its own Actor network to output the actions of a single agent according to the current state of the single agent. The user equipment unloads tasks to the UAV edge computing node according to the output actions of the single agent, and the UAV edge computing node executes tasks according to the set states of the UAV edge computing node and the conditions for state transitions;

[0012] Furthermore, the UAV edge computing system model includes several UAV edge computing nodes and several user equipments, and each UAV edge computing node corresponds to a UAV;

[0013] The user equipment is used to generate tasks and execute the tasks generated by itself or send the tasks to the UAV edge computing node for execution. Each user equipment generates only one task at the same time, and the number of tasks is equal to the number of user equipments;

[0014] The set of tasks that the UAV edge computing node needs to execute is \(K = \{k 1 ,k 2 ,...k x ...k i \}\), where \(k x \) is the \(x\)-th task that the UAV edge computing node needs to execute, \(x\) is the task number, and \(i\) is the number of tasks; each task \(k x \) includes the following information:

[0015] Computing time \(C\): The computing duration required for each unit byte to process task \(k x ;

[0016] Data size \(D\): The size of the data input by the task to the UAV edge computing node;

[0017] Priority P: The importance level of a task, measured by the remaining battery power of the user device. The lower the power, the higher the priority, which determines the weight of the unloading strategy.

[0018] The set of the UAV edge computing nodes is N = {n 1 , n 2 ,... n y ... n j}, where n y is the y-th UAV edge computing node, y is the number of the UAV edge computing node, j is the number of the UAV edge computing nodes. The characteristics of each UAV edge computing node n y include:

[0019] Computing power F: The computing power provided by the UAV edge computing node.

[0020] Cache capacity M: The amount of tasks that the UAV edge computing node can execute.

[0021] Task execution flag indicates whether task k x is executed on the UAV edge computing node n y . If then task k x is executed on the UAV edge computing node n y . If then it is not executed on the UAV edge computing node n y ;

[0022] Status flag If after a task is uploaded to the UAV edge computing node, the cache capacity of this UAV edge computing node is sufficient to execute task n y , then If not, then the task is transmitted to other idle UAV edge computing nodes for execution;

[0023] Furthermore, the UAV edge computing node adopts an information sharing mechanism: UAV edge computing nodes can send and receive information from each other, including the cache capacity M and task execution flag of the UAV edge computing node while sharing each other's status flags

[0024] Furthermore, the status of the UAV edge computing node includes:

[0025] Active state: The working state of the UAV edge computing node. Only when the UAV edge computing node is in the active state can it receive and execute the tasks assigned by the user device. The tasks assigned by the user device include the tasks directly sent by the user device and the tasks sent from other UAV edge computing nodes. At the same time, it shares the cache capacity M, task execution flag and status flag with other UAV edge computing nodes;

[0026] Sleep state: The UAV edge computing node is not currently executing any tasks, is disconnected from the user device, and only shares the cache capacity M and task execution flag with other UAV edge computing nodes, and enters the low-power mode to save energy;

[0027] Discovery state: The UAV edge computing node wakes up from the sleep state regularly and sends a communication request to discover the adjacent UAV edge computing nodes within the set surrounding area. At the same time, it shares the cache capacity M, task execution flag and status flag with other UAV edge computing nodes;

[0028] Furthermore, the conditions for the state transition include:

[0029] Condition for transitioning from the sleep state to the active state: When the remaining cache capacity of a certain UAV edge computing node is insufficient to execute the tasks generated by the user device, it sends a communication request to the adjacent UAV edge computing nodes within the set surrounding area. The UAV edge computing node with the largest cache capacity that receives the communication request will share the tasks generated by the user device and enter the active state;

[0030] Condition for transitioning from the discovery state to the active state: After maintaining the discovery state for a set time T d , it automatically enters the active state;

[0031] Condition for transitioning from the active state to the sleep state: When the cache capacity of the UAV edge computing node is insufficient, the status flag and transfers the tasks to the UAV edge computing node with the largest cache capacity within the set surrounding area, and then enters the sleep state;

[0032] Condition for transitioning from the discovery state to the sleep state: When a UAV edge computing node sends a communication request and discovers the adjacent UAV edge computing nodes within the set surrounding area, and at the same time receives a communication request from a UAV edge computing node with a larger cache capacity, it enters the sleep state;

[0033] Condition for transitioning from the sleep state to the discovery state: After maintaining the sleep state for a set time T s , it automatically enters the discovery state;

[0034] The active state transitions to the discovery state: After transitioning from the discovery state to the active state, if no task is executed within the set time T a it enters the discovery state; when there is a task to be executed by the UAV edge computing node currently, after completing the task, it enters the discovery state;

[0035] Furthermore, define the global state space of multi - agents at time slot t as S t ={E uav (t), P uav (t), d(t), D(t), D remain (t)}, where represents the set of remaining battery levels of all UAV edge computing nodes at time slot t, represents the remaining battery level of the y - th UAV edge computing node at time slot t; P uav (t)={p 1 (t), p 2 (t),..., p y (t),..., p j (t)} represents the set of position information of all UAV edge computing nodes at time slot t, p y (t) represents the position information of the y - th UAV edge computing node at time slot t; d(t)={d 1 (t), d 2 (t),..., d x (t),..., d i (t)} represents the set of position information of all user devices at time slot t, d x (t) is the position information of the x - th user device at time slot t; D(t)={D 1 (t), D 2 (t),..., D x (t),..., D i (t)} represents the set of task sizes generated by all user devices UDS at time slot t, D x (t) is the task size generated by the x - th user device at time slot t; D rem (t) represents the size of the remaining tasks of all user devices UDS at time slot t, that is, the task size generated by the user device UDS at time slot t minus the task size executed locally by the user device;

[0036] Define the state of a single agent at time slot t

[0037] Define the joint action of multi - agents at time slot t as A t={x(t), R(t), v(t), ω(t)}, where x(t) is the number of all user equipment connected to the UAV edge computing node at time slot t; R(t) is the task offloading ratio of all user equipment connected to the UAV edge computing node at time slot t, with the range [0, 1], representing the ratio of the task offloaded from the user equipment connected to the UAV edge computing node to the UAV edge computing node, v(t) is the flight speed of all UAV edge computing nodes at time slot t, and ω(t) is the flight angle of all UAV edge computing nodes at time slot t;

[0038] Define the action of a single agent at time slot t as where x y (t) is the number of the user equipment connected to the y-th UAV edge computing node at time slot t, R x (t) is the task offloading ratio of the user equipment UDS connected to the y-th UAV edge computing node at time slot t, v y (t) is the flight speed of the y-th UAV edge computing node at time slot t, ω y (t) is the flight angle of the y-th UAV edge computing node at time slot t;

[0039] Define the reward function: The reward functions of all UAV edge computing nodes are the same, which is:

[0040] r 1 (t) = r 2 (t) = … = r y (t) = … = r j (t)

[0041] = r t = r(S t , A t ) = -T sum (t)

[0042] where r y (t) is the reward function of the y-th UAV edge computing node at time slot t, r t and r(S t , A t ) is the reward function at time slot t, T sum (t) is the total task processing delay, which is the sum of the processing delays of each task;

[0043] The processing delay T(t) of each task is:

[0044]

[0045] where, is the local computing part delay of the user equipment, F UDSis the computing power of the user device UDS; is the offloading upload delay, where δ(t) is the wireless transmission rate, B is the channel bandwidth, u is the upload power of the user device, and σ 2 is the noise power, and h(t) is the channel gain between the UAV edge computing node and the task generated by the user device; is the computing delay of the UAV edge computing node, is the transmission delay between UAV edge computing nodes, and v uav (t) is the wireless transmission power between UAV edge computing nodes, and λ n is the status flag;

[0046] At time slot t, the UAV edge computing node n y and the task k generated by the user device x The channel gain is defined as where h 0 is the channel gain at unit distance, and l yx -2 (t) is the Euclidean distance between the UAV edge computing node and the user device UDS;

[0047] The Q-value function is defined as is the Q-value function of the y-th UAV edge computing node at time slot t, used to estimate the expected cumulative discounted reward that can be obtained after taking action under the state E represents the expectation, γ is the discount factor of the reward function r y (t), represents the Q-value function of the y-th UAV edge computing node at time slot t + 1;

[0048] Furthermore, the MADDPG algorithm extends the DDPG algorithm to the multi-agent field. It initializes a Critic network and an Actor network for each agent. Both the Critic network and the Actor network include an evaluation network and a target network. In the centralized training stage, each agent shares an experience pool. In the distributed execution stage, each agent uses the Actor network to select an individual agent's action based on the current state of the single agent, and then generates an experience according to the execution of the joint action of multiple agents and obtains the value of the reward function and the state of the single agent at the next time slot of each agent. The experiences of all agents are jointly stored in the shared experience pool, and during the training process of multiple agents, experiences are randomly sampled from the shared experience pool as samples for the training processes of different agents.

[0049] The second aspect of the present invention provides an edge computing task offloading system based on a group of drones, which is used to implement the edge computing task offloading method based on a group of drones, including:

[0050] A user device, which is used to generate tasks and execute the tasks generated by itself or send the tasks to a drone edge computing node for execution according to the task offloading strategy obtained by the multi-agent task offloading module;

[0051] Drones, each drone is a drone edge computing node, which is used to execute the tasks generated by the user device; the states of the drone edge computing nodes include: active state, dormant state, and discovery state;

[0052] A multi-agent task offloading module, which regards each drone edge computing node as an agent. The agent uses its own Actor network to output the actions of a single agent according to the state of the current single agent, and obtains a task offloading strategy;

[0053] The third aspect of the present invention provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the edge computing task offloading method based on a group of drones are executed;

[0054] The fourth aspect of the present invention provides a computer-readable storage medium, in which a computer program is stored. When the computer program is run by a processor, the steps of the edge computing task offloading method based on a group of drones are executed.

[0055] Compared with the prior art, the beneficial effects of the present invention are:

[0056] 1. The present invention provides an edge computing task offloading method based on a group of drones. Under the model of drone edge computing nodes and user devices, in the drone task offloading strategy, the single-drone system has a slow processing speed and poor user experience; the multi-drone system lacks communication and cooperation, resulting in resource waste. This article considers regarding the drone swarm as a whole to serve the user device. The drones cooperate with each other and share states, which can improve resource utilization and shorten the task processing delay.

[0057] 2. Enhance the adaptability to dynamic environments. The present invention introduces a GAF node conversion model, defines drones as three different states, and automatically adjusts. When the environment changes or reaches a preset condition, the system can automatically allocate resources to ensure the smooth execution of tasks. It can dynamically allocate tasks according to the real-time resource status of edge nodes, avoiding task calculation failures caused by insufficient node resources.

[0058] 3. Optimize multi - task concurrent processing. Through reasonable task priority sorting and resource allocation, ensure that high - priority tasks are processed first, and make full use of system resources. When resources are limited, dynamically schedule each task to improve the overall system performance. By optimizing the task offloading strategy, reduce the computing load of the UAV itself, lower energy consumption, and extend the flight time. Brief Description of the Drawings

[0059] Figure 1 It is a flowchart of the edge computing task offloading method based on a UAV swarm in an embodiment of the present invention;

[0060] Figure 2 It is a schematic diagram of the system model composed of a UAV and user equipment in an embodiment of the present invention. Detailed Embodiment

[0061] The following combines the drawings and embodiments to further describe the detailed implementation of the present invention in detail. The following embodiments are used to illustrate the present invention, but not to limit the scope of the present invention.

[0062] This embodiment proposes an edge computing task offloading method based on a UAV swarm, aiming to solve the load balancing and task offloading problems in a UAV - assisted distributed wireless communication network. By introducing the GAF node conversion model and defining different states of the UAVs, energy consumption is saved and the service time of the UAVs is extended. Subsequently, the DDPG algorithm is introduced to train a single UAV to obtain the optimal action. Through the MADDPG algorithm, it is extended to the multi - agent field, and the entire UAV swarm is regarded as a whole to provide services for user equipment, so as to achieve an efficient resource allocation strategy. This method can improve the task offloading efficiency, ensure that the UAV computing resources can be effectively allocated in complex scenarios, and achieve the purpose of shortening the task processing delay. The process of the specific implementation method is as Figure 1 shown, and the process is as follows:

[0063] Step 1: Establish a UAV edge computing system model; the UAV edge computing system model includes several UAV edge computing nodes and several user equipment, and each UAV edge computing node corresponds to a UAV;

[0064] In this embodiment, the UAV edge computing system model composed of the UAV edge computing node and the user equipment is as Figure 2 shown. The user equipment is initially equipped with certain Internet of Things components, which are used to generate tasks and execute the tasks generated by itself or send the tasks to the UAV edge computing node for execution. Each user equipment generates only one task at the same time, so the number of tasks is equal to the number of user equipment;

[0065] The set of tasks that the UAV edge computing node needs to execute is K = {k1 ,k 2 ,...k x ...k i}, where k x is the xth task that the drone edge computing node needs to perform, x is the task number, i is the number of tasks; each task k x Include the following information:

[0066] Computation time C: Processing task k x The computation time required for each unit byte is measured in CPU time cycles;

[0067] Data size D: the size of the data input to the edge computing node of the task, in bytes;

[0068] Priority P: The importance of the task, measured by the remaining battery power of the user device UDS. The lower the battery power, the higher the priority, which determines the weight of the offloading strategy;

[0069] The set of the UAV edge computing nodes is N = {n 1 ,n 2 ,...n y ...n j}, where n y is the yth UAV edge computing node, y is the number of the UAV edge computing node, j is the number of UAV edge computing nodes, and each UAV edge computing node n y Features include:

[0070] Computing power F: the computing power that the drone edge computing node can provide;

[0071] Cache capacity M: the amount of tasks that can be performed by the drone edge computing node;

[0072] Task execution flag Represents task k x Is it in the drone edge computing node n y If Then task k x In the drone edge computing node n y Execute on; if Then it is not in the UAV edge computing node n y Execute on;

[0073] Status Flags If the task is uploaded to the drone edge computing node, the cache capacity of the drone edge computing node is sufficient to execute task n y ,but If insufficient, Transmit the task to other idle drone edge computing nodes for execution to avoid resource waste;

[0074] The drone edge computing nodes adopt an information sharing mechanism: The drone edge computing nodes can send and receive information, including the cache capacity M and task execution flag of the drone edge computing nodes At the same time, share each other's status flags

[0075] Step 2: According to the established drone edge computing system model, define the status of the drone edge computing nodes in the drone edge computing system model and the conditions for status conversion;

[0076] The status of the drone edge computing nodes includes:

[0077] Active state: The working state of the drone edge computing node. Only the drone edge computing node in the active state can receive and execute the tasks assigned by the user device UDS. The tasks assigned by the user device UDS include the tasks directly sent by the user device and the tasks sent from other drone edge computing nodes. At the same time, it shares the cache capacity M, task execution flag and status flag with other drone edge computing nodes;

[0078] Sleep state: The drone edge computing node is not currently executing any tasks, disconnects from the user device, and only shares the cache capacity M and task execution flag with other drone edge computing nodes and enters the low power mode to save energy;

[0079] Discovery state: The drone edge computing node wakes up from the sleep state regularly and sends a communication request to discover the adjacent drone edge computing nodes within the set surrounding area. At the same time, it shares the cache capacity M, task execution flag and status flag with other drone edge computing nodes;

[0080] The conditions for the status conversion include:

[0081] Condition for the sleep state to turn to the active state: When the remaining cache capacity of a certain drone edge computing node is insufficient to execute the tasks generated by the user device, it sends a communication request to the adjacent drone edge computing nodes within the set surrounding area. The drone edge computing node with the largest cache capacity that receives the communication request will share the tasks generated by the user device and enter the active state;

[0082] Condition for the discovery state to turn to the active state: After maintaining the set time T d in the discovery state, it automatically enters the active state;

[0083] Conditions for the active state to the sleep state: When the cache capacity of the UAV edge computing node is insufficient, the status flag and transfers the task to the UAV edge computing node with the largest cache capacity within the surrounding set area, and then enters the sleep state;

[0084] Conditions for the discovery state to the sleep state: When a UAV edge computing node sends a communication request and discovers adjacent UAV edge computing nodes within the surrounding set area, and at the same time receives a communication request from a UAV edge computing node with a larger cache capacity, it enters the sleep state;

[0085] Conditions for the sleep state to convert to the discovery state: Keep the sleep state for the set time T s and then automatically enter the discovery state;

[0086] The active state transitions to the discovery state: After transitioning from the discovery state to the active state, if no task is executed within the set time T a it enters the discovery state; When the UAV edge computing node currently has a task to be executed, after completing the task, it enters the discovery state;

[0087] In a conventional multi-UAV cooperation system, each UAV is in an active state, resulting in high overall energy consumption and cost of the system. By defining the states of the UAV edge computing nodes and the conditions for state transitions, the cooperation efficiency of the UAV swarm can be improved and the energy consumption can be reduced;

[0088] Step 3: Since the UAV edge computing system model includes several UAV edge computing nodes, each UAV edge computing node is regarded as an agent, a multi-agent environment is constructed, and the global state space of the multi-agent, the state of a single agent, the joint action of the multi-agent, the action of a single agent, the reward function, and the Q-value function are defined;

[0089] Define the global state space of the multi-agent at time slot t as S t ={E uav (t),P uav (t),d(t),D(t),D remain (t)}, where represents the set of remaining battery levels of all UAV edge computing nodes at time slot t, represents the remaining battery level of the y-th UAV edge computing node at time slot t; P uav (t)={p 1 (t),p 2 (t),...,p y (t),...,p jThe set of the location information of all UAV edge computing nodes at time slot t is denoted as \(\mathcal{P}(t)\), and \(p\) y (t) represents the location information of the \(y\)-th UAV edge computing node at time slot t; \(\mathbf{d}(t)=\{d\) 1 (t), d 2 (t), \(\cdots\), d x (t), \(\cdots\), d i (t)\} represents the set of the location information of all user devices UDS at time slot t, and \(d\) x (t) is the location information of the \(x\)-th user device at time slot t; \(\mathbf{D}(t)=\{D\) 1 (t), D 2 (t), \(\cdots\), D x (t), \(\cdots\), D i (t)\} represents the set of the task sizes generated by all user devices UDS at time slot t, and \(D\) x (t) is the task size generated by the \(x\)-th user device at time slot t; \(D\) rem (t) represents the size of the remaining tasks of all user devices UDS at time slot t, that is, the task size generated by the user device UDS at time slot t minus the task size executed locally by the user device;

[0090] Define the state of a single agent at time slot t

[0091] Define the joint action of multiple agents at time slot t as \(\mathbf{A}\) t =\{x(t), R(t), v(t), \omega(t)\}, where \(x(t)\) is the number of all user devices UDS connected to the UAV edge computing node at time slot t; \(R(t)\) is the task offloading ratio of all user devices UDS connected to the UAV edge computing node at time slot t, and the range is \([0, 1]\), representing the ratio of the tasks offloaded from the user devices connected to the UAV edge computing node to the UAV edge computing node, \(v(t)\) is the flight speed of all UAV edge computing nodes at time slot t, and \(\omega(t)\) is the flight angle of all UAV edge computing nodes at time slot t;

[0092] Define the action of a single agent at time slot t as where \(x\) y (t) is the number of the user device UDS connected to the \(y\)-th UAV edge computing node at time slot t, \(R\) x (t) is the task offloading ratio of the user device UDS connected to the \(y\)-th UAV edge computing node at time slot t, \(v\) y (t) is the flight speed of the \(y\)-th UAV edge computing node at time slot t, and \(\omega\) y (t) is the flight angle of the \(y\)-th UAV edge computing node at time slot t;

[0093] Define the reward function: Since the UAV edge computing nodes in the UAV edge computing system model negotiate and cooperate with each other, the reward functions of all UAV edge computing nodes are the same, which is:

[0094] r 1 (t) = r 2 (t) = … = r y (t) = … = r j (t)

[0095] = r t = r(S t , A t ) = -T sum (t)

[0096] where r y (t) is the reward function of the y-th UAV edge computing node at time slot t, r t and r(S t , A t ) are the reward functions at time slot t, and T sum (t) is the total task processing delay, which is the sum of the processing delays of each task;

[0097] The processing delay T(t) of each task is:

[0098]

[0099] where is the local computing part delay of the user equipment UDS, and F UDS is the computing power of the user equipment UDS; is the offloading and uploading delay, where δ(t) is the wireless transmission rate, B is the channel bandwidth, u is the uploading power of the user equipment UDS, σ 2 is the noise power, and h(t) is the channel gain between the UAV edge computing node and the task generated by the user equipment; is the computing delay of the UAV edge computing node (the task is executed on the UAV edge computing node), is the transmission delay between UAV edge computing nodes, and v uav (t) is the wireless transmission power between UAV edge computing nodes, and λ n is the status flag;

[0100] At time slot t, the channel gain between the UAV edge computing node n y and the task k x generated by the user equipment is defined as where h 0 is the channel gain at unit distance, and lyx -2 (t) is the Euclidean distance between the UAV edge computing node and the user device UDS;

[0101] Define the Q-value function as is the Q-value function of the y-th UAV edge computing node at time slot t, used to estimate in the state taking the action and the expected cumulative discounted reward that can be obtained afterwards. E represents expectation, and γ is the discount factor of the reward function r y (t), represents the Q-value function of the y-th UAV edge computing node at time slot t + 1;

[0102] Step 4: According to the defined multi-agent global state space, single-agent state space, multi-agent joint action, single-agent action, reward function, and Q-value function, apply the multi-agent deep deterministic policy gradient algorithm MADDPG to centrally train several agents in the multi-agent environment to obtain trained agents;

[0103] The MADDPG algorithm extends the DDPG algorithm to the multi-agent field, initializes a Critic network and an Actor network for each agent. Both the Critic network and the Actor network include an evaluation network and a target network; in the centralized training stage, each agent shares an experience pool, so each agent can use the information of other agents to evaluate the current action value. In the distributed execution stage, each agent uses the Actor network to select a single-agent action according to the current single-agent state, and then generates an experience according to the executed multi-agent joint action and obtains the value of the reward function and the single-agent state of each agent in the next time slot; the experiences of all agents are jointly stored in the shared experience pool, and in the training process of the multi-agent, experiences are randomly sampled from the shared experience pool as samples for the training processes of different agents, and by sharing the states of other agents, the learning efficiency of the agents is improved;

[0104] During the training process of each agent:

[0105] Update the parameters of the evaluation network in the Critic network by minimizing the given loss function:

[0106]

[0107] where L is the loss function, Z is the size of the sampling batch, S y is the current state, A y is the action in the current state, e y is the target Q-value;

[0108] Update the parameters of the evaluation network in the Actor network through policy gradients to maximize the Q value of the Critic network:

[0109]

[0110] where θ y is the parameter of the Actor network of the y-th UAV edge computing node, denotes taking the expectation of the state S and joint action A sampled from the experience pool D, is the gradient of the Q value output by the Critic network with respect to the action of the y-th UAV edge computing node; is the gradient of the Actor network of the y-th UAV edge computing node with respect to its parameter θ y ; is the action generated by the Actor network, and Q(S,A) is the Q value evaluated by the Critic network;

[0111] The method for updating the parameters of the target network is:

[0112] θ target ' ← τθ+(1 - τ)θ target

[0113]

[0114] where θ and are the parameters of the evaluation networks in the current Actor network and Critic network respectively, θ target and are the parameters of the target networks in the Actor network and Critic network respectively, θ target ' and are the parameters of the target networks in the updated Actor network and Critic network respectively, and τ is the soft update coefficient;

[0115] Step 5: Each trained agent outputs the action of a single agent according to the state of the current single agent using its own Actor network. The user device offloads tasks to this UAV edge computing node according to the output action of the single agent, and the UAV edge computing node executes tasks according to the state of the UAV edge computing node and the conditions for state transition set in Step 3;

[0116] The edge computing task offloading strategy method based on UAV swarms proposed in this example significantly reduces the task execution time while improving the task processing efficiency. Compared with the ordinary DDPG algorithm, the algorithm proposed in this paper has a faster convergence speed and can better meet the timeliness requirements.

[0117] An embodiment of the present invention provides an edge computing task offloading system based on a drone swarm, which is used to implement the edge computing task offloading method based on the drone swarm, including:

[0118] A user equipment, which is used to generate tasks and execute the tasks generated by itself or send the tasks to a drone edge computing node for execution according to the task offloading strategy obtained by the multi-agent task offloading module;

[0119] Drones, each drone is a drone edge computing node, which is used to execute the tasks generated by the user equipment; the states of the drone edge computing nodes include: active state, dormant state, and discovery state;

[0120] A multi-agent task offloading module, which regards each drone edge computing node as an agent. The agent uses its own Actor network to output the action of the single agent according to the current state of the single agent, and obtains the task offloading strategy.

[0121] An embodiment of the present invention provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the edge computing task offloading method based on the drone swarm are executed;

[0122] In this embodiment, a computer-readable storage medium is further provided. The computer-readable storage medium stores a computer program, and when the computer program is run by a processor, the steps of the edge computing task offloading method based on the drone swarm are executed.

Claims

1. A method for offloading edge computing tasks based on drone groups, characterized in that: include: Establish a UAV edge computing system model; Define the states and state transition conditions of the drone edge computing nodes in the drone edge computing system model; Treat each drone edge computing node as an agent, construct a multi-agent environment, and define the global state space of multi-agent, the state of a single agent, the joint action of multiple agents, the action of a single agent, the reward function, and the Q-value function; According to the defined global state space of multi-intelligence, state space of a single agent, joint actions of multi-intelligence, actions of a single agent, reward function and Q-value function, the multi-agent deep deterministic policy gradient algorithm MADDPG is applied to centrally train several agents in the multi-agent environment to obtain trained agents; Each trained agent uses its own Actor network to output the action of the single agent according to the current state of the single agent. The user device unloads the task to the drone edge computing node according to the output action of the single agent, and the drone edge computing node performs the task according to the set drone edge computing node state and state transition conditions.

2. The edge computing task offloading method based on drone groups according to claim 1 is characterized in that: The UAV edge computing system model includes a plurality of UAV edge computing nodes and a plurality of user devices, and each UAV edge computing node corresponds to a UAV; The user device is used to generate tasks and execute the tasks generated by itself or send the tasks to the drone edge computing node for execution. Each user device only generates one task at a time, and the number of tasks is equal to the number of user devices; The set of tasks that the drone edge computing node needs to perform is K = {k1, k2, ...k x ...k i }, where k x is the xth task that the drone edge computing node needs to perform, x is the task number, i is the number of tasks; each task k x Include the following information: Computation time C: Processing task k x The computation time required for each unit byte; Data size D: the size of the data input into the edge computing node of the UAV; Priority P: The importance of the task, measured by the remaining battery power of the user's device. The lower the battery power, the higher the priority, which determines the weight of the uninstallation strategy; The set of the drone edge computing nodes is N = {n1, n2, ... n y ...n j }, where n y is the yth UAV edge computing node, y is the number of the UAV edge computing node, j is the number of UAV edge computing nodes, and each UAV edge computing node n y Features include: Computing power F: the computing power provided by the drone edge computing node; Cache capacity M: the amount of tasks that can be performed by the drone edge computing node; Task execution flag Represents task k x Is it in the drone edge computing node n y If Then task k x In the drone edge computing node n y Execute on; if Then it is not in the UAV edge computing node n y Execute on; Status Flags If the task is uploaded to the drone edge computing node, the cache capacity of the drone edge computing node is sufficient to execute task n y ,but If insufficient, The task is transmitted to other idle drone edge computing nodes for execution.

3. The edge computing task offloading method based on drone groups according to claim 2 is characterized in that: The UAV edge computing nodes adopt an information sharing mechanism: UAV edge computing nodes can send and receive information, including the cache capacity M of the UAV edge computing nodes and the task execution flag Share each other's status flags 4. The edge computing task offloading method based on drone groups according to claim 2 is characterized in that: The states of the drone edge computing nodes include: Activation state: The working state of the drone edge computing node. Only the drone edge computing node in the activated state receives and executes the tasks assigned by the user device. The tasks assigned by the user device include tasks sent directly by the user device and tasks sent from other drone edge computing nodes. At the same time, it performs cache capacity M and task execution flag with other drone edge computing nodes. and status flags sharing; Sleep state: The drone edge computing node is not currently executing any tasks, is disconnected from the user device, and only communicates with other drone edge computing nodes for cache capacity M and task execution flags sharing, enter low power mode to save energy; Discovery state: The UAV edge computing node periodically wakes up from the sleep state and sends a communication request to discover the adjacent UAV edge computing nodes in the surrounding set area, and communicates with other UAV edge computing nodes to check the cache capacity M and task execution flag. and status flags sharing.

5. The edge computing task offloading method based on drone groups according to claim 1 is characterized in that: The conditions for the state transition include: Conditions for transitioning from dormant state to active state: When the remaining cache capacity of a drone edge computing node is insufficient to execute the task generated by the user device, it sends a communication request to the adjacent drone edge computing nodes in the surrounding set area. The drone edge computing node with the largest cache capacity that receives the communication request will share the task generated by the user device and enter the active state; Conditions for the discovery state to the activation state: Maintain the set time T in the discovery state d After that, it automatically enters the activation state; Conditions for switching from active state to dormant state: When the cache capacity of the drone edge computing node is insufficient, the status flag After transmitting the task to the drone edge computing node with the largest cache capacity in the surrounding set area, it enters a dormant state; Conditions for transitioning from discovery state to sleep state: When a UAV edge computing node sends a communication request and discovers adjacent UAV edge computing nodes in the surrounding set area, and at the same time receives a communication request from a UAV edge computing node with a larger cache capacity, it enters the sleep state; The condition for switching from the sleep state to the discovery state: keep the sleep state for the set time T s After that, it automatically enters the discovery state; Activation state transition to discovery state: After the discovery state transitions to activation state, the a If there is no task to be executed within the time limit, it enters the discovery state; when the drone edge computing node currently has a task to be executed, it enters the discovery state after completing the task.

6. The edge computing task offloading method based on drone groups according to claim 1 is characterized in that: The global state space of the multi-agent system at time slot t is defined as in represents the set of remaining power of all drone edge computing nodes at time slot t, represents the remaining power of the yth drone edge computing node at time slot t; P uav (t)={p1(t),p2(t),...,p y (t),...,p j (t)} represents the set of location information of all UAV edge computing nodes at time slot t, p y (t) represents the location information of the yth UAV edge computing node at time slot t; d(t) = {d1(t), d2(t), ..., d x (t),...,d i (t)} represents the set of location information of all user equipments in time slot t, d x (t) is the location information of the x-th user equipment in time slot t; D(t) = {D1(t), D2(t), ..., D x (t),...,D i (t)} represents the set of task sizes generated by all user devices UDS in time slot t, D x (t) is the task size generated by the x-th user equipment in time slot t; D rem (t) represents the size of the remaining tasks of all user equipment UDS in time slot t, that is, the size of the task generated by the user equipment UDS in time slot t minus the size of the task executed locally by the user equipment; Define the state of a single agent at time slot t Define the joint action of multiple agents at time slot t as A t ={x(t),R(t),v(t),ω(t)}, x(t) is the number of all user devices connected to the UAV edge computing node at time slot t; R(t) is the task offloading ratio of all user devices connected to the UAV edge computing node at time slot t, ranging from [0,1], indicating the ratio of tasks offloaded from user devices connected to the UAV edge computing node to the UAV edge computing node, v(t) is the flight speed of all UAV edge computing nodes at time slot t, and ω(t) is the flight angle of all UAV edge computing nodes at time slot t; The action of a single agent in time slot t is defined as Among them, x y (t) is the number of the user equipment connected to the y-th UAV edge computing node at time slot t, R x (t) is the task offloading ratio of the user device UDS connected to the y-th UAV edge computing node at time slot t, v y (t) is the flight speed of the yth UAV edge computing node in time slot t, ω y (t) is the flight angle of the yth UAV edge computing node at time slot t; Define the reward function: The reward function of all drone edge computing nodes is the same, which is: r1(t)=r2(t)=…=r y (t)=…=r j (t) =r t =r(S t ,A t )=-T sum (t) Among them, r y (t) is the reward function of the yth UAV edge computing node in time slot t, r t and r(S t ,A t ) is the reward function at time slot t, T sum (t) is the total task processing delay, which is the sum of the processing delays of each task; The processing delay T(t) of each task is: in, The local calculation delay of the user equipment, F UDS The computing power of the user device UDS; To reduce the upload delay, Among them, δ(t) is the wireless transmission rate, B is the channel bandwidth, u is the upload power of the user equipment, σ 2 is the noise power, h(t) is the channel gain of the task generated by the UAV edge computing node and the user equipment; Calculate latency for drone edge computing nodes, is the transmission delay between the edge computing nodes of the UAV, v uav (t) is the wireless transmission power between the edge computing nodes of the UAV, λ n It is a status symbol; In time slot t, the drone edge computing node n y Task k generated by the user device x The channel gain is defined as Where h0 is the channel gain per unit distance, l yx -2 (t) is the Euclidean distance between the UAV edge computing node and the user device UDS; Define the Q value function as is the Q value function of the yth UAV edge computing node in time slot t, To estimate the state Take action The expected cumulative discounted reward that can be obtained later, E represents the expectation, and γ is the reward function r y The discount factor for (t), Represents the Q-value function of the y-th drone edge computing node at time slot t+1.

7. The edge computing task offloading method based on drone groups according to claim 1 is characterized in that: The MADDPG algorithm extends the DDPG algorithm to the multi-agent field, initializing a Critic network and an Actor network for each agent, each of which includes an evaluation network and a target network; in the centralized training phase, each agent shares an experience pool; In the distributed execution phase, each agent uses the Actor network to select the action of a single agent according to the current state of the single agent, and then generates experience by executing the joint action of multiple agents and obtaining the value of the reward function and the state of each agent in the next time slot. The experience of all agents is stored in the shared experience pool. During the training process of multiple agents, experience is randomly extracted from the shared experience pool as samples for the training process of different agents.

8. An edge computing task offloading system based on drone groups, characterized in that: The method for implementing the edge computing task offloading method based on a drone group as described in any one of claims 1 to 7 comprises: The user equipment is used to generate tasks and execute the tasks generated by itself or send the tasks to the drone edge computing node for execution according to the task offloading strategy obtained by the multi-agent task offloading module; UAVs, each UAV is a UAV edge computing node, used to execute tasks generated by user devices; the states of UAV edge computing nodes include: activation state, sleep state and discovery state; The multi-agent task offloading module regards each drone edge computing node as an agent. The agent uses its own Actor network to output the action of the single agent according to the current state of the single agent to obtain the task offloading strategy.

9. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate through the bus. When the machine-readable instructions are executed by the processor, the steps of the edge computing task offloading method based on a drone group are performed.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the edge computing task offloading method based on a drone group as described in any one of claims 1 to 7.