Task unloading method of multi-unmanned aerial vehicle assisted MEC system based on Safe MARL

By introducing Safe MARL and adaptive Lagrangian multiplier in the multi-UAV assisted MEC system, the limitations of traditional MEC systems in response to emergencies and changes in dynamic demands are solved, the system delay and energy consumption are minimized, and the flight safety and compliance of the drone are ensured.

CN120010949AActive Publication Date: 2025-05-16SOUTHEAST UNIV

Patent Information

Application Number
CN202510046008.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-16
Estimated Expiration
2045-01-13

AI Technical Summary

Technical Problem

Traditional MEC systems show limitations when dealing with emergencies and changes in dynamic demands, and single-drone-assisted MEC systems are difficult to cope with massive task processing needs, and service efficiency decreases.

Method used

The multi-UAV assisted MEC system based on Safe MARL is adopted. By introducing adaptive Lagrangian multipliers, the constraint optimization problem is converted into unconstrained optimization problem, and the UAV flight trajectory, computing task allocation and user scheduling are jointly optimized to minimize system delay and energy consumption.

Benefits of technology

It effectively reduces the total system consumption, ensures the safety and compliance of drone flight, improves the overall efficiency and security of the system, and is suitable for complex and dynamic network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010949A_ABST
    Figure CN120010949A_ABST
Patent Text Reader

Abstract

The invention discloses a Safe MARL-based task unloading method for a multi-unmanned aerial vehicle assisted MEC system, and provides efficient calculation unloading service for user equipment by cooperating with a plurality of unmanned aerial vehicles and an edge cloud deployed on a low-altitude platform. In order to cope with dynamic changes of user equipment calculation requirements and finiteness of unmanned aerial vehicle resources and consider flight safety of the unmanned aerial vehicle, the invention provides a strategy for jointly optimizing a flight path of the unmanned aerial vehicle, calculating task allocation and user scheduling so as to realize minimization of system execution delay and energy consumption and improve system performance. And meanwhile, the flight safety of the unmanned aerial vehicle is ensured. Specifically, the task unloading optimization problem is modeled into a constrained Markov decision process, and an advanced multi-agent safety reinforcement learning method is introduced to develop an MACTD3 solving algorithm. According to the method, the efficiency and the safety of the multi-unmanned aerial vehicle assisted MEC system in coping with large-scale task processing requirements are remarkably improved, and the method has wide application prospects and remarkable technical advantages.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of edge computing, and specifically relates to a task offloading method for a multi-UAV assisted MEC system based on Safe MARL. Background Art

[0002] With the rapid development of the sixth generation 6G wireless networks, computationally intensive and low-latency application scenarios such as autonomous driving and telemedicine are increasing. However, these tasks may pose great challenges to user equipment (UE) with limited computing power and battery life. To address this challenge, mobile edge computing (MEC) is considered a promising technology that allows user equipment to offload its computationally intensive applications to edge servers.

[0003] Mobile edge computing has shown great potential in promoting computationally intensive applications, but its performance is not always satisfactory. This is mainly due to the fixed location of ground MEC servers, which cannot be dynamically adjusted according to user needs. In addition, due to the fixed deployment location and high construction cost of MEC servers, they show certain limitations in responding to emergencies and dynamic demand changes. For example, when a natural disaster causes damage to the network infrastructure, the MEC server cannot recover and provide computing services in time, resulting in a shortage of computing resources and a decrease in task offloading performance.

[0004] In recent years, unmanned aerial vehicles (UAVs) have received extensive attention and research as an alternative to traditional MEC systems. UAV-assisted MEC systems provide effective computing services for service devices with their highly flexible mobility, convenient deployment methods, low cost, and line-of-sight (LoS) connection advantages. Compared with traditional MEC systems, drone-assisted MEC systems can not only provide temporary computing and communication services when traditional MEC servers are damaged, but also optimize the utilization of computing resources and improve the overall performance and reliability of the system through flexible deployment and scheduling in normal times. This innovative computing paradigm shows broad development prospects in coping with the complex and changing network environment of the future.

[0005] However, the single-drone-assisted MEC system is limited by its resources and cannot cope with the growing demand for massive task processing. And as the drone service range expands and the number of user devices increases, the service efficiency will drop sharply. Although the multi-drone-based MEC solution has higher performance, it also brings many challenges. First, users are constantly moving during the calculation process, making it difficult to obtain the optimal strategy. Second, drones need to fly from different take-off points to specific areas to provide unloading services. Different flight trajectories may lead to differences in channel quality, which in turn causes different communication delays and energy consumption. In addition, when the onboard resources are limited, the amount of computing tasks assigned to drones will also affect the computing delay and energy consumption. Finally, drones need to fly in complex urban or rural environments when performing tasks. These areas include no-fly zones stipulated by laws or policies, such as around airports, over government facilities, and other sensitive areas. In order to ensure the efficient execution of tasks and comply with relevant flight regulations, drones must strictly avoid entering these no-fly zones to ensure flight safety and mission compliance. Summary of the invention

[0006] To solve the above problems, the present invention discloses a task offloading method for a multi-UAV assisted MEC system based on Safe MARL. On the basis of the MATD3 algorithm, adaptive Lagrange multipliers are introduced to convert the constrained optimization problem into an unconstrained optimization problem, and then the UAV flight trajectory, computational task allocation and user scheduling strategies are combined to minimize the system execution delay and energy consumption, while ensuring the safety and compliance of UAV flight.

[0007] To achieve the above object, the technical solution of the present invention is as follows:

[0008] The task offloading method of the multi-UAV assisted MEC system based on Safe MARL includes the following steps:

[0009] First, design a multi-UAV-assisted MEC system.

[0010] The system consists of M mobile devices (MDs), N unmanned aerial vehicles (UAVs), and an edge cloud (EC) deployed on a stationary low-altitude platform (LAP). The user device set and the UAV set are represented as and Each drone is equipped with a small server, and the edge cloud is equipped with a high-performance server. Assume that the working time of the entire system is T, and the total time is divided into K equal-length time slots. The set of time slots is recorded as The duration of each time slot is Each user device generates a computationally intensive task at the beginning of each time slot t, which is denoted as W. m (t)=(D m (t),Cm (t)), where D m (t) is the size of the task data, C m (t) represents the number of CPU cycles required to process data. Due to limited computing power, user devices cannot complete calculations locally and need to offload computing tasks to drones. In addition, drones can not only act as computing nodes to provide computing services to user devices, but also act as relay nodes to further transmit part of the tasks of user devices to the edge cloud for processing. In order to reduce system consumption and ensure flight safety, the flight trajectory of drones must be strictly planned. It is assumed that each drone only provides services to ground user devices within its coverage area, and there is no overlap between the coverage areas of each drone. There are O no-fly zones within the drone's flight area, such as around airports, above government facilities, etc. The set of no-fly zones is denoted as Drones must avoid these no-fly zones to ensure the safety and compliance of mission execution. The entire system model can be divided into three parts: drone movement model, communication model, and calculation model.

[0011] The drone mobility model defines the following parameters:

[0012] ω n (t) = [x n (t),y n (t),H] T is the three-dimensional coordinate of UAV n at time slot t; n (t) and y n (t) represents the horizontal and vertical coordinates of UAV n at time step t; H represents the flight altitude of the UAV; l n (t)∈[0,L max ]Indicates the flight distance; is the flight angle; L max Indicates the maximum flight distance of the drone; Indicates the maximum coverage radius of the drone; φ n represents the maximum elevation angle of UAV n; v n (t) represents the flight speed of the UAV; Represents the flight energy consumption of the UAV; M UAV represents the quality of the drone; multiple user devices within the coverage area of ​​a drone are served by the same drone, M n (t) represents the number of users served by the drone in time slot t; Denoted as a service-associated variable, when UE m is served by UAV n, otherwise, To avoid interference during data transmission, each user device can only be served by one drone at any time, i.e. Safety factor illegaln (t) = {0,1} indicates whether UAV n illegally passes through the no-fly zone in time slot t. If it passes through the no-fly zone, then illegal n (t) = 1, if you successfully avoid the no-fly zone, then illegal n (t)=0.

[0013] Each variable is calculated by the following formula:

[0014]

[0015] The communication model defines the following parameters:

[0016] ω ,m (t) = [x m (t),y m (t),0] T is the user device location; mn (t) represents the distance between drone n and mobile device m; represents the probability of establishing a LoS link between UAV n and user device m at time slot t; represents the probability of establishing a non-line-of-sight link between drone n and user device m; and represent the LoS and NLoS path losses between drone n and user device m, respectively; c represents the speed of light, f c represents the carrier frequency, η LoS and η NLoS h represents the path loss coefficients of line-of-sight link and non-line-of-sight link respectively; mn (t) represents the channel gain between mobile device m and drone n; g0 represents the channel gain when the distance between drone and EC is 1m; B u is the uplink bandwidth; R mn (t) represents the data transmission rate between mobile device m and drone n in time slot t; P m is the transmit power of user equipment m; is the Gaussian white noise power; is the transmission delay between UE m and UAV n; is the received power of the UAV; The energy consumed by task transfer between user device m and drone n; ω e =[x e ,y e ,H e ] T Indicates the location of EC; d en (t) is the distance between UAV n and EC; h ne (t) is the channel gain between UAV n and EC in time slot t; Be represents the bandwidth allocated to each drone; R ne (t) is the data transmission rate between UAV n and EC; represents the transmission power of UAV n in time slot t; P max is the maximum transmission power of the drone; is the transmission delay of UAV n from user equipment m to EC; and are the proportions of tasks of user equipment m performed by EC and UAV n in time slot t, respectively; It is the energy consumption of user equipment m transmitting data to EC through drone n.

[0017] Each variable is calculated by the following formula:

[0018] d mn (t)=||ω n (t)-ω m (t)||

[0019]

[0020] d en (t)=||ω n (t)-ω e ||

[0021]

[0022] The calculation model defines the following parameters:

[0023] The computational latency of UAVn processing the task of user device m locally; f mn (t) is the computing resources allocated from drone n to mobile device m; is the computing resource of UAV n; κ is the energy consumption coefficient related to the CPU of UAV; Processing energy consumption of mobile device m for UAVn; F e is the total computing resources of EC; f me (t) the computing resources allocated to each user device by the EC; The computational delay of the partial task of mobile device m on EC after it is transferred to EC via UAV n; E n (t) represents the total energy consumption of UAV n at time step t; T n (t) is the total mission execution delay of UAV n at time step t.

[0024] Each variable is calculated by the following formula:

[0025]

[0026] Based on an in-depth analysis of the above system model, this paper constructs a constrained optimization problem for a multi-UAV assisted mobile edge computing (MEC) system. By jointly optimizing user scheduling, task allocation ratio, UAV flight trajectory and transmission power, the aim is to minimize the total processing delay and total energy consumption of the system while ensuring the flight safety of the UAVs and preventing them from colliding or entering no-fly zones during mission execution, thereby improving the overall efficiency and safety of the system. The constrained optimization problem is formalized, where D min is the minimum safe distance between drones, w1 and w2 represent the importance weights of energy consumption and execution delay respectively:

[0027]

[0028] In the above objective function, C1, C2 and C3 represent the constraints of MDs task offloading; C4 constrains UAVs not to pass through the no-fly zone during flight to ensure the compliance of task execution; C5, C6 and C7 represent the maximum flight distance constraint and working range constraint of the UAV during flight; C8 constrains the distance between UAVs to avoid collision; C9 constrains the coverage of any two UAVs not to overlap to avoid interference during data transmission; C10 constrains the total energy consumption of the UAV not to exceed its energy upper limit.

[0029] The optimization problem is constructed as a constrained Markov decision process (CMDP), and a MACTD3 algorithm based on multi-agent secure reinforcement learning is proposed to solve this problem. The detailed definitions of each element of secure reinforcement learning are as follows:

[0030] (1) State space: In the proposed multi-UAV-assisted MEC system, the state space is jointly determined by the UAVs, the real-time mobile end users, and the environment in which they are located. Therefore, the state space of UAV n can be defined as:

[0031]

[0032] where ω n (t) represents the location information of drone n, represents the remaining energy of drone n, ω m (t) represents the location information of mobile device m, W m (t) represents the task information of mobile device m. M is the total number of users, which means the task information of 1-M users is observed, so it is capitalized.

[0033] (2) Action space: The agent will perform actions based on the current state of the system and the observed environment. Each agent needs to optimize location scheduling, task allocation ratio, and upload power to maximize the system effect while avoiding unsafe behaviors. Therefore, the action space of drone n can be defined as:

[0034]

[0035] Among them l n (t), Respectively represent the flight distance and flight angle of UAV n, represents the upload power of drone n, It is represented as the proportion of tasks of user device m that are executed on UAV n.

[0036] (3) Reward function: The reward function is the feedback signal from the environment to the agent's actions, which is used to evaluate the agent's behavior and help it improve its strategy. In order to minimize system consumption and ensure that the drone can provide computing offloading services for all user equipment (UEs), the reward function is defined as follows:

[0037]

[0038] Among them, w1 and w2 are weights used to represent the importance of energy consumption and delay respectively. w1≥w2 indicates energy-sensitive scenarios, while w1<w2 is applicable to delay-sensitive scenarios. In addition, when any user device is not covered by the drone, all agents will face a penalty term ∈ represents the penalty weight.

[0039] (4) Cost function: The cost function improves the safety and compliance of the agent in the process of performing tasks. When the agent takes unsafe actions, a corresponding cost value will be generated. Therefore, it is required that the cost value of the agent cannot exceed the preset safety threshold d = 0. The agent must avoid unsafe behaviors when performing tasks. The cost function is defined as:

[0040] c n (t) = η1 ∨ η2 ∨ η3

[0041] Among them, η1, η2, and η3 are binary parameters used to measure whether the agent violates a specific safety constraint at time slot t. Specifically, η1 = 1 means that drone n enters a no-fly zone at time step t; η2 = 1 means that drone n is at risk of collision with other drones at time step t; η2 = 1 means that the coverage of drone n at time step t overlaps with the coverage of other drones. The logical OR operation ensures that as long as any safety constraint is violated, the cost value c n (t) is 1.

[0042] For the multi-agent secure reinforcement learning algorithm MACTD3 proposed in this invention, each agent n has the following network: Actor network π n , Critic Network and CostCritic Q c,n , and the corresponding target network and Among them, the Actor network is used to generate the action strategy of the agent, the Critic network is used to evaluate the cumulative reward under a given state and action, the Cost Critic network is used to evaluate the cumulative cost under a given state and action, and the target network is used to alleviate the instability during training. Since the multi-agent environment is non-stationary, which will cause the experience replay to fail, the MACTD3 algorithm uses a centralized training and distributed execution method to find the optimal joint strategy.

[0043] For each agent n, initialize its Actor network parameters Critic Network Parameters and CostCritic Network Parameters And the corresponding target network parameters and At the same time, initialize the Lagrange multiplier λ of each agent n ≥0, used to weigh the relationship between reward and cost.

[0044] At each time step t, each agent n takes its own observed state s n (t), through the Actor network π n Generate decision actions:

[0045]

[0046] Among them, a n (t) is the action selected by agent n at time t. The joint action of all agents is denoted as a(t) = (a1(t), a2(t), ..., a N (t)), where N is the total number of agents. After executing the joint action a(t), the environment feeds back the next state s(t+1), the reward r of each agent n (t) and cost c n (t). Store the experience tuple (s(t), a(t), r(t), c(t), s(t+1)) into the experience replay buffer.

[0047] Randomly sample a mini-batch of data from the experience replay buffer Where B is the mini-batch size. For each agent n, use the target Actor network Calculate the next target action and add random noise to smooth the estimate of the target value and reduce the overestimation bias:

[0048]

[0049] in, It means that the mean is 0 and the variance is σ 2 The noise is normally distributed and clipped in the interval [-ρ, ρ], where ρ is the noise clipping threshold. The joint target action is a′(j+1)=(a′1(j+1),a′2(j+1),...,a′ N (j+1)). and pass through the target Critic network and target Cost Critic network Calculate the target Q value and target cost Q value for each agent:

[0050]

[0051] Calculate the loss function based on the target Q value And update the Critic network parameters of each agent n by minimizing the loss function. The loss function is as follows:

[0052]

[0053] According to the loss function, the gradient descent method is used to update the Critic network parameter formula as follows, where β1 is the neural network learning rate

[0054]

[0055] Similarly, by minimizing the loss function L(Q c,n ) to update the CostCritic network parameters of each agent, and the loss function is as follows:

[0056]

[0057] The formula for updating the Cost Critic network parameters is as follows:

[0058]

[0059] The parameters of each agent’s Actor network are updated through the policy gradient method, and the Lagrange multiplier λ is introduced n , which encourages the agent to avoid violating the preset safety constraint d while optimizing the reward. The policy gradient can be expressed as:

[0060]

[0061] The Actor network parameters are updated as follows:

[0062]

[0063] According to the situation that the agent violates the safety constraint d, the Lagrange multiplier of each agent n is dynamically adjusted, and the update formula is as follows:

[0064]

[0065] Among them, β lag >0 is the learning rate. If agent n violates the safety constraint, its Lagrange multiplier will increase, thereby increasing the penalty for violating the constraint. If the behavior of agent n meets the constraint, its Lagrange multiplier will decrease, thereby maximizing the cumulative reward while meeting the constraint.

[0066] A soft update strategy is used to update the parameters of the target network. The existence of the target network is to alleviate the phenomenon of continuous overestimation or underestimation during network training and ensure the stability and convergence of training. The target network and its corresponding main network architecture are exactly the same. The specific soft update formula is as follows, where τ = 0.005 is the soft update coefficient:

[0067]

[0068] The beneficial effects of the present invention are:

[0069] The present invention ensures the flight safety of drones by minimizing the total system consumption, preventing drones from colliding or entering no-fly zones, and constructs the problem as a constrained optimization problem. The advanced multi-agent safety reinforcement learning method is introduced, and the MACTD3 solution algorithm is proposed. The adaptive Lagrange multiplier is introduced to transform the constrained optimization problem into an unconstrained optimization problem, and the flight trajectory and transmission power of drones, user scheduling and task allocation strategies are jointly optimized. A large number of simulation results show that compared with the existing solutions, the present invention has significant advantages in reducing task execution delays, reducing system energy consumption and improving flight safety, and is suitable for complex and dynamic network environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 The number of training rounds is 4000, each round contains 50 time steps, O = 3, N = 3, M = 20, H = 100, H e =200, when the no-fly zone radius is 90, the learning curve convergence characteristics of the three algorithms MATD3, MATD3-RS, and MACTD3 on the reward function.

[0071] Figure 2The number of training rounds is 4000, each round contains 50 time steps O = 3, N = 3, M = 20, H = 100, H e =200, and the no-fly zone radius is 90, the learning curve convergence characteristics of the three algorithms MATD3, MATD3-RS, and MACTD3 on the cost function.

[0072] Figure 3 Step O = 3, N = 3, M = 20, H = 100, H e =200, the flight trajectory of the drone under the guidance of the optimal strategy when the radius of the no-fly zone is 90.

[0073] Figure 4 O=3,N=3,M=20,H=100,H e =200, and the no-fly zone radius is 90. Comparison of the total system consumption of MATD3, MATD3-RS, and MACTD3 at different risk zone radii.

[0074] Figure 5 O=3,N=3,M=20,H=100,H e =200, and the no-fly zone radius is 90. Comparison of the total system safety cost of MATD3, MATD3-RS, and MACTD3 under different risk zone radii. DETAILED DESCRIPTION

[0075] The present invention will be further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.

[0076] As shown in the figure, the task offloading method of the multi-UAV assisted MEC system based on SafeMARL of the present invention includes the following steps:

[0077] S1: Design a multi-UAV assisted MEC system.

[0078] The system consists of M mobile user devices (UDs), N unmanned aerial vehicles (UAVs) flying at a height of H, and an edge cloud (EC) deployed on a stationary LAP. The user device set and the UAV set are represented as and Each drone is equipped with a small server, and the edge cloud is equipped with a high-performance server. Due to limited computing power, user devices cannot complete computing locally and need to offload computing tasks to drones. In addition, drones can not only act as computing nodes to provide computing services for user devices, but also as relay nodes to further transmit tasks of some user devices to the edge cloud for processing. Each drone only provides services to ground user devices within its coverage area, and no overlap is allowed between the coverage areas of each drone. There are O no-fly zones within the flight area of ​​the drone, and the set of no-fly zones is denoted as Drones must avoid these no-fly zones to ensure safe and compliant mission execution.

[0079] S2: Define the basic elements

[0080] ω n (t) = [x n (t),y n (t),H] T is the three-dimensional coordinate of UAV n at time slot t; n (t) and y n (t) represent the horizontal and vertical coordinates of UAV n at time step t; l n (t)∈[0,L max ]Indicates the flight distance; is the flight angle; L max Indicates the maximum flight distance of the drone; v n (t) represents the flight speed of the UAV; Indicates the flight energy consumption of the drone; Indicates the maximum coverage radius of the drone; φ n represents the maximum elevation angle of UAV n; v n (t) represents the flight speed of the UAV; Represents the flight energy consumption of the UAV; M UAV represents the quality of the drone; multiple user devices within the coverage area of ​​a drone are served by the same drone, M n (t) represents the number of users served by the drone in time slot t; Denoted as a service-associated variable, when UE m is served by UAV n, otherwise To avoid interference during data transmission, each user device can only be served by one drone at any time, i.e. Safety factor illegal n (t) = {0,1} indicates whether UAV n illegally passes through the no-fly zone in time slot t; ω ,m (t) = [x m (t),y m (t),0]T is the user device location; mn (t) represents the distance between drone n and mobile device m; represents the probability of establishing a LoS link between UAV n and user device m at time slot t; represents the probability of establishing a non-line-of-sight link between drone n and user device m; and denote the LoS and NLoS path losses between drone n and user device m, respectively; c denotes the speed of light, f c represents the carrier frequency, η LoS and η NLoS h represents the path loss coefficients of line-of-sight link and non-line-of-sight link respectively; mn (t) represents the channel gain between mobile device m and drone n; g0 represents the channel gain when the distance between drone and EC is 1m; B u is the uplink bandwidth; R mn (t) represents the data transmission rate between mobile device m and drone n in time slot t; P m is the transmit power of user equipment m; is the Gaussian white noise power; is the transmission delay between UE m and UAV u; is the received power of the UAV; The energy consumed by task transfer between user device m and drone u; ω e =[x e ,y e ,H e ] T Indicates the location of EC; d en (t) is the distance between UAV n and EC; h ne (t) is the channel gain between UAV n and EC in time slot t; B e represents the bandwidth allocated to each drone; R ne (t) is the data transmission rate between UAV n and EC; represents the transmission power of UAV n in time slot t; P max is the maximum transmission power of the drone; is the transmission delay of UAV n from user equipment m to EC; and are the proportions of tasks of user equipment m performed by EC and UAV n in time slot t, respectively; is the energy consumption of data transmission from user device m to EC via drone n; is the computational latency of UAV n processing the task of user device m locally; f mn(t) is the computing resources allocated from drone n to mobile device m; is the computing resource of UAV n; κ is the energy consumption coefficient related to the CPU of UAV; Energy consumption of mobile device m for UAV n; F e is the total computing resources of EC; f me (t) the computing resources allocated to each user device by the EC; The computational delay of the partial task of mobile device m on EC after it is transferred to EC via UAVn; E n (t) represents the total energy consumption of UAV n at time step t; T n (t) is the total mission execution delay of UAV n at time step t.

[0081] S3: Define how to calculate important performance indicators

[0082]

[0083]

[0084] S4: This problem is summarized as a constrained optimization problem based on CMDP, which aims to minimize the total processing delay and total energy consumption of the system while ensuring the flight safety of the drone and preventing it from colliding or entering the no-fly zone during the mission, thereby improving the overall efficiency and safety of the system. The constrained optimization problem is formalized as follows:

[0085]

[0086] C1, C2 and C3 represent the constraints for MDs task offloading; C4 constrains UAVs from passing through no-fly zones during flight to ensure compliance with task execution; C5, C6 and C7 represent the maximum flight distance constraints and working range constraints of UAVs during flight; C8 constrains the distance between UAVs to avoid collisions; C9 constrains the coverage of any two UAVs not to overlap to avoid interference during data transmission; C10 constrains the total energy consumption of UAVs not to exceed their energy limit.

[0087] S5: Based on the secure reinforcement learning framework, design secure reinforcement learning elements, namely state space, action space, reward, and cost function. The specific design is as follows:

[0088] 1) State space: In the proposed multi-UAV-assisted MEC system, the state space is jointly determined by the UAVs, the real-time mobile end users, and the environment in which they are located. Therefore, the state space of UAV n can be defined as:

[0089]

[0090] where ω n (t) represents the location information of UAV n, ω m (t) represents the location information of mobile device m, W m Represents the task information of mobile device m.

[0091] (2) Action space: The agent will perform actions based on the current state of the system and the observed environment. Each agent needs to optimize location scheduling, task allocation ratio, and upload power to maximize the system effect while avoiding unsafe behaviors. Therefore, the action space of drone n can be defined as:

[0092]

[0093] Among them l n (t), Respectively represent the flight distance and flight angle of UAV n, represents the upload power of drone n, It is represented as the proportion of tasks of user device m that are executed on UAV n.

[0094] (3) Reward function: The reward function is the feedback signal from the environment to the agent's actions, which is used to evaluate the agent's behavior and help it improve its strategy. In order to minimize system consumption and ensure that the drone can provide computing offloading services for all user equipment (UEs), the reward function is defined as follows:

[0095]

[0096] Among them, w1 and w2 are weights used to represent the importance of energy consumption and delay respectively. w1≥w2 indicates energy-sensitive scenarios, while w1<w2 is applicable to delay-sensitive scenarios. In addition, when any user device is not covered by the drone, all agents will face a penalty term ∈ represents the penalty weight.

[0097] (4) Cost function: The cost function improves the safety and compliance of the agent in the process of performing tasks. When the agent takes unsafe actions, a corresponding cost value will be generated. Therefore, it is required that the cost value of the agent cannot exceed the preset safety threshold d = 0. The agent must avoid unsafe behaviors when performing tasks. The cost function is defined as:

[0098] c n (t) = η1 ∨ η2 ∨ η3

[0099] Among them, η1, η2, and η3 are binary parameters used to measure whether the agent violates a specific safety constraint at time slot t. Specifically, η1 = 1 means that drone n has entered a no-fly zone; η2 = 1 means that drone n is at risk of collision with other drones; η2 = 1 means that the coverage of drone n overlaps with the coverage of other drones. The logical OR operation ensures that as long as any safety constraint is violated, the cost value c n (t) is 1.

[0100] S6: Use the MACTD3 algorithm to solve this problem. Each agent n has the following network: Actor network π n , Critic Network and CostCritic Q c,n , and the corresponding target network and Among them, the Actor network is used to generate the action strategy of the intelligent agent, the Critic network is used to evaluate the cumulative reward under a given state and action, the Cost Critic network is used to evaluate the cumulative cost under a given state and action, and the target network is used to alleviate the instability during training. For each intelligent agent n, initialize its Actor network parameters Critic Network Parameters and Cost Critic Network Parameters And the corresponding target network parameters and At the same time, initialize the Lagrange multiplier λ of each agent n ≥0, used to weigh the relationship between reward and cost.

[0101] At each time step t, each agent n takes its own observed state s n (t), through the Actor network π n Generate decision actions:

[0102]

[0103] Among them, a n (t) is the action selected by agent n at time t. The joint action of all agents is denoted as a(t) = (a1(t), a2(t), ..., a N (t)), where N is the total number of agents. After executing the joint action a(t), the environment feeds back the next state s(t+1), the reward r of each agent n (t) and cost c n(t). Store the experience tuple (s(t), a(t), r(t), c(t), s(t+1)) into the experience replay buffer.

[0104] Randomly sample a mini-batch of data from the experience replay buffer Where B is the mini-batch size. For each agent n, use the target Actor network Calculate the next target action and add random noise to smooth the estimate of the target value and reduce the overestimation bias:

[0105]

[0106] in, It means that the mean is 0 and the variance is σ 2 The noise is normally distributed and clipped in the interval [-ρ, ρ], where ρ is the noise clipping threshold. The joint target action is a′(j+1)=(a′1(j+1),a′2(j+1),...,a′ N (j+1)). and pass through the target Critic network and target Cost Critic network Calculate the target Q value and target cost Q value for each agent:

[0107]

[0108] Calculate the loss function based on the target Q value And update the Critic network parameters of each agent n by minimizing the loss function. The loss function is as follows:

[0109]

[0110] Similarly, by minimizing the loss function L(Q c,n ) to update the CostCritic network parameters of each agent, and the loss function is as follows:

[0111]

[0112] The parameters of each agent’s Actor network are updated through the policy gradient method, and the Lagrange multiplier λ is introduced n , which encourages the agent to avoid violating the preset safety constraint d while optimizing the reward. The policy gradient can be expressed as:

[0113]

[0114] According to the situation that the agent violates the safety constraint d, the Lagrange multiplier of each agent n is dynamically adjusted, and the update formula is as follows:

[0115]

[0116] Among them, β lag >0 is the learning rate. If agent n violates the safety constraint, its Lagrange multiplier will increase, thereby increasing the penalty for violating the constraint. If the behavior of agent n meets the constraint, its Lagrange multiplier will decrease, thereby maximizing the cumulative reward while meeting the constraint.

[0117] It should be noted that the above content only illustrates the technical idea of ​​the present invention and cannot be used to limit the protection scope of the present invention. For ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications all fall within the protection scope of the claims of the present invention.

Claims

1. The task offloading method of multi-UAV assisted MEC system based on SafeMARL is characterized by: The following steps are involved: S1: Design a multi-UAV assisted MEC system; S2: Calculate performance indicators; S3: Constrained Optimization Formalization; S4: Based on a safe multi-agent reinforcement learning framework; S5: Use the MACTD3 algorithm to solve the optimal strategy.

2. The task offloading method of the multi-UAV assisted MEC system based on SafeMARL according to claim 1 is characterized by: S1 describes the design of a drone-assisted MEC system, as follows: The system includes M mobile devices, N drones, and an edge cloud deployed on a stationary low-altitude platform; the flight trajectory of the drone must be strictly planned; each drone only provides services to ground user devices within its coverage area, and there is no overlap between the coverage areas of each drone; there are 0 no-fly zones within the flight area of ​​the drone, and the entire system model is divided into three parts: drone mobility model, communication model, and computing model.

3. The task offloading method of the multi-UAV assisted MEC system based on SafeMARL according to claim 2 is characterized in that: The drone mobility model defines the following parameters: ω n (t) = [x n (t),y n (t),H] T is the three-dimensional coordinate of UAV n at time slot t; n (t) and y n (t) represents the horizontal and vertical coordinates of UAV n at time step t; H represents the flight altitude of the UAV; l n (t)∈[0,L max ]Indicates the flight distance; is the flight angle; L max Indicates the maximum flight distance of the drone; Indicates the maximum coverage radius of the drone; φ n represents the maximum elevation angle of UAV n; v n (t) represents the flight speed of the UAV; Indicates the flight energy consumption of the drone; M UAV Represents the quality of the drone; Multiple user devices within the coverage area of ​​a certain drone are served by the same drone, M n (t) represents the number of users served by the drone in time slot t; Denoted as a service-associated variable, when UE m is served by UAV n, otherwise, To avoid interference during data transmission, each user device can only be served by one drone at any time, i.e. Safety factor illegal n (t) = {0,1} indicates whether UAV n illegally passes through the no-fly zone in time slot t. If it passes through the no-fly zone, then illegal n (t) = 1, if you successfully avoid the no-fly zone, then illegal n (t) = 0; Each variable is calculated by the following formula:

4. The task offloading method of the multi-UAV assisted MEC system based on SafeMARL according to claim 2 is characterized in that: The communication model defines the following parameters: ,m (t) = [x m (t),y m (t),0] T is the user device location; mn (t) represents the distance between drone n and mobile device m; represents the probability of establishing a LoS link between UAV n and user device m at time slot t; represents the probability of establishing a non-line-of-sight link between UAV n and user device m; and represent the LoS and NLoS path losses between drone n and user device m, respectively; c represents the speed of light, f c represents the carrier frequency, η LoS and η NLoS h represents the path loss coefficients of line-of-sight link and non-line-of-sight link respectively; mn (t) represents the channel gain between mobile device m and drone n; g0 represents the channel gain when the distance between drone and EC is 1m; B u is the uplink bandwidth; R mn (t) represents the data transmission rate between mobile device m and drone n in time slot t; P m is the transmit power of user equipment m; is the Gaussian white noise power; is the transmission delay between UE m and UAV n; is the received power of the UAV; The energy consumed by task transfer between user device m and drone n; ω e =[x e ,y e ,H e ] T Indicates the location of EC; d en (t) is the distance between UAV n and EC; h ne (t) is the channel gain between UAV n and EC in time slot t; B e represents the bandwidth allocated to each drone; R ne (t) is the data transmission rate between UAV n and EC; represents the transmission power of UAV n in time slot t; P max is the maximum transmission power of the drone; is the transmission delay of UAV n from user equipment m to EC; and are the proportions of tasks of user equipment m performed by EC and UAV n in time slot t, respectively; is the energy consumption of data transmission from user device m to EC via drone n; Each variable is calculated by the following formula: d mn (t)=||ω n (t)-ω m (t)|| 5. The task offloading method of the multi-UAV assisted MEC system based on SafeMARL according to claim 2 is characterized in that: The calculation model defines the following parameters: is the computational latency of UAV n processing the task of user device m locally; f mn (t) is the computing resources allocated from drone n to mobile device m; is the computing resource of UAV n; κ is the energy consumption coefficient related to the CPU of UAV; Energy consumption of mobile device m for UAV n; F e is the total computing resources of EC; f me (t) the computing resources allocated to each user device by the EC; The computational delay of the partial task of mobile device m on EC after it is transferred to EC via UAV n; E n (t) represents the total energy consumption of UAV n at time step t; T n (t) is the total task execution delay of UAV n at time step t; Each variable is calculated by the following formula:

6. The task offloading method of the multi-UAV assisted MEC system based on SafeMARL according to claim 1 is characterized in that: The formalization described in S3 is to summarize the problem into a constrained optimization problem based on CMDP, which aims to minimize system consumption while ensuring the flight safety of the UAV, as follows: In the above objective function, C1, C2 and C3 represent the constraints of MDs task offloading; C4 constrains UAVs not to pass through the no-fly zone during flight to ensure the compliance of task execution; C5, C6 and C7 represent the maximum flight distance constraint and working range constraint of the UAV during flight; C8 constrains the distance between UAVs to avoid collision; C9 constrains the coverage of any two UAVs not to overlap to avoid interference during data transmission; C10 constrains the total energy consumption of the UAV not to exceed its energy upper limit.

7. The task offloading method of the multi-UAV assisted MEC system based on SafeMARL according to claim 1 is characterized in that: Based on the secure multi-agent reinforcement learning framework described in S4, the four elements of secure reinforcement learning are designed, namely, state s(t), action a(t), reward r(t), and cost c(t); the specific design is as follows: 1) State space: In the proposed multi-UAV-assisted MEC system, the state space is jointly determined by the UAVs, the real-time mobile end users, and the environment in which they are located. Therefore, the state space of UAV n is defined as: where ω n (t) represents the location information of UAV n, ω m (t) represents the location information of the mobile device m, represents the remaining energy of UAV n, W m Represents task information of mobile device m; (2) Action space: The agent will perform actions based on the current state of the system and the observed environment. Each agent needs to optimize location scheduling, task allocation ratio, and upload power to maximize the system effect while avoiding unsafe behaviors. Therefore, the action space of drone n is defined as: Among them l n (t), Respectively represent the flight distance and flight angle of UAV n, represents the upload power of drone n, It is represented as the proportion of tasks of user device m executed on UAV n; (3) Reward function: The reward function is the feedback signal from the environment to the agent’s actions. It is used to evaluate the agent’s behavior and help it improve its strategy. In order to minimize system consumption and ensure that the drone can provide computing offload services for all user devices, the reward function is defined as follows: Among them, w1 and w2 are weights used to represent the importance of energy consumption and delay respectively; w1≥w2 indicates energy-sensitive scenarios, while w1<w2 is applicable to delay-sensitive situations; in addition, when any user device is not covered by the drone, all agents will face a penalty term ∈ represents the penalty weight; (4) Cost function: The cost function improves the safety and compliance of the agent in the process of performing tasks. When the agent takes unsafe actions, a corresponding cost value will be generated. Therefore, it is required that the cost value of the agent cannot exceed the preset safety threshold d = 0. The agent must avoid unsafe behaviors when performing tasks. The cost function is defined as: c n (t)=η1∨η2∨η3 Among them, η1, η2 and η3 are binary parameters used to measure whether the agent violates a specific safety constraint at time slot t; specifically, η1 = 1 means that drone n enters a no-fly zone; η2 = 1 means that drone n has a collision risk with other drones; η2 = 1 means that the coverage of drone n overlaps with the coverage of other drones; the logical OR operation ensures that as long as any safety constraint is violated, the cost value c n (t) is 1.

8. The task offloading method of the multi-UAV assisted MEC system based on SafeMARL according to claim 1 is characterized by: Using the MACTD3 algorithm described in S5, each agent n has an Actor network π n , Critic Network and Cost Evaluation Network Q c,n , and the corresponding target network and Among them, the Actor network is used to generate the action strategy of the agent, the Critic network is used to evaluate the cumulative reward under a given state and action, the Cost Critic network is used to evaluate the cumulative cost under a given state and action, and the target network is used to alleviate the instability in the training process; since the multi-agent environment is non-stationary, which will cause the experience replay to fail, the MACTD3 algorithm adopts a centralized training and distributed execution method to find the optimal joint strategy; for each agent n, initialize its Actor network parameters Critic Network Parameters and Cost Critic Network Parameters And the corresponding target network parameters and At the same time, initialize the Lagrange multiplier λ of each agent n ≥0, used to weigh the relationship between reward and cost; At each time step t, each agent n takes its own observed state s n (t), through the Actor network π n Generate decision actions: Among them, a n (t) is the action selected by agent n at time t; the joint action of all agents is denoted as a(t) = (a1(t), a2(t), ..., a N (t)), where N is the total number of agents; after executing the joint action a(t), the environment feeds back the next state s(t+1), the reward r of each agent n (t) and cost c n (t); store the experience tuple (s(t), a(t), r(t), c(t), s(t+1)) into the experience replay buffer; Randomly sample a mini-batch of data from the experience replay buffer Where B is the mini-batch size; for each agent n, using the target Actor network Calculate the next target action and add random noise to smooth the estimate of the target value and reduce the overestimation bias: in, It means that the mean is 0 and the variance is σ 2 The noise is normally distributed and clipped in the interval [-ρ, ρ], ρ is the noise clipping threshold; the joint target action is a′(j+1)=(a1′(j+1),a′2(j+1),…,a′ N (j+1)); and pass through the target Critic network and target Cost Critic network Calculate the target Q-value and target cost Q-value for each agent: Calculate the loss function based on the target Q value And update the Critic network parameters of each agent n by minimizing the loss function. The loss function is as follows: According to the loss function, the gradient descent method is used to update the Critic network parameter formula as follows, where β1 is the neural network learning rate Similarly, by minimizing the loss function L(Q c,n ) to update the CostCritic network parameters of each agent, and the loss function is as follows: The formula for updating the Cost Critic network parameters is as follows: The parameters of each agent’s Actor network are updated through the policy gradient method, and the Lagrange multiplier λ is introduced n , which encourages the agent to avoid violating the preset safety constraint d while optimizing the reward. The policy gradient is expressed as: The Actor network parameters are updated as follows: According to the situation that the agent violates the safety constraint d, the Lagrange multiplier of each agent n is dynamically adjusted, and the update formula is as follows: Among them, β lag >0 is the learning rate. If agent n violates the safety constraint, its Lagrange multiplier will increase, thereby increasing the penalty for violating the constraint. If the behavior of agent n meets the constraint, its Lagrange multiplier will decrease, thereby maximizing the cumulative reward while meeting the constraint. A soft update strategy is used to update the parameters of the target network. The existence of the target network is to alleviate the continuous overestimation or underestimation phenomenon during network training and ensure the stability and convergence of the training. The target network is exactly the same as its corresponding main network architecture. The specific soft update formula is as follows, where τ = 0.005 is the soft update coefficient:

Citation Information

Patent Citations

  • Energy-saving unloading method based on multi-unmanned aerial vehicle cooperative auxiliary edge calculation

    CN113282352A

  • Dynamic task unloading method based on multi-agent reinforcement learning

    CN116880923A

  • Multi-unmanned aerial vehicle auxiliary calculation migration method based on dual delay depth deterministic strategy gradient algorithm

    CN117149434A

  • Electronic device and method with graph generation and task set scheduling

    US20240411592A1

Cited By

  • Calculation unloading and resource allocation method and system based on unmanned aerial vehicle trajectory optimization

    CN120186683A

  • Track design and resource allocation joint optimization method oriented to unmanned aerial vehicle communication, inductance and calculation integration

    CN122120700A

  • An unmanned aerial vehicle (UAV) trajectory design and resource allocation integrated optimization method

    CN122120700B