Task offloading method for multi-uav assisted mec system based on safe marl

By optimizing the flight trajectory and task allocation of a multi-UAV MEC system using the Safe MARL and MACTD3 algorithms, the problems of fixed location and no-fly zone management in traditional MEC systems are solved, achieving efficient allocation of computing resources and security assurance, and improving the system's latency and energy consumption performance.

CN120010949BActive Publication Date: 2025-12-26SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510046008.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-12-26
Estimated Expiration
2045-01-13

AI Technical Summary

Technical Problem

Traditional MEC systems face challenges in handling compute-intensive applications, including fixed location limitations, difficulties in dynamic adjustment, uneven resource allocation, and no-fly zone management, leading to a shortage of computing resources and low task offloading efficiency.

Method used

A multi-UAV assisted MEC system based on Safe MARL is adopted. By solving the adaptive Lagrange multiplier transformation constraint optimization problem, the system combines UAV flight trajectory, computational task allocation and user scheduling, and uses the MACTD3 algorithm to optimize system latency and energy consumption, while ensuring flight safety and compliance.

Benefits of technology

It achieves efficient allocation of computing resources for multi-UAV systems in complex network environments, reduces latency and energy consumption, improves system security and task execution efficiency, and is suitable for dynamic network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010949B_ABST
    Figure CN120010949B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on Safe MARL's multi-unmanned plane assisted MEC system's task unloading method, by cooperating multiple unmanned planes and deploying on low altitude platform edge cloud, provide efficient computing offload service for user equipment.For responding to the dynamic change of user equipment computing demand and the finiteness of unmanned plane resource, while considering the flight safety of unmanned plane, the application proposes a kind of joint optimization unmanned plane flight trajectory, computing task allocation and user scheduling strategy, to realize the minimization of system execution delay and energy consumption, while ensuring the safety of unmanned plane flight.Specifically, the application models the task unloading optimization problem as a constrained Markov decision process, and introduces an advanced multi-agent safe reinforcement learning method, developing a MACTD3 solving algorithm.The method described in the application significantly improves the efficiency and safety of the multi-unmanned plane assisted MEC system in response to large-scale task processing requirements, has wide application prospect and significant technical advantage.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of edge computing, and particularly relates to a task offloading method for a multi-unmanned aerial vehicle (UAV) assisted mobile edge computing (MEC) system based on Safe multi-agent reinforcement learning (MARL). BACKGROUND

[0002] With the rapid development of the sixth generation (6G) wireless network, there are increasing application scenarios with computing-intensive and low-latency requirements, such as autonomous driving and remote medical treatment. However, these tasks may bring great challenges to user equipment (UE) with limited computing power and battery life. In order to solve this challenge, mobile edge computing (MEC) is considered as a promising technology, which allows user equipment to offload its computing-intensive applications to edge servers.

[0003] Mobile edge computing has shown great potential in promoting computing-intensive applications, but its performance is not always satisfactory. This is mainly due to the fixed location of the ground MEC server, which cannot be dynamically adjusted according to user demand. In addition, due to the fixed deployment location of the MEC server and the high construction cost, it has certain limitations in dealing with sudden events and dynamic demand changes. For example, when natural disasters occur and cause damage to network infrastructure, the MEC server cannot provide computing services in a timely manner, resulting in a shortage of computing resources and a decline in task offloading performance.

[0004] In recent years, unmanned aerial vehicles (UAVs) have been widely concerned and researched as an alternative to traditional MEC systems. The UAV-assisted MEC system, with its highly flexible mobility, convenient deployment, low cost, and line-of-sight (LoS) connection advantages, provides effective computing services for service devices. Compared with traditional MEC systems, the UAV-assisted MEC system not only provides temporary computing and communication services in the case of damage to traditional MEC servers, but also optimizes the utilization of computing resources and improves the overall performance and reliability of the system through flexible deployment and scheduling in normal times. This innovative computing paradigm has broad development prospects in dealing with complex and changing network environments in the future.

[0005] However, single-UAV-assisted MEC systems are limited in resources and difficult to cope with the growing demand for massive task processing. And as the service range of UAVs expands and the number of user devices increases, service efficiency will decline sharply. While the multi-UAV-based MEC scheme has higher performance, it also brings many challenges. First, users constantly move during the computing process, making it difficult to obtain the optimal strategy. Second, UAVs need to fly from different takeoff points to provide offloading services in a specific area. Different flight trajectories can cause differences in channel quality, leading to different communication delays and energy consumption. In addition, the amount of computing task allocation for UAVs will also affect the computing delay and energy consumption under the condition of limited on-board resources. Finally, UAVs need to fly in complex urban or rural environments when performing tasks. These areas include no-fly zones regulated by laws or policies, such as the surrounding areas of airports, the airspace above government facilities, and other sensitive areas. To ensure efficient task execution and compliance with relevant flight regulations, UAVs must strictly avoid entering these no-fly zones to ensure flight safety and task compliance. SUMMARY

[0006] To solve the above problems, the present application discloses a task offloading method for a multi-UAV-assisted MEC system based on Safe MARL. On the basis of the MATD3 algorithm, an adaptive Lagrange multiplier is introduced to convert the constrained optimization problem into an unconstrained optimization problem, and then the strategies of UAV flight trajectory, computing task allocation, and user scheduling are combined to minimize the system execution delay and energy consumption while ensuring the safety and compliance of UAV flight.

[0007] To achieve the above purpose, the technical scheme of the present application is as follows:

[0008] The task offloading method for a multi-UAV-assisted MEC system based on Safe MARL includes the following steps:

[0009] First, a multi-UAV-assisted MEC system is designed.

[0010] The system includes M mobile devices (MDs), N UAVs, and an edge cloud (EC) deployed on a static low-altitude platform (LAP). The user device set and the UAV set are represented as and Each UAV is equipped with a small server, and the edge cloud is equipped with a high-performance server. Assume that the total working time of the entire system is T, and the total time is divided into K equal time slots, with the set of time slots denoted as The duration of each time slot is Each user device will generate a compute-intensive task at the start of each time slot t, and the task is denoted as where is the size of the task data, This represents the number of CPU cycles required to process the data. Due to limited computing power, user devices cannot complete calculations locally and need to offload computing tasks to drones. Furthermore, drones can not only act as computing nodes to provide computing services to user devices, but also as relay nodes to further transmit some tasks from user devices to the edge cloud for processing. To reduce system consumption and ensure flight safety, the drone's flight path must be strictly planned. Assume that each drone only provides services to ground user devices within its coverage area, and that there is no overlap between the coverage areas of different drones. There are O no-fly zones within the drone's flight area, such as the area around airports and over government facilities; the set of no-fly zones is denoted as . Drones must avoid these no-fly zones to ensure the safety and compliance of mission execution. The entire system model can be divided into three parts: a drone mobility model, a communication model, and a computational model.

[0011] The drone mobility model defines the following parameters:

[0012] For drones In the time slot The three-dimensional coordinates; where and represents the horizontal and vertical coordinates of drone n at time step t; H represents the flight altitude of the drone. Indicate flight distance; For flight angle; Indicates the maximum flight distance of the drone; Indicates the maximum coverage radius of the drone; This represents the maximum elevation angle of drone n; Indicates the flight speed of the drone; Indicates the flight energy consumption of the drone; This represents the quality of the drone; multiple user devices within the coverage area of ​​a particular drone are served by the same drone. This indicates the number of users served by the drone in time slot t; Let this be a service-related variable, when UE m is served by UAV n, ;otherwise, To avoid interference during data transmission, each user device can only be served by one drone at any given time. Safety factor This indicates whether drone n illegally passed through the no-fly zone in time slot t. If it did, then... If the no-fly zone is successfully avoided, then .

[0013] Each variable is calculated using the following formula:

[0014]

[0015]

[0016]

[0017]

[0018]

[0019] The communication model defines the following parameters:

[0020] is the location of the user equipment; denotes a UAV and a mobile device ; is the distance between them; denotes the probability of establishing a LoS link between the UAV n and the user equipment m at time slot t; denotes the probability of establishing a non-line-of-sight (NLoS) link between the UAV n and the user equipment m; and denote the LoS and NLoS path loss between the UAV n and the user equipment m, respectively; c denotes the speed of light, denotes the carrier frequency, and denote the path loss coefficients for the line-of-sight (LoS) and non-line-of-sight (NLoS) links, respectively; denotes the channel gain between the mobile device m and the UAV n; denotes the channel gain amount when the UAV is 1 m away from the EC; is the uplink bandwidth; denotes the data transmission rate of the mobile device m and the UAV n at time slot t; is the transmit power of the user equipment m; is the Gaussian white noise power; is the transmission delay between the user equipment m and the UAV n; is the receive power of the UAV; is the energy consumed for task delivery between the user equipment m and the UAV n; denotes the location of the EC; is the distance between the UAV n and the EC; is the channel gain between the UAV n and the EC at time slot t; denotes the bandwidth allocated to each UAV; is the data transmission rate between the UAV n and the EC; denotes the transmit power of the UAV n at time slot t; is the maximum transmission power of the UAV; Transmission delay of UAV n from user device m to EC; With Proportion of tasks performed by EC and UAV n respectively for user device m at time slot t; Energy consumption of user device m for data transfer to EC through UAV n.

[0021] Each variable is calculated by the following formula:

[0022]

[0023]

[0024]

[0025]

[0026]

[0027]

[0028]

[0029]

[0030]

[0031]

[0032]

[0033]

[0034]

[0035]

[0036]

[0037] The calculation model defines the following parameters:

[0038] Computational delay of UAV n for processing user device m tasks locally; Computational resources allocated to mobile device m from UAV n; Computational resources of UAV n; Energy consumption coefficient related to UAV CPU; Energy consumption of UAV n for processing mobile device m; Total computational resources of EC; a computing resource allocated to each user equipment by the EC; a computing delay of a part of the task of the mobile device m on the EC after being transmitted to the EC by the UAV n; denotes the total energy consumption of the UAV n at time step t; denotes the total execution delay of the task of the UAV n at time step t.

[0039] Each variable is calculated by the following formula:

[0040]

[0041]

[0042]

[0043]

[0044]

[0045]

[0046]

[0047] Based on the in-depth analysis of the above system model, the present application constructs a constrained optimization problem of a multi-UAV assisted mobile edge computing (MEC) system. By jointly optimizing user scheduling, task allocation ratio, UAV flight trajectory and transmission power, the minimization of the total processing delay and the total energy consumption of the system is realized while ensuring the flight safety of the UAVs and preventing them from colliding or entering the no-fly zone during task execution, thereby improving the overall efficiency and safety of the system. The constrained optimization problem is formalized, wherein is the minimum safety distance between the UAVs, and denote the importance weights of the energy consumption and the execution delay, respectively:

[0048]

[0049] In the above objective function, C1, C2 and C3 represent the constraint conditions of MDs task offloading; C4 restricts the UAVs from passing through the no-fly zone during flight, ensuring the compliance of task execution; C5, C6 and C7 represent the maximum flight distance constraint and working range constraint of the UAVs during flight; C8 restricts the distance between the UAVs to avoid collision; C9 restricts the coverage range of any two UAVs from overlapping to avoid interference during data transmission; and C10 restricts the total energy consumption of the UAVs from exceeding the upper limit of the energy.

[0050] The optimization problem is constructed as a constrained Markov decision process (CMDP), and a MACTD3 algorithm based on multi-agent safe reinforcement learning is proposed to solve this problem. The elements of safe reinforcement learning are defined as follows:

[0051] (1) State space: In the proposed multi-UAV assisted MEC system, the state space is determined by the UAVs, real-time mobile terminal users, and the environment they are in. Therefore, the state space of UAV n can be defined as:

[0052]

[0053] where represents the position information of UAV n, represents the remaining energy of UAV n, represents the position information of mobile device m, represents the task information of mobile device m. M is the total number of users, representing the observation of 1-M user task information, so the capital is used.

[0054] (2) Action space: The agent will perform actions based on the current state of the system and the observed environment. Each agent needs to optimize the position scheduling, task allocation ratio, and upload power to maximize the system effect while avoiding unsafe behavior. Therefore, the action space of UAV n can be defined as:

[0055]

[0056] where , represent the flight distance and flight angle of UAV n, represents the upload power of UAV n, represents the proportion of the task of user device m executed on UAV n.

[0057] (3) Reward function: The reward function is the feedback signal of the environment to the agent's action, which is used to evaluate the agent's behavior and help it improve the strategy. In order to minimize the system consumption while ensuring that the UAV can provide computing offloading services for all user devices (UEs), the reward function is defined as follows:

[0058]

[0059] where, and are the weights used to represent the importance of energy consumption and delay. represents the energy-sensitive scenario, while applies to the delay-sensitive case. In addition, when any user device is not covered by the UAV, all agents will face a penalty term , This represents the penalty weight.

[0060] (4) Cost function: The cost function improves the safety and compliance of the agent in the process of performing tasks. When the agent takes an unsafe action, a corresponding cost will be generated. Therefore, the cost of the agent must not exceed the preset safety threshold. When an intelligent agent performs a task, it must avoid unsafe behaviors. The cost function is defined as:

[0061]

[0062] in, , and This is a binary parameter used to measure whether the agent violated specific security constraints in time slot t. Specifically, This indicates that drone n entered the no-fly zone at time step t; This indicates that drone n is at risk of colliding with other drones at time step t; This indicates that the coverage area of ​​drone n at time step t overlaps with the coverage areas of other drones. The logical OR operation guarantees that if any security constraint is violated, the cost will be zero. That is, 1.

[0063] For the multi-agent secure reinforcement learning algorithm MACTD3 proposed in this invention, each agent n has the following network: Actor network Critic Network and Cost Critic Network and the corresponding target network , , and The MACTD3 algorithm consists of an Actor network for generating the agent's action policy, a Critic network for evaluating the cumulative reward given a state and action, a Cost Critic network for evaluating the cumulative cost given a state and action, and a target network for mitigating instability during training. Since the multi-agent environment is non-stationary, which can lead to the failure of experience replay, the MACTD3 algorithm employs a centralized training and distributed execution approach to find the optimal joint policy.

[0064] For each agent n, initialize its Actor network parameters. Critic network parameters and CostCritic network parameters and the corresponding target network parameters 、 、 and . Meanwhile, the Lagrange multiplier of each agent is initialized to balance the trade-off between the reward and the cost.

[0065] At each time step t, each agent n generates a decision action according to its own observation state by the Actor network :

[0066]

[0067] where is the action chosen by agent n at time t. The joint action of all agents is denoted as , where N is the total number of agents. After executing the joint action , the environment feeds back the next state , the reward and the cost of each agent. The experience tuple is stored in the experience replay buffer.

[0068] A small batch of data is randomly sampled from the experience replay buffer. Where B is the batch size. For each agent n, the target action is calculated by the target Actor network and random noise is added to smooth the estimation of the target value and reduce the overestimation bias:

[0069]

[0070] where represents the noise that is subject to a normal distribution with mean 0 and variance and is clipped within the interval , and is the noise clipping threshold. The joint target action is . The target Q value and the target cost Q value of each agent are calculated by the target Critic network and the target Cost Critic network :

[0071]

[0072]

[0073] The loss function is calculated according to the target Q value, and the Critic network parameters of each agent n are updated by minimizing the loss function, which is as follows:

[0074]

[0075] The Critic network parameters are updated according to the loss function by the gradient descent method, and the formula is as follows: where is the neural network learning rate

[0076]

[0077] Similarly, the Cost Critic network parameters of each agent are updated by minimizing the loss function , and the loss function is as follows:

[0078]

[0079] The Cost Critic network parameter update formula is as follows:

[0080]

[0081] The parameters of the Actor network of each agent are updated by the policy gradient method, and the Lagrange multiplier is introduced to encourage the agent to optimize the reward while avoiding violating the preset safety constraint d. The policy gradient can be expressed as:

[0082]

[0083] The Actor network parameter update is as follows:

[0084]

[0085] According to the situation of the agent violating the safety constraint d, the Lagrange multiplier of each agent n is dynamically adjusted, and the update formula is as follows:

[0086]

[0087] where is the learning rate. If the agent n violates the safety constraint, its Lagrange multiplier will increase, thereby increasing the punishment for violating the constraint. If the behavior of agent n meets the constraint condition, its Lagrange multiplier will decrease, thereby maximizing the cumulative reward under the condition of meeting the constraint.

[0088] The soft update strategy is used to update the parameters of the target network. The existence of the target network aims to alleviate the continuous overestimation or underestimation phenomenon in the network training process, and ensure the stability and convergence of the training. The target network has the same architecture as its corresponding main network, and the specific soft update formula is as follows, where is the soft update coefficient:

[0089]

[0090]

[0091] .

[0092] The beneficial effects of the present application are:

[0093] The present application minimizes the total consumption of the system while ensuring the flight safety of the unmanned aerial vehicle, prevents the unmanned aerial vehicle from colliding or entering the no-fly zone, and constructs the problem as a constraint optimization problem. An advanced multi-agent safety reinforcement learning method is introduced, and a MACTD3 solving algorithm is proposed. The constraint optimization problem is converted into an unconstrained optimization problem by introducing an adaptive Lagrange multiplier, and the unmanned aerial vehicle flight trajectory and transmission power, user scheduling and task allocation strategy are jointly optimized. A large number of simulation results show that, compared with existing schemes, the present application has significant advantages in reducing task execution delay, reducing system energy consumption and improving flight safety, and is suitable for complex dynamic network environment. BRIEF DESCRIPTION OF DRAWINGS

[0094] Figure 1 For the training round number is 4000, each round contains 50 time steps, O=3, N=3, M=20, H=100, =200, the radius of the no-fly zone is 90, the learning curve convergence characteristics diagram of MATD3, MATD3-RS, MACTD3 three algorithms on the reward function.

[0095] Figure 2 For the training round number is 4000, each round contains 50 time steps, O=3, N=3, M=20, H=100, =200, the radius of the no-fly zone is 90, the learning curve convergence characteristics diagram of MATD3, MATD3-RS, MACTD3 three algorithms on the reward function.

[0096] Figure 3 For the training round number is 4000, each round contains 50 time steps, O=3, N=3, M=20, H=100, =200, the radius of the no-fly zone is 90, the flight trajectory diagram of the unmanned aerial vehicle under the guidance of the optimal strategy.

[0097] Figure 4 For the training round number is 4000, each round contains 50 time steps, O=3, N=3, M=20, H=100, =200, the radius of the no-fly zone is 90, the flight trajectory diagram of the unmanned aerial vehicle under the guidance of the optimal strategy.

[0098] Figure 5 For the training round number is 4000, each round contains 50 time steps, O=3, N=3, M=20, H=100, = 200, the forbidden flight zone radius is 90, and the MATD3, MATD3-RS, and MACTD3 are compared in terms of the system total safety cost under different risk zone radii. DETAILED DESCRIPTION

[0099] The present application will be further clarified by the following examples and figures, which should be understood as merely illustrative of the present application and not limiting the scope of the present application.

[0100] As shown in the figure, the task offloading method of the Safe MARL-based multi-UAV assisted MEC system according to the present application comprises the following steps:

[0101] S1: Design a multi-UAV assisted MEC system.

[0102] The system includes M mobile user devices (UDs), N unmanned aerial vehicles (UAVs) with a flight height of H, and an edge cloud (EC) deployed on a static LAP. The user device set and the UAV set are respectively denoted as and Each UAV is equipped with a small server, and the edge cloud is equipped with a high-performance server. Due to limited computing power, user devices cannot complete computing locally and need to offload computing tasks to UAVs. In addition, UAVs can not only provide computing services for user devices as computing nodes, but also further transmit tasks of some user devices to the edge cloud for processing as relay nodes. Each UAV only provides services for ground user devices within its coverage range, and there is no overlap between the coverage ranges of different UAVs. There are O forbidden flight zones in the flight area of the UAVs, and the set of forbidden flight zones is denoted as The UAVs must avoid these forbidden flight zones to ensure the safety and compliance of task execution.

[0103] S2: Define basic elements

[0104] The three-dimensional coordinates of the UAV n at time step t are denoted as ; wherein and respectively represent the horizontal and vertical coordinates of UAV n at time step t; denotes the flight distance; denotes the flight angle; denotes the maximum flight distance of the UAV; denotes the flight speed of the UAV; denotes the flight energy consumption of the UAV; denotes the maximum coverage radius of the UAV; denotes the maximum elevation angle of UAV n; denotes the flight speed of the UAV; denotes the flight speed of the UAV; Indicates the flight energy consumption of the drone; This represents the quality of the drone; multiple user devices within the coverage area of ​​a particular drone are served by the same drone. This indicates the number of users served by the drone in time slot t; Let this be a service-related variable, when UE m is served by UAV n, ,otherwise To avoid interference during data transmission, each user device can only be served by one drone at any given time. Safety factor This indicates whether drone n illegally passed through the no-fly zone in time slot t; Location of user equipment; Indicates drone With mobile devices The distance between them; This represents the probability that a Loss of Space (LoS) link is established between the drone n and the user equipment m in time slot t; This represents the probability of establishing a non-line-of-sight link between the drone n and the user equipment m; and Let LoS and NLoS represent the path loss between the drone n and the user equipment m, respectively; c represents the speed of light. Indicates the carrier frequency. and These represent the path loss coefficients for line-of-sight links and non-line-of-sight links, respectively. This represents the channel gain between the mobile device m and the drone n; This represents the channel gain when the distance between the UAV and the EC is 1m. It is the uplink bandwidth; This represents the data transmission rate between the mobile device m and the drone n within time slot t; This is the transmit power of user equipment m; This represents the power of Gaussian white noise. The transmission delay between UE m and UAV u; This refers to the receiving power of the drone; Energy consumed for mission transfer between user equipment m and drone u; Indicates the location of EC; Let n be the distance between the drone and EC. Let n be the channel gain between the UAV and EC within time slot t; This represents the bandwidth allocated to each drone; The data transmission rate between the drone n and the EC; This represents the transmit power of UAV n within time slot t; Pmax,nis the maximum transmission power of the UAV; Ptrans,nis the transmission delay of the UAV n from the user equipment m to the EC; Ptrans,nis the transmission delay of the UAV n from the user equipment m to the EC; Ptrans,nis the transmission delay of the UAV n from the user equipment m to the EC; Ptrans,nis the transmission delay of the UAV n from the user equipment m to the EC; Ptrans,nis the transmission delay of the UAV n from the user equipment m to the EC; Ptrans,nis the transmission delay of the UAV n from the user equipment m to the EC; Ptrans,nis the transmission delay of the UAV n from the user equipment m to the EC; Ptrans,nis the transmission delay of the UAV n from the user equipment m to the EC; Ptrans,nis the transmission delay of the UAV n from the user equipment m to the EC; Ptrans,nis the transmission delay of the UAV n from the user equipment m to the EC; Ptrans,nis the transmission delay of the UAV n from the user equipment m to the EC; Ptrans,nis the transmission delay of the UAV n from the user equipment m to the EC; Ptrans,nis the transmission delay of the UAV n from the user equipment m to the EC; Ptrans,nis the transmission delay of the UAV n from the user equipment m to the EC;

[0105] S3: Define the calculation method of important performance indicators

[0106]

[0107]

[0108]

[0109]

[0110]

[0111]

[0112]

[0113]

[0114]

[0115]

[0116]

[0117]

[0118] S4: The problem is summarized as a constraint optimization problem based on CMDP, aiming to minimize the total processing delay and total energy consumption while ensuring the flight safety of UAVs, preventing collisions or entering no-fly zones during task execution, thereby improving the overall efficiency and safety of the system. The constraint optimization problem is formalized as follows:

[0119]

[0120] C1, C2 and C3 represent the constraints of MDs task offloading; C4 restricts the UAVs from passing through the no-fly zone during flight, ensuring compliance with task execution; C5, C6 and C7 represent the maximum flight distance constraint and working range constraint of UAVs during flight; C8 restricts the distance between UAVs to avoid collisions; C9 restricts the coverage range of any two UAVs from overlapping to avoid interference during data transmission; C10 restricts the total energy consumption of UAVs from exceeding the upper limit of energy.

[0121] S5: Based on the safety reinforcement learning framework, the safety reinforcement learning elements are designed, including state space, action space, reward, and cost function. The specific design is as follows:

[0122] 1) State space: In the proposed multi-UAV assisted MEC system, the state space is determined by the UAVs, real-time mobile terminal users and the environment they are in, so the state space of UAV n can be defined as:

[0123]

[0124] where represents the position information of UAV n, represents the position information of mobile device m, represents the task information of mobile device m.

[0125] (2) Action space: The agent will perform actions based on the current state of the system and the observed environment. Each agent needs to optimize the position scheduling, task allocation ratio and upload power to maximize the system effect while avoiding unsafe behavior, so the action space of UAV n can be defined as:

[0126]

[0127] where , represent the flight distance and flight angle of UAV n, represent the upload power of UAV n, represent the proportion of task for user device m executed on UAV n.

[0128] (3) Reward function: The reward function is the feedback signal of the environment to the agent's action, which is used to evaluate the agent's behavior and help it improve the policy. In order to achieve the minimization of system consumption while ensuring that the UAV can provide computational offloading services for all user equipment (UEs), the defined reward function is as follows:

[0129]

[0130] wherein and are the weights used to represent the importance of energy consumption and delay, respectively. represents an energy-sensitive scenario, while applies to a delay-sensitive case. In addition, when any user equipment is not covered by the UAV, all agents will face a penalty term , representing the penalty weight.

[0131] (4) Cost function: The cost function improves the safety and compliance of the agent during the execution of the task. When the agent takes unsafe actions, the corresponding cost value will be generated, so the cost value of the agent cannot exceed the preset safety threshold , and the agent must avoid unsafe behavior when executing the task. The cost function is defined as:

[0132]

[0133] wherein , and are binary parameters used to measure whether the agent violates a specific safety constraint at time slot t. Specifically, indicates that the UAV n has entered the no-fly zone; indicates that the UAV n has a collision risk with other UAVs; indicates that the coverage range of the UAV n overlaps with the coverage range of other UAVs. The logical OR operation ensures that as long as any of the safety constraints is violated, the cost value is 1.

[0134] S6: Use the MACTD3 algorithm to solve the problem, and each agent n has the following network: Actor network , Critic network , , cost evaluation network (Cost Critic) , and the corresponding target network , , , and . Among them, the Actor network is used to generate the action policy of the agent, the Critic network is used to evaluate the cumulative reward under the given state and action, the Cost Critic network is used to evaluate the cumulative cost under the given state and action, and the target network is used to alleviate the instability in the training process. For each agent n, initialize its Actor network parameters , Critic network parameters , , Cost Critic network parameters , and the corresponding target network parameters , , and . At the same time, the Lagrange multiplier of each agent is initialized to balance the relationship between reward and cost.

[0135] At each time step t, each agent n generates a decision action according to its own observation state through the Actor network :

[0136]

[0137] where is the action selected by agent n at time t. The joint action of all agents is denoted as , where N is the total number of agents. After executing the joint action , the environment feeds back the next state , the reward and the cost of each agent. The experience tuple is stored in the experience replay buffer.

[0138] A small batch of data is randomly sampled from the experience replay buffer. Where B is the batch size. For each agent n, the target Actor network is used to calculate the target action for the next step, and random noise is added to smooth the estimation of the target value and reduce the overestimation bias:

[0139]

[0140] where represents the noise subject to a normal distribution with mean 0 and variance and clipped within the interval , and is the noise clipping threshold. The joint target action is . And the target Critic network And target Cost Critic network Compute the target Q value and target cost Q value of each agent:

[0141]

[0142]

[0143] According to the target Q value, the loss function is calculated , and the Critic network parameters of each agent n are updated by minimizing the loss function, and the loss function is as follows:

[0144]

[0145] Similarly, the Cost Critic network parameters of each agent are updated by minimizing the loss function , and the loss function is as follows:

[0146]

[0147] The parameters of the Actor network of each agent are updated by the policy gradient method, and the Lagrange multiplier is introduced to encourage the agent to optimize the reward while avoiding violating the preset safety constraint d, and the policy gradient can be expressed as:

[0148] According to the situation of the agent violating the safety constraint d, the Lagrange multiplier of each agent n is dynamically adjusted, and the update formula is as follows:

[0149]

[0150] Wherein, is the learning rate, if the agent n violates the safety constraint, the Lagrange multiplier will increase, so as to increase the punishment for violating the constraint, if the behavior of the agent n meets the constraint condition, the Lagrange multiplier will be reduced, so as to obtain the maximum cumulative reward under the condition of meeting the constraint.

[0151] It should be noted that the above content only illustrates the technical idea of the present application, and cannot be used to limit the protection scope of the present application. For ordinary skilled persons in the technical field, they can make some improvements and refinements without departing from the principles of the present application, and these improvements and refinements fall within the protection scope of the claims of the present application.

Claims

1. A task offloading method for a multi-UAV assisted MEC system based on Safe MARL, characterized in that: Comprising the following steps: S1: design a multi-UAV assisted MEC system, specifically as follows: The system includes M mobile devices, N UAVs, and an edge cloud EC deployed on a static low-altitude platform; each UAV only provides services for the ground mobile devices within its coverage range, and there is no overlap between the coverage ranges of the UAVs; there are 0 no-fly zones within the flight area of the UAVs, and the entire system model is divided into three parts: UAV movement model, communication model, and computing model; S2: calculate performance indicators; S3: constraint optimization formalization; S4: based on the Safe MARL framework, design the four elements of safe reinforcement learning, namely state s(t), action a(t), reward r(t), and cost c(t); the specific design is as follows: 1) State space: In the proposed multi-UAV assisted MEC system, the state space is determined by the UAVs, real-time mobile end users, and the environment in which they exist. The state space of UAV n is defined as: ; wherein represents position information of the UAV n, represents residual energy of the UAV n, represents position information of the M mobile devices, represents task information of the M mobile devices; (2) Action space: The action space of UAV n is defined as: ; wherein , respectively represent the flight distance and the flight angle of the UAV n, represent the upload power of the UAV n, represent the proportion of the task of the mobile device m executed on the UAV n; (3) Reward function: The reward function is the feedback signal of the environment to the agent, which is used to evaluate the agent's behavior and help it improve its strategy. The defined reward function is as follows: ; wherein, and are weights for representing the importance of energy consumption and delay, respectively; represents an energy consumption sensitive scenario, while is suitable for a delay sensitive case; furthermore, when any mobile device is not covered by a drone, all drones will face a penalty term , represents a penalty weight; (4) Cost function: When the agent takes unsafe actions, the corresponding cost value will be generated, and the cost value of the agent cannot exceed the preset safety threshold , the agent must avoid unsafe behavior when performing tasks; the cost function is defined as: ; wherein, , and are binary parameters that measure whether the drone n has violated a particular safety constraint at time slot t; in particular, denotes that drone n has entered a no-fly zone; denotes that drone n is at risk of collision with another drone; denotes that the coverage area of drone n overlaps with the coverage area of another drone; the logical OR operation ensures that the value of is 1 as long as there is a violation of any of the safety constraints. S5: solving the optimal strategy using MACTD3 algorithm, each UAV n has an Actor network , a Critic network , a Cost Critic network , and a corresponding target network , , , and ; wherein the Actor network is used to generate the action policy of the agent, the Critic network is used to evaluate the cumulative reward under a given state and action, the Cost Critic network is used to evaluate the cumulative cost under a given state and action, and the target network is used to alleviate instability in the training process; the MACTD3 algorithm adopts a centralized training and distributed execution method to find the optimal joint strategy.

2. The task offloading method for the Safe MARL-based multi-UAV assisted MEC system according to claim 1, wherein: The UAV movement model defines the following parameters: for a UAV at a time slot in a three-dimensional coordinate; wherein and respectively represent the horizontal longitudinal coordinates of the UAV n at the time slot t; H represents the flight height of the UAV; represents the flight distance; is the flight angle; represents the maximum coverage radius of the UAV; represents the maximum elevation angle of the UAV n; represents the flight speed of the UAV; represents the flight energy consumption of the UAV; represents the mass of the UAV; Each variable is calculated by the following formula: ; ; ; ; 。 3. The task offloading method for the Safe MARL-based multi-UAV assisted MEC system according to claim 2, wherein: The communication model defines the following parameters: Location of mobile device; Indicates drone With mobile devices The distance between them; This represents the probability that a Loss of Space (LoS) link is established between drone n and mobile device m in time slot t; This represents the probability of establishing a non-line-of-sight link between the drone n and the mobile device m; and These represent the LoS and NLoS path losses between the drone n and the mobile device m, respectively; c represents the speed of light. Indicates the carrier frequency. and These represent the path loss coefficients for line-of-sight links and non-line-of-sight links, respectively. This represents the channel gain between the mobile device m and the drone n; This represents the channel gain when the distance between the UAV and the EC is 1m. It is the uplink bandwidth; This represents the data transmission rate between the mobile device m and the drone n within time slot t; This is the transmit power of the mobile device m; The power of Gaussian white noise; The transmission delay between the mobile device m and the drone n; This refers to the receiving power of the drone; The energy consumed for task transfer between mobile device m and drone n; Indicates the location of EC; Let n be the distance between the drone and EC. Let n be the channel gain between the UAV and EC within time slot t; This represents the bandwidth allocated to each drone; The data transmission rate between the drone n and EC; This represents the transmit power of UAV n within time slot t; This is the maximum transmission power of the drone; Let n be the transmission delay from the drone n to the mobile device m and then to the EC. and These represent the proportions of tasks performed by the EC and the drone n in time slot t, respectively, for tasks performed by mobile device m. Energy consumption for mobile device m to transmit data to EC via drone n; Indicates the maximum flight distance of the drone; multiple mobile devices within the coverage area of ​​a particular drone are served by the same drone. This indicates the number of users served by the drone in time slot t; Each variable is calculated by the following formula: ; ; ; ; ; ; ; ; ; ; ; ; ; ; 。 4. The task offloading method for the Safe MARL-based multi-UAV assisted MEC system according to claim 3, characterized in that: The computational model defines the following parameters: the computational delay for the drone n to process the mobile device m task locally; the computational resources allocated to the mobile device m from the drone n; the computational resources of the drone n; the energy consumption coefficient related to the drone CPU; the energy consumption of the drone n to process the mobile device m; the total computational resources of the EC; the computational resources allocated to each mobile device by the EC; the computational delay of the partial task of the mobile device m after being transmitted to the EC by the drone n on the EC; denotes the total energy consumption of the drone n at time slot t; the total execution delay of the task of the drone n at time slot t; denoted as a service association variable when the mobile device m is served by the drone n, ; Otherwise, ; to avoid interference during data transmission, each mobile device can only be served by one UAV at any time, i.e. ; Each variable is calculated by the following formula: ; ; ; ; ; ; 。 5. The method of claim 4, wherein the Safe MARL based multi-UAV assisted MEC system task offloading method is characterized by: The formalization described in S3 is to reduce the problem to a constraint optimization problem based on CMDP, aiming to minimize system consumption while ensuring the flight safety of UAVs. The specific objective function is as follows: ; In the above objective function, C1, C2, and C3 represent the constraints of mobile device task offloading; C4 constrains the UAVs from passing through no-fly zones during flight, ensuring compliance with task execution; C5, C6, and C7 represent the maximum flight distance constraints and working range constraints for UAVs during flight; C8 constrains the distance between UAVs to avoid collisions; C9 constrains the coverage range of any two UAVs to avoid interference during data transmission; and C10 constrains the total energy consumption of UAVs to not exceed their energy upper limit; Safety factor represents whether the UAV n illegally passes the no-fly zone in the time slot t, if passing the no-fly zone, if successfully avoiding the no-fly zone, .

6. The task offloading method for the Safe MARL-based multi-UAV assisted MEC system according to claim 5, wherein: Step S5 is as follows: For each drone n, initialize its Actor network parameters. Critic network parameters and CostCritic network parameters and the corresponding target network parameters , , and Simultaneously, initialize the Lagrange multipliers for each UAV. This is used to balance the relationship between rewards and costs; At each time slot t, each drone n generates a decision action based on its own observation state , by the Actor network : ; wherein, is the action selected by drone n at time slot t; the joint action of all drones is denoted as where N is the total number of drones; the joint action is performed After that, the environment feeds back the next state , the reward of each drone and the cost The experience tuple is stored into the experience replay buffer; randomly sample a mini-batch of data from the experience replay buffer ; where B is the mini-batch size; for each drone n, utilize the target Actor network compute the next target action and add random noise to smooth the estimate of the target value and reduce overestimation bias: ; wherein, represents noise that is subject to a normal distribution with mean 0 and variance and is clipped within the interval is a noise clipping threshold; the joint target action is ; and the target Critic network and the target Cost Critic network compute the target Q value and the target cost Q' value for each drone.​ ; ; According to the target Q value, a loss function is calculated And the Critic network parameters of each drone n are updated by minimizing the loss function, and the loss function is as follows: ; According to the loss function, the Critic network parameter is updated by the gradient descent method, and the formula is as follows, wherein is the neural network learning rate ; Similarly, the Cost Critic network parameters for each drone are updated by minimizing a loss function , which is given by: ; The formula for updating the Cost Critic network parameters is as follows: ; The parameters of the Actor network for each UAV are updated using the policy gradient method, and Lagrange multipliers are introduced. This prompts the drone to optimize rewards while avoiding violating the preset safety constraint d. The policy gradient is represented as: ; The Actor network parameter update is as follows: ; According to the situation of UAVs violating safety constraints d, the Lagrange multiplier of each UAV n is dynamically adjusted, and the update formula is as follows: ; wherein, is the learning rate, if the UAV n violates the safety constraint, its Lagrange multiplier will increase, thus increasing the punishment for violating the constraint, if the behavior of the UAV n meets the constraint condition, its Lagrange multiplier will decrease, thus obtaining the maximum cumulative reward in the case of meeting the constraint; The parameter of the target network is updated by using a soft update strategy, and the target network is used to alleviate the continuous overestimation or underestimation in the network training process, and ensure the stability and convergence of the training. The target network has the same architecture as the corresponding main network, and the specific soft update formula is as follows, wherein is a soft update coefficient. ; ; 。

Citation Information

Patent Citations

  • Dynamic task unloading method based on multi-agent reinforcement learning

    CN116880923A

  • Multi-unmanned aerial vehicle auxiliary calculation migration method based on dual delay depth deterministic strategy gradient algorithm

    CN117149434A