Dual-uav edge computing system security offloading method based on multi-agent reinforcement learning
By optimizing the 3D trajectory and dynamic resource allocation of dual UAVs using the MADDPG algorithm, the risk of data leakage caused by multiple ground-based eavesdropping in multi-UAV edge computing systems is resolved, achieving a balance between security and performance and improving user satisfaction.
Patent Information
- Application Number
- CN202410208483.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-26
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2044-02-26
AI Technical Summary
In multi-drone-assisted edge computing systems, ground users face a high risk of data leakage due to multiple eavesdroppers. Furthermore, traditional optimization methods struggle to effectively address the complexity of resource allocation and task scheduling, as well as environmental uncertainties, leading to a decline in system performance.
The deterministic policy gradient (MADDPG) algorithm based on multi-agent deep reinforcement learning is adopted to jointly optimize the 3D trajectory and dynamic resource allocation of dual UAVs. Through collaborative learning and adaptive adjustment, the computational cost for system users is reduced and user satisfaction is improved.
While ensuring the security of user uninstallation data, it effectively reduces the normalized energy consumption and latency of system users, improves user satisfaction, and adapts to changing real-world application scenarios.
Smart Images

Figure CN118139013B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of multi-unmanned aerial vehicle assisted mobile edge computing, in particular to an offloading method for guaranteeing user offloading data security based on multi-agent deep reinforcement learning. BACKGROUND
[0002] With the rapid development of Internet of Things and 5G technology, the popularity of sensors and intelligent devices, and the increasing demand for real-time and low-latency, cloud computing has high latency and network congestion problems when processing massive data. Mobile edge computing (MEC) is a new computing model that originated from the challenges and changes in demand for traditional cloud computing. It brings computing resources close to data sources and terminal devices, enabling local data processing and analysis, thereby reducing the need for data transmission to the cloud and improving system response speed. This computing model takes advantage of the flexibility and flexibility of edge nodes to better meet the requirements of real-time, privacy protection and network resource utilization efficiency, providing an effective solution to the challenges faced by current cloud computing. However, in remote or mountainous areas, terminal devices face many difficulties. This is because it is difficult to obtain stable and reliable MEC servers and infrastructure coverage in these areas, resulting in severe limitations in computing and communication. This situation makes it extremely difficult to achieve efficient data processing and instant communication in these areas, becoming a problem that needs to be solved.
[0003] To solve the above problems, unmanned aerial vehicles have been widely used to assist MEC systems in performing computing-intensive tasks and providing edge computing services for remote areas due to their flexible deployment and wide coverage. Specifically, MEC servers are mounted on unmanned aerial vehicles, and users offload data to unmanned aerial vehicles for processing. Unmanned aerial vehicles act as flying MEC servers and can provide flexible services to ground users.
[0004] Although the multi-unmanned aerial vehicle assisted MEC system can effectively reduce the latency and energy consumption of users, the broadcast nature of the wireless channel makes the line-of-sight link communication between the unmanned aerial vehicle and the ground user also benefit from the malicious eavesdropping on the ground, especially when there are multiple eavesdropping collusions and non-collusions on the ground. Cooperation and information sharing among malicious eavesdroppers pose a serious risk of security breaches to the system. To solve the problem of data leakage in wireless communication systems, physical layer security technology is widely used in wireless communication systems. Unlike traditional encryption methods, physical layer security technology improves the resistance of communication systems to external attacks by utilizing the physical principles of nature, providing a higher level of security for information transmission.
[0005] In view of the complexity of resource allocation and task scheduling in MEC networks and the uncertainty of the environment, traditional optimization methods are usually difficult to effectively solve such problems. Using multi-agent deep reinforcement learning to solve the optimization problem of multi-UAV assisted edge computing system can effectively cope with the demand of complex and dynamic environment. Multi-intelligent deep reinforcement learning can learn and adaptively adjust to effectively solve the optimization problem in UAV cooperative edge computing. This method can flexibly cope with changes in task demand and network conditions in real-time scenarios by learning and optimizing decisions, improving system efficiency and performance. The cooperative effect of multi-intelligence can better tap resources in multi-UAV assisted edge computing, realize task cooperative execution, and optimize the overall performance of the system, making it more suitable for changing real-world application scenarios. SUMMARY
[0006] The present application considers the user offloading data security transmission problem of a dual-UAV assisted edge computing system in the presence of multiple ground eavesdropping, where one UAV carries a MEC server to provide computing services for ground users, and the other UAV acts as a jammer to transmit interference signals to the ground eavesdropping. The purpose is to improve the demand satisfaction of system users under the premise of ensuring the security of user offloading data. By jointly optimizing the 3D trajectory of the dual-UAV, user transmit power and computing frequency to reduce the computing cost of system users, the demand satisfaction of users is improved. Because the problem is a non-convex multi-joint problem, it is difficult to solve using traditional optimization algorithms. Therefore, the present application proposes a multi-agent deep deterministic policy gradient (MADDPG) algorithm based on deep reinforcement learning to solve the 3D trajectory of the dual-UAV and the dynamic resource allocation scheme. The purpose of the present application is to provide a dual-UAV assisted edge computing system UAV 3D trajectory and dynamic resource optimization method based on the MADDPG algorithm, considering the case where multiple ground eavesdropping steals user information, reducing the computing cost of system users under the premise of ensuring the security of user offloading data, thereby improving the demand satisfaction of users.
[0007] A dual-UAV edge computing system security offloading method based on multi-agent reinforcement learning, comprising the following steps:
[0008] Step S1: Construct a dual-UAV assisted edge computing model considering the presence of multiple ground eavesdropping, where one UAV acts as an aerial edge computing server to provide edge computing services for ground user devices, and the other UAV acts as an aerial jammer to transmit interference signals to the ground multiple eavesdropping.
[0009] Step S2: Based on the system model provided in step S1, the average computing cost of the system user is calculated under the premise of ensuring the security of the system user offloading data. The inverse of the computing cost is taken as the user demand satisfaction degree, and the user demand satisfaction degree is considered comprehensively considering the security of the user offloading data, the computing delay and the energy consumption. The optimization problem is constructed by taking maximizing the user demand satisfaction degree as the optimization goal.
[0010] Step S3: The optimization problem is modeled as a Markov decision process, including the setting of the system state space, the action space and the reward function.
[0011] Step S4: Due to the continuity of the action and the large dimension of the solution, the MADDPG algorithm in deep reinforcement learning is used to jointly optimize the 3D trajectory and dynamic resource allocation strategy of the dual unmanned aerial vehicle to reduce the computing cost of the system user, thereby improving the demand satisfaction degree of the user.
[0012] Further, the dual unmanned aerial vehicle assisted edge computing model includes K ground users U k unload data to the server unmanned aerial vehicle S through a time division mode, and M potential eavesdropping users E m constantly attempt to steal the information offloaded by the user, and a friendly unmanned aerial vehicle J assists S in suppressing eavesdropping. m Send an interference signal to suppress eavesdropping.
[0013] Further, the user demand satisfaction degree in step S2 is:
[0014]
[0015] represents the normalized energy consumption, represents the normalized delay, x k is set to 1, and a k represents the energy consumption control weight.
[0016] Further, the optimization problem in step S2 is
[0017]
[0018]
[0019]
[0020]
[0021]
[0022] C5:E S (N)≥0,E J (N)≥0,
[0023]
[0024]
[0025]
[0026] where C1-C3 represent the limits on user transmit power, user computation frequency and safety offloading threshold, respectively; C4 is the constraint on UAV computation capability; C5 is the energy consumption limit on UAV; C6-C8 represent the constraints on UAV flight speed, flight height and collision avoidance, respectively, where RS(k) is the user demand satisfaction, p k (n) is the user transmit power, f k (n) is the user computation frequency; is the user instantaneous safety reachable speed, is the user safety offloading threshold; is the data amount offloaded by the nth time slot user, is the maximum computation frequency of the server UAV, C S is the CPU computation frequency required by the server UAV to compute 1-bit data, δ t is the time slot length, E S (N) and E J (N) represent the remaining energy of the server UAV S and the interfering UAV J, respectively, v S (n) and v J (n) are the flight speed of S and J, respectively, and represent the maximum flight speed of S and J, respectively; z S (n) and z J (n) are the flight height of S and J, respectively, z min is the minimum flight height, z max is the maximum flight height; q S (n) and q J (n) are the positions of S and J, respectively, d min is the minimum safety distance to prevent collision between two UAVs.
[0027] Further, the modeling of the optimization problem as a Markov decision process in step S3 comprises:
[0028] the state set of the server UAV S is the state set of the friendly UAV J is the state set of the system is:
[0029]
[0030] q S (n), qJ (n) denote the coordinate position of S and J, respectively, E S (n), E J (n) denote the residual energy of the nth time slot S and J, respectively, denote the instantaneous safe rate of the system, L k (n) denote the flight speed of the nth time slot U k unprocessed data. denote the channel gain between J and E m .
[0031] The action set of the server drone S is The action set of the friendly drone J is The action set of the system is:
[0032]
[0033] v S (n) and v J (n) denote the flight speed of S and J, respectively, θ S (n) and θ J (n) denote the polar angle of the nth time slot S and J, respectively, and denote the horizontal angle of S and J, respectively.f k (n) is U k The frequency is calculated in the nth time slot, p k (n) denotes the transmission power of the nth time slot user;
[0034] The reward function of the server drone S is:
[0035]
[0036] where, is the optimization target, r off (n) is the reward for the user to offload data, r p,S (n) is the punishment S receives for violating the constraints; where, r off (n) is represented as:
[0037]
[0038] where, κ f is a positive integer that adjusts the reward r off (n), denote the instantaneous safe rate of the system, δ t denotes the time slot length;
[0039] r p,S (n) is defined as:
[0040] r p,S(n) = - λ ac (n) κ ac,S - λ rc (n) κ rc,S + κ er,S (n)
[0041] where κ ac,S , κ rc,S , κ er,S (n) represent the punishment of the collision constraint, the computation capability constraint and the residual energy constraint of the UAV S respectively, λ ac (n) and λ rc (n) are binary coefficients, κ er,S (n) is used to judge the sparse reward related to the residual energy of S after all the user data are processed, and is defined as:
[0042]
[0043] where ζ is a positive number used to adjust κ er,S (n), and E S (n) represents the residual energy of S in the nth time slot; the reward function of the friendly UAV J is:
[0044]
[0045] where κ g is a positive integer used to adjust the total eavesdropping rate, r k (n) represents the instantaneous achievable rate between U m and E p,J , and r p,J (n) is the punishment of J for violating the constraint condition, and is represented as:
[0046] r ac (n) = - λ ac,J (n) κ er,J + κ ac,J (n).
[0047] where κ er,J , κ er,J (n) are the collision punishment and the energy residual punishment of J respectively, and κ J (n) is defined as:
[0048]
[0049] where ζ er,J is a positive number used to adjust κ S (n).
[0050] The overall reward of the system r(n) can be defined as
[0051] r(n) = [r J(n), r J (n)].
[0052] Further, the MADDPG algorithm in step S4 includes the following steps:
[0053] (1) Initialization: for each agent, initialize its policy network, action network, target network and experience replay buffer;
[0054] (2) Experience sampling: each agent interacts with the environment according to the action output by its policy network, collects experience and stores it in the experience replay buffer;
[0055] (3) Training: each agent randomly samples a batch of experience from the experience replay buffer, calculates the gradient of its action value function and updates its policy network and action network;
[0056] (4) Target network update: every certain time step, the target network of each agent is updated softly, that is, the parameters of the target network slowly approach the parameters of the action network;
[0057] (5) Shared experience: when each agent is sampling experience, it shares the experience of other agents, thereby accelerating the learning process.
[0058] (6) Coordinated strategy: in each training, the policy network of each agent can consider the actions of other agents, thereby realizing the collaborative decision-making among multiple agents.
[0059] The objective function of each agent in the MADDPG algorithm is defined as:
[0060]
[0061] wherein J(θ i ) is the performance measure of the policy θ i of agent i, γ is the discount factor, t is the time step, is the immediate reward obtained by agent i at time step t, E[.] is the expectation operation on the cumulative reward;
[0062] The policy gradient of each agent can be calculated by the following formula:
[0063]
[0064] wherein is the gradient of the policy parameter θ i of agent i, E o,a~D [.] is the expectation operation on the sampled state o and action a in the experience pool D, and D is the experience pool storing the state value and action of the agent in the process of interacting with the environment, Policy function of agent i i Gradient with respect to its parameters theta, Action value function of agent i, representing the expected cumulative reward under a given state o and action a according to the current policy pi, Gradient of the action a i , a i = pi i is the action selected by the agent under the policy pi i under the state o i , represents the action taken by the agent according to its own policy.
[0065] A computer readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the steps of the multi-agent reinforcement learning based dual-UAV edge computing system security offloading method described above.
[0066] A communication system comprising a plurality of users, two UAVs, one of which serves as an aerial edge computing server to provide edge computing services for ground user equipment, and the other of which serves as an aerial jammer to transmit interference signals to ground multi-eavesdropping, the UAVs having a processor and a memory, the processor being configured to execute executable instructions stored in the memory to implement the method described above.
[0067] The present application adopts the MADDPG algorithm to jointly optimize the 3D trajectory and dynamic resource allocation strategy of the dual-UAV to reduce the computing cost of the system users, thereby improving the demand satisfaction of the users.
[0068] After sufficient training, when the cumulative reward of the system tends to be stable, the training process is stopped; the training results are deployed with the UAV base station platform to guide the UAV system to quickly and efficiently perform tasks in practice, so as to minimize the latency and energy consumption of the system users, thereby improving the demand satisfaction of the users.
[0069] The present application has the following beneficial technical effects relative to the prior art:
[0070] 1. The multi-agent reinforcement learning based dual-UAV edge computing system security offloading method provided by the present application considers the security problem of user offloading data in the presence of multiple malicious eavesdroppers on the ground, while ensuring the security of user offloading data, the normalized energy consumption and latency of the system users are reduced as much as possible, and at the same time, the user can set the corresponding energy consumption weight and latency weight according to his own needs, thereby improving the demand satisfaction of the system users.
[0071] 2. This invention considers both conspiratorial and non-conspiratorial eavesdropping scenarios in a drone-assisted edge computing system, designing dual-drone 3D trajectory and user resource optimization schemes for each of the two security scenarios. In the conspiratorial eavesdropping scenario, we assume all eavesdroppers share information, and the system eavesdropping rate is the sum of the rates of all eavesdroppers. An optimization scheme is designed under the worst-case system security condition to ensure the security of user data uninstallation. In the non-conspiratorial eavesdropping scenario, the eavesdropping rate of the most capable eavesdropper is selected as the system eavesdropping rate. Again, an optimization scheme is designed under the worst-case system security condition to ensure the security of user data uninstallation.
[0072] 3. This invention addresses the control problem of high-dimensional, continuous action space in edge computing systems assisted by dual UAVs. This invention employs the MADDPG algorithm, which optimizes through cooperation among agents to obtain an effective user resource allocation scheme and a dual UAV 3D trajectory strategy. Attached Figure Description
[0073] Figure 1 Diagram of a secure communication model for edge computing assisted by dual drones;
[0074] Figure 2 This is a diagram of the MADDPG algorithm structure. Detailed Implementation
[0075] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0076] A secure offloading method for a dual-UAV edge computing system based on multi-agent reinforcement learning includes the following steps:
[0077] Step S1: Construct a secure communication model for a dual-UAV-assisted edge computing system.
[0078] like Figure 1 As shown, to achieve secure offloading functionality in a dual-UAV edge computing system based on multi-agent reinforcement learning, this step constructs a dual-UAV-assisted edge computing system. One UAV carries an MEC server to provide computing services to ground users, while the other UAV acts as a jammer, transmitting jamming signals to multiple ground-based eavesdroppers. Through the collaborative cooperation of the two UAVs, while ensuring the security of system user data, the system's normalized energy consumption and latency are minimized by designing the 3D trajectories of the two UAVs, user transmission power, and computing frequency, thereby improving user satisfaction. In the training process, the proposed secure offloading method for a dual-UAV edge computing system based on multi-agent reinforcement learning uses a simulated virtual urban environment, where the virtual UAV processes the offloading data from ground users within this virtual urban environment. Therefore, before constructing the deep reinforcement learning network framework, it is necessary to model the virtual urban environment first.
[0079] likeFigure 1 A secure edge computing system with dual-UAV assistance is considered, where K ground users (U k , k = 1,..., K) offload data to a MEC server drone S through a time-division pattern to reduce their own latency and energy consumption. M potential eavesdroppers (E m , m = 1,..., M) constantly attempt to steal the information offloaded by users. To weaken the eavesdropping ability, a friendly jammer drone J is set to assist S to send interference signals to suppress eavesdropping. All devices in the system are equipped with a single antenna. m
[0080] The task data L k of each user needs to be processed within the flight time T t of S, T t is evenly divided into N time slots, each with a length of δ t = T t / N, assuming δ t is small enough so that the position of S and J in each time slot is considered fixed. We use three-dimensional Cartesian coordinates to represent the position of S and J, denoted as q S (n) = [x S (n), y S (n), z S (n)] T and q J (n) = [x J (n), y J (n), z J (n)] T , respectively. The positions of ground users U k and eavesdroppers E m are denoted as and respectively. In the formula, x represents the horizontal coordinate, y represents the vertical coordinate, and z represents the flight height of the drone.
[0081] Assuming that the link between the user and the drone adopts a probabilistic line-of-sight link model, the line-of-sight link probability is represented as:
[0082]
[0083] where d is the distance between user U k and S, z S (n) is the flight height of S in the nth time slot, and η a and η b are environment-dependent hyperparameters.
[0084] The average path loss between U k and S is:
[0085]
[0086] where, is the probability of line-of-sight link transmission, and are the path loss of line-of-sight and non-line-of-sight links, respectively. The free space path loss is given by where, f c is the carrier frequency, c is the speed of light, η LoS and η NLoS are the additional path loss of line-of-sight and non-line-of-sight links, respectively.
[0087] The channel gain between the nth time slot U k and S is given by:
[0088]
[0089] Similarly, the channel gain between J and E m can be calculated as
[0090] The channel between U k and the eavesdropper is modeled as an independent Rayleigh fading model, thus the channel gain between U k and the eavesdropper E m is given by:
[0091]
[0092] where, is the distance between U k and the eavesdropper E m , β0is the unit channel gain, a is the path loss exponent, ξ e is the Rayleigh fading coefficient, which follows an exponential distribution with mean 1.
[0093] Let P J > 0 be the transmit power of J, p k (n) be the transmit power of the nth time slot user, p k (n) should satisfy:
[0094] 0 < p k (n) < P k max (5)
[0095] where, P k max is the maximum transmit power of U k .
[0096] Thus, the instantaneous achievable rate between U k and S is given by, in unit of (bps / Hz):
[0097]
[0098] where B is the channel bandwidth, is the additive white Gaussian noise at S. U k and E m The instantaneous achievable rate between U
[0099]
[0100] where, is the additive white Gaussian noise at E.
[0101] When there is no collusion among the eavesdroppers, the instantaneous secure achievable rate of the system is:
[0102]
[0103] When there is collusion among the eavesdroppers, the instantaneous secure achievable rate of the system is:
[0104]
[0105] where, [x] + = max{x, 0}.
[0106] The instantaneous secure rate of the system can be summarized as:
[0107]
[0108] In order to protect the security of the offloading data of users, the secure rate should satisfy:
[0109]
[0110] where, is the offloading threshold, only when the instantaneous secure rate of a user is greater than or equal to the offloading threshold, the user can offload data.
[0111] Suppose we select the user with the maximum instantaneous secure rate to communicate with S in each time slot, and use θ k (n) to represent the scheduling coefficient between U k and S, θ k (n) is defined as:
[0112]
[0113] Local computation
[0114] In order to save the system energy consumption and shorten the delay as much as possible, U k and S both use dynamic frequency and voltage adjustment technology, suppose f k (n) is the frequency of U k in the nth time slot, Ck For U k CPU cycles required to compute 1 bit of data, For U k The effective capacitance coefficient of each time slot U k The data calculated locally is:
[0115]
[0116] f k (n) should satisfy:
[0117] 0≤f k (n)≤F k max , (14)
[0118] Where F k max is the maximum calculation frequency of U k .
[0119] The energy consumption of each time slot U k The local calculation energy consumption is:
[0120]
[0121] Safe uninstall
[0122] When the user's safety rate is greater than 0, the user will unload data to S for processing, then the data and energy consumption of the nth time slot user for safe uninstall are:
[0123]
[0124]
[0125] Assuming that the unloaded data is irrelevant, there is no need to reconstruct the data of the calculation result at S every time slot. Considering that the data amount of the actual calculation result is much smaller than the data amount of the safe uninstall, the invention omits the latency and energy consumption of returning the calculation result.
[0126] Therefore, the total amount of data calculated and the energy consumption of each time slot U k are:
[0127]
[0128]
[0129] The nth time slot U k The unprocessed data is represented as:
[0130]
[0131] Where Lk (0) = L k .
[0132] Define T k (n) is the state of the user's remaining data in the nth time slot, denoted as:
[0133]
[0134] The execution delay of user U k is:
[0135]
[0136] The user U k computes the energy consumption as:
[0137]
[0138] The energy consumption of S is mainly composed of three parts, namely propulsion energy consumption, computing energy consumption and communication energy consumption. Since the communication energy consumption is usually two orders of magnitude smaller than the propulsion energy consumption, the present application ignores the communication energy consumption, and the propulsion power is represented as:
[0139]
[0140] Wherein, v(n) is the real-time speed of the unmanned aerial vehicle, P0 and P i are two constants, respectively representing inherent blade type power and induced power. U tip represents the blade tip speed, v0 is the average rotor induced speed when hovering, d0, s, p, A respectively represent the fuselage drag ratio, the wind wheel solidity, the air density and the rotor disc area.
[0141] The propulsion energy consumption of S and J in the nth time slot is respectively:
[0142]
[0143]
[0144] Wherein, and are the propulsion power of S and J in the nth time slot.
[0145] Assuming that the data unloaded by each time slot user can be processed by S, the computing capacity of S should satisfy:
[0146]
[0147] Wherein, is the maximum computing capacity of S, C S is the CPU cycle number required by S to calculate 1 bit of data. S is U kThe calculation frequency of distribution is:
[0148]
[0149] The calculation energy consumption produced by each time slot S is:
[0150]
[0151] The residual energy of the nth time slot S is expressed as:
[0152]
[0153] Wherein, is the battery capacity of S,
[0154] The residual energy of the nth time slot J is expressed as:
[0155]
[0156] Similarly, is the battery capacity of J,
[0157] In order to ensure that the calculation tasks of all users can be processed before the energy of S and J is exhausted, the energy of S and J should satisfy:
[0158] E S (N)≥0,E J (N)≥0. (31)
[0159] The present application adopts spherical coordinates to describe the speed and flight direction of S and J, wherein v represents the flight speed, θ is the polar angle, is the horizontal angle.
[0160] Let v S (n) and v J (n) represent the flight speed of S and J respectively, then v S (n) and v J (n) should satisfy:
[0161]
[0162] Wherein, and represent the maximum flight speed of S and J respectively.
[0163] Meanwhile, the flight height of S and J should satisfy:
[0164] Z min ≤z S (n),z J (n)≤Z max, (33)
[0165] wherein Z min represents the minimum height of the UAV flight, Z max represents the maximum height of the UAV flight.
[0166] In order to avoid the collision of the UAV, a minimum safety distance d min , d min should be satisfied:
[0167] ||q S (n)-q J (n)||≥d min , (34)
[0168] In the system, the user adopts a partial offloading strategy.
[0169] In the present application, the communication link between the ground node and the UAV is a probabilistic line-of-sight link model, and the channel gain between the user and the server UAV S and the channel gain between the interference UAV J and the eavesdropping can be calculated by equations (1)-(3), respectively. The channel between the user and the eavesdropping is modeled as a Rayleigh fading model, and the channel gain is calculated by equation (4). After calculating the channel gain of each link, the transmission rate between the user and the server UAV S and the transmission rate between the user and the eavesdropping are calculated by equations (6) and (7), respectively, and the difference between the two is the instantaneous achievable safety rate of the user. The eavesdropping rate under the eavesdropping collusion and non-collusion conditions is calculated by equations (8) and (9), respectively, and is summarized as shown in equation (10). The instantaneous safety achievable rate of the user should satisfy the condition equation (11) to offload data. The user and the UAV communicate in a scheduling manner, and the user with the maximum instantaneous safety rate is selected to communicate with the server UAV S in each time slot, and the user selection rule is shown in equation (12).
[0170] In the edge computing model, the local computing and offloading computing of the user will be carried out at the same time, and the data amount and the local computing energy consumption of the user in each time slot are calculated according to equations (13)-(15), and the offloaded data amount and the offloading energy consumption of the user in each time slot are calculated according to equations (16) and (17). The delay and energy consumption of the user are calculated according to equations (18)-(22).
[0171] In the present application, rotary-wing UAVs are used and the energy consumption of each UAV is considered to ensure that all user data can be processed before the UAV runs out of energy. The energy consumption of UAV S is mainly composed of two parts: calculation energy consumption and flight energy consumption. The calculation energy consumption can be obtained by formula (28), and the flight energy consumption can be obtained by formula (23) and (24). UAV J only considers flight energy consumption, which can be calculated by formula (23) and (25). Formulas (31)-(34) correspond to UAV energy consumption, speed, flight height and collision constraint, respectively.
[0172] Step S2: Determine the system optimization target: maximize user demand satisfaction as the optimization target under the premise of ensuring user offloading data security.
[0173] On the basis of ensuring offloading security, the user's business demand for reducing latency and saving energy is considered. Due to the large difference in units and dimensions of time and energy, simple weighted combination cannot provide a more fair preference index for balancing the reduction of latency and energy saving. Therefore, it is necessary to solve the dimensional difference between them and provide on-demand services.
[0174] When the user performs partial security offloading, the normalized energy consumption and the normalized latency are used to reduce the dimensional difference of energy consumption and latency, and can intuitively reflect the energy saving degree and latency reduction degree brought by secure partial computation offloading, which are respectively defined as:
[0175]
[0176]
[0177] wherein, and are the U k with the maximum computation frequency F k max The energy consumption and latency of performing complete local computation are defined as:
[0178]
[0179]
[0180] According to the above analysis, the user demand satisfaction (RS) index is further defined, including providing security protection for users, reducing latency and saving energy. The demand satisfaction index is defined as:
[0181]
[0182] During the calculation and unloading process, security is a fundamental requirement for all users; therefore, each user U k x k Set to 1. 0≤α k ≤1 represents the corresponding energy consumption control weight, which is given by the user based on their latency and energy consumption preferences. Therefore, maximizing the satisfaction of each user's needs is equivalent to minimizing the weighted normalized latency and energy consumption cost for the user to complete the task while ensuring safety.
[0183] By jointly optimizing the 3D trajectories of S and J, and the user transmit power p k (n) and the calculated frequency f k (n), maximizing the total user satisfaction, let The optimization problem can then be expressed as:
[0184]
[0185] Where C1 to C3 represent the restrictions on user transmit power, user computing frequency, and safety offload threshold, respectively; C4 is the constraint on the UAV's computing power; C5 is the constraint on the UAV's energy consumption; and C6 to C8 represent the constraints on the UAV's flight speed, flight altitude, and collision avoidance, respectively.
[0186] This invention uses normalized latency and energy consumption to reduce the dimensional differences between latency and energy consumption. The normalization process is shown in equations (35)-(38), and the reciprocal of the weighted sum of normalized latency and energy consumption is defined as the user satisfaction, as shown in equation (39). Considering the above factors, the optimization problem of the constructed UAV-edge computing model is shown in equation (40).
[0187] Step S3: Transform the optimization problem of the dual-UAV-assisted edge computing model into a Markov decision process. This process includes a system state set, a system dynamic resource allocation and trajectory action set, and an agent reward function set.
[0188] The state set of drone S is The state set of drone J is Therefore, the system's state set is:
[0189]
[0190] The action set of drone S is The action set of drone J is The system's action set is then:
[0191]
[0192] The reward function for drone S is:
[0193]
[0194] wherein, r off (n) is the reward for the user to unload data, r p,S (n) is the punishment for the S to violate the constraint condition. Wherein, r off (n) is represented as:
[0195]
[0196] wherein, κf is the positive integer for adjusting r off (n) reward.
[0197] r p,S (n) is defined as:
[0198] r p,S (n) = -λ ac (n) κ ac,S -λ rc (n) κ rc,S + κ er,S (n) (45)
[0199] wherein, κ ac,S , κ rc,S , κ er,S (n) respectively represent the punishment for the unmanned aerial vehicle collision constraint (40C8), the S computing capability constraint (40C4), and the S residual energy constraint (40C5). λ ac (n) and λ rc (n) are binary coefficients, λ ac (n) = 1 indicates that the constraint (40C8) is not satisfied. Similarly, λ rc (n) is a binary coefficient related to the S computing resource, κ er,S (n) is a sparse reward related to the energy remaining of the S after all user data is processed, and is defined as:
[0200]
[0201] wherein, ζ is a normal number for adjusting κ er,S (n). When E S (n) > 0, a positive reward is given to the agent S, and when E S (n) < 0, a negative punishment is given to the agent. The agent in the present application is the unmanned aerial vehicle.
[0202] The reward function of the unmanned aerial vehicle J is:
[0203]
[0204] wherein, κg r is a positive integer to adjust the total eavesdropping rate. p,J (n) is the penalty that J violates the constraint, which is expressed as:
[0205] r p,J (n) = -λ ac (n)κ ac,J + κ er,J (n). (48)
[0206] where κ ac,J , κ er,J (n) are the collision penalty and the energy surplus penalty of J, respectively. κ er,J (n) is defined as:
[0207]
[0208] where ζ J is a positive constant to adjust κ er,J (n). When E J (n) > 0, it gives a positive reward to the agent, and when E J (n) < 0, it gives a negative penalty to the agent.
[0209] The system-wide reward r(n) can be defined as:
[0210] r(n) = [r S (n), r J (n)]. (50)
[0211] First, each agent perceives the environment state s(n) to obtain information about the current situation. Then, each agent selects an action a(n) based on its current state value s(n). The agent executes the selected action a(n) and interacts with the environment, which includes applying the action a(n) to the environment, observing the feedback of the environment, and obtaining the corresponding reward or penalty r(n). The environment jumps to the next state s(n+1). (s(n), a(n), r(n), s(n+1)) is stored in the experience pool as experience and is randomly sampled to train the action neural network and the Q neural network.
[0212] The specific implementation method of this step is to convert the optimization problem into a <S, A, R> triple, where the state space of the UAV S is the position and remaining energy of S, the user safe offloading rate, and the user remaining unprocessed data. The state space of the UAV J is the position and remaining energy of J, the user safe offloading rate, and the channel gain between J and the eavesdropper. The action space of the UAV S is the vector velocity of S, the user transmission power, and the calculation frequency. The action space of the UAV J is the vector velocity of J. The reward function designs of S and J are formula (43) and formula (47), respectively.
[0213] Step S4: solving the optimization problem using the MADDPG algorithm, the specific process is:
[0214] (1) Initialization: for each agent, initialize its policy network, action network, target network and experience replay buffer;
[0215] (2) Experience sampling: each agent interacts with the environment according to the action output by its policy network, collects experience and stores it in the experience replay buffer.
[0216] (3) Training: each agent randomly samples a batch of experience from the experience replay buffer, calculates the gradient of its action-value function and updates its policy network and action network.
[0217] (4) Target network update: every certain steps, the target network of each agent is updated softly, that is, the parameters of the target network gradually approach the parameters of the action network.
[0218] (5) Shared experience: when sampling experience, each agent can share the experience of other agents, thereby accelerating the learning process.
[0219] (6) Coordination strategy: in each training, the policy network of each agent can consider the actions of other agents, thereby realizing the collaborative decision-making among multiple agents.
[0220] As shown in Figure 2 , in the framework of MADDPG, each agent needs to learn an Actor network π i (a i |o i ) and a Q network Among them, the Q network receives the environment information X composed of the observations of all agents and the actions (a1,…a N ) of all agents as input in the training stage, to centrally calculate the action-value function of the agent; while the Actor network takes the local observation of the agent as input, and outputs its action. Since each agent learns a Q value individually, each agent can design its own reward.
[0221] The MADDPG algorithm is based on the following basic constraints:
[0222] Each agent can only calculate the action value according to its local observation during the execution of the action;
[0223] Without considering the differential equation model of environmental dynamics;
[0224] Without considering the specific communication protocol or method between agents.
[0225] Based on the above constraints, objective functions are designed for each module and gradients are calculated.
[0226] In the diagram, π = (π1, ... π) N Let θ represent the policies of N agents, defined by θ = (θ1, ..., θ2). N The N Actor networks are used to fit the data.
[0227] In MADDPG, the objective function for each agent is defined as follows:
[0228]
[0229] Wherein, J(θ) i Let θ be the policy of agent i. i The performance metric is the expected cumulative reward that agent i will receive in the environment. γ is a discount factor, 0 ≤ γ ≤ 1. When γ is close to 1, the agent focuses more on future rewards; when γ is close to 0, the agent focuses more on immediate rewards. t is the time step. Let E[t] be the immediate reward obtained by agent i at time step t. Let E[.] be the expectation of the cumulative reward.
[0230] The policy gradient for each agent can be calculated by the following formula:
[0231]
[0232] in, This is the policy parameter θ of agent i. i The gradient, E o,a~D [.] This describes the desired operation performed on the sampled state o and action a in the experience pool D. D is the experience pool, which stores the state values and actions of the agent during its interaction with the environment. Let π be the policy function of agent i. i The gradient relative to its parameter θ. Let be the action value function of agent i, representing the expected cumulative reward obtained according to the current policy π given state o and action a. Indicates action a i The gradient, a i =π i This is in Strategy π i According to state o i The chosen action represents the action taken by the agent according to its policy. Therefore, the entire formula represents sampling state o and action a from the experience pool D, and then calculating the expected gradient of agent i's policy under the current policy parameters. This expectation is based on the action value function. For the policy parameter θ i gradient and policy function π i(o i ) the product of the gradient with respect to the parameters, which can be used to update the policy parameters to maximize the expected cumulative reward.
[0233] Through the iteration of the above steps, the MADDPG algorithm can gradually optimize the policy network and action network of each agent, achieve collaborative decision-making among multiple agents, and thus obtain better performance.
[0234] After sufficient training, the policy network realizes dynamic resource allocation and UAV trajectory optimization. Once the cumulative reward value reaches the maximum and tends to be stable, the training stops. At this time, the complete policy network is directly deployed to the UAV base station platform to guide the system to efficiently perform tasks in practice to minimize user latency and energy consumption.
Claims
1. A secure offloading method for a dual-UAV edge computing system based on multi-agent reinforcement learning, characterized in that, Includes the following steps: Step S1: Construct a dual-UAV-assisted edge computing model considering the presence of multiple ground-based eavesdropping incidents; the dual-UAV-assisted edge computing model includes... ground users drones to the server via time-division mode Unload data, ground A potential eavesdropping user A friendly drone constantly attempts to steal user uninstallation information. Support Towards Sending jamming signals to suppress eavesdropping; including drones As an aerial edge computing server, drones provide edge computing services to ground user equipment. As an airborne jammer, it transmits jamming signals to the ground for multiple eavesdropping operations. Step S2: Based on the system model provided in Step S1, and with the premise of ensuring the security of system users' data uninstallation, calculate the average computing cost of system users, and take the reciprocal of the computing cost as the user satisfaction. When considering user satisfaction, the security of user data uninstallation, computing latency and energy consumption are comprehensively considered. The optimization problem is constructed with maximizing user satisfaction as the optimization objective. The user satisfaction level is: This represents the normalized energy consumption. This represents the normalized delay. Set to 1, Indicates the weight of energy consumption control; The optimization problem is: in These represent the restrictions on user transmit power, user computing frequency, and safety offload threshold, respectively. To constrain the computing power of drones; To limit the energy consumption of drones; These represent the constraints on the drone's flight speed, flight altitude, and collision avoidance, respectively. To meet user needs and satisfaction, For user transmit power, Calculate the frequency for the user; For users, instantaneous and secure achievable speeds, To ensure user safety, a threshold for uninstallation is set. For the first The amount of data offloaded by time-slot users This represents the maximum computing frequency of the server-side drone. The CPU computing frequency required to compute 1 bit of data for a server-side drone. The time slot length, and These represent server drones. and jamming drones The remaining energy, and They are respectively and Flight speed, and They represent and Maximum flight speed; and They are respectively and Flight altitude Minimum flight altitude Maximum flight altitude; and They are respectively and Location, Minimum safe distance to prevent collisions between two drones; Step S3: Model the optimization problem as a Markov decision process, including the setting of the system state space, action space, and reward function; Step S4: Use the MADDPG algorithm in deep reinforcement learning to jointly optimize the 3D trajectory and dynamic resource allocation strategy of the two UAVs to reduce the computational cost for system users.
2. The secure offloading method for a dual-UAV edge computing system based on multi-agent reinforcement learning according to claim 1, characterized in that: Step S3, which involves modeling the optimization problem as a Markov decision process, includes: Server drones The state set is Friendly drones The state set is The system's state set is: , They represent and The coordinates of the location , They represent the first Time slot and The remaining energy, Indicates the instantaneous safe rate of the system. Indicates the first Time slot Unprocessed data express and Channel gain between; Server drones Action set as Friendly drones Action set as The system's action set is then: and They represent the first Time slot and Flight speed, and They represent the first Time slot and polar angle, and They represent and Horizontal angle, for In the Time slot calculation frequency, Indicates the first Transmit power of time-slot users; Server drones The reward function is: in, To optimize the objective, Rewards for users who uninstall data. for Penalties for violating constraints; among which, Represented as: in, To adjust The positive integer of the reward Indicates the instantaneous safe rate of the system. Indicates the time slot length; Defined as: in, Representing drones Collision constraints Computational constraints Penalty for residual energy constraints and Binary coefficients Used to determine when all user data has been processed. The sparse reward related to energy surplus is defined as: in, It is used for adjustment positive numbers, Indicates the first Time slot Remaining energy; friendly drones The reward function is: in, To adjust the total eavesdropping rate to a positive integer, express and The instantaneous achievable rate between for The penalty for violating the constraints is represented as follows: in, They are respectively Collision penalty and energy surplus penalty, Defined as: in, It is used for adjustment ; positive numbers; Overall system rewards Defineable 。 3. The secure offloading method for a dual-UAV edge computing system based on multi-agent reinforcement learning according to claim 1, characterized in that: The MADDPG algorithm described in step S4 includes the following steps: (1) Initialization: For each agent, initialize its policy network, action network, target network and experience replay buffer; (2) Experience sampling: Each agent interacts with the environment according to the actions output by its policy network, collects experience and stores it in the experience replay buffer; (3) Training: Each agent randomly samples a batch of experiences from the experience replay buffer, calculates the gradient of its action value function, and updates its policy network and action network; (4) Target network update: At regular intervals, the target network of each agent is softly updated, that is, the parameters of the target network are gradually brought closer to the parameters of the action network. (5) Sharing experience: When each agent performs experience sampling, it shares the experience of other agents, thereby accelerating the learning process; (6) Coordination strategy: In each training session, each agent's policy network can consider the actions of other agents, thereby achieving collaborative decision-making among multiple agents.
4. The secure offloading method for a dual-UAV edge computing system based on multi-agent reinforcement learning according to claim 3, characterized in that: The objective function for each agent in the MADDPG algorithm is defined as follows: in, For intelligent agents strategy Performance metrics, As a discount factor, For time step, In time step At that time, intelligent agent Instant rewards received To calculate the expected value of the cumulative reward; The policy gradient for each agent can be calculated by the following formula: in, It is an intelligent agent strategy parameters gradient, For experience pool Mid-sampled state and actions Perform the desired operation. It is an experience pool that stores the state values and actions of the agent during its interaction with the environment. For intelligent agents strategy function Relative to its parameters gradient, For intelligent agents The action value function represents the value of an action given a state. and actions Next, based on the current strategy The expected cumulative return obtained, Indicates the action gradient, In strategy According to the state The chosen action represents the action taken by the agent according to its own strategy.
5. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the secure offloading method for a dual-UAV edge computing system based on multi-agent reinforcement learning as described in any one of claims 1 to 4.
6. A communication system comprising several users and two unmanned aerial vehicles (UAVs), characterized in that: One of the drones acts as an aerial edge computing server, providing edge computing services to ground user equipment, while the other drone acts as an aerial jammer, transmitting jamming signals to multiple ground eavesdroppers. The drones have a processor and a memory, and the processor, when executing executable instructions stored in the memory, implements the method described in any one of claims 1 to 4.
Citation Information
Patent Citations
Task security unloading strategy determination method based on air-ground collaborative edge computing network
CN112104494A
High-safety unloading energy efficiency method using physical layer safety technology in UAV-MEC environment
CN114362877A