A multi-unmanned aerial vehicle assisted MEC system deployment offloading strategy joint optimization method

By constructing a dynamic model and using deep reinforcement learning methods for a multi-UAV assisted edge computing system, the flight trajectory and task allocation of UAVs were optimized, solving the UAV overload problem caused by excessive concentration of tasks on ground terminals, and achieving low-energy and high-efficiency task processing.

CN116980852BActive Publication Date: 2026-04-28NANJING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING UNIV OF POSTS & TELECOMM
Filing Date
2023-05-26
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In existing technologies, multi-UAV assisted edge computing systems suffer from the problem of excessive concentration of ground terminal tasks leading to UAV overload, and fail to effectively optimize the UAV's three-dimensional flight trajectory to reduce energy consumption and latency.

Method used

A dynamic edge computing system model with multiple UAVs as multiple ground terminals is established, and an optimal offloading strategy and trajectory optimization method are constructed. The flight trajectory of the UAVs is optimized through an agent model and deep reinforcement learning to minimize the system energy consumption.

Benefits of technology

Under the premise of fairness and low energy consumption, the task allocation and flight trajectory of UAVs were optimized, reducing system energy consumption and improving the system's task processing capability and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116980852B_ABST
    Figure CN116980852B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of computer wireless communication, and discloses a multi-unmanned aerial vehicle assisted MEC system deployment and unloading strategy joint optimization method, which comprises the following steps: a dynamic multi-unmanned aerial vehicle model for real-time communication and data transmission is established for a multi-ground terminal edge computing system; in the case of a given unmanned aerial vehicle flight trajectory, an optimal unloading strategy is constructed; according to the given unmanned aerial vehicle flight trajectory and the optimal unloading strategy, a ground terminal unloading matching decision is constructed; and according to the optimal unloading strategy and the ground terminal unloading matching decision, the multi-unmanned aerial vehicle trajectory is optimized. The optimization method has good convergence, can reduce optimization variables, optimize three-dimensional unmanned aerial vehicle trajectories, and is more in line with real scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer wireless communication technology, and in particular to a joint optimization method for deployment and offloading strategies of multi-UAV assisted MEC systems. Background Technology

[0002] Mobile edge computing (MEC) deploys servers at the network edge and is considered an effective solution for fixed MEC server deployments. In MEC, mobile devices can offload their tasks to servers closer to them. Compared to mobile cloud computing, MEC involves shorter transmission distances, less transmission time, and less energy consumption. By configuring ubiquitous computing resources near user devices, mobile edge computing (MEC) can serve compute-intensive and latency-sensitive services, thereby supporting a large number of devices and processing large amounts of data in a timely manner.

[0003] In recent years, unmanned aerial vehicles (UAVs) have received widespread attention in wireless communication. UAVs have been extensively studied and are considered a viable method for supplementing wireless communication networks. The development of UAVs has, to some extent, overcome the limitations of flight time and battery power. For example, UAVs have been applied in areas with limited communication infrastructure, such as developing countries or mountainous regions, as well as in earthquake response, emergency rescue, and battlefield communications. Recently, a literature has investigated a UAV MEC wireless system in which the MEC server is installed on the UAV (i.e., a flying edge cloud). This system offers two advantages: 1) Due to the higher altitude of the flying edge server, it can provide better line-of-sight links to mobile users with a higher probability; 2) Due to the flexible deployment of UAVs, the transmission distance can be further shortened; 3) Their hovering stability and Loss of Space (LoS) transmission characteristics provide reliable and low-latency communication links for ground terminals.

[0004] In multi-UAV-assisted wireless communication systems, UAVs typically function as airborne base stations (BS) or airborne mobile terminals. When a UAV is used as an airborne base station, the ground terminal communicates with the UAV via a Loss of Specs (LoS) link. However, the large amount of data transmission between the ground terminal and the UAV can lead to channel congestion. Furthermore, UAV coverage is limited. When UAVs are used as airborne mobile terminals, the increased number of UAVs can overload cellular network bandwidth. Additionally, UAVs will compete with ground terminals for limited spectrum resources.

[0005] For multi-UAV assisted edge computing systems, existing work mainly focuses on computation offloading, resource allocation, and UAV trajectory optimization. Computation offloading can offload tasks to nearby MEC servers to improve service quality, and is usually optimized in conjunction with task scheduling and load balancing. Meanwhile, resource allocation optimization can rationally allocate computing resources to ground terminals, reducing resource waste. UAV trajectory optimization can reduce latency, save energy, and improve communication throughput, bringing better service quality to ground terminals. Summary of the Invention

[0006] The purpose of this section is to outline some aspects of embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.

[0007] In view of the aforementioned existing problems, the present invention is proposed.

[0008] Therefore, the technical problem solved by the present invention is that the existing technology ignores the problem of overloading of a single UAV due to excessive concentration of ground terminal tasks, and ignores the problem of optimizing the three-dimensional flight trajectory of multiple UAVs in real-world scenarios, under the premise of handling as many tasks as possible and minimizing energy consumption and latency.

[0009] To address the aforementioned technical problems, this invention provides the following technical solution: a joint optimization method for deployment and offloading strategies of multi-UAV assisted MEC systems, comprising:

[0010] Establish a dynamic multi-UAV-to-multi-ground-terminal edge computing system model for real-time communication and data transmission;

[0011] Given the drone's flight trajectory, construct the optimal unloading strategy;

[0012] Based on the given UAV flight trajectory and the optimal unloading strategy, a ground terminal unloading matching decision is constructed.

[0013] Based on the optimal unloading strategy and the ground terminal unloading matching decision, the trajectories of multiple UAVs are optimized.

[0014] As a preferred embodiment of the joint optimization method for deployment and offloading strategy of multi-UAV assisted MEC system described in this invention, the edge computing system includes K ground terminals and M UAVs;

[0015] All drones are equipped with small MEC servers for communication and computing;

[0016] Define all ground terminals as randomly distributed in {X}size ,Y size On the plane of {X, 0}, all drones are in {X} size ,Y size Flying in the three-dimensional space of H, all UAVs complete the flight missions generated by the connected ground terminal within T time slots.

[0017] As a preferred embodiment of the joint optimization method for deployment and offloading strategy of multi-UAV assisted MEC system described in this invention, wherein: in each time slot, the UAV completes the task generated by the connected ground terminal;

[0018] Defined in the next time slot, the position and task of the ground terminal are randomly updated within a certain range. The ground terminal reselects the optimal UAV based on its position and task. After a series of time slots and task processing, the UAV flies from the starting point to the ending point, completing the trajectory design.

[0019] In the t-th time slot, the position of UAV m is Where X m (t) represents the X-axis position of the UAV m, and Y... m (t) represents the Y-axis position of the UAV m, Z... m (t)} represents the Z-axis position of the UAV m;

[0020] The location of ground terminal k is

[0021] The data volume of ground terminal task k is D k (t);

[0022] The number of CPU cycles required for ground terminal k to process 1 bit of data is F. k (t), the task of ground terminal k is

[0023] As a preferred embodiment of the joint optimization method for deployment and offloading strategies of multi-UAV-assisted MEC systems described in this invention, the total energy consumption of the UAV-assisted edge computing system based on fairness in time slot t is expressed as:

[0024]

[0025] in, Let K be the energy consumption for transmission between the ground terminal k and the UAV m. The energy consumption for calculating the ground terminal k. Let m be the computational energy consumption of the drone. Let m be the flight energy consumption of the drone, I(t) be the fairness index among drones, and ω be the weight of the drone's flight energy consumption.

[0026] As a preferred embodiment of the joint optimization method for deployment and offloading strategies of multi-UAV assisted MEC systems described in this invention, the total energy consumption optimization problem of the system based on fairness is expressed as:

[0027]

[0028] stC1:

[0029] C2:

[0030] C3:

[0031] C4:

[0032] C5:

[0033] in, The objective function is to minimize the total system energy consumption for the UAV to complete the mission. C1 is the range constraint for the movement of the UAV and the ground terminal. C2 is the data unloading ratio constraint for the unloading strategy. C3 is the safe distance constraint for the UAV. C4 is the unloading matching decision between the ground terminal and the UAV, where each ground terminal needs to select and match one UAV for unloading. C5 is the fairness index constraint between UAVs. The closer it is to 1, the fairer it is.

[0034] As a preferred embodiment of the joint optimization method for deployment and offloading strategies of multi-UAV assisted MEC systems described in this invention, the optimal offloading strategy is a task offloading ratio, including:

[0035] Based on the given UAV trajectory Θ in time slot t, the unloading matching decision of the fixed ground terminal. Analyzing the concavity and convexity of the optimal strategy Ψ, the total energy consumption of the system simplifies to:

[0036]

[0037]

[0038] in, As a fairness factor among drones, For power coefficient, and These are the energy consumption coefficients for transmission and computing at the ground terminal and computing at the UAV terminal, respectively. This refers to the flight power factor of the drone. and These are the latency coefficients for transmission and computation at the ground terminal and computation at the UAV terminal, respectively.

[0039] When the total energy consumption of the system is minimized

[0040]

[0041]

[0042]

[0043] The optimal unloading strategy for transferring data from ground terminal k to UAV m is expressed as:

[0044]

[0045] As a preferred embodiment of the joint optimization method for deployment and offloading strategies of multi-UAV assisted MEC systems described in this invention, the ground terminal offloading matching decision includes:

[0046] Given the UAV trajectory Θ in time slot t, and by obtaining the optimal unloading strategy Ψ of the ground terminal k, the unloading matching decision of the ground terminal is analyzed. At this time, the total energy consumption of the system is simplified to:

[0047]

[0048]

[0049] Where r is the Euclidean distance d k,m The data transmission rate function of (t), and These are constants following the fixed unloading strategy Ψ and the UAV trajectory Θ, respectively.

[0050] The offloading matching decision for K ground terminals is defined as follows:

[0051]

[0052] The offloading matching decision for ground terminals other than ground terminal k is represented as follows:

[0053]

[0054] Match the unloading decision of the ground terminal To initialize, select the nearest drone and repeat the following steps:

[0055] For ground terminal k (k∈[1,K]), keeping the unloading matching decisions of other ground terminals unchanged, calculate the current optimal choice Δ for ground terminal k. k and

[0056] when When, modify Δ k For the optimal selection of ground terminal k, update Until no ground terminal offers a better alternative;

[0057] When in Nash equilibrium, the ground terminal's offloading matching decision reaches its best, and E is also minimized in time slot t.

[0058] As a preferred embodiment of the joint optimization method for deployment and offloading strategies of multi-UAV assisted MEC systems described in this invention, each UAV is considered as an intelligent agent, and the environment model is described as follows: Define a Markov decision process, including state, behavior, transition probability, reward, and initial state;

[0059] The state consists of the state of each intelligent agent and the ground terminal, including: the position of the drone, the position of the ground terminal, the amount of data of the ground terminal, and the number of CPU cycles required for the ground terminal to match the drone m to calculate 1 bit of data;

[0060] The state of agent m can be represented as:

[0061]

[0062] The behavior includes defining the horizontal deflection angle and the vertical deflection angle as the behavior of each agent, expressed as:

[0063]

[0064] in,

[0065] The behavior m of the agent is standardized and represented as follows:

[0066]

[0067]

[0068] Among them, the horizontal deflection angles of the three drones The ranges are as follows: [0,π], Vertical deflection angles of the three drones

[0069] The transition probability is expressed as:

[0070]

[0071] This indicates that based on the behavior a = [a1, ..., a...] M From state s = [s1, ... s M To the next state s′=[s1′,…s′ MThe transition probability of ];

[0072] The reward includes, based on the objective function, the sum of energy consumption of all agents in T time slots, which is defined as the reward under the premise of ensuring fairness and the optimal unloading strategy.

[0073] The reward is represented as follows:

[0074]

[0075] Punishment and reward are represented as follows:

[0076]

[0077] The initial state assumes that each drone completes a flight path from the starting point to the end point, and then returns to the starting point for training, until the reward converges.

[0078] As a preferred embodiment of the joint optimization method for deployment and offloading strategies of multi-UAV assisted MEC systems described in this invention, the method employs a framework of centralized training and distributed execution, including:

[0079] The transition probability is expressed as:

[0080] P(s′|s,a,μ)=P(s′|s,a)=P(s′|s,a,μ′)

[0081] Where μ = [μ1,…,μ] M [] represents the deterministic policy of M agents in an actor policy network, where μ′ = [μ′1, ..., μ′] M ] represents the deterministic policy of M agents in the target policy network, using θ = [θ1, ..., θ2]. M ] represents the parameter of the deterministic policy μ in the actor policy network;

[0082] The cumulative expected reward of agent m is expressed as:

[0083]

[0084] Where D represents the experience replay buffer, including {s,a,r,s′,done}, and r=[r1,…,r M ] is the set of rewards for all agents, and γ represents the reward discount factor;

[0085] The policy gradient of a deterministic policy μ is expressed as:

[0086]

[0087] in, This represents the output of a centralized action-value function of a critic network, whose inputs are the states and actions of all agents. The quality of the actor network's output policy is evaluated by updating the critic policy network by minimizing the loss function.

[0088] The loss function is expressed as:

[0089]

[0090] Among them, the target value a′=[μ′1(s1′),…,μ′ M (s′ M [)] is the set of behaviors of M intelligent agents. Let θ' represent a target network based on a set of deterministic policies μ', with delay parameters θ' = [θ1', ..., θ'']. M ];

[0091] The delay parameter θ′ is updated in the following way:

[0092] θ′ m ←τθ m +(1-τ)θ′ m

[0093] Where τ is the soft update coefficient;

[0094] Training a set of U different policies is represented as follows: For agent m, the cumulative expected reward is updated as follows:

[0095]

[0096] The policy gradient is updated as follows:

[0097]

[0098] As a preferred embodiment of the joint optimization method for deployment and offloading strategies of multi-UAV assisted MEC systems described in this invention, the optimization of multi-UAV trajectories includes:

[0099] Initialize the agent policy network μ and the target policy network μ′, initialize the replay memory for storing the agent's experience to rpm, and generate a random process. For action exploration, for each agent m, an action is selected according to the Markov decision process. The states s and actions a of all agents are input, and the agent's reward r and next state s′ are obtained. (s, a, r, s′, done) is stored in the experience storage pool rpm. A random batch is sampled from the experience replay pool. Set the target value and minimize the loss function to update the critic network;

[0100] The target value is expressed as:

[0101]

[0102] The minimized loss function is expressed as:

[0103]

[0104] The gradient of the updated actor policy network is represented as:

[0105]

[0106] The delay parameter of the target network for each agent m is softly updated as θ′, and the reward value for each drone reaching the destination is represented as:

[0107]

[0108] When drone m flies out of the boundary or flies within the safe distance, the reward value is updated to:

[0109]

[0110] Repeat the optimization steps until the maximum number of iterations is reached, at which point the optimal drone trajectory with the lowest energy consumption is obtained.

[0111] The beneficial effects of this invention are as follows: This invention provides a joint optimization method for deployment and offloading strategies of multi-UAV assisted MEC systems. By using a dynamic multi-UAV assisted edge computing system, auxiliary UAVs share the computing tasks of the main UAV, completing system tasks within the ground terminal's task tolerance time with greater fairness and lower energy consumption. The optimization method proposed in this invention has good convergence, can reduce optimization variables, optimize 3D UAV trajectories, and better reflect real-world scenarios. Attached Figure Description

[0112] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0113] Figure 1 This is an overall flowchart of the joint optimization method for deployment and offloading strategies of multi-UAV assisted MEC systems according to an embodiment of the present invention;

[0114] Figure 2This is a system model diagram of the joint optimization method for deployment and offloading strategies of a multi-UAV assisted MEC system according to an embodiment of the present invention;

[0115] Figure 3 This is a comparative test diagram of the joint optimization method for deployment and offloading strategies of a multi-UAV assisted MEC system according to an embodiment of the present invention. Detailed Implementation

[0116] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0117] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0118] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0119] This invention is described in detail with reference to the schematic diagrams. When detailing the embodiments of this invention, for ease of explanation, the cross-sectional views illustrating the device structure may be partially enlarged, not adhering to the usual scale. Furthermore, the schematic diagrams are merely examples and should not be construed as limiting the scope of protection of this invention. In actual fabrication, the three-dimensional spatial dimensions of length, width, and depth should be included.

[0120] Furthermore, in the description of this invention, it should be noted that the terms "upper," "lower," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used solely for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In addition, the terms "first," "second," or "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0121] Unless otherwise explicitly specified and limited, the terms "installation," "connection," and "joining" in this invention should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; similarly, they can refer to mechanical connections, electrical connections, or direct connections, or indirect connections through an intermediate medium, or internal connections between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0122] Example 1

[0123] Reference Figure 1 —2, is the first embodiment of the present invention. This embodiment provides a joint optimization method for deployment and offloading strategies of multi-UAV assisted MEC systems, characterized in that it includes:

[0124] S1: Establish a dynamic multi-UAV to multi-ground terminal edge computing system model for real-time communication and data transmission;

[0125] Furthermore, such as Figure 2 As shown, a dynamic multi-UAV / multi-ground-terminal edge computing system model is established for real-time communication and data transmission, including K ground terminals, set as follows: And M drones, grouped together All drones are equipped with small MEC servers for communication and computing;

[0126] Define all ground terminals as randomly distributed in {X} size ,Y size On the plane of {X, 0}, all drones are in {X} size ,Y size Flying in a three-dimensional space of H, all UAVs complete flight tasks generated by connected ground terminals within T time slots, and the set is...

[0127] It should be noted that there are three types of drones: main drone 1, main drone 2, and auxiliary drone. The main drone is responsible for communication and computation with most ground terminals and has a fixed start and end point. The auxiliary drone serves a small number of ground terminals, relieving the load on the main drone and achieving better load fairness among all drones. The main drone and auxiliary drone have the same structure, but their respective service targets and flight trajectories are different.

[0128] Furthermore, within each time slot, the drone completes the task generated by the connected ground terminal;

[0129] Defined in the next time slot, the position and task of the ground terminal are randomly updated within a certain range. The ground terminal reselects the optimal UAV based on its position and task. After a series of time slots and task processing, the UAV flies from the starting point to the ending point, completing the trajectory design.

[0130] In the t-th time slot, the position of UAV m is

[0131] The location of ground terminal k is

[0132] The data volume of ground terminal task k is D k (t);

[0133] The number of CPU cycles required for ground terminal k to process 1 bit of data is F. k (t), the task of ground terminal k is

[0134] Furthermore, the data transmission rate between the ground terminal k and the UAV m is expressed as:

[0135]

[0136] Where B is the channel bandwidth, P k For each mobile user, the transmit power is given; β0 is the channel gain at the reference distance; G0 is a positive constant; and N0 is the noise power spectral density.

[0137] The transmission delay when ground terminal k communicates with UAV m is expressed as:

[0138]

[0139] In time slot t, The task offloading ratio for communication between ground terminal k and UAV m;

[0140] The transmission energy consumption between the ground terminal k and the UAV m is expressed as:

[0141]

[0142] The computation delay of ground terminal k is expressed as:

[0143]

[0144] Among them, f k (t) represents the calculation frequency of the ground terminal k;

[0145] The computational energy consumption of ground terminal k is expressed as:

[0146]

[0147] Among them, K gd This represents the CPU capacitance coefficient of the ground terminal;

[0148] The computational delay of the drone m is expressed as:

[0149]

[0150] Among them, f k,m (t) represents the computing resources of the UAV m to compute the ground terminal k;

[0151] The computational energy consumption of the drone m is expressed as:

[0152]

[0153] Among them, K uav The CPU capacitance coefficient of the drone;

[0154] The flight delay of UAV m in time slot t is expressed as:

[0155]

[0156] Wherein, K′ m The total number of ground terminals serving UAV m, P fly (t) represents the flight power of the UAV;

[0157] The flight energy consumption of drone m is expressed as:

[0158]

[0159] The average workload of the UAV m connected to the ground terminal is expressed as:

[0160]

[0161] The fairness index among drones is expressed as:

[0162]

[0163] The total energy consumption of a fairness-based drone-assisted edge computing system in time slot t is expressed as:

[0164]

[0165] Where ω represents the weight of the UAV's flight energy consumption.

[0166] S2: Given the drone's flight trajectory, construct the optimal unloading strategy;

[0167] Furthermore, the total energy consumption optimization problem of a fairness-based system can be expressed as:

[0168]

[0169] stC1:

[0170] C2:

[0171] C3:

[0172] C4:

[0173] C5:

[0174] in, The objective function is to minimize the total system energy consumption for the UAV to complete the mission. C1 is the range constraint for the movement of the UAV and the ground terminal. C2 is the data unloading ratio constraint for the unloading strategy. C3 is the safe distance constraint for the UAV. C4 is the unloading matching decision between the ground terminal and the UAV, where each ground terminal needs to select and match one UAV for unloading. C5 is the fairness index constraint between UAVs. The closer it is to 1, the fairer it is.

[0175] Furthermore, the optimal unloading strategy is the task load unloading ratio, including:

[0176] Based on the given UAV trajectory Θ in time slot t, the unloading matching decision of the fixed ground terminal. Analyzing the concavity and convexity of the optimal strategy Ψ, the total energy consumption of the system simplifies to:

[0177]

[0178]

[0179] in, As a fairness factor among drones, For power coefficient, and These are the energy consumption coefficients for transmission and computing at the ground terminal and computing at the UAV terminal, respectively. This refers to the flight power factor of the drone. and These are the latency coefficients for transmission and computation at the ground terminal and computation at the UAV terminal, respectively.

[0180] It should be noted that E(t) is a linear function of Ψ and can be differentiated at extrema. and They are about An increasing or decreasing function.

[0181] When the total energy consumption of the system is minimized

[0182]

[0183]

[0184]

[0185] The optimal unloading strategy for transferring data from ground terminal k to UAV m is expressed as:

[0186]

[0187] S3: Based on the given UAV flight trajectory and the optimal unloading strategy, construct a ground terminal unloading matching decision;

[0188] Furthermore, the ground terminal offloading and matching decision includes:

[0189] Given the UAV trajectory Θ in time slot t, and by obtaining the optimal unloading strategy Ψ of the ground terminal k, the unloading matching decision of the ground terminal is analyzed. At this time, the total energy consumption of the system is simplified to:

[0190]

[0191]

[0192] Where r is the Euclidean distance d k,m The data transmission rate function of (t), and These are constants following the fixed unloading strategy Ψ and the UAV trajectory Θ, respectively.

[0193] It should be noted that the ground terminal's offloading matching decision is reflected in ensuring that all ground terminals can choose to offload the mission to the appropriate drone, and how many ground terminals choose drone 1, drone 2, or auxiliary drone.

[0194] Since the number of ground terminals is constant in communication interactions, that is... The total number of elements is constant, but the combination of elements is variable. Meanwhile, the Euclidean distance d(t) from the ground terminal to the UAV is an important factor affecting the unloading and matching decision of the ground terminal.

[0195] Furthermore, the offloading matching decision for K ground terminals is defined as follows:

[0196]

[0197] The offloading matching decision for ground terminals other than ground terminal k is represented as follows:

[0198]

[0199] It should be noted that for any ground terminal, when in Nash equilibrium, if Δ k Changes have occurred, and If the energy consumption value E remains unchanged, then the energy consumption value E will not decrease. This is because if the offloading matching decisions of other ground terminals remain unchanged, then no matter how a ground terminal changes its offloading matching decision, it cannot break the Nash equilibrium.

[0200] Furthermore, the decision to offload ground terminals will be matched. To initialize, select the nearest drone and repeat the following steps:

[0201] For ground terminal k (k∈[1,K]), keeping the unloading matching decisions of other ground terminals unchanged, calculate the current optimal choice Δ for ground terminal k. k and

[0202] when When, modify Δ k For the optimal selection of ground terminal k, update Until no ground terminal offers a better alternative;

[0203] When in Nash equilibrium, the ground terminal's offloading matching decision reaches its best, and E is also minimized in time slot t.

[0204] S4: Optimize the trajectories of multiple UAVs based on the optimal unloading strategy and the ground terminal unloading matching decision.

[0205] Furthermore, treating each drone as an intelligent agent, the environment model is described as follows: Define a Markov decision process, including state, behavior, transition probability, reward, and initial state;

[0206] The state consists of the state of each agent and the ground terminal, including: the position of the drone, the position of the ground terminal, the amount of data of the ground terminal, and the number of CPU cycles required for the ground terminal to match the drone m to calculate 1 bit of data;

[0207] The state of agent m can be represented as:

[0208]

[0209] It should be noted that the above four states change in different time slots, which indicates that the ground terminal is moving and generating new tasks, which is more in line with the real-world scenario.

[0210] The behavior includes defining the horizontal and vertical deflection angles as the behavior of each agent, expressed as:

[0211]

[0212] in,

[0213] The behavior m of the agent is standardized and represented as follows:

[0214]

[0215]

[0216] Among them, the horizontal deflection angles of the three drones The ranges are as follows: [0,π], Vertical deflection angles of the three drones

[0217] The transition probability is expressed as:

[0218]

[0219] This indicates that based on the behavior a = [a1, ..., a...] M From state s = [s1, ... s M To the next state s′=[s1′,…s′ M The transition probability of ];

[0220] The reward is defined as the total energy consumption of all agents in T time slots, based on the objective function, while ensuring fairness and the optimal offloading strategy.

[0221] It should be noted that, in order to reflect the rationality of the reward, a negative value of energy consumption is defined as the reward.

[0222] The reward is expressed as follows:

[0223]

[0224] Punishment and reward are represented as follows:

[0225]

[0226] The initial state assumes that each drone completes a flight path from the starting point to the end point, and then returns to the starting point for training, until the reward converges.

[0227] Furthermore, a framework employing centralized training and distributed execution includes:

[0228] The transition probability is expressed as:

[0229] P(s′|s,a,μ)=P(s′|s,a)=P(s′|s,a,μ′)

[0230] Where μ = [μ1, ..., μ] M [] represents the deterministic policy of M agents in an actor policy network, where μ′ = [μ′1, ..., μ′] M ] represents the deterministic policy of M agents in the target policy network, using θ = [θ1, ..., θ2]. M ] represents the parameter of the deterministic policy μ in the actor policy network;

[0231] The cumulative expected reward of agent m is expressed as:

[0232]

[0233] Where D represents the experience replay buffer, including {s,a,r,s′,done}, and r=[r1,…,r M ] is the set of rewards for all agents, and γ represents the reward discount factor;

[0234] The policy gradient of a deterministic policy μ is expressed as:

[0235]

[0236] in, This represents the output of a centralized action-value function of a critic network, whose inputs are the states and actions of all agents. The quality of the actor network's output policy is evaluated by updating the critic policy network by minimizing the loss function.

[0237] The loss function is expressed as:

[0238]

[0239] Among them, the target value a′=[μ′1(s1′),…,μ′ M (s′ M [)] is the set of behaviors of M intelligent agents. Let θ' represent a target network based on a set of deterministic policies μ', with delay parameters θ' = [θ1', ..., θ'']. M ];

[0240] The delay parameter θ′ is updated as follows:

[0241] θ′ m ←τθ m +(1-τ)θ′ m

[0242] Where τ is the soft update coefficient;

[0243] Training a set of U different policies is represented as follows: For agent m, the cumulative expected reward is updated as follows:

[0244]

[0245] The policy gradient is updated as follows:

[0246]

[0247] It should be noted that during the training phase, each agent's critic network collects the states and behaviors of all agents and generates Q-values, but each agent's actor network makes decisions based on its own partial states. The critic network is extended to learn the policies of other agents, so each agent executes a function that approximates the policies of other agents.

[0248] Furthermore, optimize multi-drone trajectories, including:

[0249] Initialize the agent policy network μ and the target policy network μ′, initialize the replay memory for storing the agent's experience to rpm, and generate a random process. For action exploration, for each agent m, an action is selected according to the Markov decision process. The states s and actions a of all agents are input, and the agent's reward r and next state s′ are obtained. (s, a, r, s′, done) is stored in the experience storage pool rpm. A random batch is sampled from the experience replay pool. Set the target value and minimize the loss function to update the critic network;

[0250] The target value is expressed as:

[0251]

[0252] The loss function is minimized as follows:

[0253]

[0254] The gradient of the updated actor policy network is represented as:

[0255]

[0256] The delay parameter of the target network for each agent m is softly updated as θ′, and the reward value for each drone reaching the destination is represented as:

[0257]

[0258] When drone m flies out of the boundary or flies within the safe distance, the reward value is updated to:

[0259]

[0260] Repeat the above optimization steps until the maximum number of iterations is reached, at which point the optimal drone trajectory with the lowest energy consumption is obtained.

[0261] Example 2

[0262] Reference Figure 3 This is one embodiment of the present invention, which provides a joint optimization method for deployment and offloading strategies of multi-UAV assisted MEC systems. To verify the beneficial effects of the present invention, comparative experiments are conducted for scientific demonstration.

[0263] like Figure 3 As shown, by comparing with existing UAV control methods, the UAV controlled by the deployment and unloading strategy designed in this invention optimizes the ground terminal unloading decision, the proportion of ground terminal unloading tasks, and the UAV flight trajectory, enabling it to handle a greater number of tasks. In practical applications, the system energy consumption can be minimized simply by using the aforementioned deep reinforcement learning optimization algorithm, demonstrating strong practicality.

[0264] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A joint optimization method for deployment and offloading strategies of multi-UAV assisted MEC systems, characterized in that, include: Establish a dynamic multi-UAV-to-multi-ground-terminal edge computing system model for real-time communication and data transmission; Given the drone's flight trajectory, construct the optimal unloading strategy; Based on the given UAV flight trajectory and the optimal unloading strategy, a ground terminal unloading matching decision is constructed. Based on the optimal unloading strategy and the ground terminal unloading matching decision, optimize the trajectory of multiple UAVs; The optimal unloading strategy is the task load unloading ratio, including: According to the given time slot Chinese drone trajectory Offloading matching decision for fixed ground terminals Analyze the optimal strategy Given the concavity and convexity of the system, the total energy consumption of the system simplifies to: , in, As a fairness index among drones, For power coefficient, , and These are the energy consumption coefficients for transmission and computing at the ground terminal and computing at the UAV terminal, respectively. This refers to the flight power factor of the drone. , and These are the latency coefficients for transmission and computation at the ground terminal and computation at the UAV terminal, respectively. When the total energy consumption of the system is minimized , , , Ground terminal Unload to drone The optimal unloading strategy is expressed as: , The ground terminal unloading matching decision includes: Given in Time-slot drone trajectory Then obtain the ground terminal Optimal Unloading Strategy To analyze the unloading and matching decisions of the ground terminal, the total energy consumption of the system can be simplified as follows: , , in, It concerns Euclidean distance. The data transmission rate function, and Fixed uninstallation strategies and drone trajectory The constant after; definition The offloading matching decision for each ground terminal is represented as follows: , In addition to ground terminals The offloading matching decision for other ground terminals is represented as follows: , Match the unloading decision of the ground terminal To initialize, select the nearest drone and repeat the following steps: For ground terminals Keeping the unloading matching decisions of other ground terminals unchanged, calculate the ground terminal The current optimal choice and ; when When, modify For ground terminal The optimal choice, update This will continue until no ground terminal offers a better alternative. When in Nash equilibrium, the ground terminal's offloading matching decision reaches its optimal state, and in the time slot... The middle also reached minimize; Treating each drone as an intelligent agent, the environment model is described as follows: Define a Markov decision process, including state, behavior, transition probability, reward, and initial state; The state consists of the state of each intelligent agent and the ground terminal, including: the position of the drone, the position of the ground terminal, the amount of data of the ground terminal, and the number of CPU cycles required for the ground terminal to match the drone m to calculate 1 bit of data; intelligent agent The state is represented as: , The behavior includes defining the horizontal deflection angle and the vertical deflection angle as the behavior of each agent, expressed as: , in, ; The behavior of the intelligent agent Standardized operations are represented as follows: , , Among them, the horizontal deflection angles of the three drones The ranges are as follows: , , Vertical deflection angles of the three drones ; The transition probability is expressed as: , Indicates based on behavior From state To the next state The transition probability; The reward includes, based on the objective function, ensuring fairness and the optimal unloading strategy, all agents... The total energy consumption in each time slot is defined as the reward; The reward is represented as follows: , Punishment and reward are represented as follows: , The initial state assumes that each drone completes a flight path from the starting point to the end point, and then returns to the starting point for training, until the reward converges.

2. The joint optimization method for deployment and offloading strategies of multi-UAV assisted MEC systems as described in claim 1, characterized in that: The edge computing system includes ground terminals and One drone; All drones are equipped with small MEC servers for communication and computing; Define all ground terminals as randomly distributed in On the plane, all drones are Flying in three-dimensional space, all drones in The flight mission generated by the connected ground terminal is completed within a time slot.

3. The joint optimization method for deployment and offloading strategies of multi-UAV assisted MEC systems as described in claim 2, characterized in that: In each time slot, the drone completes the task generated by the connected ground terminal; Defined in the next time slot, the position and task of the ground terminal are randomly updated within a certain range. The ground terminal reselects the optimal UAV based on its position and task. After a series of time slots and task processing, the UAV flies from the starting point to the ending point, completing the trajectory design. In the In each time slot, the drone The position is ; Ground terminal The position is ; Ground terminal The data volume of the task is ; Ground terminal The number of CPU cycles required to process 1 bit of data is Ground terminal The task is .

4. The joint optimization method for deployment and offloading strategies of multi-UAV assisted MEC systems as described in claim 3, characterized in that: In the time slot The total energy consumption of a drone-assisted edge computing system based on fairness is expressed as: , in, For ground terminal With drones Transmission energy consumption, For ground terminal Computational energy consumption, For drones Computational energy consumption, For drones Flight energy consumption, As a fairness index among drones, The weighting of drone flight energy consumption.

5. The joint optimization method for deployment and offloading strategies of multi-UAV assisted MEC systems as described in claim 4, characterized in that: The total energy consumption optimization problem of a fairness-based system is expressed as: , in, , , The objective function is to minimize the total system energy consumption for the UAV to complete the mission. C1 is the range constraint for the movement of the UAV and the ground terminal. C2 is the data unloading ratio constraint for the unloading strategy. C3 is the safe distance constraint for the UAV. C4 is the unloading matching decision between the ground terminal and the UAV, where each ground terminal needs to select and match one UAV for unloading. C5 is the fairness index constraint between UAVs. The closer it is to 1, the fairer it is.

6. The joint optimization method for deployment and offloading strategies of multi-UAV assisted MEC systems as described in claim 5, characterized in that: A framework employing centralized training and distributed execution includes: The transition probability is expressed as: , in, In the agent strategy network Deterministic policies of individual agents, In the target policy network A deterministic strategy for each agent, using Deterministic policies in an actor policy network Parameters; intelligent agent The cumulative expected reward is expressed as: , in, Represents the experience replay buffer, including , It is the set of rewards for all intelligent agents. Indicates the reward discount factor; Deterministic strategy The policy gradient is expressed as: , in, This represents the output of a centralized action-value function of a critic network, whose inputs are the states and actions of all agents. The quality of the actor network's output policy is evaluated by updating the critic policy network by minimizing the loss function. ; The loss function is expressed as: , Among them, the target value , yes A set of behaviors of an intelligent agent Indicates a deterministic strategy The target network of the set, whose delay parameter is ; The delay parameter The update method is as follows: , in, This is a soft update coefficient; Training a by The set of different strategies is represented as For intelligent agents The cumulative expected reward has been updated as follows: , The policy gradient is updated as follows: 。 7. The joint optimization method for deployment and offloading strategies of multi-UAV assisted MEC systems as described in claim 1 or 6, characterized in that: Optimize multi-drone trajectories, including: Initialize the actor policy network and target policy network Initialize the replay memory used to store the agent's experience as follows: Generate a random process Used for action exploration, for each agent The action is selected according to the Markov decision process, and the states of all agents are input. and actions To obtain rewards from intelligent agents and the next state ,Will Stored in the experience storage pool In the middle, a random batch is sampled from the experience replay pool. Set the target value and minimize the loss function to update the critic network; The target value is expressed as: , The minimized loss function is expressed as: , The gradient of the updated actor policy network is represented as: , Soft update for each agent The latency parameter of the target network is The reward value for each drone reaching the finish line is expressed as: , When drones If the flight goes outside the boundary or stays within the safe distance, the reward value is updated as follows: , Repeat the optimization steps until the maximum number of iterations is reached, at which point the optimal drone trajectory with the lowest energy consumption is obtained.

Citation Information

Patent Citations

  • Joint trajectory, unloading and resource allocation optimization method in multi-unmanned aerial vehicle assisted MEC system

    CN115833907A