Multi-unmanned aerial vehicle auxiliary edge computing resource and trajectory optimization method
Through multi-UAV assisted edge computing resources and trajectory optimization methods, the problem of unstable communication links in IoT data collection is solved, efficient and reliable data transmission is achieved, and it adapts to the real-time requirements in complex environments, especially in user-intensive and heterogeneous channel scenarios.
Patent Information
- Application Number
- CN202511041275.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-10-10
AI Technical Summary
Traditional multi-hop routing for IoT data collection results in unstable communication links in complex environments, leading to increased transmission delays and packet loss. Energy-constrained IoT devices struggle to maintain long-distance communications, and existing algorithms lack the ability to adaptively optimize unstable links, especially in remote areas or areas lacking infrastructure, making them unable to meet real-time requirements.
A multi-UAV assisted edge computing resource and trajectory optimization method is adopted. By establishing an IoT data collection network model, combining non-orthogonal multiple access technology and information age constraints, the Markov decision process and multi-agent flexible actor-critic algorithm are used to optimize the power allocation and trajectory of UAVs, achieving efficient and reliable data transmission.
It significantly improves data collection efficiency, dynamically adapts to the uneven distribution requirements in user-dense scenarios, improves spectrum efficiency and fairness for edge users, ensures the real-time nature of key information, and is particularly suitable for large-scale IoT scenarios.
Smart Images

Figure CN120769307A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of mobile edge computing, in particular to a multi-unmanned aerial vehicle assisted edge computing resource and trajectory optimization method. BACKGROUND
[0002] The Internet of Things terminal device continuously generates massive real-time data through a distributed sensor network, and its diversified data stream provides support for transformative applications such as autonomous driving, industrial automation, and precision agriculture. The traditional multi-hop routing data collection method is limited by complex environments and device energy consumption, often leading to unstable communication links. Therefore, it is of great significance to build an efficient and reliable Internet of Things data collection system.
[0003] Traditional IoT data collection relies on multi-hop routing, but in complex environments, unstable links can lead to increased transmission delay and packet loss rate. Energy-constrained IoT devices cannot maintain long-distance communication, and the dynamic maintenance cost of multi-hop paths is high, which cannot meet the real-time requirements. The challenge of direct transmission to the base station has not been completely solved, especially in remote or infrastructure-poor areas, and existing algorithms lack adaptive optimization capabilities for unstable links.
[0004] As the infrastructure of smart cities, the Internet of Things relies on distributed sensor networks to achieve real-time monitoring of environmental data, but faces challenges such as limited transmission power and scarce uplink channel resources. Although the unmanned aerial vehicle (UAV) assisted mobile edge computing (MEC) system alleviates the load pressure of ground base stations through three-dimensional deployment, its resource allocation strategy still has defects. Existing algorithms mostly assume uniform user distribution or static channel conditions, and cannot effectively cope with the spatio-temporal dynamics caused by high user density, heterogeneous channels, and fast UAV movement. SUMMARY
[0005] The technical problem to be solved by the present application is to overcome the shortcomings of the prior art, and to provide a multi-unmanned aerial vehicle assisted edge computing resource and trajectory optimization method, comprising the following steps:
[0006] S1: Establish an Internet of Things data collection network model, including a communication sub-model, a transmission sub-model, and a movement sub-model, involving multiple Internet of Things devices, multiple unmanned aerial vehicles equipped with mobile edge computing servers, and a base station. The unmanned aerial vehicles collect data from multiple Internet of Things devices and unload the data to the base station;
[0007] The communication sub-model is used to represent the data transmission relationship between the unmanned aerial vehicles, the Internet of Things devices, and the base station, and to calculate the data transmission rate of the unmanned aerial vehicles. The transmission sub-model is used to calculate the time delay and energy consumption of data transmission between the unmanned aerial vehicles, the Internet of Things devices, and the base station. The movement sub-model is used to calculate the flight energy consumption of the unmanned aerial vehicles;
[0008] S1.1: Set the parameters of the IoT data collection network model, including the number of drones, the number of IoT devices, the range of movement, and the location coordinates of drones, IoT devices, and base stations within the range of movement;
[0009] S1.2: Establish a communication sub-model and use it to characterize the data transmission relationship between the drone, IoT device, and base station. Calculate the data transmission rate between the drone and the base station and between the IoT device and the drone.
[0010] A communication sub-model is established based on the air-ground path loss model. In the data uplink, the line-of-sight connection probability, non-line-of-sight connection probability, line-of-sight connection path loss, and non-line-of-sight connection path loss between the IoT device and the drone are calculated. The average path loss and channel gain between the IoT device and the drone are also calculated.
[0011] Calculate the average path loss between the drone and the base station and the transmission rate between the drone and the base station;
[0012] Use constraints to limit the data upload service relationship between IoT devices and drones, and obtain the drone's receiving signal;
[0013] Establish a decoding order decision criterion to determine the accumulated inter-cluster interference by identifying the optimal decoding order and eliminate the intra-cluster interference between IoT devices;
[0014] Calculate the signal-to-interference-to-noise ratio between the IoT device and the drone to obtain the data transmission rate between the IoT device and the drone;
[0015] S1.3: Establish a transmission sub-model to calculate the latency and energy consumption of data transmission between the drone, IoT device, and base station. The transmission sub-model includes a data upload unit, a data offload unit, and a local computing unit. The data upload unit is used to calculate the latency and energy consumption of data uploaded by the IoT device to the drone. The data offload unit is used to calculate the latency and energy consumption of the drone offloading collected data to the base station. The local computing unit is used to assign a local computing method to the IoT device data and calculate the latency and energy consumption of the drone's local computing.
[0016] The specific method by which the local computing unit allocates local computing methods for the data of IoT devices is as follows:
[0017] Calculate the information age of IoT devices at time slot t;
[0018] Set the maximum information age and assign local calculation methods to the data of IoT devices based on the maximum information age, specifically:
[0019] The information age A of the kth IoT device k(t) Exceeds the maximum information age A MAX When , the data of the kth IoT device is distributed to the UAV equipped with the mobile edge computing server for local calculation; otherwise, the data of the kth IoT device is offloaded to the base station for calculation;
[0020] S1.4: Establish a mobile sub-model and use it to calculate the flight energy consumption of the UAV:
[0021] S1.5: Calculate the total energy consumption of the IoT data collection network based on the energy consumption of the drone uploading data, the energy consumption of the drone offloading collected data to the base station, the energy consumption of the drone local computing, and the energy consumption of the drone flying.
[0022] S2: Set decision variables based on the IoT data collection network model, and set constraints and constraint targets based on the decision variables;
[0023] S2.1: Set decision variables based on the IoT data collection network model, including: a service indicator variable, a power allocation variable, and a motion control variable; the service indicator variable is used to represent the association between the IoT device and the drone; the power allocation variable is used to represent the transmission power of the IoT device to the drone; and the motion control variable is used to represent the flight speed and flight angle of the drone;
[0024] S2.2: Set constraints and constraint objectives based on decision variables;
[0025] Set constraints based on service indicator variables, power allocation variables, and motion control variables, including:
[0026] Constraints for limiting the number of drones that can serve each IoT device;
[0027] Constraints on the decoding order of non-orthogonal multiple access for eliminating intra-cluster interference between IoT devices;
[0028] Constraints for limiting the power consumption of each drone;
[0029] Constraints used to limit the age of information;
[0030] Constraints used to avoid collisions between drones;
[0031] Constraints used to limit the range of drone activities;
[0032] Constraints used to constrain the drone's flight speed and angle;
[0033] The constraint objective is set to minimize the total energy consumption of the IoT data collection network model;
[0034] S3: Model the constraint target as a Markov decision process, establish a multi-agent flexible actor-critic algorithm based on K-means clustering, jointly optimize the power allocation and trajectory of the UAV, and obtain the power allocation result of the UAV and the optimized UAV trajectory;
[0035] S3.1: Model the UAV as an agent in the Markov decision process, set the state set, action set and reward function of the UAV;
[0036] S3.2: Distribute the Internet of Things devices to the UAV based on K-means clustering;
[0037] Taking the set of Internet of Things device positions as input, clustering users through channel broadcasting, and dividing the Internet of Things devices into several clusters equal to the number of UAVs based on Euclidean distance;
[0038] Determine the cluster center of each cluster by iteratively calculating the minimum sum of squared errors within the cluster, and assign each cluster to the nearest UAV; when the load of the mobile edge computing server equipped by the UAV exceeds the threshold, dynamically reassign the Internet of Things device farthest from the cluster center to other UAVs;
[0039] S3.3: Construct a multi-agent flexible actor-critic algorithm to jointly optimize the power allocation and trajectory of the UAV, and obtain the power allocation result of the UAV and the optimized UAV trajectory;
[0040] A centralized training and decentralized execution framework is used to construct a multi-agent flexible actor-critic algorithm, including: multiple policy subnetworks, an experience replay pool, multiple critic subnetworks and an entropy adjustment subnetwork;
[0041] The policy subnetwork is used to output a continuous action a t based on the state s t of the UAV; the experience replay buffer is used to store transition tuples <S, A, P, R, >; the critic subnetwork is used to optimize the target Q value based on the mean square error loss function and the Adam optimizer; and the entropy adjustment subnetwork is used to dynamically adjust the temperature coefficient a of the policy subnetwork;
[0042] The multi-agent flexible actor-critic algorithm is used to jointly optimize the power allocation and trajectory of the UAV, and obtain the power allocation result of the UAV and the optimized UAV trajectory.
[0043] The beneficial effects of the above technical solution are as follows: the present invention provides a multi-UAV-assisted edge computing resource and trajectory optimization method, which uses unmanned aerial vehicles (UAVs) to assist mobile edge computing (MEC) and, by coordinating with ground base stations, significantly improves data collection efficiency. In densely populated scenarios, ground base stations often struggle to serve edge users due to limited bandwidth. However, UAVs, with their three-dimensional maneuverability, can dynamically adapt to the needs of unevenly distributed users.
[0044] Non-orthogonal multiple access (NOMA) technology effectively addresses channel heterogeneity challenges through superposition coding and successive interference cancellation (SIC). Deploying NOMA in drone MEC systems can improve spectrum efficiency while also enhancing fairness for edge users. This technology is particularly well-suited for large-scale IoT scenarios with significant spatial heterogeneity.
[0045] Timeliness of Information (AoI) is a core metric that quantifies the time delay between data generation and reception. Compared to traditional latency metrics, it better reflects the freshness of information in dynamic network environments. In drone-assisted MEC networks, AoI optimization is particularly important due to the mobility and spatiotemporal dynamics of terminals. By prioritizing the transmission of high-AoI data, the system ensures the real-time delivery of critical information, which is particularly valuable in emergency scenarios such as forest fire monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 A schematic diagram of multi-UAV assisted mobile edge computing provided by an embodiment of the present invention;
[0047] Figure 2 A schematic diagram of the calculation method allocation based on the maximum information age provided in an embodiment of the present invention;
[0048] Figure 3 A multi-agent flexible actor-critic algorithm based on K-means clustering provided by an embodiment of the present invention;
[0049] Figure 4 A schematic diagram of the convergence performance of the multi-agent flexible actor-critic algorithm based on K-means clustering under different learning rates provided by an embodiment of the present invention;
[0050] Figure 5 A graph showing the average rewards of different algorithms as a function of training rounds provided by the embodiments of the present invention;
[0051] Figure 6 A comparison chart showing the impact of power allocation on total energy consumption in NOMA and OMA scenarios, provided by an embodiment of the present invention;
[0052] Figure 7 A graph showing the evolution of information age over time slots when information age constraints are applied, provided in an embodiment of the present invention;
[0053] Figure 8 This is a graph showing the evolution of information age over time slots when there is no information age constraint provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0054] The following embodiments of the present invention are described in further detail with reference to the accompanying drawings and examples. The following examples are used to illustrate the present invention but are not intended to limit the scope of the present invention.
[0055] To address the problems of limited transmission power and scarce uplink channel resources, this embodiment provides a multi-unmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC) data acquisition system framework. By combining the UAV-assisted MEC system with non-orthogonal multiple access (NOMA) technology, it effectively improves data transmission efficiency and network fairness. On this basis, a joint power allocation and trajectory optimization method based on NOMA dynamic decoding order and age of information (AoI) constraints is established to ensure data freshness. The joint power allocation and UAV trajectory optimization problem is modeled as a Markov decision process (MDP). In view of the high-dimensional non-convex nature of the problem, a KMASAC algorithm based on improved K-means and multi-agent flexible actor-critic (MASAC) is proposed to minimize the total energy consumption of the system. In the KMASAC framework, the K-means algorithm first optimizes the clustering centers to achieve Internet of Things device (IoTD) clustering and ensures quality of service (QoS) through device reallocation; then, the MASAC algorithm is used to jointly optimize the multi-UAV trajectory and power allocation.
[0056] A multi-UAV assisted edge computing resource and trajectory optimization method of this embodiment involves multiple IoT devices, multiple UAVs equipped with mobile edge computing servers and base stations, such as Figure 1 As shown, the following steps are included:
[0057] S1: Establish an IoT data collection network model, including a communication sub-model, a transmission sub-model, and a mobility sub-model. This involves multiple IoT devices, multiple drones equipped with mobile edge computing servers, and base stations. The drones collect data from multiple IoT devices and offload the data to the base stations.
[0058] The communication sub-model is used to characterize the data transmission relationship between the drone, IoT device and base station, and calculate the data transmission rate of the drone; the transmission sub-model is used to calculate the delay and energy consumption of data transmission between the drone, IoT device and base station; the mobility sub-model is used to calculate the flight energy consumption of the drone;
[0059] S1.1: Set the parameters of the IoT data collection network model, including the number of drones, the number of IoT devices, the range of movement, and the location coordinates of drones, IoT devices, and base stations within the range of movement;
[0060] Assume that the drone set in the IoT data collection network model is U = {U1,…,U M}, the IoT device set is K = {K1,…,K N}, where M is the number of drones in the IoT data collection network, and N is the number of IoT devices in the IoT data collection network;
[0061] Assume that all IoT devices and drones are in three-dimensional space {X size ,Y size ,H}, in time slot t∈{1,2,…,T}, the position coordinate of the u-th UAV is l u (t) = {x u (t),y u (t),H u}, the location coordinates of the kth IoT device are l k (t) = {x k (t),y k (t),0}, the location coordinates of base station B are l B ={x B (t),y B (t),0}, where x u (t)∈X size is the x-axis coordinate of the u-th drone, y u (t)∈Y size is the y-axis coordinate of the u-th UAV, H u ∈H is the height of the u-th UAV, x k (t)∈X size is the x-axis coordinate of the kth IoT device, y k (t)∈Y size is the y-axis coordinate of the kth IoT device, x B (t)∈X size is the x-axis coordinate of base station B, y B (t)∈Y size is the y-axis coordinate of base station B; the UAV's mobile edge computing server provides services to IoT devices in a time division multiple access manner;
[0062] S1.2: Establish a communication sub-model and use it to characterize the data transmission relationship between the drone, IoT device, and base station. Calculate the data transmission rate between the drone and the base station and between the IoT device and the drone:
[0063] A communication sub-model is established based on the air-ground path loss model. In the data uplink, the line-of-sight connection probability, non-line-of-sight connection probability, line-of-sight connection path loss, and non-line-of-sight connection path loss between the IoT device and the UAV are calculated respectively. In the t time slot, the line-of-sight connection probability between the k-th IoT device and the u-th UAV is calculated. As shown in the following formula:
[0064]
[0065] Among them, η a and η b is a constant related to the physical environment in which wireless signals propagate, θ k,u is the elevation angle between the kth IoT device and the uth UAV;
[0066] The probability of non-line-of-sight connection between the kth IoT device and the uth drone in time slot t As shown in the following formula:
[0067]
[0068] Path loss of line-of-sight connection in time slot t and path loss for non-line-of-sight connections As shown in the following formula:
[0069]
[0070] Among them, η LOS is the additional path loss for line-of-sight connections, η NLOS is the additional path loss for the non-line-of-sight path, is the free space path loss, as shown in the following formula:
[0071]
[0072] Among them, d k,u (t) is the Euclidean distance between the kth IoT device and the uth drone, f s is the system frequency, c is the speed of light;
[0073] Based on the path loss of line-of-sight connection and non-line-of-sight connection, the average path loss between the drone and the base station and the transmission rate between the drone and the base station are calculated, and the average path loss L between the kth IoT device and the uth drone is calculated. k,u (t) is shown in the following formula:
[0074]
[0075] According to the average path loss L between the kth IoT device and the uth drone k,u(t), the channel gain g between the kth IoT device and the uth drone k,u (t) is shown in the following formula:
[0076]
[0077] Among them, H k,u (t) is the channel transfer function between the kth IoT device and the uth UAV in time slot t;
[0078] Calculate the average path loss between the UAV and the base station and the transmission rate between the UAV and the base station, the average path loss between the u-th UAV and the base station L u The calculation method of (t) is the same as the average path loss L between the kth IoT device and the uth drone. k,u (t) is similar and will not be elaborated here; the transmission rate between the u-th drone and the base station is R u (t) is shown in the following formula:
[0079]
[0080] Where B is the bandwidth, δ0 is Gaussian white noise, and p u is the transmission power of the UAV;
[0081] Use constraints to limit the data upload service relationship between IoT devices and drones, and get the drone's receiving signal. The receiving signal y of the u-th drone is u,k (t) is shown in the following formula:
[0082]
[0083] in, is a constraint term, x k,u (t) is the signal transmitted by the k-th IoT device to the u-th drone, P k,u (t) is the power distribution value, σ u,k (t) is Gaussian noise, is the accumulated inter-cluster interference, is the intra-cluster interference;
[0084] Constraints middle, is a binary variable used to characterize the data upload relationship between the kth IoT device and the uth drone. Let the constraint To limit the kth IoT device to only have one drone providing data upload service;
[0085] Since IoT devices and drones have time-varying position characteristics, the decoding order must be dynamically determined in each time slot to ensure the effective implementation of SIC. The judgment criterion G for the decoding order is established.u,k (t), as shown in the following formula:
[0086]
[0087] Determine the accumulated inter-cluster interference by identifying the optimal decoding order And eliminate intra-cluster interference between IoT devices. In this embodiment, consider that the cluster served by the u-th drone contains k IoT devices. To successfully eliminate the intra-cluster interference of all IoT devices except the k-th IoT device, the optimal decoding order must meet the following conditions:
[0088] G 1,u (t)≤G 2,u (t)≤…≤G k,u (t)
[0089] After eliminating the signals of IoT devices other than the kth IoT device, the intra-cluster interference As shown in the following formula:
[0090]
[0091] Calculate the signal-to-interference-and-noise ratio between the IoT device and the drone, and obtain the data transmission rate between the IoT device and the drone, and the signal-to-interference-and-noise ratio γ between the kth IoT device and the uth drone. u,k (t) is shown in the following formula:
[0092]
[0093] The data transmission rate R between the kth IoT device and the uth drone u,k (t) is shown in the following formula:
[0094] R u,k (t) = Blog2(1+γ u,k (t) 2 )
[0095] S1.3: Establish a transmission sub-model to calculate the latency and energy consumption of data transmission between the drone, IoT device, and base station. The transmission sub-model includes a data upload unit, a data offload unit, and a local computing unit. The data upload unit is used to calculate the latency and energy consumption of data uploaded by the IoT device to the drone. The data offload unit is used to calculate the latency and energy consumption of the drone offloading collected data to the base station. The local computing unit is used to assign a local computing method to the IoT device data and calculate the latency and energy consumption of the drone's local computing.
[0096] The IoT device uploads all data to the drone through the NOMA link, and uses the data upload unit to calculate the transmission delay and energy consumption of the IoT device uploading data to the drone. The transmission delay of the k-th IoT device uploading data to the u-th drone is and energy consumption They are:
[0097]
[0098] Use the data unloading unit to calculate the delay and energy consumption of the drone unloading the collected data to the base station. The delay of the u-th drone unloading the collected data to the base station and energy consumption They are:
[0099]
[0100] Among them, the local computing unit distributes local computing methods for the data of IoT devices. The specific method for calculating the latency and energy consumption of the drone's local computing is as follows:
[0101] The age of information (AoI) of the data received by the drone is used as a key performance indicator to characterize the timeliness of the drone's sampled information update. The age of information increases linearly with time from the initial age until the next update data packet sent by the IoT device to the drone is received;
[0102] The information age A(τ) of IoT devices at time τ is set to:
[0103] A(τ)=τ-o(τ)
[0104] Wherein, o(τ) is the generation time of the base station receiving the sampling data;
[0105] Calculate the information age of IoT devices at time slot t, the information age A of the kth IoT device k (t) is shown in the following formula:
[0106]
[0107] in, The flight time of the drone;
[0108] Set the maximum information age A MAX , based on the maximum information age A MAX Assign local computation to the data of the kth IoT device, such as Figure 2 As shown, specifically:
[0109] The information age A of the kth IoT device k (t) Exceeds the maximum information age A MAXWhen , the data of the kth IoT device is distributed to the UAV equipped with the mobile edge computing server for local calculation; otherwise, the data of the kth IoT device is offloaded to the base station for calculation;
[0110] Calculate the latency and energy consumption of the local computation of the UAV, the latency of the local computation of the u-th UAV and energy consumption As shown in the following formula:
[0111]
[0112] S1.4: Establish a mobile sub-model and use it to calculate the flight energy consumption of the UAV:
[0113] The position of the UAV in time slot t+1 is calculated based on the position of the UAV in time slot t. The position coordinates of the UAV in time slot t are l u (t) = {x u (t),y u (t),H u}, the position of the u-th UAV at time slot t+1 is shown in the following formula:
[0114] l u (t+1)={x u (t)+v u (t)t fly cosβ(t),y u (t)+v u (t)t fly cosβ(t),H u}
[0115] Among them, β(t)∈[0,2π] is the flight angle, v u (t)∈[0,v max ] is the flight speed of the u-th UAV, t fly is the flight time of the drone;
[0116] Calculate the flight energy consumption of the UAV, the flight energy consumption of the u-th UAV As shown below:
[0117]
[0118] Among them, Q u is the mass of the u-th drone, is the flight time of the u-th drone;
[0119] S1.5: Calculate the total energy consumption of the IoT data collection network model based on the energy consumption of the drone uploading data, the energy consumption of the drone offloading collected data to the base station, the energy consumption of the drone local computing, and the energy consumption of the drone flying.
[0120] The total energy consumption E(t) of the IoT data collection network is shown in the following formula:
[0121]
[0122] S2: Set decision variables, and set constraints and constraint targets based on the decision variables;
[0123] S2.1: Set decision variables, including service indicator variables, power allocation variables, and motion control variables;
[0124] Service indicator variables Used to represent the relationship between IoT devices and drones;
[0125] Power allocation variable P = {P k,u (t),0 <k<K,0<u<U,0<t<T},用于表示物联网设备到无人机的发射功率;
[0126] Motion control variable U={v u (t),β u (t),0 <u<U,0<t<T},用于表示无人机的飞行速度v u (t) and flight angle β u (t);
[0127] S2.2: Set constraints and constraint objectives based on decision variables;
[0128] Constraints set based on service indicator variables, power allocation variables, and motion control variables include:
[0129]
[0130] C2:G u,k (t)≥G u,j (t)
[0131]
[0132] C4:A k (t)≤A MAX
[0133] C5:l i (t)≠l j (t)i,j∈U
[0134] C6:l u (t), l k (t)∈{X size ,Y size ,H}
[0135] C7:vu (t)∈[0,v max ],β u (t)∈[0,2π]
[0136] Among them, constraint C1 is used to ensure that each IoT device u∈U is served by only one UAV; constraint C2 is the non-orthogonal multiple access (NOMA) decoding order constraint to ensure the successful implementation of continuous interference cancellation; constraint C3 is the power constraint, requiring that the power consumption of each UAV must not exceed its maximum power limit; constraint C4 is the information age constraint; constraint C5 is the UAV anti-collision constraint; constraint C6 is the UAV activity range restriction constraint; constraint C7 is the UAV flight speed and flight angle constraint;
[0137] The constraint goal is to minimize the total energy consumption of the IoT data collection network, as shown in the following formula:
[0138]
[0139] S3: Model the constraint objectives as a Markov decision process and establish a multi-agent flexible actor-critic algorithm based on K-means clustering to jointly optimize the power allocation and trajectory of the UAV, obtaining the UAV power allocation result and the optimized UAV trajectory;
[0140] In this example, minimizing the total energy consumption of the IoT data collection network is a complex non-convex optimization problem that is difficult to solve directly. Using a solution based on deep reinforcement learning (DRL), we combined improved K-means clustering with the multi-agent flexible actor-critic (MASAC) method to develop a joint algorithm based on K-means clustering and the multi-agent flexible actor-critic method to optimize the power allocation and trajectory of multiple drones.
[0141] In the information age-aware IoT data collection network, drones are regarded as intelligent agents, IoT devices and base stations are regarded as the environment, and each intelligent agent needs to dynamically decide its position, speed, flight angle, transmission power and service indicator variables to minimize the total energy consumption of the system. The intelligent agent continuously interacts with the environment and updates the current state based on the historical state and action. Since the state information of the intelligent agent contains all relevant historical data, the current state is sufficient to determine future actions. Therefore, the constraint objective can be effectively modeled as a Markov decision process (MDP). A typical Markov decision process includes an intelligent agent and an environment, which can be defined as a five-tuple<S,A,P,R,γ> , where S is the state space, A is the action space, P(s i+1 |s i ,a i ) indicates the execution of action a i ∈A from the current state s i∈S transfers to s i+1 ∈S, R is the agent reward function, and γ∈[0,1] is the discount factor;
[0142] S3.1: Model the drone as an agent in a Markov decision process, setting the drone's state set, action set, and reward function;
[0143] Based on the drone, IoT devices, and environment, the state set s(t) of the drone at the t-th time slot is set as shown in the following formula:
[0144] s(t)= <l(t),G(t),A(t),D remain (t)>
[0145] Among them, l(t) is the position coordinate of the drone, G(t) is the channel gain of the IoT device, A(t) is the information age of the IoT device, and D remain (t) is the amount of remaining data to be collected in the t-th time slot;
[0146] According to the state set s(t) of the UAV in the tth time slot and the environment, the action a(t) of the agent in the tth time slot is set, including the service indicator variable of the tth time slot Transmitting power P(t), UAV’s flight angle β u (t) and the flight speed v of the UAV u (t), the range of the action set a(t) of the UAV in the t-th time slot is set based on the constraints, including:
[0147]
[0148] P(t)∈[0,P MAX ]
[0149] v u (t)∈[0,v max ]
[0150] β u (t)∈[0,2π]
[0151] With the goal of constraining the target and avoiding UAV collisions, the reward function r(t) of the t-th time slot is set as shown in the following formula:
[0152] r(t)=-E(t)-μ
[0153] Among them, μ is the collision penalty coefficient;
[0154] S3.2: Assign IoT devices to drones based on K-means clustering method;
[0155] The IoT device location set L = {l1,…,l k} as input, user clustering is performed through channel broadcasting, and IoT devices are divided into u clusters U = {1, .., u} equal to the number of drones based on Euclidean distance, ensuring intra-cluster connectivity; by iteratively calculating the minimization of intra-cluster squared error and determining the cluster center, each cluster is assigned to the drone closest to the cluster center;
[0156] The intra-cluster squared error and J are shown in the following formula:
[0157]
[0158] Among them, c i is the cluster center of the i-th cluster;
[0159] Considering the capacity limitations of the mobile edge computing servers equipped with drones, when the load of the mobile edge computing servers equipped with drones exceeds the threshold, the IoT devices farthest from the cluster center are dynamically reallocated to other drones to ensure service quality and system stability;
[0160] S3.3: Construct a multi-agent flexible actor-critic algorithm to jointly optimize the power allocation and trajectory of the UAV, and obtain the UAV power allocation result and the optimized UAV trajectory;
[0161] In this embodiment, for the non-convex high-dimensional optimization problem of constraint objectives, a multi-agent flexible actor-critic algorithm based on K-means clustering is established, such as Figure 3 As shown in the figure, by fusing enhanced K-Means with the Multi-Agent Flexible Actor-Critic (MASAC) algorithm, an efficient solution is achieved based on a two-stage optimization framework.
[0162] Multiple UAVs are modeled as a distributed agent system. The Markov Decision Process (MDP) framework is used to formally describe the state space, action space, and reward function of the UAVs. On this basis, a multi-agent flexible actor-critic algorithm is constructed to achieve autonomous decision optimization of each UAV. The specific method is as follows:
[0163] To address the highly dynamic network characteristics caused by time-varying channel states and random device movement, a multi-agent flexible actor-critic algorithm is constructed using a centralized training and decentralized execution framework, including multiple policy sub-networks, an experience replay pool, multiple critic sub-networks, and an entropy regulation sub-network.
[0164] The policy sub-network is used to calculate the state of the drone according to the t Output continuous action a t , by minimizing the KL divergence to achieve strategy improvement, the loss function J of the f-th strategy sub-network π,f (θ) is shown in the following formula:
[0165]
[0166] Where α is the temperature coefficient, is the function used to normalize the policy distribution;
[0167] Function to normalize policy distribution The value of depends on the state of the agent, but has no effect on the gradient of the policy network parameters. The entropy network has the ability to automatically adjust the temperature coefficient α to maintain a balance between exploration and exploitation.
[0168] Experience replay buffer is used to store transfer tuples<S,A,P,R,> ,By randomly sampling training samples, it is possible to reuse the previous data for training, thereby improving the sample utilization and breaking the correlation between samples;
[0169] The critic subnetwork is used to optimize the target Q value based on the mean square error loss function and Adam optimizer. The loss function J of the g-th critic subnetwork is Q,g (w) is shown in the following formula:
[0170]
[0171] Among them, w is a parameter, Calculate the expected mean, Q i (·) is the g-th critic network in state s t and action a t The predicted Q value is is the target Q value, and D is the experience replay buffer, which is used to collect training data and provide training samples;
[0172] The entropy regulation sub-network is used to dynamically adjust the temperature coefficient α to balance exploration and utilization. The loss function J of the entropy regulation sub-network is h (α) is shown in the following formula:
[0173]
[0174] in, is the pre-set minimum policy entropy threshold;
[0175] In order to make the exploration behavior of the agent more consistent with the continuous control logic of the physical world, by introducing time correlation and mean regression characteristics, Ornstein-Uhlenbeck noise, which can generate time-series-related exploration, is used as the exploration noise of the strategy sub-network.
[0176] Because the multi-agent flexible actor-critic algorithm adopts an online learning strategy, the drone can continuously optimize its strategy during mission execution and adapt to environmental changes. By jointly optimizing power and trajectory, the drone dynamically adjusts its transmission power according to the optimization strategy, reduces interference and extends flight time, and flies along the optimized three-dimensional path, avoiding obstacles and maintaining the best communication link.
[0177] In this embodiment, numerical simulation is used to further illustrate the power allocation and flight trajectory optimization algorithm in the UAV-assisted mobile edge computing data collection system. The simulation parameter settings include:
[0178] Set a two-dimensional square area as the drone's service range, set the experience replay buffer pool capacity to 10,000 records, set the mini-batch size to 128, set the neural network hidden layer and output layer to use the ReLU activation function, and apply the Adam optimizer for training;
[0179] In this embodiment, the influence of key parameters on the convergence of the multi-agent flexible actor-critic algorithm based on K-means clustering is analyzed by setting different learning rates. The convergence performance diagram of the multi-agent flexible actor-critic algorithm based on K-means clustering under different learning rates is shown in FIG. Figure 4 As shown, the parameter is lr act =0.001,lr cri =0.003,lr alpha = 0.0003, the multi-agent flexible actor-critic algorithm based on K-means clustering has a faster convergence speed and higher stability. This is because the high lr alpha It will lead to excessive exploration of the strategy and poor performance; and too large lr act with lr cri This will make the parameter update step too large, making it difficult to achieve fine optimization and eventually converge to a suboptimal solution. act =0.001,lr cri =0.003,lr alpha =0.0003 as the optimal parameter combination.
[0180] This example evaluates the performance of a multi-agent flexible actor-critic algorithm based on K-means clustering in multiple scenarios and compares it with a baseline solution. For comprehensive comparison, three multi-agent deep reinforcement learning baseline algorithms were selected, including:
[0181] (1) Multi-Agent Deep Deterministic Policy Gradient (MADDPG): This algorithm uses a centralized training and decentralized execution framework, allowing each agent to share policy information during training, thereby effectively coordinating multi-agent behavior. This method is suitable for competitive or collaborative multi-agent environments, and can not only solve the problem of environmental non-stationarity, but also improve policy stability and convergence.
[0182] (2) Multi-Agent Deep Q-Network (MADQN): This method trains an independent Q-network for each agent, enabling it to learn collaborative or competitive strategies in complex environments. Through a shared experience replay mechanism, agents can learn from each other's learning experiences, thereby improving training efficiency.
[0183] (3) Multi-Agent Flexible Executor-Critic (MASAC): This algorithm promotes exploration by maximizing policy entropy, enabling agents to learn more robust collaborative strategies in complex environments. It combines centralized training with a decentralized execution framework, allowing information sharing during training, thereby improving learning efficiency and policy stability.
[0184] In the scenario where 3 drones serve 18 IoT devices, the average rewards of different algorithms change with the number of training rounds as shown in the following figure. Figure 5 As shown in the figure, the experimental results show that the K-means clustering method and the multi-agent flexible actor-critic algorithm proposed in this embodiment have the fastest convergence speed and the highest system average reward; MASAC has the second best convergence performance; the reward value of MADDPG is significantly lower than KMASAC, indicating that it is not effective enough in scenarios with dense IoT devices; and MADQN has a convergence speed that lags significantly behind the proposed algorithm due to problems such as the curse of dimensionality and training instability, which limits its application potential in large-scale complex multi-agent environments.
[0185] In this embodiment, the OMA solution is used to compare with the NOMA solution implemented in this embodiment during the data upload phase. The remaining parameters are exactly the same as those of the proposed system. In the OMA solution, the bandwidth allocated to each IoT device is 1 / U of that in the NOMA case. When interference factors are ignored, the data transmission rate of the OMA solution is shown in the following formula:
[0186]
[0187] The comparison of the impact of power allocation on total energy consumption in NOMA and OMA scenarios is shown in the figure below. Figure 6 As shown in Figure 2, when no power allocation is performed, data is transmitted at maximum power. The results show that NOMA consistently outperforms OMA in all scenarios. Under both NOMA and OMA conditions, the power allocation strategy based on the KMASAC algorithm can significantly improve performance.
[0188] The embodiment compares the evolution of information age with and without information age constraint. Under the condition of T=30s, three randomly selected Internet of Things devices are selected. The evolution curve of information age with information age constraint is shown in FIG. 8, and the evolution curve of information age without information age constraint is shown in FIG. 9. When the information age of the collected data reaches the critical threshold A_MAX, the mobile edge computing server will immediately trigger the computing task. By comparing FIG. 8 and FIG. 9, it can be observed that the multi-unmanned aerial vehicle assisted edge computing information age perception network based on reinforcement learning in the embodiment can maintain a lower information age value, and the collected data can be processed more quickly. Figure 7 Figure 8 Figure 7 Figure 8
[0189] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some or all of the technical features. Such modifications or substitutions do not cause the corresponding technical solutions to deviate from the scope defined by the present application.
Claims
1. A multi-UAV assisted edge computing resource and trajectory optimization method, characterized in that: The following steps are involved: S1: Establish an IoT data collection network model, including a communication sub-model, a transmission sub-model, and a mobility sub-model. This involves multiple IoT devices, multiple drones equipped with mobile edge computing servers, and base stations. The drones collect data from multiple IoT devices and offload the data to the base stations. The communication sub-model is used to characterize the data transmission relationship between the drone, IoT device and base station, and calculate the data transmission rate of the drone; The transmission sub-model is used to calculate the delay and energy consumption of data transmission between the drone, IoT device and base station; the mobility sub-model is used to calculate the flight energy consumption of the drone; S2: Set decision variables based on the IoT data collection network model, and set constraints and constraint targets based on the decision variables; S3: The constraint objectives are modeled as a Markov decision process, and a multi-agent flexible actor-critic algorithm based on K-means clustering is established to jointly optimize the power allocation and trajectory of the UAV, obtaining the UAV power allocation result and the optimized UAV trajectory.
2. A multi-UAV assisted edge computing resource and trajectory optimization method according to claim 1, characterized in that: Said S1 comprises: S1.1: Set the parameters of the IoT data collection network model, including the number of drones, the number of IoT devices, the range of movement, and the location coordinates of drones, IoT devices, and base stations within the range of movement; S1.2: Establish a communication sub-model and use it to characterize the data transmission relationship between the drone, IoT device, and base station. Calculate the data transmission rate between the drone and the base station and between the IoT device and the drone. S1.3: Establish a transmission sub-model to calculate the latency and energy consumption of data transmission between the drone, IoT device, and base station. The transmission sub-model includes a data upload unit, a data offload unit, and a local computing unit. The data upload unit is used to calculate the latency and energy consumption of data uploaded by the IoT device to the drone. The data offload unit is used to calculate the latency and energy consumption of the drone offloading collected data to the base station. The local computing unit is used to assign a local computing method to the IoT device data and calculate the latency and energy consumption of the drone's local computing. S1.4: Establish a mobile sub-model and use it to calculate the flight energy consumption of the UAV: S1.5: Calculate the total energy consumption of the IoT data collection network based on the energy consumption of the drone uploading data, the energy consumption of the drone offloading collected data to the base station, the energy consumption of the drone local computing, and the energy consumption of the drone flying.
3. A multi-UAV assisted edge computing resource and trajectory optimization method according to claim 2, characterized in that: The specific method of S1.2 is: A communication sub-model is established based on the air-ground path loss model. In the data uplink, the line-of-sight connection probability, non-line-of-sight connection probability, line-of-sight connection path loss, and non-line-of-sight connection path loss between the IoT device and the drone are calculated. The average path loss and channel gain between the IoT device and the drone are also calculated. Calculate the average path loss between the drone and the base station and the transmission rate between the drone and the base station; Use constraints to limit the data upload service relationship between IoT devices and drones, and obtain the drone's receiving signal; Establish a decoding order decision criterion to determine the accumulated inter-cluster interference by identifying the optimal decoding order and eliminate the intra-cluster interference between IoT devices; Calculate the signal-to-interference-and-noise ratio between the IoT device and the drone to obtain the data transmission rate between the IoT device and the drone.
4. A multi-UAV assisted edge computing resource and trajectory optimization method according to claim 2, characterized in that: The specific method of allocating local computing mode for data of IoT devices by the local computing unit in S1.3 is as follows: Calculate the information age of IoT devices at time slot t; Set the maximum information age and assign local calculation methods to the data of IoT devices based on the maximum information age, specifically: The information age A of the kth IoT device k (t) Exceeds the maximum information age A MAX When the kth IoT device’s data is allocated to the UAV equipped with the mobile edge computing server for local calculation, otherwise, the kth IoT device’s data is offloaded to the base station for calculation.
5. The multi-UAV assisted edge computing resource and trajectory optimization method according to claim 1 is characterized in that: The S2 includes: S2.1 Set decision variables based on the IoT data collection network model, including: a service indicator variable, a power allocation variable, and a motion control variable; the service indicator variable is used to represent the association between the IoT device and the drone; the power allocation variable is used to represent the transmission power of the IoT device to the drone; and the motion control variable is used to represent the flight speed and flight angle of the drone; S2.2: Set constraints and constraint objectives based on decision variables.
6. A multi-UAV assisted edge computing resource and trajectory optimization method according to claim 5, characterized in that: The S2.2 is specifically: Set constraints based on service indicator variables, power allocation variables, and motion control variables, including: Constraints for limiting the number of drones that can serve each IoT device; Constraints on the decoding order of non-orthogonal multiple access for eliminating intra-cluster interference between IoT devices; Constraints for limiting the power consumption of each drone; Constraints used to limit the age of information; Constraints used to avoid collisions between drones; Constraints used to limit the range of drone activities; Constraints used to constrain the drone's flight speed and angle; The constraint objective is set to minimize the total energy consumption of the IoT data collection network model.
7. The multi-UAV assisted edge computing resource and trajectory optimization method according to claim 1, characterized in that: The S3 includes: S3.1: Model the drone as an agent in a Markov decision process, setting the drone's state set, action set, and reward function; S3.2: Assign IoT devices to drones based on K-means clustering; S3.3: Construct a multi-agent flexible actor-critic algorithm to jointly optimize the power allocation and trajectory of the UAV, and obtain the UAV power allocation result and the optimized UAV trajectory.
8. The multi-UAV assisted edge computing resource and trajectory optimization method according to claim 7, characterized in that: The specific method of S3.2 is: Taking the location set of IoT devices as input, user clustering is performed through channel broadcasting, and IoT devices are divided into several clusters equal to the number of drones based on Euclidean distance; By iteratively minimizing the intra-cluster square error and determining the cluster center of each cluster, each cluster is assigned to the drone closest to the cluster center; when the load of the mobile edge computing server equipped with the drone exceeds the threshold, the IoT device farthest from the cluster center is dynamically reallocated to other drones.
9. The multi-UAV assisted edge computing resource and trajectory optimization method according to claim 7, characterized in that: The specific method of S3.3 is: A multi-agent flexible actor-critic algorithm is constructed using a centralized training and decentralized execution framework, including multiple policy sub-networks, an experience replay pool, multiple critic sub-networks, and an entropy regulation sub-network. The policy sub-network is used to calculate the state of the drone according to the t Output continuous action a t ; Experience replay buffer is used to store transfer tuples<S,A,P,R,> The critic sub-network is used to optimize the target Q value based on the mean square error loss function and the Adam optimizer; the entropy regulation sub-network is used to dynamically adjust the temperature coefficient α of the strategy sub-network; The multi-agent flexible actor-critic algorithm is used to jointly optimize the power allocation and trajectory of the UAV, and the UAV power allocation result and the optimized UAV trajectory are obtained.