Multi-UAV deployment and collaborative offloading method for the Internet of Things

By optimizing UAV deployment and offloading through constrained K-Means clustering and MARL strategies, the problems of uneven deployment and unbalanced load of UAVs in IoT systems are solved, coverage is improved, latency and energy consumption are reduced, and the system adapts to the changes of large-scale IoT systems.

CN119300048BActive Publication Date: 2025-09-26FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411328230.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-24
Publication Date
2025-09-26
Estimated Expiration
2044-09-24

AI Technical Summary

Technical Problem

In existing IoT systems, the uneven deployment and unreasonable offloading of UAVs lead to inefficient resources and unbalanced loads. The complexity of large-scale systems is difficult to solve efficiently, affecting service coverage and latency.

Method used

A UAV deployment scheme based on constrained K-Means clustering and a multi-UAV collaborative offloading strategy based on MARL are adopted. By adaptively adjusting UAV positions and collaborative offloading decisions, UAV deployment and computation offloading are optimized, thereby improving coverage balance and resource utilization.

Benefits of technology

It achieves better UAV deployment performance, improves service coverage and coverage balance, reduces the total task processing latency and UAV energy consumption, and adapts to the changes of large-scale IoT systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119300048B_ABST
    Figure CN119300048B_ABST
Patent Text Reader

Abstract

This paper provides a multi-UAV deployment and collaborative offloading method for the Internet of Things. First, the original joint optimization problem is transformed into a UAV deployment subproblem and a computation offloading subproblem. Next, a UAV deployment scheme based on constrained K-Means clustering is proposed for the UAV deployment subproblem. By introducing UAV coverage constraints into K-Means clustering, the UAV deployment locations are adaptively adjusted to improve the coverage and coverage balance of the computation offloading service in the system. Finally, a multi-UAV collaborative computation offloading strategy based on MARL is proposed for the computation offloading subproblem. Through a centralized training and decentralized execution model, the proposed strategy achieves near-optimal computation offloading and UAV collaboration. Application of this technical solution can achieve better UAV deployment performance, improving service coverage while improving coverage balance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Internet of Things, and in particular to a multi-UAV deployment and collaborative unloading method for the Internet of Things. Background Art

[0002] With the rapid development of the Internet of Things (IoT) and communication technologies, it is becoming increasingly prevalent in scenarios such as smart cities, Industry 4.0, and smart agriculture. Hundreds of millions of sensors and smart devices have formed large-scale IoT systems, giving rise to many emerging intelligent applications (such as facial recognition and virtual reality). However, limited by battery capacity and hardware performance, IoT devices struggle to meet the high real-time requirements of these compute-intensive intelligent applications. While classic cloud computing provides abundant computing and storage resources, long-distance communications significantly increase data transmission latency, severely impacting Quality of Service (QoS). By shifting resources to the edge of the network, the emergence of Mobile Edge Computing (MEC) can effectively reduce data transmission latency, alleviate the bandwidth load on the core network, and thereby improve the performance of IoT systems and user experience.

[0003] The communication of existing IoT systems mainly relies on ground infrastructure (such as base stations). However, when facing remote areas (such as forests and oceans) and emergency situations (such as disaster relief and military exercises), the ground infrastructure may not be able to meet the communication needs. As an aerial mobile device, unmanned aerial vehicles (UAVs) can be used in more diverse application scenarios. Thanks to the mobility and communication capabilities of UAVs, they can assist MEC to realize the construction of aerial networks and provide more flexible computing offloading services for IoT devices. However, due to the uneven distribution of IoT devices and the unreasonable deployment of UAVs, the load between UAVs may be unbalanced, resulting in increased system latency and excessive energy consumption of UAVs. Therefore, how to design an efficient multi-UAVs deployment and collaborative offloading method still faces the following challenges:

[0004] (1) Uneven distribution of IoT devices leads to low resource efficiency of UAVs. The distribution of IoT devices is generally diverse and uneven, but UAVs can only provide computing offload services for IoT devices within a limited coverage area. Therefore, unreasonable UAV deployment plans will seriously affect the number of IoT devices that can be served and also lead to large differences in service coverage in different regions.

[0005] (2) Diverse offloading requests lead to significant differences in load between UAVs. The number of requests and computing requirements of IoT devices within the coverage area of ​​different UAVs vary greatly, which leads to large differences in load conditions on different UAVs and significant differences in resource utilization between different UAVs.

[0006] (3) Complex problems caused by large-scale IoT systems are difficult to solve efficiently. The offloading requests from a large number of IoT devices and the real-time collaboration between UAVs significantly increase the complexity of the problem, making it difficult to solve efficiently. Existing solutions usually use control theory or iterative algorithms, resulting in excessive overhead. Some research works using classic centralized reinforcement learning are plagued by problems such as explosion of action dimensions and low exploration efficiency. Summary of the Invention

[0007] In light of this, the present invention aims to provide a multi-UAV deployment and collaborative offloading method for the Internet of Things. First, it achieves better UAV deployment performance, improving both service coverage and coverage balance. Second, it efficiently performs computation offloading and UAV collaboration to reduce overall task processing latency and UAV energy consumption.

[0008] To achieve the above objectives, the present invention adopts the following technical solutions: a multi-UAV deployment and collaborative offloading method for the Internet of Things. First, the original joint optimization problem is transformed into a UAV deployment subproblem and a computation offloading subproblem. Then, for the UAV deployment subproblem, a UAV deployment scheme based on constrained K-Means clustering is proposed. By introducing UAV coverage range constraints into K-Means clustering, the deployment positions of UAVs are adaptively adjusted to improve the coverage and coverage balance of computation offloading services in the system. Finally, for the computation offloading subproblem, a multi-UAV collaborative computation offloading strategy based on MARL is proposed. Through the centralized training and decentralized execution mode, the proposed strategy achieves a near-optimal computation offloading and UAV collaboration strategy.

[0009] In a preferred embodiment, a UAV deployment and collaborative offloading system for IoT devices is included, wherein the UAV deployment and collaborative offloading system for IoT devices is composed of UAVs and IoT devices; IoT devices request computing offloading services from UAVs via wireless links, denoted as a set U = {u1, u2, ..., u n}, the horizontal coordinates are represented as {w1,w2,...,w n UAVs provide computing offloading services to IoT devices within their coverage area, which is recorded as the set UAV = {uav1, uav2, ..., uav m}, the horizontal coordinates are expressed as {q1,q2,...,q m};The horizontal height of UAVs is H, then u i(1≤i≤n) With UAV j(1≤j≤m) The distance between them is:

[0010]

[0011] Collaborative offloading is required between multiple UAVs. UAVs forward tasks from IoT devices to their collaborative UAVs for execution. The system time slot is represented as t∈{1,2,...,T}. At the beginning of each time slot, the IoT device chooses to execute its computing request locally or offload it to the UAVs connected to it. If the IoT device is not within the coverage of any UAVs, its computing request is executed locally.

[0012] In a preferred embodiment, at time slot t, the i The task is defined as a triple <D i [t],S i [t],c i >, where D i [t] represents the amount of input data for the task, S i [t] represents the computational complexity of the task, c i represents the associated indicator variable; when u i Connect to UAV j When c i =j; when u i When unable to connect to an available UAV, c i = 0; tasks from IoT devices can be executed locally or offloaded to UAVs; u i The offloading decision at time slot t is denoted as α i [t]∈{0,1}; when u i Executing the mission locally or failing to connect to an available UAV (c i =0), α i [t]=0; when u i When offloading the task to a UAV for execution, α i [t]=1; specifically, it is divided into the following two calculation modes:

[0013] (1) Local computing mode

[0014] When the task is executed locally, its latency is:

[0015]

[0016] Among them, f i Indicates u i computing power;

[0017] (2) UAV calculation mode

[0018] IoT devices access UAVs via OFDMA; when u i The task is offloaded to uav j When the input data is uploaded first; according to Shannon's theorem, u i With UAV j The data transmission rate between is:

[0019]

[0020] Among them, B u represents the channel bandwidth between IoT devices and UAVs, P i Represents u i The transmission power; Where ρ0 represents the channel gain per unit distance, and N0 represents the noise power spectral density;

[0021] Therefore, u i Offload tasks to uav j The transmission delay is:

[0022]

[0023] Next, UAV j Need to process concurrently j [t] tasks; therefore, from u i The mission in UAV j The computational delay on is:

[0024]

[0025] in, f max Indicates the computing power of a UAV;

[0026] Tasks on high-load UAVs can be forwarded to low-load UAVs to achieve load balancing; specifically, uav j The cooperative offloading decision at time slot t is defined as β ijk [t]∈{0,1}; when uav j Decided to put u i The task is forwarded to uav k When β ijk [t]=1; otherwise, β ijk [t]=0; a task is forwarded at most once, denoted as When the task is from uav j Forward to uav k When cooperative offloading is performed, the data transmission rate between the two is:

[0027]

[0028] Among them, B uav represents the channel bandwidth between UAVs, P uav represents the transmission power between UAVs, UAV j With UAV k The distance between them; therefore, u i The mission from UAV j Forward to uav k The transmission delay is:

[0029]

[0030] The forwarded task is in uav k The computational delay on is:

[0031]

[0032] Execute u i The task delay is:

[0033]

[0034] Processing i The total delay of the task is:

[0035]

[0036] In summary, the total delay for processing all tasks in time slot t is:

[0037]

[0038] In a preferred embodiment, when tasks are offloaded to UAVs for processing, three aspects of energy consumption are generated, including the energy consumption of UAVs executing tasks, the energy consumption of forwarding tasks between UAVs, and the energy consumption of UAVs maintaining a hovering state;

[0039] uav j Execute u i The energy consumption of the task is:

[0040]

[0041] Where κ represents the effective capacitance coefficient;

[0042] uav j will u i Task forwarding to uav k The energy consumption is:

[0043]

[0044] The energy consumption of UAVs to maintain the hovering state in each time slot is e; therefore, the total energy consumption of all UAVs in time slot t is:

[0045]

[0046] in, Indicates uav j The set of tasks to be performed, Indicates that from uav j The set of tasks forwarded;

[0047] Due to the limited power of UAVs, it is impossible to provide computing offload services for IoT devices all the time; when the power is lower than the threshold b min When the UAVs are in the state of emergency, they will no longer provide computing offload services and return to the home station for battery recharge;

[0048] Based on the model proposed above, the optimization objective is to minimize the weighted sum of the total task processing delay and the UAV energy consumption. The optimization problem is formally defined as:

[0049]

[0050] Among them, λ1 and λ2 represent the weights of delay and energy consumption respectively, b j [t] indicates uav j The remaining power at time slot t, b min represents the minimum amount of power that UAVs must retain; C1 and C2 represent the value range constraints for offloading decisions; C3 represents the number of times a task is forwarded between UAVs; C4 represents the distance constraint between the IoT device and its associated UAV, where R is the coverage radius of the UAVs; C5 represents the value range constraint of the associated indicator variable; constraint C6 indicates that the computing resources used by UAVs to execute tasks cannot exceed their maximum computing resource constraints; C7 represents the minimum amount of power that UAVs need to retain when providing computing offloading services. min Used for returning to the hop and continuing the power supply.

[0051] In a preferred embodiment, it is further divided into two sub-problems: UAV deployment and computational offloading. First, UAV deployment optimization is performed to improve the coverage and coverage balance of UAVs computational offloading services. After the UAVs are deployed, computational offloading optimization is performed to reduce the total task processing latency and UAV energy consumption.

[0052] In a preferred embodiment, for the UAV deployment problem, IoT devices are divided into multiple clusters, and UAVs are deployed at the centroid of each cluster to maximize service coverage. Service coverage and coverage balance are defined as performance indicators of UAV deployment, and a UAV deployment scheme based on constrained K-Means clustering is designed. Specifically, clusters are divided based on UAV coverage, so that the cluster size does not exceed the UAV coverage range. At the same time, the balance of the number of IoT devices within the coverage range of each UAV is considered when dividing the clusters.

[0053] In a preferred embodiment, in each iteration, the positions of the UAVs are first randomly initialized, and the cluster set change flag is initialized to check whether the cluster division has reached convergence. Next, the IoT devices within the coverage area of ​​the UAVs are assigned to the corresponding UAV cluster. Specifically, unlike the classic K-Means algorithm, which assigns data points based on the principle of minimum distance to the centroid, the allocation of IoT devices considers whether the IoT device is within the coverage radius of the UAV. When an IoT device is covered by multiple UAVs at the same time, the nearest UAV is selected and the change is marked, thereby updating the number of UAVs covered. Next, the current coverage rate and the load level of all UAVs are calculated, which are defined as:

[0054]

[0055] When the coverage is close to equilibrium, F1≈F2≈...≈F m ; The mean square error is used to evaluate the coverage balance, which is defined as:

[0056]

[0057] in,

[0058] Next, the coverage rate and coverage balance are normalized to the same numerical range; the mean of each cluster is used as the new centroid and UAVs are deployed; through iteration, the deployment location of UAVs will gradually tend to areas with a dense distribution of IoT devices, enabling them to serve more IoT devices; after the iteration is completed, the solution with the smallest weighted sum of coverage rate and coverage balance is selected for UAV deployment.

[0059] In a preferred embodiment, after solving the UAV deployment problem, the computation offloading problem is further separated from P1 and expressed as:

[0060]

[0061] A MARL-based multi-UAV collaborative computation offloading strategy is proposed. The proposed multi-UAV collaborative offloading system is regarded as an environment. Multiple agents interact with the environment simultaneously and select corresponding offloading actions. After receiving the reward signal from the environment, each actor network is updated in a distributed manner. The critic network adopts a centralized update method. Accordingly, the state space, action space, and reward function are defined as follows:

[0062] State space: The state space includes task attributes, computing capabilities of IoT devices, and i , data transmission power P of IoT devices i , the remaining power of UAV b[t]; therefore, at time slot t, the state observed by agent j is expressed as:

[0063]

[0064] Among them, 1≤j≤m, and Represents the set consisting of the data volume and computational complexity of all tasks respectively;

[0065] Action space: The action space contains the offloading decision α[t] of the IoT device and the collaborative offloading decision β[t] of the UAVs. α[t] and β[t] are integrated, and the integrated offloading decision is recorded as η[t]∈{0,1,2,...,m}. Therefore, at time slot t, the action of agent j is expressed as:

[0066]

[0067] in, represents the offloading decision made by agent j for all tasks; η i [t]=0 means u i No task is generated or the agent decides to execute the task locally; η i [t]=k, k∈{1,2,...,m} means the agent decides to offload the task to the uav k Execute on;

[0068] Reward function: The optimization objective of P2 is to minimize the weighted sum of the total task processing delay and the UAV energy consumption; therefore, at time slot t, the reward function is expressed as:

[0069] r[t]=-(λ1N(T u_tot [t])+λ2N(E uav_tot [t])), (22)

[0070] Here, λ1 and λ2 represent the weights of delay and energy consumption, respectively; N is a normalization function that normalizes delay and energy consumption to the same numerical range.

[0071] In a preferred embodiment, the value network acts as a critic and uses multi-step temporal difference to estimate the advantage function of the state-action pair; the policy network acts as an actor to update the policy according to the advantage function; the obtained UAVs coordinates q and the associated indicator variable c will be used as the input of Algorithm 2; specifically, the value and policy networks are first initialized; in each round of iteration, each agent obtains the initial state of the environment, initializes the round end flag and the interaction data record table, where the interaction data record table is used to collect all data of the interaction between the agent and the environment in the current round; in each round, each agent selects the appropriate unloading action according to its strategy, and converts η into the unloading decision α of the corresponding IoT device and the collaborative unloading decision β of the UAVs; then, α and β are executed and the reward, state and round end flag are updated; the round ends when the power of all UAVs reaches the minimum threshold, that is: when When , done = True; then, collect the interaction data of each agent and update its state; then, use the interaction data collected in the current round to train the agent; specifically, extract the interaction data and splice the state to centrally train the value network, whose generalized advantage function is defined as:

[0072]

[0073] Among them, γ is the discount factor and introduces a hyperparameter When λ = 0, it is the advantage obtained by one step of differentiation; when λ = 1, it is the complete average of the advantages obtained by each step of differentiation; by using the advantage function, the variance of the policy gradient estimation can be significantly reduced and the stability of the training process can be improved;

[0074] Next, the strategy obtained in the previous round is used as the old strategy for this round, and the action probability distribution of each agent is calculated; then, the same batch of data is used to update the network parameters multiple times; unlike other DRL algorithms, the previously sampled data is discarded after the network parameters are updated. The proposed method converts it into a different-strategy-based training process by setting new and old strategies; specifically, data is collected by interacting with the old strategy and the environment, and then the new strategy is trained using this data; according to the principle of importance sampling, the data sampled by the old strategy can be used multiple times, and gradient ascent can be performed multiple times for policy updates; after each update of the policy network parameters, the action probability distribution of each agent and the loss function of the policy network are calculated; in particular, clipping is used to limit the amplitude of each policy update to ensure that the policy does not change drastically during training. Finally, the loss function of the value network is calculated and the parameters of the policy and value networks are updated using gradient descent.

[0075] Compared with existing technologies, this invention has the following advantages: The proposed multi-UAV deployment and collaborative offloading method for the Internet of Things (IoT) achieves superior UAV deployment performance in various scenarios, improving both service coverage and coverage balance. Furthermore, the multi-UAV deployment and collaborative offloading method for the IoT enables more optimal offloading decisions. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] Figure 1 This is a schematic diagram of a UAV deployment and collaborative offloading system for IoT devices according to a preferred embodiment of the present invention;

[0077] Figure 2 This is a schematic diagram of multi-UAV collaborative offloading based on MARL in a preferred embodiment of the present invention;

[0078] Figure 3 Schematic diagrams of three scenarios with different IoT device distributions according to a preferred embodiment of the present invention, where (a) is scenario 1, (b) is scenario 2, and (c) is scenario 3;

[0079] Figure 4 Schematic diagrams of UAV deployment in different scenarios of a preferred embodiment of the present invention, where (a) is scenario 1, (b) is scenario 2, and (c) is scenario 3;

[0080] Figure 5 Schematic diagram comparing the UAV deployment performance of different methods according to a preferred embodiment of the present invention, wherein (a) is the weighted sum of coverage rate and coverage balance, (b) is the coverage rate, and (c) is the coverage balance;

[0081] Figure 6 This is a schematic diagram of average time slot rewards for different methods according to a preferred embodiment of the present invention.

[0082] Figure 7 Schematic diagram of average time slot delay of different methods in a preferred embodiment of the present invention.

[0083] Figure 8 This is a schematic diagram of average time slot UAV energy consumption in different methods of a preferred embodiment of the present invention.

[0084] Figure 9 The preferred embodiment of the present invention is different from max Below is a diagram of the average time slot rewards for different methods.

[0085] Figure 10 Figure 1 is a schematic diagram of the loads of different UAVs according to a preferred embodiment of the present invention.

[0086] Figure 11 Schematic diagram of load balancing degree of different methods according to a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0087] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0088] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present application belongs.

[0089] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form, and it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or their combinations.

[0090] refer to Figure 1-11 This paper proposes MUCO, a novel multi-UAV deployment and collaborative offloading method for IoT systems. First, the original joint optimization problem is transformed into two sub-problems: UAV deployment and computation offloading. Next, a UAV deployment scheme based on constrained K-Means clustering is proposed for the UAV deployment sub-problem. By incorporating UAV coverage constraints into K-Means clustering, the proposed scheme adaptively adjusts UAV deployment locations to improve the coverage and coverage balance of computation offloading services in the system. Finally, a multi-UAV collaborative computation offloading strategy based on MARL is proposed for the computation offloading sub-problem. Through centralized training and decentralized execution, the proposed strategy achieves near-optimal computation offloading and UAV collaboration. In particular, the distributed task execution approach is better suited for large-scale IoT systems. Extensive experiments validate the effectiveness of the proposed MUCO method, which enables adaptive UAV deployment and collaborative offloading based on the distribution of IoT devices. Compared with baseline methods, the proposed MUCO method achieves superior UAV deployment performance in various scenarios, improving both service coverage and coverage balance. In addition, the MUCO method can make better offloading decisions and achieve lower weighted sum of latency and energy consumption under the UAV power constraint.

[0091] System model and problem definition

[0092] like Figure 1 As shown in the figure, the proposed UAV deployment and collaborative offloading system for IoT devices consists of UAVs and IoT devices. IoT devices request computing offloading services from UAVs through wireless links, denoted as the set U = {u1,u2,...,u n}, the horizontal coordinates are represented as {w1,w2,...,w n UAVs can provide computing offloading services to IoT devices within their coverage area, which is denoted as the set UAV = {uav1, uav2, ..., uav m}, the horizontal coordinates are expressed as {q1,q2,...,q m}. The horizontal height of UAVs is H, then u i(1≤i≤n) With UAV j(1≤j≤m) The distance between them is:

[0093]

[0094] To improve system resource utilization, multiple UAVs need to collaborate on offloading. Specifically, UAVs can forward tasks from IoT devices to their collaborating UAVs for execution. At the start of each system time slot, denoted as t∈{1,2,...,T}, an IoT device can choose to execute its computational requests locally or offload them to connected UAVs. If an IoT device is not within the coverage area of ​​any UAV, its computational requests are executed locally.

[0095] Computation and Communication Model

[0096] At time slot t, from u i The task is defined as a triple <D i [t],S i [t],c i >, where D i [t] represents the amount of input data for the task, S i [t] represents the computational complexity of the task (i.e., the number of computational resources required per unit of computational task input: cycle / bit), c i represents the associated indicator variable. i Connect to UAV j When c i =j; when u i When unable to connect to an available UAV, c i = 0. Tasks from IoT devices can be executed locally or offloaded to UAVs. i The offloading decision at time slot t is denoted as α i [t]∈{0,1}. When u i Executing the mission locally or failing to connect to an available UAV (c i =0), α i [t]=0; when u i When offloading the task to a UAV for execution, α i [t] = 1. Specifically, it can be divided into the following two calculation modes.

[0097] (1) Local computing mode

[0098] When the task is executed locally, its latency is:

[0099]

[0100] Among them, f i Indicates u i computing power.

[0101] (2) UAV calculation mode

[0102] IoT devices access UAVs through OFDMA. i The task is offloaded to uav j When u i With UAV j The data transmission rate between is:

[0103]

[0104] Among them, B u represents the channel bandwidth between IoT devices and UAVs, P i Represents u i transmission power. Where ρ0 represents the channel gain per unit distance, and N0 represents the noise power spectral density.

[0105] Therefore, u i Offload tasks to uav j The transmission delay is:

[0106]

[0107] Next, UAV j Need to process concurrently j [t] tasks. Therefore, from u i The mission in UAV j The computational delay on is:

[0108]

[0109] in, f max Indicates the computing power of a UAV.

[0110] To improve the resource utilization efficiency of the system, tasks on high-load UAVs can be forwarded to low-load UAVs to achieve load balancing. j The cooperative offloading decision at time slot t is defined as β ijk[t]∈{0,1}. When uav j Decided to put u i The task is forwarded to uav k When β ijk [t]=1; otherwise, β ijk [t]=0. A task is forwarded at most once, which is recorded as When the task is from uav j Forward to uav k When cooperative offloading is performed, the data transmission rate between the two is:

[0111]

[0112] Among them, B uav represents the channel bandwidth between UAVs, P uav represents the transmission power between UAVs, UAV j With UAV k Therefore, u i The mission from UAV j Forward to uav k The transmission delay is:

[0113]

[0114] The forwarded task is in uav k The computational delay on is:

[0115]

[0116] Compared with the input of the task, the amount of data of the calculation result is usually small, so the delay of returning the result can be ignored. i The task delay is:

[0117]

[0118] Furthermore, processing u i The total delay of the task is:

[0119]

[0120] In summary, the total delay for processing all tasks in time slot t is:

[0121]

[0122] Energy consumption model

[0123] Unmanned aerial vehicles (UAVs) consume significant energy when providing computational offloading services. Due to their limited battery capacity, UAVs need to improve their energy efficiency to better complete tasks from IoT devices. When tasks are offloaded to UAVs, energy consumption can arise in three areas: the energy consumed by the UAVs executing the task, the energy consumed by forwarding tasks between UAVs, and the energy consumed by the UAVs maintaining their hovering state.

[0124] uav j Execute u i The energy consumption of the task is:

[0125]

[0126] Here, κ represents the effective capacitance coefficient.

[0127] uav j will u i Task forwarding to uav k The energy consumption is:

[0128]

[0129] The energy consumption of UAVs to maintain the hovering state in each time slot is e. Therefore, the total energy consumption of all UAVs in time slot t is:

[0130]

[0131] in, Indicates uav j The set of tasks to be performed, Indicates that from uav j The set of tasks forwarded.

[0132] Due to the limited power of UAVs, they cannot always provide computing offloading services for IoT devices. min When the UAVs are in the air, they will no longer provide computing offload services and return to the home station for battery recharging.

[0133] Based on the above proposed model, the optimization goal of the present invention is to minimize the weighted sum of the total task processing delay and the UAV energy consumption. Therefore, the optimization problem is formally defined as:

[0134]

[0135] Among them, λ1 and λ2 represent the weights of delay and energy consumption respectively, b j [t] indicates uav j The remaining power at time slot t, b minIndicates the minimum amount of power that UAVs must retain. C1 and C2 represent the value range constraints for offloading decisions. C3 represents the number of times a task is forwarded between UAVs. C4 represents the distance constraint between the IoT device and its associated UAV, where R is the coverage radius of the UAVs. C5 represents the value range constraint of the associated indicator variable. Constraint C6 indicates that the computing resources used by UAVs to execute tasks cannot exceed their maximum computing resource constraints. C7 indicates that UAVs need to retain the minimum amount of power b when providing computing offloading services. min Used for returning to the hop and continuing the power supply.

[0136] In order to solve P1 more efficiently, the present invention further splits it into two sub-problems: UAV deployment and computational offloading. First, UAV deployment optimization is performed to improve the coverage and coverage balance of UAVs computational offloading services. IoT devices that are not within the coverage range of UAVs can only perform tasks locally, and a large amount of local computing will cause excessive latency, so improving coverage helps reduce system latency. In addition, the uneven distribution of IoT devices leads to unbalanced loads on different UAVs. If the coverage balance can be improved, the number of tasks offloaded to each UAV will be more evenly distributed, thereby reducing the number of task forwardings between UAVs and reducing transmission latency. After the UAVs are deployed, computational offloading optimization is performed to obtain the best offloading and collaborative offloading decisions.

[0137] The proposed MUCO method

[0138] This paper proposes a novel Multi-UAV deployment and collaborative offloading (MUCO) method for IoT systems. First, a UAV deployment scheme based on constrained K-Means clustering is designed to improve service coverage and coverage balance. Next, a multi-UAV collaborative offloading strategy based on MARL is developed to reduce total task processing latency and UAV energy consumption.

[0139] UAV deployment based on constrained K-Means clustering

[0140] Considering the superior partitioning clustering effect of the K-Means algorithm, this paper introduces it for the UAV deployment problem. For the UAV deployment problem, IoT devices can be divided into multiple clusters, and then UAVs are deployed at the centroid of each cluster to maximize service coverage. The size and density of each cluster obtained by the classic K-Means algorithm may vary significantly, and a single UAV may not be able to cover a cluster well. Therefore, directly deploying UAVs at the centroid may result in a large difference in the number of IoT devices served by each UAV, which will seriously affect the quality of UAV deployment.

[0141] To address this issue, this paper defines service coverage and coverage balance as performance metrics for UAV deployment, and then designs a UAV deployment solution based on constrained K-Means clustering. Specifically, clusters are divided based on UAV coverage, ensuring that cluster size does not exceed the UAV coverage area. Furthermore, cluster division takes into account the balance of the number of IoT devices within each UAV's coverage area. The main steps are shown in Algorithm 1.

[0142]

[0143]

[0144] In each round of iteration, the positions of the UAVs are first randomly initialized, and the cluster set change identifier is initialized (lines 3-4), which is used to check whether the cluster division has reached convergence. Then, the IoT devices within the coverage range of the UAVs are assigned to the clusters of the corresponding UAVs (lines 7-15). Specifically, unlike the classic K-Means algorithm that distributes data points according to the principle of minimum distance from the center of mass, the present invention considers whether the IoT device is within the coverage radius of the UAVs when allocating IoT devices (line 9). When an IoT device is covered by multiple UAVs at the same time, the nearest UAV will be selected and the change will be identified (lines 10-12), and the number of UAVs covered will be updated (line 13). Next, the current coverage rate and the load level of all UAVs are calculated (lines 16-17), which are defined as:

[0145]

[0146] When the coverage is close to equilibrium, F1≈F2≈...≈F m The mean square error is used to evaluate the coverage balance (line 18), which is defined as:

[0147]

[0148] in,

[0149] Next, the coverage rate and coverage balance are normalized to the same numerical range (line 20). The mean of each cluster is used as the new centroid and UAVs are deployed (lines 21-23). ​​Through iteration, the deployment locations of UAVs will gradually move toward areas with a high density of IoT devices, enabling them to serve more IoT devices. After the iteration is complete, the solution with the smallest weighted sum of coverage rate and coverage balance is selected for UAV deployment (line 26).

[0150] Multi-UAV cooperative offloading based on MARL

[0151] After solving the UAV deployment problem, the computation offloading problem can be further separated from P1, which is expressed as:

[0152]

[0153] Most existing work on computation offloading uses centralized DRL, which can make reasonable decisions in small action spaces. However, as the scale of IoT systems continues to grow, this approach leads to dimensionality explosion and unacceptable training complexity. The emergence of MARL offers a new approach to addressing this problem. Multiple agents can independently handle a large number of offload requests from users, reducing training complexity and accelerating decision-making, better meeting the stringent real-time requirements of IoT systems. Furthermore, the combined strategy of centralized training and distributed execution promotes collaboration between agents while maintaining state consistency, improving system performance. Specifically, centralized training enables multiple agents to learn from a shared global perspective, accelerating their learning process by sharing experience during training. Distributed execution enables each agent to make decisions based on individual perception and local information, enabling rapid response and adaptation to the environment. This significantly reduces the size of each agent's state and action space, simplifying training complexity.

[0154] Based on the above analysis, this paper proposes a multi-UAV collaborative computing offloading strategy based on MARL to optimize system latency and energy consumption. Figure 2 As shown in the figure, the proposed multi-UAV cooperative offloading system is considered an environment. Multiple agents simultaneously interact with the environment and select corresponding offloading actions. To reduce the training complexity of each agent, each actor network is updated in a distributed manner after receiving reward signals from the environment. To better grasp the overall system state information, the critic network adopts a centralized update method. Accordingly, the state space, action space, and reward function are defined as follows.

[0155] State space: The state space includes task attributes, computing capabilities of IoT devices, and i , data transmission power P of IoT devices i , the remaining power of the UAV b[t]. Therefore, at time slot t, the state observed by agent j (1≤j≤m) is expressed as:

[0156]

[0157] in, and Represents the set consisting of the data volume and computational complexity of all tasks respectively.

[0158] Action space: The action space contains the offloading decision α[t] of the IoT device and the collaborative offloading decision β[t] of the UAVs. To improve training efficiency, α[t] and β[t] are integrated, and the integrated offloading decision is recorded as η[t]∈{0,1,2,...,m}. Therefore, at time slot t, the action of agent j is expressed as:

[0159]

[0160] in, represents the offloading decision made by agent j for all tasks. i [t]=0 means u i No task is generated or the agent decides to execute the task locally. i [t]=k, k∈{1,2,...,m} means the agent decides to offload the task to the uav k Execute on.

[0161] Reward function: The optimization objective of P2 is to minimize the weighted sum of the total task processing delay and the UAV energy consumption. Therefore, at time slot t, the reward function is expressed as:

[0162] r[t]=-(λ1N(T u_tot [t])+λ2N(E uav_tot [t])), (22)

[0163] Where λ1 and λ2 represent the weights of latency and energy consumption, respectively. N is a normalization function that normalizes latency and energy consumption to the same numerical range.

[0164] The main steps of the proposed MARL-based multi-UAV cooperative offloading strategy are shown in Algorithm 2.

[0165]

[0166]

[0167] Algorithm 2 is designed based on the actor-critic DRL framework. Among them, the value network acts as a critic and uses multi-step temporal difference to estimate the advantage function of the state-action pair; the policy network acts as an actor to update the policy according to the advantage function. The UAVs coordinates q and the associated indicator variable c obtained by Algorithm 1 will be used as the input of Algorithm 2. Specifically, the value and policy networks are first initialized (line 1). In each round of iteration, each agent obtains the initial state of the environment (line 3), initializes the round end flag and the interaction data record table (lines 4-5), where the interaction data record table is used to collect all data of the interaction between the agent and the environment in the current round. In each round, each agent selects a suitable unloading action according to its strategy (line 7), and converts η into the unloading decision α of the corresponding IoT device and the collaborative unloading decision β of the UAVs (line 8). Then, execute α and β and update the reward, state and round end flag (line 9). The round ends when the power of all UAVs reaches the minimum threshold, that is: when When , done = True. Subsequently, the interaction data of each agent is collected and its state is updated (lines 10-11). Next, the agent is trained using the interaction data collected in the current round (lines 13-22). Specifically, the interaction data is extracted and the state is concatenated for centralized training of the value network (lines 13-15), whose generalized advantage function is defined as:

[0168]

[0169] Among them, γ is the discount factor and introduces a hyperparameter When λ = 0, it is the advantage obtained by one step of differentiation; when λ = 1, it is the complete average of the advantages obtained by each step of differentiation. By using the advantage function, the variance of the policy gradient estimate can be significantly reduced and the stability of the training process can be improved.

[0170] Next, the policy obtained in the previous round is used as the old policy for this round, and the action probability distribution of each agent is calculated (line 16). Subsequently, the same batch of data is used to update the network parameters multiple times (lines 17-22). Unlike other DRL algorithms, which discard previously sampled data after updating network parameters, the proposed method transforms this into an off-policy training process by setting up a new and old policy. Specifically, data is collected through the interaction between the old policy and the environment, and this data is then used to train the new policy. Based on the principle of importance sampling, the data sampled by the old policy can be reused multiple times, and gradient ascent can be performed multiple times to update the policy. This design significantly reduces the time overhead of data sampling and improves the utilization of the same batch of data. After each update of the policy network parameters, the action probability distribution of each agent and the policy network loss function are calculated (lines 18-19). In particular, clipping is used to limit the magnitude of each policy update, ensuring that the policy does not change drastically during training, thereby avoiding instability caused by excessive updates. Finally, the value network loss function is calculated, and the parameters of the policy and value networks are updated using gradient descent (lines 20-21).

[0171] Method evaluation

[0172] The simulation experiment was carried out on a workstation equipped with Intel (R) Xeon (R) Silver 4208 CPU and NVIDIA GetForce GTX 3090 GPU, with a CPU clock frequency of 2.10 GHz and 96 GB of memory. The proposed system model and method were implemented in Python, with the main simulation parameter settings shown in Table 1. To better simulate the dynamic nature of the environment, the task traffic was set to vary over time slots (i.e., different IoT devices generated tasks in each time slot), and the number and attributes of tasks also varied over time slots. The number of time slots in a round was determined by the time it took for all UAVs to reach a minimum power threshold from full charge.

[0173] Table 1 Parameter settings

[0174]

[0175]

[0176] In the simulation experiment, all IoT devices are distributed in an area of ​​800×800m 2 The number of IoT devices n = 150 and the number of UAVs m = 4. Figure 3 As shown, three scenarios with different IoT device distributions are considered, with the following settings:

[0177] (1) Scenario 1: Figure 3 As shown in (a), the scenario contains a local area with a high density of IoT devices. About 90% of IoT devices are concentrated in this local area, and the remaining about 10% of IoT devices are randomly distributed in other areas.

[0178] (2) Scenario 2: Figure 3 As shown in (b), the scenario contains a local area with a high density of IoT devices. About 60% of IoT devices are concentrated in this local area, and the remaining 40% of IoT devices are randomly distributed in other areas.

[0179] (3) Scenario 3: Figure 3 As shown in (c), the scene contains two local areas with a high density of IoT devices. Approximately 50% and 30% of IoT devices are concentrated in these two local areas, respectively, and the remaining approximately 20% of IoT devices are randomly distributed in other areas.

[0180] To verify the superiority of the proposed MUCO method, it is compared with the following UAV deployment and computation offloading methods:

[0181] (1) RD: Randomly select m locations from the locations of IoT devices as the deployment locations of UAVs.

[0182] (2) KD: IoT devices are regarded as clusters to be divided, and the number of UAVs m is regarded as the number of clusters. Then, the classic K-Means is used to divide IoT devices into corresponding clusters, and UAVs are deployed at the centroid of each cluster.

[0183] (3) Local: All tasks of IoT devices are executed locally.

[0184] (4) MEC: The tasks of IoT devices are all offloaded to UAVs for execution, and UAVs collaborate randomly with each other.

[0185] (5) PPO: Integrate the state and action spaces of all agents and use centralized PPO to make offloading decisions.

[0186] (6) DDQN: Use DDQN to make offloading decisions.

[0187] (7)SA: Use simulated annealing algorithm to make unloading decisions.

[0188] first, Figure 4The proposed MUCO method was used to demonstrate UAV deployment in different scenarios. Experimental results show that in all three scenarios, the proposed MUCO method effectively locates areas with dense IoT device distribution and achieves relatively balanced service coverage. Within the limited UAV coverage area, the MUCO method can cover areas with higher IoT device density, improving coverage. Compared to Scenario 2, ideal UAV deployment points are easier to find in Scenarios 1 and 3. This is because IoT devices are densely distributed in areas with very few devices elsewhere. In this scenario, UAVs with limited coverage can cover more IoT devices. In contrast, in Scenario 2, IoT devices are more sparsely distributed, resulting in more IoT devices being beyond the reach of UAVs. In this scenario, the MUCO method achieves good UAV deployment results, even within the constraints of the limited number of UAVs and their coverage area.

[0189] Then, the UAV deployment performance of each method in different scenarios was compared, including indicators such as coverage rate and coverage balance. Figure 5As shown, the proposed MUCO method achieves the minimum weighted sum of coverage and coverage balance in all scenarios, and generally outperforms the other two methods in both coverage and coverage balance. This is because the proposed MUCO method combines UAV coverage and clustering methods to better understand IoT device density for different IoT device distributions, thereby obtaining the optimal UAV deployment strategy. Compared to the RD method, the MUCO and KD methods achieve better coverage in all scenarios. This is because the RD method randomly selects UAV deployment locations and fails to understand the distribution of IoT devices. This may result in UAVs being deployed in areas with sparse IoT device distribution, resulting in poor coverage. The MUCO and KD methods achieve similar coverage in scenarios 1 and 3, but MUCO performs better in scenario 2. This is because the MUCO method considers UAV coverage constraints when dividing clusters, effectively controlling the range of each cluster while increasing the density. While the KD method can divide IoT devices into multiple clusters, it cannot control the range of each cluster, resulting in UAVs not being able to fully cover all IoT devices in each cluster. At the same time, the density of IoT device distribution within each cluster varies significantly, resulting in UAVs deployed in larger, sparser cluster centroids only covering a small number of IoT devices. Consequently, the KD method exhibits lower coverage in Scenario 2. Furthermore, in Scenario 2, the number and density of IoT devices within each cluster divided by the KD method vary significantly, leading to severe coverage imbalance. The RD method achieves slightly better coverage balance than the MUCO method in Scenario 2. This is because, in this scenario, it is easier to find several UAV deployment locations with similar distribution sparseness. Compared to the KD and RD methods, the proposed MUCO method achieves an average improvement in coverage and coverage balance of approximately 23.82% and 28.13% respectively across different scenarios.

[0190] Based on the UAV deployment results, the multi-UAV cooperative offloading performance of the proposed MUCO method was further tested. First, the average time slot rewards of different methods were compared. Figure 6As shown, the SA, MEC, and Local methods use single-step decisions and do not involve a learning process. MEC and Local methods perform inferior to other methods because their approach to task execution is too simplistic and fails to fully consider system state and task characteristics. For example, due to the limited computing power of UAVs, offloading all tasks to them would overload them, significantly reducing the computing resources allocated to each task and leading to severe performance degradation. Compared to the DDQN method, the MUCO, PPO, and SA methods achieve higher rewards. This is because the DDQN method cannot effectively cope with complex and changing environments and is prone to getting stuck in local optima. The PPO method integrates the states and actions of each agent, resulting in an overly large state and action space, making training more difficult and slowing convergence. However, the PPO method can outperform the SA method after a certain number of training rounds. The SA method continuously iterates over the same time slot to find the optimal solution for that time slot, but it cannot perceive environmental changes and find the optimal solution, and is prone to getting stuck in local optima. Compared to the PPO and SA methods, the proposed MUCO method effectively improves the average time slot reward, demonstrating the best performance among all methods. This is because the MUCO method uses a distributed execution approach, which reduces the state and action space that each agent processes, reducing training pressure and accelerating convergence. Furthermore, the centralized training approach used by the MUCO method enables it to fully perceive the overall state of the environment, thereby better optimizing the agent's offloading decisions.

[0191] Then, the average slot delays under different methods are compared. Figure 7 As shown in the figure, as the number of training rounds increases, the average slot delay of the MEC and Local methods remains basically unchanged and is at the highest value. Since all tasks are offloaded to the UAV for execution, this results in excessively high delay, making the performance of the MEC method similar to that of the Local method. In addition, the MEC method adopts a random collaboration approach, which achieves load balancing to a certain extent, but the effect is limited. The SA method reduces the average slot delay to a certain extent, but it no longer changes with the increase in training rounds. As the number of training rounds increases, the average slot delay of the other methods will gradually decrease. However, the average slot delay that the proposed MUCO method can finally achieve is the lowest, indicating that this method can make better offloading decisions, thereby reducing task processing delay.

[0192] Secondly, the average time slot UAV energy consumption of different methods is compared. According to formula (12), when the average allocation of computing resources is adopted, the fewer tasks a UAV processes, the more computing resources can be allocated to each task, and the higher the average time slot UAV energy consumption; conversely, the average time slot UAV energy consumption is lower. Figure 8As shown in the figure, the average time slot UAV energy consumption of the MEC method is at a low value. This is because it offloads all tasks to UAVs for execution. Each UAV needs to process a large number of tasks, so its average time slot UAV energy consumption is low. As the number of training rounds increases, the average time slot UAV energy consumption of the MUCO, PPO and DDQN methods gradually increases. This is because these three methods gradually learn more optimized strategies and then reasonably select tasks to be offloaded to UAVs for execution, which better balances the load between IoT devices and UAVs and reduces system latency. Although the proposed MUCO method is slightly higher than other methods in average time slot UAV energy consumption, the reward it obtains is significantly higher than other methods (such as Figure 6 At the same time, the advantage in reducing system latency is also significant.

[0193] Then, different f max The average time slot rewards of different methods are as follows. Figure 9 As shown, with f max As f increases, the rewards of all methods show an upward trend. This is because when the UAV has more computing resources, more computing resources are allocated to each task, which reduces the computing latency. max As f increases, the rate of increase of the rewards of each method gradually slows down. This is because the increase of computing resources allocated to the task also leads to an increase in computing energy consumption. max When f is low, the reward gap between the MEC method and other methods is large. This is because the MEC method offloads all tasks to UAVs for execution, and each task can only be allocated to limited computing resources, resulting in increased computing latency. max When f is high, the rewards of the MEC method are comparable to those of other methods. This is because when computing resources are sufficient, the MEC method can allocate enough computing resources to the task. Compared with other methods, the proposed MUCO method has a higher reward at different f max The highest reward can be obtained under all conditions, which verifies the superiority of the MUCO method.

[0194] The decision-making time of different methods was then compared. As shown in Table 2, the SA method requires a large number of iterations to find the optimal solution, and its decision-making time is approximately 10 times that of the other methods. In contrast, the MUCO, PPO, and DDQN methods are able to make decisions quickly based on the current environmental state, better meeting the real-time requirements of large-scale IoT systems. Compared to the PPO and DDQN methods, the proposed MUCO method achieves a shorter decision-making time. This is because the MUCO method uses a distributed execution approach, and the state and action space required of each agent is smaller than those of the PPO and DDQN methods.

[0195] Table 2 Decision time of different methods (ms)

[0196] method Decision Time MUCO 2.01 PPO 7.99 DDQN 8.00 SA 43.00

[0197] Finally, we evaluated the load on different UAVs within a time slot and the load balancing degree of different methods. The load of a UAV refers to the number of tasks offloaded to and calculated by the UAV. The load balancing degree is calculated as:

[0198]

[0199] Among them, ld j represents the load of the j-th UAV, Indicates the average load of all UAVs. The smaller the load balance value, the closer the load of each UAV is.

[0200] like Figure 10 As shown in Figure 2, when the proposed MUCO method is used for collaborative offloading, the loads of each UAV are closest, and the loads between UAVs are more balanced compared to other methods. Figure 11 As shown in Figure 3, the proposed MUCO method can achieve lower load balancing degree values ​​compared with other methods. The above results reflect that the proposed MUCO method can effectively utilize and coordinate multiple UAVs to achieve better load balancing effect.

[0201] Process or method of use

[0202] (1) MUCO performs adaptive UAV deployment based on the distribution of ground IoT devices.

[0203] (2) In each time slot, each MARL agent collects the task information generated by all IoT devices within the UAV coverage area, including the number of tasks and the computational complexity of the tasks.

[0204] (3) Each MARL agent generates task offloading and forwarding decisions based on the task information generated by all IoT devices, the computing power of the corresponding IoT devices, the transmission power of the corresponding IoT devices, and the remaining power of the UAV.

[0205] (4) UAV allocates computing resources to existing tasks for parallel computing based on an even distribution of computing resources.

[0206] (5) During the computation offloading process, each MARL agent continuously records the state of each time slot, the actions taken, the rewards obtained, and the new state entered, and continuously optimizes its own performance based on the above information.

Claims

1. A multi-UAV deployment and collaborative offloading method for the Internet of Things, characterized by: First, the original joint optimization problem is transformed into a UAV deployment subproblem and a computation offloading subproblem. Then, a UAV deployment scheme based on constrained K-Means clustering is proposed for the UAV deployment subproblem. By introducing UAV coverage constraints into K-Means clustering, the UAV deployment locations are adaptively adjusted to improve the coverage and coverage balance of the computation offloading service in the system. Finally, a multi-UAV collaborative computation offloading strategy based on MARL is proposed for the computation offloading subproblem. Through a centralized training and decentralized execution model, the proposed strategy achieves near-optimal computation offloading and UAV collaboration strategies. The UAV deployment and collaborative offloading system for IoT devices is composed of UAVs and IoT devices. IoT devices request computing offloading services from UAVs through wireless links, which is denoted as the set U = {u1,u2,...,u n }, the horizontal coordinates are represented as {w1,w2,...,w n UAVs provide computing offloading services to IoT devices within their coverage area, which is recorded as the set UAV = {uav1, uav2, ..., uav m }, the horizontal coordinates are expressed as {q1,q2,...,q m };The horizontal height of UAVs is H, then u i(1≤i≤n) With UAV j(1≤j≤m) The distance between them is: Collaborative offloading is required between multiple UAVs. UAVs forward tasks from IoT devices to their collaborative UAVs for execution. The system time slot is represented as t∈{1,2,...,T}. At the beginning of each time slot, the IoT device chooses to execute its computing request locally or offload it to the UAVs connected to it. If the IoT device is not within the coverage of any UAV, its computing request is executed locally. At time slot t, from u i The task is defined as a triple <D i [t],S i [t],c i >, where D i [t] represents the amount of input data for the task, S i [t] represents the computational complexity of the task, c i represents the associated indicator variable; when u i Connect to UAV j When c i =j; when u i When unable to connect to an available UAV, c i = 0; tasks from IoT devices can be executed locally or offloaded to UAVs; u i The offloading decision at time slot t is denoted as α i [t]∈{0,1}; When u i When executing a mission locally or failing to connect to an available UAV, α i [t]=0; when u i When offloading the task to a UAV for execution, α i [t]=1; specifically, it is divided into the following two calculation modes: (1) Local computing mode When the task is executed locally, its latency is: Among them, f i Indicates u i computing power; (2) UAV calculation mode IoT devices access UAVs via OFDMA; when u i The task is offloaded to uav j When the input data is uploaded first; according to Shannon's theorem, u i With UAV j The data transmission rate between is: Among them, B u represents the channel bandwidth between IoT devices and UAVs, P i Represents u i transmission power; Where ρ0 represents the channel gain per unit distance, and N0 represents the noise power spectral density; Therefore, u i Offload tasks to uav j The transmission delay is: Next, UAV j Need to process concurrently j [t] tasks; therefore, from u i The mission in UAV j The computational delay on is: in, f max Indicates the computing power of a UAV; Tasks on high-load UAVs can be forwarded to low-load UAVs to achieve load balancing; specifically, uav j The cooperative offloading decision at time slot t is defined as β ijk [t]∈{0,1}; when uav j Decided to put u i The task is forwarded to uav k When β ijk [t]=1; otherwise, β ijk [t]=0; a task is forwarded at most once, denoted as When the task is from uav j Forward to uav k When cooperative offloading is performed, the data transmission rate between the two is: Among them, B uav represents the channel bandwidth between UAVs, P uav represents the transmission power between UAVs, UAV j With UAV k The distance between them; therefore, u i The mission from UAV j Forward to uav k The transmission delay is: The forwarded task is in uav k The computational delay on is: Execute u i The task delay is: Processing i The total delay of the task is: T i total [t]=(1-α i [t])T i u_c [t]+α i [t](T i u_t [t]+T i uav_total [t]) (10) In summary, the total delay for processing all tasks in time slot t is: When tasks are offloaded to UAVs for processing, three aspects of energy consumption are generated, including the energy consumption of UAVs executing tasks, the energy consumption of forwarding tasks between UAVs, and the energy consumption of UAVs maintaining the hovering state; uav j Execute u i The energy consumption of the task is: Where κ represents the effective capacitance coefficient; uav j will u i Task forwarding to uav k The energy consumption is: The energy consumption of UAVs to maintain the hovering state in each time slot is e; therefore, the total energy consumption of all UAVs in time slot t is: in, Indicates uav j The set of tasks to be performed, Indicates that from uav j The set of tasks forwarded; Due to the limited power of UAVs, it is impossible to provide computing offload services for IoT devices all the time; when the power is lower than the threshold b min When the UAVs are in the state of emergency, they will no longer provide computing offload services and return to the home station for battery recharge; The optimization goal is to minimize the weighted sum of the total task processing delay and the UAV energy consumption; the optimization problem is formally defined as: Among them, λ1 and λ2 represent the weights of delay and energy consumption respectively, b j [t] indicates uav j The remaining power at time slot t, b min represents the minimum amount of power that UAVs must retain; C1 and C2 represent the value range constraints for offloading decisions; C3 represents the number of times a task is forwarded between UAVs; C4 represents the distance constraint between the IoT device and its associated UAV, where R is the coverage radius of the UAVs; C5 represents the value range constraint of the associated indicator variable; constraint C6 indicates that the computing resources used by UAVs to execute tasks cannot exceed their maximum computing resource constraints; C7 represents the minimum amount of power that UAVs need to retain when providing computing offloading services. min Used for returning to the destination and continuing the power supply; For the UAV deployment problem, IoT devices are divided into multiple clusters, and UAVs are deployed at the centroid of each cluster to maximize service coverage. Service coverage and coverage balance are defined as performance indicators for UAV deployment, and a UAV deployment scheme based on constrained K-Means clustering is designed. Specifically, clusters are divided based on UAV coverage, ensuring that the cluster size does not exceed the UAV coverage area. At the same time, the cluster division takes into account the balance of the number of IoT devices within each UAV coverage area. In each iteration, the positions of the UAVs are first randomly initialized, and the cluster set change flag is initialized to check whether the cluster division has reached convergence. Then, the IoT devices within the coverage range of the UAVs are assigned to the corresponding UAV cluster. Specifically, unlike the classic K-Means algorithm, which assigns data points based on the principle of minimum distance to the centroid, the allocation of IoT devices takes into account whether the IoT device is within the coverage radius of the UAV. When an IoT device is covered by multiple UAVs at the same time, the nearest UAV is selected and the change is marked, thereby updating the number of UAVs covered. Then, the current coverage rate and the load level of all UAVs are calculated, which are defined as: When the coverage is close to equilibrium, F1≈F2≈...≈F m ; The mean square error is used to evaluate the coverage balance, which is defined as: in, Next, the coverage rate and coverage balance are normalized to the same numerical range. The mean of each cluster is used as the new centroid and UAVs are deployed. Through iteration, the deployment locations of UAVs will gradually move towards areas with a dense distribution of IoT devices, enabling them to serve more IoT devices. After the iteration, the solution with the smallest weighted sum of coverage rate and coverage balance is selected for UAV deployment. After solving the UAV deployment problem, the computation offloading problem is further separated from P1 and is expressed as follows: A MARL-based multi-UAV collaborative computation offloading strategy is proposed. The proposed multi-UAV collaborative offloading system is regarded as an environment. Multiple agents interact with the environment simultaneously and select corresponding offloading actions. After receiving the reward signal from the environment, each actor network is updated in a distributed manner. The critic network adopts a centralized update method. Accordingly, the state space, action space, and reward function are defined as follows: State space: The state space includes task attributes, computing capabilities of IoT devices, and i , data transmission power P of IoT devices i , the remaining power of UAV b[t]; therefore, at time slot t, the state observed by agent j is expressed as: Among them, 1≤j≤m, and Represents the set consisting of the data volume and computational complexity of all tasks respectively; Action space: The action space contains the offloading decision α[t] of the IoT device and the collaborative offloading decision β[t] of the UAVs. α[t] and β[t] are integrated, and the integrated offloading decision is recorded as η[t]∈{0,1,2,...,m}. Therefore, at time slot t, the action of agent j is expressed as: in, represents the offloading decision made by agent j for all tasks; η i [t]=0 means u i No task is generated or the agent decides to execute the task locally; η i [t]=k, k∈{1,2,...,m} means the agent decides to offload the task to the uav k Execute on; Reward function: The optimization objective of P2 is to minimize the weighted sum of the total task processing delay and the UAV energy consumption; therefore, at time slot t, the reward function is expressed as: r[t]=-(λ1N(T u_tot [t])+λ2N(E uav_tot [t])), (22) Here, λ1 and λ2 represent the weights of delay and energy consumption, respectively; N is a normalization function that normalizes delay and energy consumption to the same numerical range.

2. The method for multi-UAV deployment and collaborative offloading for the Internet of Things according to claim 1 is characterized in that: The original joint optimization problem is further split into two sub-problems: UAV deployment and computation offloading. First, UAV deployment optimization is performed to improve the coverage and coverage balance of UAVs computation offloading services. After UAVs are deployed, computation offloading optimization is performed to reduce the total task processing latency and UAV energy consumption.

3. The method for multi-UAV deployment and collaborative offloading for the Internet of Things according to claim 1 is characterized in that: The value network acts as a critic and uses multi-step temporal difference to estimate the advantage function of the state-action pair; the policy network acts as an actor to update the policy according to the advantage function; the obtained UAVs coordinates q and the associated indicator variable c will be used as the input of Algorithm 2; specifically, the value and policy networks are first initialized; in each round of iteration, each agent obtains the initial state of the environment, initializes the round end mark and the interaction data record table, where the interaction data record table is used to collect all data of the interaction between the agent and the environment in the current round; in each round, each agent selects the appropriate unloading action according to its strategy, and converts η into the unloading decision α of the corresponding IoT device and the collaborative unloading decision β of the UAVs; then, α and β are executed and the reward, state and round end mark are updated; the round ends when the power of all UAVs reaches the minimum threshold, that is: when When , done = True; then, collect the interaction data of each agent and update its state; then, use the interaction data collected in the current round to train the agent; specifically, extract the interaction data and splice the state to centrally train the value network, whose generalized advantage function is defined as: Among them, γ is the discount factor, and a hyperparameter δ is introduced t+l =-V θ (s[t+l])+r[t+l]+γV θ (s[t+l+1]),λ∈[0,1]; when λ=0, it is the advantage obtained by one-step difference; when λ=1, it is the complete average of the advantages obtained by each step difference; by using the advantage function, the variance of the policy gradient estimation can be significantly reduced and the stability of the training process can be improved; Next, the strategy obtained in the previous round is used as the old strategy for this round, and the action probability distribution of each agent is calculated; then, the same batch of data is used to update the network parameters multiple times; unlike other DRL algorithms, the previously sampled data is discarded after updating the network parameters. The proposed method converts it into a training process based on different strategies by setting new and old strategies; specifically, data is collected by interacting with the old strategy and the environment, and then the new strategy is trained using these data; according to the principle of importance sampling, the data sampled by the old strategy can be used multiple times, and gradient ascent can be performed multiple times for policy updates; after each update of the policy network parameters, the action probability distribution of each agent and the loss function of the policy network are calculated; clipping is used to limit the amplitude of each policy update to ensure that the strategy does not change drastically during training. Finally, the loss function of the value network is calculated and the parameters of the policy and value networks are updated using gradient descent.

Citation Information

Patent Citations

  • Model training method and system and electronic equipment

    CN116362327A

  • Safe unloading method for dual-unmanned aerial vehicle edge computing system based on multi-agent reinforcement learning

    CN118139013A