Method for path planning and channel selection for multi-uav assisted internet of things data collection

By employing multi-agent reinforcement learning based on the ITPCS-DC algorithm, the optimization problems of trajectory planning and channel selection in multi-UAV-assisted IoT data collection were solved, achieving the minimization of AoI and the improvement of network utility, while ensuring the freshness and stability of data transmission.

CN118968819BActive Publication Date: 2025-10-21NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410944746.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-15
Publication Date
2025-10-21
Estimated Expiration
2044-07-15

AI Technical Summary

Technical Problem

In scenarios involving multiple drones assisting IoT data collection, existing technologies struggle to effectively optimize drone trajectory planning and channel selection, resulting in ineffective control of information staleness (AoI), which may lead to erroneous control decisions and network inefficiency.

Method used

The ITPCS-DC algorithm is adopted, which transforms the problem into a POMDP problem. Combined with MADRL and SAC algorithms, the UAV objective function is designed to achieve joint optimization of UAV trajectory planning and channel selection, avoid local optima, and utilize multi-agent reinforcement learning optimization strategy to maximize cumulative reward and entropy and reduce AoI.

Benefits of technology

It effectively reduces the AoI of data collection from IoT devices, improves network communication efficiency and drone flight efficiency, avoids local optima traps, and enhances the freshness of data transmission and network stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118968819B_ABST
    Figure CN118968819B_ABST
Patent Text Reader

Abstract

The application discloses a method for path planning and channel selection for multi-unmanned aerial vehicle (UAV) assisted Internet of Things (IoT) data collection, introduces an age of information (AoI) to measure the freshness of information collection, and optimizes the target by jointly optimizing multi-UAV path planning and channel selection to minimize the AoI of IoT data collection.Three-dimensional interference in the system environment comes from UAVs and ground jamming machines, and considering the instability of the environment, a multi-agent reinforcement learning algorithm based on the random strategy characteristics of SAC is proposed, which is named ITPCS-DC, i.e., intelligent joint trajectory planning and channel selection for data collection.The algorithm can enable the agent to avoid falling into local optimization in a multi-dimensional complex interference environment.Simulation results show that ITPCS-DC is superior to other benchmark algorithms in terms of cumulative reward, channel switching cost, average AoI and path length.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of flight control, and in particular relates to a method for track planning and channel selection for multi-UAV-assisted Internet of Things data collection. Background Art

[0002] With the development of 6G networks, the Internet of Things (IoT) has gained widespread adoption, but this has also brought various challenges to service providers and network operators, such as managing the connection of massive IoT devices, ensuring security and privacy, processing and storage of data, and managing energy consumption. Therefore, IoT devices require various technologies to facilitate data processing, collection, and dissemination, particularly in emerging healthcare, smart homes, smart transportation, and security applications such as anomaly detection. Traditional cellular base stations can serve as IoT data collectors, but they are susceptible to interference from terrain, buildings, and other factors, preventing them from providing the stable signals required by IoT applications. Drones, on the other hand, can be flexibly deployed as mobile base stations, enabling real-time adjustments based on the actual needs of IoT applications to improve network efficiency and service quality. Therefore, drones are considered a powerful tool for efficient IoT data collection and transmission. However, in latency-sensitive IoT applications such as forest fire monitoring and railway inspection, outdated information from drones can lead to erroneous control decisions, resulting in serious accidents. To ensure the freshness of received data, we propose Age of Information (AoI) as a new performance metric. Data collection based on AoI can ensure the freshness of information in IoT networks.

[0003] To effectively improve information freshness, spectrum resource allocation and trajectory planning are crucial in multi-UAV data collection scenarios. Due to the coupled nature of trajectory planning and resource allocation, the optimization problem in UAV-assisted wireless communication networks is non-convex. Numerous studies have demonstrated that MADRL can adapt to the complexity and diversity of various non-convex problems, demonstrating significant advantages in adaptability to dynamic environments and real-time resource allocation. When applied to multi-UAV collaborative assisted communication, MADRL is modeled as a sequential Markov decision process (MDP) problem. Through reward design, the difficult-to-optimize problem is transformed into maximizing cumulative rewards. After sufficient training, well-trained deep neural networks (DNNs) are used to make decisions. When MADRL solves the joint optimization problem, the interdependencies and constraints between actions within the complex action space make it difficult to optimize the learning policy and may lead to local optima. Summary of the Invention

[0004] In view of the shortcomings of the existing technology, this paper aims to provide a method for trajectory planning and channel selection for multi-UAV-assisted IoT data collection. To minimize the AoI (AoI) of the system, this method formulates an optimization problem by jointly optimizing the UAV trajectory planning and channel selection. Because this system problem is a complex combinatorial optimization problem and difficult to solve, it is converted into a POMDP problem, and the objective function corresponding to the UAVs is designed. Because the three-dimensional interference in this system comes from UAVs and ground jammers, and considering the instability of this environment, we propose an ITPCS-DC algorithm based on MADRL. This algorithm utilizes the SAC (Surrounding Accumulation) property of maximizing cumulative reward and entropy to prevent the agent from falling into local optima.

[0005] In order to achieve the above technical objectives, the technical solution adopted by the present invention is:

[0006] A method for trajectory planning and channel selection for multi-UAV-assisted IoT data collection includes the following steps:

[0007] Step 1: Set up a UAV-assisted IoT network environment with several UAV base stations, several IoT devices, several channels, and several ground jammers. The positions are randomly assigned in each round, where the number of channels is less than the number of UAV base stations. Each UAV base station can only select one channel to access in each time slot and collect data for the IoT devices in need through trajectory planning. At the same time, avoid collisions and avoid exceeding the set boundary area during flight.

[0008] Step 2: Consider an environment with J ground jammers, where U rotor drones act as mobile base stations to provide data collection services for G IoT devices. In a three-dimensional jamming environment, the drones fly within a predetermined altitude and speed range. The position of drone i at time t is IoT devices can move randomly, and the position of IoT device j at time t is expressed as The position of the ground jammer k at time t is expressed as But it is not mobile. There are C channels in the network, and C<U. x, y, and z are the x-axis coordinates, y-axis coordinates, and z-axis coordinates in the spatial coordinate system respectively. T is the time limit.

[0009] Step 3: Since the navigation trajectories of the UAVs are inconsistent, when the distance between the UAVs is less than the interference threshold distance d UU When the same channel is selected, spectrum conflicts occur between drones. Similarly, when the distance between the drone and the ground jammer is less than the interference threshold distance d UJWhen the same channel is selected, the UAV and the ground jammer will have a spectrum conflict. Because the position of the UAV changes in real time, there will be a complex interference relationship between the UAVs. Therefore, the system is a three-dimensional dynamic interference model.

[0010] Step 4: The channel selected by drone i to collect IoT data is determined based on factors such as environmental conditions, drone location, IoT device location, and interference. Then, the channel selection information is sent to the IoT device through the downlink channel. Finally, IoT device j will send a Q-sized channel according to the channel. j,i The data at (t) is uploaded to UAV i. The length of each moment t is recorded as τ seconds. This also means that the uplink and downlink of the UAV must share the same channel during the data collection process. In order to avoid interference, time division multiple access (TDMA) technology is used to downlink the UAV control information and upload the IoT device data.

[0011] Step 5: Due to the environment, building density and height, and the elevation angle between the IoT device and the drone, the air-to-ground channel is affected by line-of-sight (LoS) and non-line-of-sight (NLoS) links, as well as small-scale multipath fading. Ignoring small-scale fading, the loss of the air-to-ground A2G channel model at time t is expressed as:

[0012]

[0013] Among them, d i,j (t) is the propagation distance between drone i and IoT device j at time t, α UG =3 is the path loss factor of the A2G channel, η NLoS =20dB is the additional attenuation factor of the NLoS link,

[0014] Step 6: The LoS link connectivity probability between drone i and IoT device j at time t is expressed as:

[0015]

[0016] where hd i,j (t)=||(x i (t), y i (t))-(x j (t), y j (t))|| represents the horizontal distance from drone i to IoT device j, and is a constant related to the propagation environment type, and the probability of NLoS link connectivity is expressed as

[0017] Step 7: In this system, the uplink and downlink between drone i and IoT device j share the same channel via TDMA. Therefore, the downlink channel power gain and uplink channel power gain Expressed as:

[0018]

[0019] Step 8: At time t, the signal-to-interference-and-noise ratio (SINR) of the uplink from IoT device j to drone i is:

[0020]

[0021] Among them, p j =0.1W is the transmission power of IoT device j, N0=-120dBm=10 -15 W is the variance of the Gaussian noise at the receiver,

[0022] Step 9: If UAV i and UAV n select the same channel at time t, then I i,n (t) represents the interference of UAV i by UAV n at time t. If UAV i and ground jammer k choose the same channel at time t, I i,k (t) is the interference of UAV i by ground jammer k at time t, which can be expressed as:

[0023]

[0024]

[0025] where p n =1W is the transmission power of the jamming drone n, d i,n (t) represents the distance at which UAV i is interfered with by UAV n at time t, α UU =2 is the path loss factor of the air-to-air A2A channel, p k =1W is the transmission power of ground jammer k, d i,k (t) represents the distance at which UAV i is interfered by ground jammer k at time t,

[0026] Step 10: For transmission quality, set a signal-to-noise ratio threshold k g,u =20dB. If the signal-to-interference-noise ratio is greater than this threshold, the IoT device data transmission is considered successful. Therefore, the constraint on the signal-to-noise ratio at the drone i receiver is given as:

[0027]

[0028] Step 11: Given a bandwidth B, the transmission rate from IoT device j to drone i at time t is expressed as:

[0029]

[0030] Where B = 1MHz is the channel bandwidth,

[0031] Step 12: At the receiver of UAV i, from the interference cancellation perspective, the network weighted interference is expressed as:

[0032]

[0033] Step 13: In a multi-channel UAV communication system, UAV i can reduce interference with other interference sources by selecting different channels. However, frequent channel switching not only leads to a decrease in throughput, but also causes unnecessary energy loss and even communication interruption. Therefore, the communication utility of the network is expressed as:

[0034]

[0035] Where C is the channel hopping cost, f(c i (t), c i (t-1)) represents the channel selection c of drone i at the current moment i (t) and the channel selection c at the previous moment i (t-1) is the same, that is:

[0036]

[0037] Step 14: Therefore, the network communication utility of the entire task is expressed as:

[0038]

[0039] Step 15: In addition to considering the network communication utility, we also consider minimizing the track distance of each UAV in the process of flying to the target point to reduce flight energy consumption. Therefore, the track distance of the entire process is expressed as:

[0040]

[0041] Step 16: At the same time, in order to consider the safety of the drones during flight, set the flight safety distance d between drones. safe , therefore, the risk factor of the whole process is:

[0042]

[0043]

[0044] Step 17: Therefore, the network-wide task utility is expressed as:

[0045]

[0046] Step 18: Use AoI to measure the timeliness of drones collecting IoT device data. At time t, the AoI of drone i collecting data packets from IoT device j is:

[0047]

[0048] in, is the moment when the data packet is generated, (x) + =max{0,x}, when hour, This means that the data of IoT device j has not been collected yet.

[0049] Step 19: For ease of analysis, the AoI of IoT device j is the time required to upload data to drone i. This time is related to the upload rate, in other words, the distance between IoT device i and drone i, the channel status, etc. If the distance between IoT device j and drone i is close and the channel status is good, the upload rate is higher and the time required for data upload is shorter. Conversely, the AoI of IoT device j is larger. Therefore, the AoI of IoT device j uploading a data packet to drone i at time t is expressed as:

[0050]

[0051] Q j,i (t+1)=Q j,i (t)-R j,i (t)*τ (19)

[0052] Among them, Q j,i (t) represents the remaining transmission data amount uploaded by IoT device j to drone i at time t, Q j,i (0) = 10 Mbits, R j,i (t) is the transmission rate from IoT device j to drone i at time t, R j,i (t) and each element within the whole network task utility D and S are closely related, for example: Frequent channel switching in [1] will lead to a decrease in throughput; the smaller the distance between the drone and the IoT device in [D], the greater the throughput; the higher the risk factor in [S], the lower the throughput.

[0053] Step 20: Set the goal to minimize the total AoI of all IoT devices’ uploaded data by jointly optimizing the drone’s trajectory planning and channel selection.

[0054]

[0055]

[0056]

[0057]

[0058]

[0059] Among them, u i,j represents the trajectory of drone i serving IoT device j, c i,j represents the channel selection of drone i serving IoT device j, is the initial position of UAV i serving IoT j, and Equation (20b) indicates that the UAV starts to move from the initial position; V is the flight speed of the UAV, δ t is the time interval, so Equation (20c) indicates that the position state of UAV i serving IoT j at time t+1 depends on the position state at time t and δ t Flight speed V during the time interval; δ d is the safe distance between UAV i serving IoT j and UAV n serving IoT k. Formula (20d) indicates that the distance between any two UAVs at time t must be greater than or equal to the safe distance; c i,j (t) is the channel selection of UAV i serving IoT j at time t. Equation (20e) indicates that the channel selection of UAV i is not 0 at any time. P0 needs to minimize AoI by jointly optimizing trajectory planning and channel selection.

[0060] Step 21: Define the key elements of reinforcement learning in a multi-UAV environment: observation space, action space, and reward function.

[0061] Step 22: Build an ITPCS-DC framework to solve the model by:

[0062] Build and train the MADRL algorithm network model for joint trajectory planning and channel selection for multi-UAV-assisted IoT data collection;

[0063] Step 22-1: Multi-agent reinforcement learning is used to solve the problem modeled as Markov games. The agents are drones. The Markov games of U drones are defined by a tuple (S, A, R, P, γ), where S represents the state of the environment, A represents the set of actions of all agents, R represents the set of rewards obtained by all agents, P represents the state transition probability, and γ represents the reward discount factor. At each moment, the state of the environment is s(t), and each agent can only receive local observations o i (t) = b i (s(t)), and select action a based on local observation i(t)=π i (o i (t)), where b i and π i Represents the observation function and strategy of agent i. After selecting an action, agent i will obtain a reward r(t) = {r1(t), r2(t), ..., r U (t)}, r U (t) is the reward obtained by the U-th agent, and then the environment is transformed according to the state transfer function p i Transition to the next state s(t+1),

[0064] The goal of the drone is to minimize the AoI of IoT device data collection through trajectory planning and channel selection. Therefore, the ITPCS-DC algorithm is combined with the MAAC architecture. At the same time, a random strategy is used based on SAC to maximize the cumulative reward and entropy. The ITPCS-DC algorithm combined with the MAAC architecture contains 5 networks in each agent: an actor network For distributed execution, where Represents the weight of the actor network, the input of the actor network is the local observation o of agent i i , the output is action a i ; Four critic networks are used for centralized training, including the state value estimation V network and the state-action value estimation Q network Among them, s t , a t Respectively represent the observations and actions of all drones at time t, represents the network weight, Represents the Q network weight. The ITPCS-DC algorithm also reduces the oscillation of the training process by setting the experience replay pool and the target network. At each time t, the corresponding experience tuple (o(t), a(t), r(t), o(t+1)) is stored in a size of Experience replay pool In the process, if the experience replay pool is full, the new experience tuple will replace the old experience tuple, and batch sampling from the experience replay pool is used to train the actor and critic networks. The random samples break the correlation between sequence samples and reduce training oscillations. In addition, the V network and the Q network have corresponding target networks that share the same architecture with the online network.

[0065] Step 22-2: Use SAC to maximize the cumulative reward value and entropy to make the strategy as random as possible;

[0066] Step 22-3: Use ITPCS-DC algorithm to update the multi-UAV control network.

[0067] Step 22-4: Repeat steps 22-1 to 22-3, and stop training when the set number of training steps for one round is reached;

[0068] Step 22-5: Select an untrained UAV mission environment from the U-frame UAV mission environment created in step 1 and load it. Repeat steps 22-1 to 22-4 until the training is completed after the set number of rounds are loaded.

[0069] Step 22-6: Use the trained MADRL algorithm model for joint multi-UAV trajectory planning and channel selection to minimize AoI by collecting data from multiple UAVs for IoT devices.

[0070] To optimize the technical solution, further improvements include:

[0071] In step 21, the key elements of reinforcement learning in a multi-UAV environment are defined: observation space, action space, and reward function. The specific method is:

[0072] Step 22-1: Define the observation space

[0073] Step 22-1-1: The drones in this system need to sense the locations of each IoT device to plan their trajectories, minimizing the overall flight distance while avoiding collisions between drones. The downlink throughput provided by the drones to IoT devices is related to the channel status. Since the number of channels in the system is smaller than the number of drones, drones need to sense the channel selections of other drones and ground jammers to avoid interference risks within a certain distance.

[0074] 1) UAV i obtains the location of all UAVs through GPS positioning {u i (t)} i∈U , the location of all IoT devices and the location of the ground jammer {b k (t)} k∈J ,

[0075] 2) UAV i obtains the channel selection of other UAV n through information interaction {c n (t)} n∈U,n≠i And the channel selection of ground jammer k {c k (t)} k∈J ,

[0076] 3) The remaining data volume of the currently served IoT device j observed by drone i {Q j,i (t)} j∈G,i∈U ,

[0077] In summary, the observation space of UAV i at time t is expressed as:

[0078]

[0079] Step 22-2: Define the action space

[0080] Step 22-2-1: The drone provides data collection services for IoT devices in need through trajectory planning and channel selection. The actions in the trajectory planning process are processed by changing the speed in each direction, and the actions in the channel selection process are processed by switching channels.

[0081] 1) Track planning:

[0082] We normalized the physical movement of drone i in the system model at time t in the xy plane and the positive z axis direction, and further obtained the velocity v in the xy plane. xy (t) and the velocity v in the positive z-axis direction z (t), where and are the maximum velocities in the xy plane and the positive z axis, respectively. Therefore, the physical movement action space of drone i is expressed as:

[0083]

[0084] Therefore, the relationship between the position state of drone i and its physical movement action is expressed as:

[0085] (x i (t), y i (t))=(x i (t-1), y i (t-1))+v xy (t-1)*δ(t-1) (23)

[0086] z i (t) = z i (t-1)+v z (t-1)*δ(t-1) (24)

[0087] Where δ(t-1) represents the time interval length of time slot t-1,

[0088] 2) Channel selection

[0089] The channel selection state of drone i for collecting data from IoT device j at time t is represented as c i,j (t) If there is interference in the same channel, it may be necessary to switch channels. Right now

[0090]

[0091] Therefore, the relationship between the channel selection state and the channel switching action is expressed as:

[0092]

[0093] In summary, the action space of drone i is expressed as

[0094]

[0095] Step 22-3: Define the reward function

[0096] In order to minimize the value of the total AOI of collected data by jointly optimizing UAV trajectory planning and channel selection, according to the setting of the objective function, the reward is designed as follows:

[0097] 1) The transmission rate reward from IoT device j to drone i at time t is expressed as:

[0098]

[0099] in Is a positive constant used to adjust the reward portion of the transmission rate.

[0100] 2) Network communication utility U from IoT device j to drone i at time t j,i The reward of (t) is expressed as:

[0101]

[0102] in Is a positive constant used to adjust the reward portion obtained from network communication utility.

[0103] 3) The flight distance D of drone i to the IoT device j in need at time t i,j The reward of (t) is expressed as:

[0104]

[0105] in Is a positive constant used to adjust the reward portion obtained by the track distance.

[0106] 4) The danger S that the flight distance between any two drones i and n at time t is less than the safe distance i,n The reward of (t) is expressed as:

[0107]

[0108] in is a positive constant that adjusts the reward component of the risk,

[0109] Therefore, the total reward obtained by drone i at any time t is expressed as:

[0110]

[0111] In step 22-2, SAC is used to maximize the cumulative reward value and entropy, making the strategy as random as possible. The specific method is:

[0112] Step 22-2-1: Since the purpose of SAC is to maximize the cumulative reward value and entropy and make the strategy as random as possible, the objective function of SAC is:

[0113]

[0114] Where T is the total number of time steps, is expectation, r(o t , a t ) is the agent’s observation at time t as o t Make an action t The environmental reward obtained under π is the strategy π (o t , a t ), H(·) represents the entropy value, β is a hyperparameter used to control the randomness of the optimal strategy and weigh the importance of entropy for rewards, π(·|o t )Observe as o t Agent strategy.

[0115] Step 22-2-2: Therefore, the SAC formula of the optimal strategy is defined as:

[0116]

[0117] where γ t is the reward discount factor, H(π(·|o t ))=E[-logπ(·|o t )] is the observed state o t The entropy of the policy distribution under

[0118] Step 22-2-3: The Q value of SAC is calculated based on the entropy-modified Bellman equation. The Q value function Q(s t , a t ) is defined as:

[0119]

[0120] Among them, s t+1 From the experience pool The sampling is obtained, r(st , a t ) is the total state s at time t t -Action a t The global rewards obtained.

[0121] Step 22-2-4: Therefore, the state value function V(s t ) is defined as follows:

[0122]

[0123] In step 22-3, the specific method of using the ITPCS-DC algorithm to update the multi-UAV control network is as follows:

[0124] The control network of each drone consists of five networks: actor network, Q network, target Q network, V network and target V network.

[0125] Step 22-3-1: According to the SAC algorithm, the V function and Q function of agent i are related, and a separate network can be used to estimate the V function for stable training. Therefore, the loss of the V network is expressed as:

[0126]

[0127] in, is the local observation state of agent i at time t Make the move strategy.

[0128] Step 22-3-2: Introducing the target network To improve the stability of Q network training, the Q network of agent i is updated by minimizing the mean squared error loss:

[0129]

[0130] in

[0131]

[0132] in, Denotes the state value estimate obtained by the target V network.

[0133] Step 22-3-3: The parameters of the V network and Q network are updated using stochastic gradient V as follows:

[0134]

[0135]

[0136] Step 22-3-4: For the optimization of the actor network, the parameters of the policy function are updated by minimizing the loss function of Kullback-Leibler divergence, which is described as:

[0137]

[0138] in, Calculate the difference between distributions a and b, Denotes that agent i observes The target policy value at time .

[0139] Step 22-3-5: Reparameterize the policy function by using neural network transformation

[0140]

[0141] Among them, ∈ t is the standard normal distribution from agent i The sampled input noise vector,

[0142] Step 22-3-6: Then rewrite formula (42) as:

[0143]

[0144] Step 23-3-7: Approximate the gradient of the actor network parameters as follows:

[0145]

[0146] Step 22-3-8: Finally, the parameters of the target V network and the target Q network are updated through soft update.

[0147]

[0148]

[0149] Where τ<<1, represents the parameters of the target V network of agent i, represents the parameters of the target Q network of agent i.

[0150] The beneficial effects of the present invention are:

[0151] The present invention discloses a MADRL algorithm (ITPCS-DC) for joint optimization of trajectory planning and channel selection for multi-UAV assisted IoT data collection. The system introduces AoI to measure the freshness of information collection, and the optimization goal is to minimize the AoI of IoT data collection by jointly optimizing multi-UAV trajectory planning and channel selection. The three-dimensional interference in the system environment comes from UAVs and ground jammers. Taking into account the instability of the environment, an algorithm based on multi-agent reinforcement learning is proposed using the random strategy characteristics of SAC, named ITPCS-DC (Intelligent Joint Trajectory Planning and Channel Selection for Data Collection). The algorithm can enable the agent to avoid falling into local optimality in a multi-dimensional complex interference environment. Simulation results show that ITPCS-DC outperforms other benchmark algorithms in terms of cumulative reward, channel switching cost, average AoI and trajectory length. BRIEF DESCRIPTION OF THE DRAWINGS

[0152] Figure 1 Schematic diagram of the implementation steps for jointly optimizing drone trajectory planning and channel selection in a multi-drone-assisted IoT data collection system.

[0153] Figure 2 This is a schematic diagram of a scenario in which drone trajectory planning and channel selection are jointly optimized in a multi-drone-assisted IoT network data collection system.

[0154] Figure 3 It is a schematic diagram of air-to-ground channel propagation in a dense urban environment according to the present invention.

[0155] Figure 4 Schematic diagram of the safe distance between drones of the present invention.

[0156] Figure 5 It is a structural diagram of the present invention based on the ITPCS-DC algorithm.

[0157] Figure 6 It is a flowchart for constructing a network model for joint optimization of drone trajectory planning and channel selection in a multi-drone assisted IoT network data collection system.

[0158] Figure 7 It is the cumulative reward convergence diagram of the trajectory planning and channel selection test results of multiple UAVs.

[0159] Figure 8 It is the average track length convergence diagram of the track planning and channel selection test results of multiple UAVs.

[0160] Figure 9 This is the average AoI convergence diagram of the trajectory planning and channel selection test results of multiple UAVs. DETAILED DESCRIPTION

[0161] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is described and illustrated below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely used to explain this application and are not intended to limit this application. Based on the embodiments provided in this application, all other embodiments obtained by those of ordinary skill in the art without making any creative efforts are within the scope of protection of this application.

[0162] like Figure 1 As shown, the present invention proposes an ITPCS-DC algorithm for trajectory planning and channel selection for multi-UAV assisted IoT data collection, which includes the following steps:

[0163] Step 1: Based on the multi-UAV collaborative communication environment, establish a three-dimensional interference environment for multiple UAVs to collect data for IoT devices;

[0164] like Figure 2 As shown in the figure, a drone-assisted IoT network environment consists of several drone base stations, several IoT devices, several channels, and several ground jammers. Positions are randomly assigned in each round. The number of channels is smaller than the number of drone base stations. Each drone base station can only select one channel to access during each time slot and collect data for the IoT devices in need through trajectory planning. Furthermore, during flight, collisions and exceeding the specified boundary area are avoided.

[0165] Step 2: Describe the optimization problem as a combinatorial optimization problem, transform it into a POMDP problem, and design the objective function corresponding to the UAV. To solve the optimization problem of this system, we propose an ITPCS-DC algorithm to solve it.

[0166] Step 3: Consider an environment with J ground jammers, where U rotorcraft drones act as mobile base stations to provide data collection services for G IoT devices. In a three-dimensional jamming environment, drones can fly within a certain range of altitude and speed. The position of drone i at time t is IoT devices can move randomly, and the position of IoT device j at time t can be expressed as The position of the ground jammer k at time t can be expressed as But it does not have mobility. There are C channels in the network, and C<U.

[0167] Step 4: Since the navigation trajectories of the UAVs are inconsistent, when the distance between the UAVs is less than the interference threshold distance d UU When the same channel is selected, spectrum conflicts will occur between drones. Similarly, when the distance between the drone and the ground jammer is less than the interference threshold distance d UJWhen the same channel is selected, the UAV and the ground jammer will have a spectrum conflict. Because the position of the UAV changes in real time, there will be complex interference relationships between the UAVs. Therefore, the system is a three-dimensional dynamic interference model.

[0168] Step 5: The channel selected by drone i to collect IoT data is determined based on factors such as the environmental conditions, drone location, IoT device location, and interference. Then, the channel selection information is sent to the IoT device through the downlink channel. Finally, IoT device j will send a Q-sized channel to the IoT device based on the channel. j,i The data at (t) is uploaded to UAV i, with the duration of each moment t being τ seconds. This also means that the uplink and downlink of the UAV must share the same channel during the data collection process. To avoid interference, we use time division multiple access (TDMA) technology to downlink UAV control information and upload IoT device data.

[0169] Step 6: Due to the environment, building density and height, and the elevation angle between the IoT device and the drone, the air-to-ground channel is affected by line-of-sight (LoS) and non-line-of-sight (NLoS) links and small-scale multipath fading. Usually, we ignore small-scale fading because it is too weak compared to the line-of-sight and non-line-of-sight components. At time t, the loss of the air-to-ground (A2G) channel model can be expressed as:

[0170]

[0171] Among them, d i,j (t) is the propagation distance between drone i and IoT device j at time t, α UG =3 is the path loss factor of the A2G channel, η NLoS =20dB is the additional attenuation factor of the NLoS link.

[0172] Step 7: Figure 3 As shown in Figure 2, since the line-of-sight connection probability depends on the geographical environment, the location of the IoT device and the drone, the LoS link connectivity probability between drone i and IoT device j at time t can be expressed as:

[0173]

[0174] where hd i,j (t)=||(x i (t), y i (t))-(x j (t), y j (t))|| represents the horizontal distance from drone i to IoT device j, and is a constant related to the propagation environment type, and the probability of NLoS link connectivity is expressed as

[0175] Step 8: In this system, the uplink and downlink between drone i and IoT device j share the same channel via TDMA. Therefore, the downlink channel power gain and uplink channel power gain It can be expressed as:

[0176]

[0177] Step 9: At time t, the signal-to-interference-and-noise ratio (SINR) of the uplink from IoT device j to drone i is:

[0178]

[0179] Among them, p j =0.1W is the transmission power of IoT device j. N0 = -120dBm = 10 -15 W is the variance of the Gaussian noise at the receiver.

[0180] Step 10: If UAV i and UAV n select the same channel at time t, then I i,n (t) represents the interference of UAV i by UAV n at time t. If UAV i and ground jammer k select the same channel at time t, I i,k (t) is the interference of UAV i by ground jammer k at time t, which can be expressed as:

[0181]

[0182]

[0183] where p n =1W is the transmission power of the jamming drone n, d i,n (t) represents the distance at which UAV i is interfered with by UAV n at time t, α UU =2 is the path loss factor of the A2A channel. k =1W is the transmission power of ground jammer k, d i,k (t) represents the distance at which UAV i is interfered by ground jammer k at time t.

[0184] Step 11: For transmission quality, we set a signal-to-noise ratio threshold k g,u =20dB. If the signal-to-interference-noise ratio is greater than this threshold, the IoT device data transmission is considered successful. Therefore, the constraint on the signal-to-noise ratio at the drone i receiver is given as:

[0185]

[0186] Step 12: Given a bandwidth B, the transmission rate from IoT device j to drone i at time t can be expressed as:

[0187]

[0188] Where B=1MHz is the channel bandwidth.

[0189] Step 13: At the receiver of UAV i, from the interference cancellation perspective, the network weighted interference can be expressed as:

[0190]

[0191] Step 14: In a multi-channel UAV communication system, UAV i can reduce interference with other interference sources by selecting different channels. However, frequent channel switching not only leads to a decrease in throughput, but also causes unnecessary energy loss and even communication interruption. Therefore, the communication utility of the network can be expressed as:

[0192]

[0193] Where C is the channel hopping cost, f(c i (t), c i (t-1)) represents the channel selection c of drone i at the current moment i (t) and the channel selection c at the previous moment i (t-1) is the same, that is:

[0194]

[0195] Step 15: Therefore, the network communication utility of the entire task can be expressed as:

[0196]

[0197] Step 16: In addition to considering the network communication utility, this paper also considers minimizing the track distance of each UAV during its flight to the target point to reduce flight energy consumption. Therefore, the track distance of the entire process can be expressed as:

[0198]

[0199] Step 17: At the same time, in order to consider the safety of the drones during flight, set the flight safety distance d between drones. safe , therefore, the risk factor of the whole process is:

[0200]

[0201]

[0202] Step 18: Therefore, the utility of the whole network task in this paper can be expressed as:

[0203]

[0204] Step 19: We use AoI to measure the timeliness of the drone’s collection of IoT device data. The AoI of the data packet collected by drone i from IoT device j at time t is:

[0205]

[0206] in, is the moment when the data packet is generated, (x) + =max{0,x}. When hour, This means that the data of IoT device j has not been collected yet. It is obvious that the AoI of a data packet increases with time.

[0207] Step 20: For ease of analysis, the AoI of IoT device j is the time required to upload data to drone i. This time is related to the upload rate, in other words, the distance between IoT device i and drone i, the channel conditions, and other factors. If the distance between IoT device j and drone i is close and the channel conditions are good, the upload rate is higher and the time required for data upload is shorter. Conversely, the AoI of IoT device j is greater. Therefore, mathematically, the AoI of IoT device j uploading a data packet to drone i at time t can be expressed as:

[0208]

[0209] Q j,i (t+1)=Q j,i (t)-R j,i (t)*τ (19)

[0210] Among them, Q j,i (t) can be expressed as the remaining transmission data amount uploaded by IoT device j to drone i at time t, Q j,i (0) = 10Mbits. R j,i (t) and each element within the whole network task utility D and S are closely related, for example: Frequent channel switching in [1] will lead to a decrease in throughput; the smaller the distance between the drone and the IoT device in [1], the greater the throughput; the higher the risk factor in [1], the lower the throughput.

[0211] Step 21: The goal in this paper is to minimize the total AoI of all IoT devices’ uploaded data by jointly optimizing the UAV’s trajectory planning and channel selection.

[0212]

[0213]

[0214]

[0215]

[0216]

[0217] Among them, u i,j represents the trajectory of drone i serving IoT device j, c i,j represents the channel selection of drone i serving IoT device j, is the initial position of UAV i serving IoT j, and Equation (20b) indicates that the UAV starts to move from the initial position; V is the flight speed of the UAV, δ t is the time interval, so Equation (20c) indicates that the position state of UAV i serving IoT j at time t+1 depends on the position state at time t and δ t Flight speed V during the time interval; δ d is the safe distance between UAV i serving IoT j and UAV n serving IoT k. Formula (20d) indicates that the distance between any two UAVs at time t must be greater than or equal to the safe distance; c i,j (t) is the channel selection of drone i serving IoT j at time t. Equation (20e) indicates that the channel selection of drone i is non-zero at any time. P0 requires minimizing the AoI by jointly optimizing trajectory planning and channel selection, a well-known NP-hard problem. P0 is difficult to solve using traditional optimization methods. Fortunately, MADRL can effectively address complex optimization problems and explore efficient and reliable solutions from a large policy space.

[0218] Step 22: Based on the above description, define the key elements of reinforcement learning in a multi-UAV environment: observation space, action space, and reward function.

[0219] Step 22-1: Observe the Space

[0220] Step 22-1-1: The drones in this system need to sense the location of each IoT device to plan their trajectories, minimizing the overall flight distance while avoiding collisions between drones. The downlink throughput provided by the drones to the IoT devices is related to the channel state. Because the number of channels in the system is smaller than the number of drones, drones need to sense the channel selection of other drones and ground jammers to avoid interference risks within a certain distance.

[0221] 1) UAV i obtains the location of all UAVs through GPS positioning {u i (t)} i∈U , the location of all IoT devices and the location of the ground jammer {b k (t)} k∈J .

[0222] 2) UAV i obtains the channel selection of other UAV n through information interaction {c n (t)} n∈U,n≠i And the channel selection of ground jammer k {c k (t)} k∈J .

[0223] 3) The remaining data volume of the currently served IoT device j observed by drone i {Q j,i (t)} j∈G,i∈U .

[0224] In summary, the observation space of UAV i at time t can be expressed as:

[0225]

[0226] Step 22-2: Action Space

[0227] Step 22-2-1: The drone provides data collection services for IoT devices through trajectory planning and channel selection. Movement during trajectory planning is handled by changing speed in each direction, while movement during channel selection is handled by switching channels.

[0228] 1) Track planning:

[0229] We normalized the physical movement of drone i in the system model at time t in the xy plane and the positive z axis direction, and further obtained the velocity v in the xy plane. xy (t) and the velocity v in the positive z-axis direction z (t). and are the maximum velocities in the xy plane and the positive z axis, respectively. Therefore, the physical movement action space of drone i can be expressed as:

[0230]

[0231] Therefore, the relationship between the position state of drone i and its physical movement action can be expressed as:

[0232] (x i (t), y i (t))=(x i(t-1), y i (t-1))+v xy (t-1)*δ(t-1) (23)

[0233] z i (t) = z i (t-1)+v z (t-1)*δ(t-1) (24)

[0234] Wherein δ(t-1) represents the time interval length of time slot t-1.

[0235] 2) Channel selection

[0236] The channel selection state of drone i for collecting data from IoT device j at time t is represented as c i,j (t) If there is interference in the same channel, it may be necessary to switch channels. Right now

[0237]

[0238] Therefore, the relationship between the channel selection state and the channel switching action can be expressed as:

[0239]

[0240] In summary, the action space of drone i can be expressed as

[0241]

[0242] Step 22-3: Reward Function

[0243] In order to minimize the total AOI value of collected data by jointly optimizing UAV trajectory planning and channel selection, according to the setting of the objective function, the reward design is as follows:

[0244] 1. The transmission rate reward from IoT device j to drone i at time t is expressed as:

[0245]

[0246] in Is a positive constant used to adjust the reward portion of the transmission rate.

[0247] 2. Network communication utility U from IoT device j to drone i at time t j,i The reward of (t) is expressed as:

[0248]

[0249] in Is a positive constant used to adjust the reward portion obtained from network communication utility.

[0250] 3. The flight distance D of drone i to the IoT device j in need at time t i,j The reward of (t) is expressed as:

[0251]

[0252] in Is a positive constant used to adjust the reward portion obtained by track distance.

[0253] 4. If Figure 4 As shown in the figure, a safe distance is established between drones. At time t, the flight distance between any two drones i and n is less than the danger distance S i,n The reward of (t) is expressed as:

[0254]

[0255] in is a positive constant that adjusts the reward component of the risk.

[0256] Therefore, the total reward obtained by drone i at any time t can be expressed as:

[0257]

[0258] Step 23: Based on the description that the optimization problem in this scenario is a combinatorial optimization problem, an ITPCS-DC framework is established to solve the model.

[0259] like Figure 5 and Figure 6 As shown, the MADRL algorithm network model for joint trajectory planning and channel selection for multi-UAV assisted IoT data collection is constructed and trained;

[0260] Step 23-1: Multi-agent reinforcement learning can be used to solve problems modeled as Markov games. In this system, the agents are drones, and the Markov games of U drones can be defined by a tuple (S, A, R, P, γ), where S represents the state of the environment, A represents the set of actions of all agents, R represents the set of rewards received by all agents, P represents the state transition probability, and γ represents the reward discount factor. At each moment, the state of the environment is s(t), and each agent can only receive local observations o i (t) = b i (s(t)), and select action a based on local observation i (t)=π i (o i (t)), where bi and π i Represents the observation function and strategy of agent i. After selecting an action, agent i will receive a reward r(t) = {r1(t), r2(t), ..., r U (t)}, and then the environment is transformed according to the state transfer function p i Transition to the next state s(t+1).

[0261] In this system, the goal of the drone is to minimize the AoI of IoT device data collection through trajectory planning and channel selection. Due to the inherent non-stationarity of the multi-agent environment and the three-dimensional interference added by the environment, these may cause the agent to fall into a local optimal solution. We propose the ITPCS-DC algorithm, which combines the MAAC architecture and uses a random strategy based on SAC to maximize the cumulative reward and entropy. Each agent in the algorithm contains five networks: an actor network For distributed execution, where Represents the weight of the actor network. The input of the actor network is the local observation o of agent i i , the output is action a i ; Four critic networks are used for centralized training, including the state value estimation V network and the state-action value estimation Q network Among them, s t , a t Respectively represent the observations and actions of all drones at time t, represents the network weight, Represents the Q network weight. The ITPCS-DC algorithm also reduces the oscillation of the training process by setting the experience replay pool and the target network. At each time t, the corresponding experience tuple (o(t), a(t), r(t), o(t+1)) is stored in a size of Experience replay pool In

[15] , if the experience replay pool is full, new experience tuples replace old ones. The actor and critic networks are trained by batch sampling from the experience replay pool. The random samples break the correlation between sequential samples and reduce training oscillations. Furthermore, both the V and Q networks have corresponding target networks that share the same architecture as the online network.

[0262] Step 23-2: The role of SAC

[0263] Step 23-2-1: SAC is an algorithm for continuous action space reinforcement learning and is one of the most advanced algorithms in the field of deep reinforcement learning. By adding entropy to the objective function, this algorithm greatly improves the algorithm's exploration ability and robustness. As can be seen, the goal of SAC is to maximize the cumulative reward value and entropy, making the strategy as random as possible. Therefore, the objective function of SAC is:

[0264]

[0265] Where T is the total number of time steps, is expectation, r(o t , a t ) is the agent’s observation at time t as o t Make an action t The environmental reward obtained under π is the strategy π (o t , a t ), H(·) represents the entropy value, β is a hyperparameter used to control the randomness of the optimal strategy and weigh the importance of entropy for rewards, π(·|o t )Observe as o t Agent strategy.

[0266] Step 23-2-2: Therefore, the SAC formula of the optimal strategy is defined as:

[0267]

[0268] where γ t is the reward discount factor,

[0269] H(π(·|o t ))=E[-logπ(·|o t )] is the observed state o t The entropy of the policy distribution under .

[0270] Step 23-2-3: The Q value of SAC can be calculated based on the entropy-modified Bellman equation. The Q value function Q(s t , a t ) is defined as:

[0271]

[0272] Among them, s t+1 From the experience pool The sampling is obtained, r(s t , a t ) is the total state s at time t t -Action a t The global rewards obtained.

[0273] Step 23-2-4: Therefore, the state value function V(s t ) is defined as follows:

[0274]

[0275] Step 23-3: Use the ITPCS-DC algorithm to update the multi-UAV control network.

[0276] The control network of each drone consists of five networks: actor network, Q network, target Q network, V network and target V network.

[0277] Step 23-3-1: According to the SAC algorithm, the V function and Q function of agent i are related. Using a separate network to estimate the V function can stabilize training. Therefore, the loss of the V network can be expressed as:

[0278]

[0279] in, is the local observation state of agent i at time t Make the move strategy.

[0280] Step 23-3-2: Like DQN, we introduce the target network To improve the stability of Q network training. Therefore, the Q network of agent i can be updated by minimizing the mean squared error (MSE) loss:

[0281]

[0282]

[0283] in, Denotes the state value estimate obtained by the target V network.

[0284] Step 23-3-3: The parameters of the V network and Q network can be adjusted using stochastic gradient The updates are:

[0285]

[0286]

[0287] Step 23-3-4: For the optimization of the actor network, the parameters of the policy function can be updated by minimizing the loss function of Kullback-Leibler (KL) divergence, which can be described as:

[0288]

[0289] in, Calculate the difference between distributions a and b, Denotes that agent i observes The target policy value at time .

[0290] Step 23-3-5: Reparameterize the policy function by using neural network transformation

[0291]

[0292] Among them, ∈ t is the standard normal distribution from agent i The sampled input noise vector,

[0293] Step 23-3-6: Then rewrite formula (42) as:

[0294]

[0295] Step 23-3-7: Approximate the gradient of the actor network parameters as follows:

[0296]

[0297] Step 23-3-8: Finally, update the parameters of the target V network and the target Q network through soft update.

[0298]

[0299]

[0300] Where τ<<1, represents the parameters of the target V network of agent i, represents the parameters of the target Q network of agent i.

[0301] Step 23-4: Repeat steps 23-1 to 23-3, and stop training when the set number of training steps for one round is reached;

[0302] Step 23-5: Select an untrained UAV mission environment from the U-frame UAV mission environment created in step 1 and load it. Repeat steps 23-1 to 23-4 until the training is completed after the set number of rounds are loaded.

[0303] Step 23-6: Use the trained MADRL algorithm model for joint multi-UAV trajectory planning and channel selection to minimize AoI by collecting data from multiple UAVs for IoT devices.

[0304] Example:

[0305] In this example, the initial positions of all drones, IoT devices, and jammers in each episode are randomly distributed within a 3km*3km service area, with the origin of the coordinate system at the center of the area. Therefore, the two-dimensional horizontal and vertical coordinates range from (-1500 to +1500). The upper limit of the flight altitude of all drones is The lower limit is The height of IoT devices is fixed. The maximum flight speed of the drone is set to The maximum acceleration is The safe distance between drones is 5m, and the interference threshold distance is d UU =d UJ =800m. The total frequency bandwidth of each UAV in this system is B uav =0.5MHz, the transmission power of the UAV and ground jammer is p n =p k =1W, the transmission power of IoT devices is p j =0.1W, noise power spectrum density n0 = -120dBm. In addition, the system considers a dense urban environment, and the corresponding channel-related parameter η a =12.08,η b =0.11. The path loss factors of A2G channel and A2A channel are α UG =3 and α UU = 2. The additional attenuation factor of the NLoS link is η NLoS =20dB. Signal-to-noise ratio threshold k g,u =20dB. The channel hopping cost is C=1.5*10 -6 The initial transmission data volume of IoT device j is Q j,i (0) = 10Mbits.

[0306] Hyperparameters: In the experimental simulation, the total number of episodes is set to 50,000, the step size of each round is 100, and the time slot of each step is δ t = 0.5s, mini-batch size = 1024. The learning rate of the actor network is 5*10 -4 , and the learning rate of V network and Q network is 5*10 -3 The target network is updated with τ = 0.01, discount factor γ = 0.999, and the size of the experience buffer is 2*10 5 Once the data in the experience pool exceeds the maximum value, the original experience data will be lost.

[0307] Network architecture: Each drone in this system has the network structure of the proposed algorithm ITPCS-DC. It is worth noting that the actor network, V network architecture and Q network have the same hidden layer with 100, 150, 100 and 50 neurons respectively. The actor network of drone i inputs the drone’s own observation Output the velocity in the xy plane and the positive z axis and channel switching actions The activation function of the Actor network output layer is the softmax function, and the activation functions of other layers are all ReLU functions; the V network of drone m inputs the observations {s(t)} of all drones t∈T , the output is the corresponding state value estimate {V i (t)} i∈U,t∈T The Q network of drone i inputs the observations and actions {s(t), a(t)} of all drones t∈T , outputting a state-action value estimate Q-value to evaluate the learning policy of drone i. The activation function of all neural network layers in the V and Q networks is the Reinforced Luminance (ReLU) function. The number of neurons in the input layer of all networks is related to the number of drones, IoT devices, and jammers in the scenario and can be flexibly configured as needed.

[0308] The present invention randomly initializes the positions of three UAVs, three IoT devices, and one ground jammer in a dynamic environment model in three-dimensional space. In this three-dimensional interference environment, the UAVs minimize the AoI for IoT data collection by jointly optimizing trajectory planning and channel selection.

[0309] like Figure 7 The figure shows the convergence of the cumulative rewards obtained by the total drones of the ITPCS-DC algorithm for jointly optimizing drone trajectory planning and channel selection in a multi-drone assisted IoT network data collection system and other benchmark algorithms. It can be concluded that the ITPCS-DC algorithm can obtain more cumulative rewards than other benchmark algorithms, and the generated strategy is more stable and converges the fastest.

[0310] like Figure 8 The figure shows the convergence of the average track length obtained by the total drones of the ITPCS-DC algorithm for jointly optimizing drone track planning and channel selection in the multi-drone assisted IoT network data collection system and other benchmark algorithms. It can be concluded that the ITPCS-DC algorithm has a shorter average track length than other benchmark algorithms and can more effectively plan effective tracks to further optimize the learning strategy.

[0311] like Figure 9The figure shows a comparison of the average AoI obtained by the ITPCS-DC algorithm for joint optimization of drone trajectory planning and channel selection in a multi-drone-assisted IoT network data collection system and other benchmark algorithms under different drone numbers. It can be concluded that the ITPCS-DC algorithm achieves the smallest average AoI under any number of drones compared to other benchmark algorithms, reflecting that the ITPCS-DC algorithm has higher information freshness in drone-assisted IoT data collection.

[0312] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions based on the principles of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should be considered within the scope of protection of the present invention.

Claims

1. A method for trajectory planning and channel selection for multi-UAV-assisted IoT data collection, characterized by: The following steps are involved: Step 1: Set up a UAV-assisted IoT network environment with several UAV base stations, several IoT devices, several channels, and several ground jammers. The positions are randomly assigned in each round, where the number of channels is less than the number of UAV base stations. Each UAV base station can only select one channel to access in each time slot and collect data for the IoT devices in need through trajectory planning. At the same time, avoid collisions and avoid exceeding the set boundary area during flight. Step 2: Considering the scenario of U rotary-wing UAVs acting as mobile base stations to provide data collection services for G Internet of Things (IoT) devices in an environment with J ground jammers. In the three-dimensional interference environment, the UAVs fly within a predetermined altitude range and speed range. The position of UAV i at time t is 0 ≤ t ≤ T, i ∈ U. The IoT devices can move randomly. The position of IoT device j at time t is denoted as 0 ≤ t ≤ T, j ∈ G. The position of ground jammer k at time t is denoted as 0 ≤ t ≤ T, k ∈ J. However, it does not have mobility. There are C channels in the network, and C < U. x, y, and z are the x-axis coordinate, y-axis coordinate, and z-axis coordinate in the spatial coordinate system respectively, and T is the upper limit of time. Step 3: Since the navigation trajectories of the UAVs are inconsistent, when the distance between the UAVs is less than the interference threshold distance d UU When the same channel is selected, spectrum conflicts occur between drones. Similarly, when the distance between the drone and the ground jammer is less than the interference threshold distance d UJ When the same channel is selected, the UAV and the ground jammer will have a spectrum conflict. Due to the real-time change of the UAV position, a complex interference relationship will appear between the UAVs, thus forming a three-dimensional dynamic interference model. Step 4: The channel selected by drone i to collect IoT data is determined based on the environmental conditions, drone location, IoT device location, and interference factors. Then, the channel selection information is sent to the IoT device through the downlink channel. Finally, IoT device j will select a channel of size Q according to the channel. j,i The data of (t) is uploaded to UAV i. The length of each moment t is recorded as τ seconds. This also means that the uplink and downlink of the UAV must share the same channel during the data collection process. In order to avoid interference, time division multiplexing technology is used to downlink the UAV control information and upload the IoT device data. Step 5: Due to the environment, building density and height, and the elevation angle between the IoT device and the drone, the air-to-ground channel is affected by line-of-sight (LoS) and non-line-of-sight (NLoS) links, as well as small-scale multipath fading. Ignoring small-scale fading, the loss of the air-to-ground A2G channel model at time t is expressed as: Among them, d i,j (t) is the propagation distance between drone i and IoT device j at time t, α UG =3 is the path loss factor of the A2G channel, η NLoS =20dB is the additional attenuation factor of the NLoS link, Step 6: The LoS link connectivity probability between drone i and IoT device j at time t is expressed as: where hd i,j (t)=‖(x i (t),y i (t))-(x j (t),y j (t))‖ represents the horizontal distance from drone i to IoT device j, and is a constant related to the propagation environment type, and the probability of NLoS link connectivity is expressed as Step 7: The uplink and downlink between drone i and IoT device j share the same channel via TDMA. Therefore, the downlink channel power gain and uplink channel power gain Expressed as: Step 8: The signal-to-noise ratio of the uplink from IoT device j to drone i at time t is: Among them, p j =0.1W is the transmission power of IoT device j, N0=-120dBm=10 -15 W is the variance of the Gaussian noise at the receiver, Step 9: If UAV i and UAV n select the same channel at time t, then I i,n (t) represents the interference of UAV i by UAV n at time t. If UAV i and ground jammer k choose the same channel at time t, I i,k (t) is the interference of UAV i by ground jammer k at time t, which can be expressed as: where p n =1W is the transmission power of the jamming drone n, d i,n (t) represents the distance at which UAV i is interfered with by UAV n at time t, α UU =2 is the path loss factor of the air-to-air A2A channel, p k =1W is the transmission power of ground jammer k, d i,k (t) represents the distance at which UAV i is interfered by ground jammer k at time t, Step 10: For transmission quality, set a signal-to-noise ratio threshold k g,u =20dB. If the signal-to-noise ratio is greater than this threshold, the IoT device data transmission is considered successful. Therefore, the constraint on the signal-to-noise ratio at the drone i receiver is given as: Step 11: Given a bandwidth B, the transmission rate from IoT device j to drone i at time t is expressed as: Where B = 1MHz is the channel bandwidth, Step 12: At the receiver of UAV i, from the interference cancellation perspective, the network weighted interference is expressed as: Step 13: In a multi-channel UAV communication system, UAV i can reduce interference with other interference sources by selecting different channels. However, frequent channel switching not only leads to a decrease in throughput, but also causes unnecessary energy loss and even communication interruption. Therefore, the communication utility of the network is expressed as: Where C is the channel hopping cost, f(c i (t),c i (t-1)) represents the channel selection c of drone i at the current moment i (t) and the channel selection c at the previous moment i (t-1) is the same, that is: Step 14: Therefore, the network communication utility of the entire task is expressed as: Step 15: In addition to considering the network communication utility, we also consider minimizing the track distance of each UAV in the process of flying to the target point to reduce flight energy consumption. Therefore, the track distance of the entire process is expressed as: Step 16: At the same time, in order to consider the safety of the drones during flight, set the flight safety distance d between drones. safe , therefore, the risk factor of the whole process is: Step 17: Therefore, the network-wide task utility is expressed as: Step 18: Use AoI to measure the timeliness of drones collecting IoT device data. At time t, the AoI of drone i collecting data packets from IoT device j is: in, is the moment when the data packet is generated, (x) + =max{0,x}, when hour, This means that the data of IoT device j has not been collected yet. Step 19: For ease of analysis, the AoI of IoT device j is the time required to upload data to drone i. This time is related to the upload rate, in other words, the distance between IoT device j and drone i and the channel status. If the distance between IoT device j and drone i is close and the channel status is good, the upload rate is higher and the time required for data upload is shorter. Conversely, the AoI of IoT device j is larger. Therefore, the AoI of IoT device j uploading a data packet to drone i at time t is expressed as: Q j,i (t+1)=Q j,i (t)-R j,i (t)*τ (19) Among them, Q j,i (t) represents the remaining transmission data amount uploaded by IoT device j to drone i at time t, Q j,i (0) = 10 Mbits, R j,i (t) is the transmission rate from IoT device j to drone i at time t, Step 20: Set the goal to minimize the total AoI of all IoT devices’ uploaded data by jointly optimizing the drone’s trajectory planning and channel selection. Among them, u i,j represents the trajectory of drone i serving IoT device j, c i,j represents the channel selection of drone i serving IoT device j, is the initial position of UAV i serving IoT j, and Equation (20b) indicates that the UAV starts to move from the initial position; V is the flight speed of the UAV, δ t is the time interval, so Equation (20c) indicates that the position state of UAV i serving IoT j at time t+1 depends on the position state at time t and δ t Flight speed V during the time interval; δ d is the safe distance between UAV i serving IoT j and UAV n serving IoT k. Formula (20d) indicates that the distance between any two UAVs at time t must be greater than or equal to the safe distance; c i,j (t) is the channel selection of UAV i serving IoT j at time t. Equation (20e) indicates that the channel selection of UAV i is not 0 at any time. P0 needs to minimize AoI by jointly optimizing trajectory planning and channel selection. Step 21: Define the key elements of reinforcement learning in a multi-UAV environment: observation space, action space, and reward function. Step 22: Build an ITPCS-DC framework to solve the model by: Build and train the MADRL algorithm network model for joint trajectory planning and channel selection for multi-UAV-assisted IoT data collection; Step 22-1: Multi-agent reinforcement learning is used to solve the problem modeled as Markov games. The agents are drones. The Markov games of U drones are defined by a tuple (S, A, R, P, γ), where S represents the state of the environment, A represents the set of actions of all agents, R represents the set of rewards obtained by all agents, P represents the state transition probability, and γ represents the reward discount factor. At each moment, the state of the environment is s(t), and each agent can only receive local observations o i (t) = b i (s(t)), and select action a based on local observation i (t)=π i (o i (t)), where b i and π i Represents the observation function and strategy of agent i. After selecting an action, agent i will obtain a reward r(t) = {r1(t), r2(t), …, r U (t)}, r U (t) is the reward obtained by the U-th agent, and then the environment is transformed according to the state transfer function p i Transition to the next state s(t+1), The goal of the drone is to minimize the AoI of IoT device data collection through trajectory planning and channel selection. Therefore, the ITPCS-DC algorithm is combined with the MAAC architecture. At the same time, a random strategy is used based on SAC to maximize the cumulative reward and entropy. The ITPCS-DC algorithm combined with the MAAC architecture contains 5 networks in each agent: an actor network For distributed execution, where Represents the weight of the actor network, the input of the actor network is the local observation o of agent i i , the output is action a i ; Four critic networks are used for centralized training, including the state value estimation V network and the state-action value estimation Q network Among them, s t ,a t Respectively represent the observations and actions of all drones at time t, represents the network weight, Represents the Q network weight. The ITPCS-DC algorithm also reduces the oscillation of the training process by setting the experience replay pool and the target network. At each time t, the corresponding experience tuple (o(t), a(t), r(t), o(t+1)) is stored in a size of Experience replay pool In the process, if the experience replay pool is full, the new experience tuple will replace the old experience tuple, and batch sampling from the experience replay pool is used to train the actor and critic networks. The random samples break the correlation between sequence samples and reduce training oscillations. In addition, the V network and the Q network have corresponding target networks that share the same architecture with the online network. Step 22-2: Use SAC to maximize the cumulative reward value and entropy to make the strategy random; Step 22-3: Use ITPCS-DC algorithm to update the multi-UAV control network. Step 22-4: Repeat steps 22-1 to 22-3, and stop training when the set number of training steps for one round is reached; Step 22-5: Select an untrained UAV mission environment from the U-frame UAV mission environment created in step 1 and load it. Repeat steps 22-1 to 22-4 until the training is completed after the set number of rounds are loaded. Step 22-6: Use the trained MADRL algorithm model for joint multi-UAV trajectory planning and channel selection to minimize AoI by collecting data from multiple UAVs for IoT devices.

2. The method for trajectory planning and channel selection for multi-UAV-assisted IoT data collection according to claim 1, characterized in that: In step 21, the key elements of reinforcement learning in a multi-UAV environment are defined: observation space, action space, and reward function. The specific method is: Step 22-1: Define the observation space Step 22-1-1: The drones in this system need to sense the locations of each IoT device to plan their trajectories, minimizing the overall flight distance while avoiding collisions between drones. The downlink throughput provided by the drones to IoT devices is related to the channel status. Since the number of channels in the system is smaller than the number of drones, drones need to sense the channel selections of other drones and ground jammers to avoid interference risks within a certain distance. 1) UAV i obtains the location of all UAVs through GPS positioning {u i (t)} i∈U , the location of all IoT devices and the location of the ground jammer {b k (t)} k∈J , 2) UAV i obtains the channel selection of other UAV n through information interaction {c n (t)} n∈U,n≠i And the channel selection of ground jammer k {c k (t)} k∈J , 3) The remaining data volume of the currently served IoT device j observed by drone i {Q j,i (t)} j∈G,i∈U , In summary, the observation space of UAV i at time t is expressed as: Step 22-2: Define the action space Step 22-2-1: The drone provides data collection services for IoT devices in need through trajectory planning and channel selection. The actions in the trajectory planning process are processed by changing the speed in each direction, and the actions in the channel selection process are processed by switching channels. 1) Track planning: The physical movement of the drone i in the model at time t in the xy plane and the positive z axis direction is normalized respectively, and the velocity v in the xy plane is further obtained. xy (t) and the velocity v in the positive z-axis direction z (t), where and are the maximum velocities in the xy plane and the positive z axis, respectively. Therefore, the physical movement action space of drone i is expressed as: Therefore, the relationship between the position state of drone i and its physical movement action is expressed as: (x i (t),y i (t))=(x i (t-1),y i (t-1))+v xy (t-1)*δ(t-1) (23) z i (t)=z i (t-1)+v z (t-1)*δ(t-1) (24) Where δ(t-1) represents the time interval length of time slot t-1, 2) Channel selection The channel selection state of drone i for collecting data from IoT device j at time t is represented as c i,j (t) If there is interference in the same channel, it may be necessary to switch channels. Right now Therefore, the relationship between the channel selection state and the channel switching action is expressed as: In summary, the action space of drone i is expressed as Step 22-3: Define the reward function In order to minimize the total AoI value of the collected data by jointly optimizing the UAV trajectory planning and channel selection, according to the setting of the objective function, the reward is designed as follows: 1) The transmission rate R from IoT device j to drone i at time t j,i (t) Reward is expressed as: in Is a positive constant used to adjust the reward portion of the transmission rate. 2) Network communication utility U from IoT device j to drone i at time t j,i The reward of (t) is expressed as: in Is a positive constant used to adjust the reward portion obtained from network communication utility. 3) The flight distance D of drone i to the IoT device j in need at time t i,j The reward of (t) is expressed as: in Is a positive constant used to adjust the reward portion obtained by the track distance. 4) The danger S that the flight distance between any two drones i and n at time t is less than the safe distance i,n The reward of (t) is expressed as: in is a positive constant that adjusts the reward component of the risk, Therefore, the total reward obtained by drone i at any time t is expressed as:

3. The method for trajectory planning and channel selection for multi-UAV assisted IoT data collection according to claim 1, characterized in that: In step 22-2, SAC is used to maximize the cumulative reward value and entropy, and the specific method of making the strategy random is: Step 22-2-1: Since the purpose of SAC is to maximize the cumulative reward value and entropy and make the strategy random, the objective function of SAC is: Where T is the total number of time steps, is expectation, r(o t ,a t ) is the agent’s observation at time t as o t Make an action t The environmental reward obtained under π is the strategy π (o t ,a t ), H(·) represents the entropy value, β is a hyperparameter used to control the randomness of the optimal strategy and weigh the importance of entropy for rewards, π(·|o t )Observe as o t The agent strategy, Step 22-2-2: Therefore, the SAC formula of the optimal strategy is defined as: where γ t is the reward discount factor, H(π(·|o t ))=E[-logπ(·|o t )] is the observed state o t The entropy of the policy distribution under Step 22-2-3: The Q value of SAC is calculated based on the entropy-modified Bellman equation. The Q value function Q(s t ,a t ) is defined as: Among them, s t+1 From the experience pool The sampling is obtained, r(s t ,a t ) is the total state s at time t t -Action a t The global rewards obtained under Step 22-2-4: Therefore, the state value function V(s t ) is defined as follows:

4. The method for trajectory planning and channel selection for multi-UAV-assisted IoT data collection according to claim 3, characterized in that: In step 22-3, the specific method of using the ITPCS-DC algorithm to update the multi-UAV control network is as follows: The control network of each drone consists of five networks: actor network, Q network, target Q network, V network and target V network. Step 22-3-1: According to the SAC algorithm, the V function and Q function of agent i are related, and a separate network can be used to estimate the V function for stable training. Therefore, the loss of the V network is expressed as: in, is the local observation state of agent i at time t Make the move strategy, Step 22-3-2: Introducing the target network To improve the stability of Q network training, the Q network of agent i is updated by minimizing the mean squared error loss: in in, represents the state value estimate obtained by the target V network, Step 22-3-3: Use stochastic gradient to adjust the parameters of V network and Q network The updates are: Step 22-3-4: For the optimization of the actor network, the parameters of the policy function are updated by minimizing the loss function of Kullback-Leibler divergence, which is described as: in, Calculate the difference between distributions a and b, Denotes that agent i observes The target policy value when Step 22-3-5: Reparameterize the policy function by adopting neural network transformation: Among them, ∈ t is the standard normal distribution from agent i The sampled input noise vector, Step 22-3-6: Then rewrite formula (42) as: Step 23-3-7: Approximate the gradient of the actor network parameters as follows: Step 22-3-8: Finally, the parameters of the target V network and the target Q network are updated through soft update. Where τ<<1, represents the parameters of the target V network of agent i, represents the parameters of the target Q network of agent i.

Citation Information

Patent Citations

  • Multi-unmanned aerial vehicle intelligent route planning method for data acquisition

    CN116795138A

  • Internet of Things information age and power optimization method based on unmanned aerial vehicle

    CN117615386A