Unmanned aerial vehicle trajectory optimization and bandwidth allocation method based on hierarchical deep reinforcement learning
By adopting a layered deep reinforcement learning method in the drone-assisted IoT node data collection scenario, optimizing the drone trajectory and bandwidth allocation, the problem of low data collection efficiency is solved, and efficient data collection under energy, speed and bandwidth constraints are achieved.
Patent Information
- Application Number
- CN202510211107.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-10
AI Technical Summary
In the drone-assisted IoT node data collection scenario, it is difficult to effectively optimize the drone trajectory and bandwidth allocation, resulting in low data collection efficiency under the constraints of energy, speed and bandwidth.
Using a method based on hierarchical deep reinforcement learning, a semi-Markov decision-making process is constructed, and data collection is optimized by adjusting the location and bandwidth allocation ratio of the drone, and data collection is maximized from the environmental terrestrial IoT nodes.
Under the constraints of energy, speed and bandwidth, the efficiency of drones in IoT node data collection is significantly improved, achieving faster convergence speed and less network computing.
Smart Images

Figure CN120128895A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of unmanned aerial vehicle (UAV) assisted communication, and particularly relates to a method for UAV trajectory optimization and bandwidth allocation based on hierarchical deep reinforcement learning. Background Art
[0002] With the development of communication technologies, UAVs have been increasingly widely used in the field of wireless communication. UAVs have the characteristics of high mobility, easy deployment, and low cost. Therefore, using a UAV assisted communication system can efficiently help complete tasks such as farmland mapping, data collection, traffic patrol, and emergency communication. By using IoT nodes to monitor and collect environmental data, and then using the UAV as an aerial mobile base station to collect the data sensed by these IoT nodes, the human consumption can be greatly reduced and the work efficiency can be improved. At the same time, compared with traditional ground communication systems, using UAVs can establish a line-of-sight (LoS) channel with a higher probability and approach distributed IoT nodes, providing higher quality communication while effectively reducing the transmission energy consumption of IoT nodes. In addition, since many IoT nodes are located in areas with relatively high Internet access costs, such as uninhabited areas and mountains, using UAVs can also effectively reduce costs. Of course, when UAVs are used for data collection, they also face many challenges such as tight power supply, obstacles, and unknown interferers.
[0003] Currently, many explorations have been carried out in the related field of UAV assisted IoT node data collection. Since most of the optimization model problems constructed in this field scenario are highly non-convex problems, the solution is often very difficult. Traditional methods include using convex optimization methods to transform and solve the problems, but there are many limitations, such as a significant increase in the complexity of the transformed problems, difficulty in balancing multi-objective requirements, and poor adaptability to dynamic environments, which easily lead to suboptimal solutions being introduced into the problems. Of course, some heuristic algorithms are also used to solve the problems, such as ant colony algorithms and genetic algorithms. However, these algorithms have a high computational complexity for high-dimensional and multi-constraint problems, and in a dynamic environment, the real-time adjustment ability of the algorithms is insufficient, making it difficult to cope with environmental changes and easily falling into local optimal solutions.
[0004] In recent years, people have gradually studied the use of deep reinforcement learning algorithms to assist in UAV data collection. Different from traditional methods, deep reinforcement learning usually models the problem as a Markov decision process (MDP). Among them, the UAV is used as the agent, the IoT node state, the UAV state, etc. constitute the environment, and the objective function is transformed into the reward obtained by the agent from the environment after taking actions. The entire algorithm aims to maximize the cumulative reward, enabling the agent to continuously iteratively learn in the interaction with the environment. Therefore, without focusing on the mathematical model of the highly non-convex optimization problem of scenario construction, effective solutions can be obtained. The existing related work mainly focuses on aspects such as UAV trajectory optimization, UAV trajectory and bandwidth allocation optimization, UAV and node-related energy consumption optimization, UAV trajectory and IoT node data collection information age (AoI) optimization, etc. However, if the deep reinforcement learning method is applied to reality, then the excessively high training data requirements, high-dimensional action state space, or sparse rewards will pose severe challenges. How to efficiently process tasks, accelerate the convergence speed of the network, and reduce the training parameters of the network to adapt to low-computing-power UAVs is very important. Based on the above problems, the present invention designs a TBH-DDPG algorithm for the scenario of UAV-assisted IoT node data collection, and by optimizing the UAV trajectory and bandwidth allocation ratio, under the constraints of energy, speed, and bandwidth, realizes the maximization of data collection from ground IoT nodes in the environment. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a UAV trajectory optimization and bandwidth allocation method based on hierarchical deep reinforcement learning for the deficiencies of the above-mentioned existing technologies. Using hierarchical deep reinforcement learning technology, for the scenario of UAV-assisted IoT node data collection, by optimizing the UAV trajectory and bandwidth allocation ratio, under the constraints of energy, speed, and bandwidth, realizes the maximization of data collection from ground IoT nodes in the environment.
[0006] To achieve the above technical objectives, the technical solutions adopted by the present invention are as follows:
[0007] A UAV trajectory optimization and bandwidth allocation method based on hierarchical deep reinforcement learning, including the steps of:
[0008] (1) Construct a UAV-assisted IoT node data collection system model;
[0009] (2) Based on the constructed system model, establish an optimization problem of maximizing the data collection amount by adjusting the position and bandwidth allocation ratio of the UAV;
[0010] (3) Model the optimization problem as a semi-Markov decision process, and use a hierarchical deep reinforcement learning method to train an agent that can autonomously adjust the position and bandwidth allocation ratio of the UAV.
[0011] To optimize the above technical solutions, the specific measures also include:
[0012] The UAV-assisted IoT node data collection system model described in step (1) includes a UAV-assisted IoT node data collection scenario model, a LoS / NLoS communication channel model, and a UAV energy consumption model;
[0013] The UAV-assisted IoT node data collection scenario model considers a rasterized square area of size Y*Y, which contains three types of obstacle areas: areas that only affect the communication area, areas that only affect the flight area, and areas that affect both the flight and communication areas; there are a total of I ground IoT nodes in the scenario, and the position of the i-th IoT node is represented by p i ={x i ,y i ,0}, where i ∈ {1, 2,..., I}; the scenario includes J ground jammers, and they are unknown to the UAV. The position of the j-th jammer is represented by p j ={x j ,y j ,0}, where j ∈ {1, 2,..., J}; the UAV takes off from the takeoff area and flies at a fixed height H. The time for the entire mission is represented by T; the mission time T is equally divided into N flight periods, and the n-th flight period is represented by δ n , where n ∈ {1, 2,..., N}; the real-time position of the UAV can be represented by p u ={x u (δ n ),y u (δ n ),H}; assume that the speed of the UAV within each flight period is 0 or v, i.e., v u ∈ {0, v}; the single flight period is equally divided into M communication time slots, and the m-th communication time slot in the n-th flight period is denoted as δ n,m , where m ∈ {1, 2,..., M}; when performing bandwidth allocation communication within a flight period, the position of the UAV is simplified, and the position during bandwidth allocation communication is regarded as the position of the UAV in the next flight period; the data volume of IoT nodes in the scenario will increase in real time. Assume that the data growth of each IoT node is the same. After each communication time slot δ n,m , the IoT node will increase a fixed data volume d inc ; of course, the data volume stored in each IoT node has an upper limit, denoted as D max ;
[0014] The LoS / NLoS communication channel model considers the shadow effect, log-normal fading, and the influence of jammers; in time slot δ n,m, the channel gain g between the UAV and the i-th IoT node i is:
[0015]
[0016] where d i (δ n,m ) represents the distance between the UAV and IoT node i at time slot δ n,m ; α LoS and α NLoS represent the transmission path losses of the line-of-sight and non-line-of-sight channels respectively; η LoS and η NLoS represent the shadowing effects of the line-of-sight and non-line-of-sight channels, which follow normal distributions and
[0017] A jammer model is designed, assuming that the main lobe direction of the jammer is always perpendicular to the ground; the main lobe gain G m exists within the beamwidth θ m , and the gains in other directions are ignored; the relationship between the isotropic antenna gain G s and the main lobe gain G m is defined as:
[0018]
[0019] where d m is the distance from the UAV to the center line of the main beam, and P is the transmission power; the antenna gain G k (d m,k ) of the k-th jammer is expressed as:
[0020]
[0021] where θ m,k represents the beamwidth of the k-th jammer; Combining the above, the SINR n,m between the UAV and IoT node i at time slot δ i (δ n,m ) is expressed as:
[0022]
[0023] where P i represents the transmission power of IoT node i, S N represents the noise power spectral density, and P k represents the transmission power of the jammer;
[0024] The transmission rate C n,m between the UAV and the i-th IoT node at time slot δ i(δ n,m ) is expressed as:
[0025] C i (δ n,m ) = b i (δ n,m ) log 2 (1 + SINR i (δ n,m ))
[0026] where b i (δ n,m ) represents the bandwidth size allocated to IoT node i in time slot δ n,m ; the sum of the bandwidths allocated to each IoT node in any communication time slot is less than or equal to the total bandwidth size B, that is:
[0027]
[0028] The UAV energy consumption model only considers the hover energy consumption and propulsion energy consumption of the UAV. The UAV propulsion power P UAV is
[0029]
[0030] where P 0 represents the blade section power, u tip represents the UAV rotor tip speed, P 1 represents the induced power, v 0 represents the average rotor induced speed when the UAV is hovering, v u represents the UAV flight speed, d 0 represents the fuselage drag ratio, ρ represents the air density, s 0 represents the rotor robustness, A r represents the rotor disk area; when the UAV flight speed is 0, P UAV (0) = P 0 + P 1 ; assuming the amount of data transmitted in a communication time slot is When the remaining data volume of IoT node i is less than C i (δ n,m )δ n,m , only the remaining data volume of IoT node i can be obtained, which can be expressed as:
[0031]
[0032] where δ n,m is the m-th communication time slot in the n-th flight period, C i (δ n,m)For the transmission rate of the drone and the i-th IoT node in time slot δ n,m , ξ sets a threshold for normal communication.
[0033] The optimization problem of maximizing the data collection volume described in step (2) is as follows:
[0034]
[0035]
[0036] The first constraint is the energy constraint. At any time, the energy of the drone should be greater than 0 to ensure that the drone can reach the designated landing point normally after completing data collection; the second constraint is the speed constraint, that is, the speed of the drone in each time slot is 0 or v; the last one is the bandwidth constraint, that is, the sum of the bandwidths allocated to each IoT node in each communication time slot is always less than or equal to the total bandwidth B.
[0037] The system scenario map of the semi-Markov decision process described in step (3) is composed of 5 maps, namely the no-fly zone map, the communication obstacle zone map, the takeoff and landing map, the IoT node map and the jammer map, specifically represented as [Y, Y, 5]; after obtaining the map and the drone state information each time, the map is centered with the drone as the center, and the centered map is processed in two ways. One way is to perform global pooling on the map, and the other way is to perform central cropping on the map; a convolutional neural network that does not participate in learning is used to extract features from the two ways; the output of the network is flattened and added to the remaining battery information of the drone to obtain the input state of the algorithm.
[0038] Step (3) models the optimization problem as a semi-Markov decision process and consists of <S, A, P r , R>, adding an option space to solve the problem of inconsistent action time slot content; where the state space S = {m r , m g , m k , m i , m j , E u , p u}, where the first 5 elements respectively represent the no-fly zone, the communication obstacle zone, the takeoff and landing zone, the IoT node and the jammer map matrix, E u represents the remaining battery power of the drone, and p u represents the position of the drone; the action space A = {a f , a b}, where a f represents the flight action, and a b represents the bandwidth allocation action; the flight action a f ∈ {an , a e , a s , a w , a h , a l}, where a n represents flying one grid north, a e represents flying one grid east, a s represents flying one grid south, a w represents flying one grid west, a h represents hovering, a l represents landing; the bandwidth allocation action a b = <b 1 , b 2 ,..., b i ,..., b I >, where b i represents the bandwidth size allocated to the i-th IoT node; P r represents the probability of state transition after executing the action; the option space O = {o f}, the execution method of the option is specified by the triple <W, π, β>, where W represents the initial state of the option, π represents the set of policies under the option, and β represents the end condition of the option; define o f ∈ {o n , o e , o s , o w , o h , o l}, where the subscripts of the elements in the set correspond to the subscripts of the flight action a f ; W is defined as the state space S; the policy set is defined as π = <a f , a b1 , a b2 ,..., a bm ,..., a bM >, where a f corresponds to the selected o f , a bM is the bandwidth allocation action in the M-th communication time slot, and the termination condition β is that all actions in π are completed in sequence; R represents the reward obtained after executing an action in a state, which consists of two parts. The lower-layer reward r n,m obtained in the communication time slot δ lo (δ n,m ) is:
[0039] r lo (δ n,m ) = r collection (δ n,m ) + r loss(δ n,m )
[0040] where r collection (δ n,m ) represents the positive data collection reward, and r loss (δ n,m ) represents the negative data loss reward; the data collection volume reward is given as a fixed percentage of the data volume collected each time data is collected. Specifically:
[0041] r collection (δ n,m ) = ò cen ·r cen (δ n,m )
[0042] where ò cen is the proportionality coefficient of the data collection reward, and is the data volume collected within a single communication time slot; D i (δ n,m ) is the remaining data volume of node i in the m-th communication time slot under the n-th flight time slot, and D i (δ n,m+1 ) is the remaining data volume of node i in the (m + 1)-th communication time slot under the n-th flight time slot; the data loss reward is a fixed penalty value given when the data storage volume exceeds the storage capacity limit of the IoT node. Specifically:
[0043]
[0044] where r ls represents the penalty constant value;
[0045] The upper-layer reward r n (δ up ) obtained in the flight time slot δ n includes all communication time slot rewards and rewards obtained by performing flight actions in this flight time slot. Specifically:
[0046] r up (δ n ) = r f (δ n ) + r lo (δ n,1 ) + … + r lo (δ n,m ) + … r lo (δ n,M )
[0047] where r f (δ n ) = r collision (δ n ) + r nland(δ n ) + r return (δ n ), are the negative collision reward, negative return reward, and negative crash reward respectively; the collision reward is a fixed penalty value given when the drone enters the no-fly zone, and the specific expression is:
[0048]
[0049] Among them, r csn represents the penalty constant value, and p u (δ n ) represents the real-time position of the drone in the nth flight period;
[0050] The return reward is a penalty given when the drone's battery level is less than the threshold, and the smaller the battery level and the farther away from the landing point, the greater the penalty. The specific expression is:
[0051]
[0052] Among them, d re (δ n ) represents the distance of the current drone to the takeoff and landing area, ò re represents the return penalty coefficient, E(δ n ) represents the remaining battery level of the current drone, and E tsd represents the return penalty threshold; finally, there is the landing reward, which is a penalty given when the drone's battery runs out and it has not landed yet. The specific expression is:
[0053]
[0054] Among them, ò re1 and ò re2 represent the penalty coefficients for not landing.
[0055] The hierarchical deep reinforcement learning adopts the actor and critic network architectures. The actor network formulates a strategy based on the input state s, denoted as π(s; θ π ), where θ π represents the network parameters. The critic network evaluates the value of the action a selected in the state s, denoted as Q(s, a; θ Q ); The memory bank and target network techniques are used to stabilize the learning process. The memory bank stores the states, actions, rewards, and next states during the interaction between the agent and the environment. Each time, a batch of samples of size batch size is randomly sampled from it for training the network; The target network creates the same parameters as the original network, π'(s; θ π' ) and Q'(s, a; θ Q' ), and gradually updates the network through soft update; The expression for soft update is:
[0056] θ' = εθ + (1 - ε)θ'
[0057] Where ε represents the update rate, 0 < ε < 1, θ is the network parameter, and θ' is the target network parameter;
[0058] The loss function of the actor network is:
[0059]
[0060] Where is the average value calculated after taking out the batch size, s n is the state at time n, θ Q is the critic network parameter;
[0061] The loss function of the critic network is:
[0062]
[0063] Where Ta n = (r t + γQ'(s t+1 , π'(s t+1 ; θ π' )); θ Q' ) represents the update target, a n is the action selected at time n, r n is the reward obtained at time n, and γ is the reward discount;
[0064] The lower - layer actor network outputs the proportional coefficient of the bandwidth allocation for each IoT node. Through the softmax layer, the sum is 1, and the result is denoted as: λ b = <λ b1 , λ b2 ,..., λ bi ,..., λ bI >, and the bandwidth allocation size for each IoT node is a b = B * λ b .
[0065] The hierarchical deep reinforcement learning initializes the network parameters, the environmental map, the position and the initial power of the drone using the TBH - DDPG algorithm based on DDPG; entering the flight period, the forward propagation of the state map is performed to obtain the input state s(δ n ); the position change of the drone is simplified, and the position for bandwidth - allocation communication is regarded as the position at the start of the next flight period; the upper - layer network obtains o n f (δ(δ n ) and a f(δ n ) After that, the UAV directly executes action a f (δ n ), and at the same time obtains the reward r for the flight action f (δ n ), as well as the next state map; enters the communication time slot, and propagates the state map forward to obtain s(δ n,m ); the lower-layer network obtains the bandwidth allocation action a n,m through s(δ b (δ n,m )); the UAV executes bandwidth allocation communication and obtains the reward r lo (δ n,m ) and the next state map; at the same time, adds this communication reward to the upper-layer network reward r up (δ n ); the next state map is propagated forward again to obtain s(δ n,m+1 ), and the communication time slot ends in a loop; the upper-layer memory bank stores s(δ n ), a f (δ n ), r up (δ n ), s(δ n+1 ); the lower-layer memory bank stores <s(δ n,m ), a b (δ n,m ), r lo (δ n,m ), s(δ n,m+1 )>; when a certain value is stored, the network is updated.
[0066] The present invention has the following beneficial effects:
[0067] The present invention proposes a method for UAV trajectory optimization and bandwidth allocation based on hierarchical deep reinforcement learning, considering a system environment with multiple obstacle regions, unknown interference, and the real-time growth of IoT node data. It focuses on the optimization of UAV trajectories and bandwidth allocation to maximize data collection from IoT nodes. It adopts hierarchical deep learning technology, proposes the problem of maximizing the data collection volume of UAV-assisted IoT nodes, and optimizes the UAV trajectory and bandwidth allocation ratio under the constraints of energy, speed, and bandwidth. Considering that there are multiple communications when the UAV executes flight actions, it is transformed into an SMDP model, and a hierarchical deep reinforcement learning algorithm is proposed. The analysis and simulation results both show that the proposed scheme achieves a faster convergence speed and less network computation. Description of the Drawings
[0068] Figure 1 It is a system model scenario diagram related to the embodiment of the present invention;
[0069] Figure 2 Schematic diagram of time slot division involved in the embodiment of the present invention;
[0070] Figure 3 Schematic diagram of the jammer model involved in the embodiment of the present invention;
[0071] Figure 4 Abstracted system scenario diagram involved in the embodiment of the present invention;
[0072] Figure 5 Overall framework diagram of the hierarchical deep reinforcement learning algorithm involved in the embodiment of the present invention;
[0073] Figure 6 Comparison of reward convergence curves of different algorithms involved in the embodiment of the present invention;
[0074] Figure 7 Comparison of average collision times of different algorithms in different scenarios involved in the embodiment of the present invention; Detailed implementation manners
[0075] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0076] Although the steps in the present invention are arranged with reference numerals, they are not used to limit the order of the steps. Unless the order of the steps is clearly stated or the execution of a certain step requires other steps as a basis, the relative order of the steps can be adjusted. It can be understood that the term "and / or" used in the present invention covers any and all possible combinations of one or more of the related listed items.
[0077] This embodiment is based on Figure 1 the UAV-assisted IoT node data collection model scenario shown in Figure 1 As shown in the system model scenario diagram, the UAV trajectory optimization and bandwidth allocation method based on hierarchical deep reinforcement learning of the present invention includes the steps of:
[0078] (1) Construct a UAV-assisted IoT node data collection system model;
[0079] The UAV-assisted IoT node data collection system model includes a UAV-assisted IoT node data collection scenario model, a LoS / NLoS communication channel model, and a UAV energy consumption model.
[0080] 1) UAV-assisted IoT node data collection scenario model;
[0081] As shown in Figure 1The system model scenario is shown. The scenario model of data collection by UAV-assisted IoT nodes considers a rasterized square area of size Y * Y, which contains three types of obstacle areas: areas that only affect the communication area, areas that only affect the flight area, and areas that affect both the flight and communication areas. There are a total of I ground IoT nodes in the scenario. The position of the i-th IoT node is represented by p i ={x i ,y i ,0}, where i ∈ {1, 2,..., I}. The scenario includes J ground jammers, which are unknown to the UAV. The position of the j-th jammer is represented by p j ={x j ,y j ,0}, where j ∈ {1, 2,..., J}. The UAV takes off from the takeoff area and flies at a fixed height H. The time for the entire mission is represented by T. The mission time T is equally divided into N flight epochs. The n-th flight epoch is denoted as δ n , where n ∈ {1, 2,..., N}. The real-time position of the UAV can be represented by p u ={x u (δ n ),y u (δ n ),H}. It is assumed that the speed of the UAV in each flight epoch is either 0 or v, i.e., v u ∈ {0, v}. Considering that within a single flight epoch, the UAV can perform multiple bandwidth allocation communications, the single flight epoch is equally divided into M communication time slots. The m-th communication time slot in the n-th flight epoch is denoted as δ n,m , where m ∈ {1, 2,..., M}. In any time slot, the UAV uses frequency division multiple access (FDMA) for communication. The time slot division method for the entire mission is as shown in Figure 2 (the time slot division schematic diagram). When performing bandwidth allocation communication within a flight epoch, the position of the UAV is simplified, and the position during bandwidth allocation communication is regarded as the position of the UAV in the next flight epoch. The data volume of the IoT nodes in the scenario grows in real time. It is assumed that the data growth amounts of each IoT node are the same. After each communication time slot δ n,m , the IoT node will increase a fixed data volume d inc . The data volume stored in each IoT node has an upper limit, denoted as D max .
[0082] 2) LoS / NLoS communication channel model;
[0083] The LoS / NLoS communication channel model takes into account the shadow effect, log-normal fading, and the influence of jammers. In the time slot δ n,m, the channel gain g between the UAV and the i-th IoT node i is:
[0084]
[0085] where d i (δ n,m ) represents the distance between the UAV and IoT node i in time slot δ n,m . α LoS and α NLoS represent the transmission path losses of the line-of-sight and non-line-of-sight channels respectively. η LoS and η NLoS represent the shadow effects of the line-of-sight and non-line-of-sight channels, which respectively follow normal distributions and
[0086] As Figure 3 shown in the schematic diagram of the jammer model, it is assumed that the main lobe direction of the jammer is always perpendicular to the ground. The main lobe gain G m exists within the beamwidth θ m , and the gain in other directions is ignored. The relationship between the isotropic antenna gain G s and the main lobe gain G m is defined as:
[0087]
[0088] where d m is the distance from the UAV to the center line of the main beam, and P is the transmission power. The antenna gain G k (d m,k ) of the k-th jammer is expressed as:
[0089]
[0090] where θ m,k represents the beamwidth of the k-th jammer. Combining the above, the expression for the SINR n,m between the UAV and IoT node i in time slot δ i (δ n,m ) can be obtained as:
[0091]
[0092] where P i represents the transmission power of IoT node i, S N represents the noise power spectral density, and P k represents the transmission power of the jammer.
[0093] The UAV and the i-th IoT node in time slot δ n,mTransmission rate C i (δ n,m ) is expressed as:
[0094] C i (δ n,m ) = b i (δ n,m ) log 2 (1 + SINR i (δ n,m ))
[0095] where b i (δ n,m ) represents the bandwidth size allocated to IoT node i in time slot δ n,m . The sum of the bandwidths allocated to each IoT node in any communication time slot is less than or equal to the total bandwidth size B, that is:
[0096]
[0097] 3) UAV energy consumption model;
[0098] The UAV energy consumption model only considers the hovering energy consumption and propulsion energy consumption of the UAV. The UAV propulsion power P UAV is
[0099]
[0100] where P 0 represents the blade section power, u tip represents the UAV rotor tip speed, P 1 represents the induced power, v 0 represents the average rotor induced speed when the UAV is hovering, d 0 represents the fuselage drag ratio, ρ represents the air density, s 0 represents the rotor robustness, A r represents the rotor disk area. When the flight speed of the UAV is 0, P UAV (0) = P 0 + P 1 . Assuming that the amount of data transmitted in a communication time slot is When the remaining data volume of IoT node i is less than C i (δ n,m )δ n,m , only the remaining data volume of IoT node i can be obtained, which can be expressed as:
[0101]
[0102] where δ n,mis the m-th communication time slot in the n-th flight period, C i (δ n,m ) is the transmission rate between the UAV and the i-th IoT node in time slot δ n,m , and ξ is a threshold set for normal communication.
[0103] (2) Based on the constructed system model, an optimization problem is established to maximize the data collection volume by adjusting the position and bandwidth allocation ratio of the UAV;
[0104] By optimizing the trajectory and bandwidth allocation ratio of the UAV, under the constraints of energy, speed and bandwidth, maximize the data collection from the environmental ground IoT nodes. The optimization problem of maximizing the data collection volume is specifically as follows:
[0105]
[0106] The first constraint condition is the energy constraint. At any time, the energy of the UAV should be greater than 0 to ensure that the UAV can reach the designated landing point normally after completing data collection. The second constraint condition is the speed constraint, that is, the speed of the UAV in each time slot is 0 or v. The last one is the bandwidth constraint, that is, the sum of the bandwidths allocated to each IoT node in each communication time slot is always less than or equal to the total bandwidth B.
[0107] (3) Model the optimization problem as a semi-Markov decision process, and use a hierarchical deep reinforcement learning method to train an agent that can autonomously adjust the position and bandwidth allocation ratio of the UAV;
[0108] As Figure 4 shown is the abstracted system scenario map. The green area is the communication obstacle area, which affects the LoS / NLoS channel between the UAV and the IoT node, that is, the transmission power of the IoT node. The system scenario map is composed of 5 maps, namely the no-fly zone map, the communication obstacle area map, the takeoff and landing map, the IoT node map and the jammer map. The yellow area is formed by the superposition of the no-fly zone and the communication obstacle area. These 5 maps form a three-dimensional matrix, specifically represented as [Y, Y, 5]. Among them, the third dimension of the no-fly zone, the communication obstacle area and the takeoff and landing map matrix is represented by 1 and 0 respectively for existence or non-existence. The third dimension of the IoT node map is filled with the remaining data volume of the IoT node, and the third dimension of the jammer map is filled with the calculated ratio of the interference power to the noise power.
[0109] After obtaining the map and drone status information each time, center the map with the drone as the center and process the centered map in two paths. One path performs global pooling on the map, and the other path performs central cropping on the map. Use a convolutional neural network that does not participate in learning to extract features from the two paths. Flatten the output of the network and add it to the remaining battery information of the drone to obtain the input state of the algorithm.
[0110] The optimization problem is modeled as a semi-Markov decision process and consists of <S, A, P r , R>. The option space is increased to solve the problem of inconsistent action slot content. The state space S = {m r , m g , m k , m i , m j , E u , p u}, where the first five elements represent the no-fly zone, communication obstacle area, takeoff and landing area, IoT node, and jammer map matrix respectively. Eu represents the remaining battery of the drone, and pu represents the position of the drone. The action space A = {a f , a b}, where a f represents the flight action, and a b represents the bandwidth allocation action. The flight action a f ∈ {a n , a e , a s , a w , a h , a l}, where a n represents flying one grid north, a e represents flying one grid east, a s represents flying one grid south, a w represents flying one grid west, a h represents hovering, and a l represents landing. The bandwidth allocation action a b = <b 1 , b 2 ,..., b i ,..., b I >, where b i represents the bandwidth size allocated to the i-th IoT node. P r represents the probability of state transition after executing the action. The option space O = {o f}, and the execution method of the option is specified by the triple <W, π, β>, where W represents the initial state of the option, π represents the set of policies under the option, and β represents the end condition of the option. Define of ∈{o n ,o e ,o s ,o w ,o h ,o l}, where the subscript of the element in the set corresponds to the subscript of the flight action a f . W is defined as the state space S. The policy set is defined as π = <a f , a b1 , a b2 ,..., a bm ,..., a bM >, where a f corresponds to the selected o f , a bM is the bandwidth allocation action in the M-th communication time slot, and the termination condition β is that all the actions in π are completed in sequence. R represents the reward obtained after performing an action in a state, which consists of two parts. The lower-layer reward r n,m obtained in the communication time slot δ lo (δ n,m ) is:
[0111] r lo (δ n,m ) = r collection (δ n,m ) + r loss (δ n,m )
[0112] where r collection (δ n,m ) represents the positive data collection reward, and r loss (δ n,m ) represents the negative data loss reward. The data collection volume reward is given by multiplying the amount of data collected by a fixed ratio every time data collection is performed. Specifically:
[0113] r collection (δ n,m ) = ò cen ·r cen (δ n,m )
[0114] where ò cen is the proportionality coefficient of the data collection reward, and is the amount of data collected in a single communication time slot, D i (δ n,m ) is the remaining data volume of node i in the m-th communication time slot at the n-th flight time slot, D i (δ n,m+1) is the remaining data volume of node i in the (m + 1)-th communication time slot of the n-th flight time slot. The data loss reward is a fixed penalty value given when the data storage volume exceeds the upper limit of the IoT node's storage volume. Specifically:
[0115]
[0116] where r ls represents the penalty constant value. The upper-layer reward r n obtained in the flight time slot δ up (δ n ) includes all the communication time slot rewards and the rewards obtained by performing flight actions in this flight time slot. Specifically:
[0117] r up (δ n ) = r f (δ n ) + r lo (δ n,1 ) + … + r lo (δ n,m ) + … r lo (δ n,M )
[0118] where r f (δ n ) = r collision (δ n ) + r nland (δ n ) + r return (δ n ), which are the negative collision reward, negative return reward, and negative crash reward respectively. The collision reward is a fixed penalty value given when the drone enters the no-fly zone. The specific expression is:
[0119]
[0120] where, r csn represents the penalty constant value. The return reward is a penalty given when the drone's battery level is less than the threshold, and the smaller the battery level and the farther away from the landing point, the greater the penalty. The specific expression is:
[0121]
[0122] where, d re (δ n ) represents the distance from the current drone to the takeoff and landing area, ò re represents the return penalty coefficient, E(δ n ) represents the remaining battery level of the current drone, E tsdIndicates the return penalty threshold. Finally, there is the landing reward, which is a penalty given when the drone's power runs out and it has not landed yet. The specific expression is:
[0123]
[0124] where ò re1 and ò re2 represent the penalty coefficients for not landing.
[0125] During a single option process, the upper - layer network outputs the upper - layer option o n based on the observed state s(δ f ) and obtains the flight action a f , and at the same time records the reward r f (δ n ) obtained by the agent executing the flight action. Then, the lower - layer network outputs the bandwidth allocation action based on the state s(δ n,1 ) after the action is executed and o f , and obtains the reward of the lower - layer network. After the bandwidth allocation action of the last communication time slot is executed, this option ends, and at the same time, the rewards of all lower - layer actions are added to the reward of the flight action to obtain the reward of the upper - layer network.
[0126] Hierarchical deep reinforcement learning adopts the actor and critic network architectures. The actor network formulates a policy based on the input state s, denoted as π(s; θ π ), where θ π represents the network parameters. The critic network evaluates the value of the action a selected in the state s, denoted as Q(s,a; θ Q ). The replay buffer and target network techniques are used to stabilize the learning process. The replay buffer stores the states, actions, rewards, and next states during the interaction between the agent and the environment. Each time, a batch of samples of size batchsize is randomly sampled from it for training the network. The target network creates the same parameters as the original network π'(s; θ π' ) and Q'(s,a; θ Q' ), and the updated network is gradually updated through soft update. The expression of soft update is:
[0127] θ' = εθ+(1 - ε)θ'
[0128] where ε represents the update rate, 0 < ε < 1, θ is the network parameter, and θ' is the target network parameter.
[0129] The loss function of the actor network is:
[0130]
[0131] where, is the average value calculated after taking out the batch size, s n is the state at time n, θ Q are the parameters of the critic network.
[0132] The loss function of the critic network is:
[0133]
[0134] where Ta n =(r t +γQ'(s t+1 ,π'(s t+1 ;θ π' );θ Q' )) represents the updated target, a n is the action selected at time n, r n is the reward obtained at time n, and γ is the discount of the reward.
[0135] The lower-layer actor network outputs the proportionality coefficients for the bandwidth allocation of each IoT node. Through the softmax layer, the sum is 1, and the result is denoted as: λ b =<λ b1 ,λ b2 ,…,λ bi ,...,λ bI >, and the bandwidth allocation size for each IoT node is a b =B*λ b .
[0136] As Figure 5 shown is the overall framework diagram of the hierarchical deep reinforcement learning algorithm. The hierarchical deep reinforcement learning adopts the TBH-DDPG algorithm based on DDPG, initializing the network parameters, the environmental map, the position of the drone, and the initial power.
[0137] Entering the flight period, the forward propagation of the state map is performed to obtain the input state s(δ n ). Simplifying the position change of the drone, the position for bandwidth allocation communication is regarded as the position at the start of the next flight period. After the upper-layer network obtains o n (δ f )(δ n ) and a f (δ n ), the drone directly executes the action a f (δ n ), and at the same time obtains the reward r f (δ n ) of the flight action and the next state map.
[0138] Enter the communication time slot, and perform forward propagation on the state map to obtain s(δ n,m ). The lower-layer network obtains the bandwidth allocation action a n,m through s(δ b )(δ n,m ). The UAV executes bandwidth allocation communication and obtains the reward r lo (δ n,m ) and the next state map. At the same time, add this communication reward to the upper-layer network reward r up (δ n ). The next state map is then propagated forward to obtain s(δ n,m+1 ), and the communication time slot ends after continuous looping. The upper-layer memory bank stores <s(δ n ), a f (δ n ), r up (δ n ), s(δ n+1 )>, and the lower-layer memory bank stores s(δ n,m ), a b (δ n,m ), r lo (δ n,m ), s(δ n,m+1 )>. When the storage reaches a certain value, the network is updated.
[0139] The simulation analysis is as follows:
[0140] Figure 6 Figure shows the comparison of the reward convergence curves of different algorithms. Each algorithm was run 5 times with the same parameters. The solid lines with markers represent the average rewards of the 5 experiments, and the shaded areas represent the range of the mean plus or minus 50% variance. In the convergence stage, the performance of the proposed algorithm is similar to that of the non-hierarchical algorithm TBJN-DDPG based on DDPG, but its convergence speed is increased by 44%. In contrast, the hierarchical algorithm TBH-SAC based on SAC is slightly lower in performance than the proposed algorithm and TBJN-DDPG, but its convergence speed is comparable to that of the proposed algorithm. The reward of the proposed hierarchical algorithm in the lower layer is only the reward for data collection and data loss, and the optimized action output is only bandwidth allocation. Although the reward in the upper layer includes the reward in the lower layer, the optimized action output is only flight option. Compared with the non-hierarchical algorithm that includes all rewards and optimizes both flight actions and bandwidth allocation actions at the same time, the proposed algorithm can effectively shorten the convergence time of the algorithm compared with the non-hierarchical algorithm. In addition, the performance of the algorithms that can perform autonomous bandwidth allocation is better than that of a greedy time-division multiplexing algorithm TDMA-DDPG based on DDPG, and can achieve data collection for multiple IoT nodes, thus significantly improving the bandwidth utilization efficiency.
[0141] Figure 7Shows the comparison of the obstacle avoidance effects of different algorithms in different scenarios. The obstacle avoidance effects of each algorithm were tested in 3 different scenarios. In each scenario, the trained algorithm was tested 500 times, and the average number of collisions in 500 tests was taken as the result. It can be seen that the proposed algorithm and TBJN-DDPG can effectively avoid collisions and complete the task in different scenarios. However, there are still some fewer collisions in TBH-SAC and TDMA-DDPG, but the average number of collisions is less than 1.
[0142] Table 1 shows the influence of the number of network layers involved in the embodiments of the present invention on the algorithm.
[0143] Table 1
[0144]
[0145] As shown in Table 1, it can be observed that the convergence rounds of the non-hierarchical algorithm TBJN-DDPG range from 2050 to 2350 under different hidden layer sizes, while the convergence rounds of the hierarchical algorithm TBH-DDPG range from 1150 to 1400, and its convergence speed has been significantly improved. At the same time, under similar rewards, the network hidden layer size of the hierarchical algorithm only needs 128, while the network layer of the non-hierarchical algorithm needs 1024, and the finally saved network size is approximately reduced by 6 times. In addition, this paper also counted the FLOPS required for the actor network in the two algorithms. The TBJN-DDPG algorithm only contains one actor network and only predicts once per flight cycle; while the TBH-DDPG algorithm contains two actor networks, where the upper actor predicts once per flight cycle and the lower actor network predicts four times. It can be seen that when obtaining similar rewards, the actor network FLOPS of the hierarchical algorithm proposed in this paper is reduced by 58.05%. It can be obtained from this that the hierarchical algorithm proposed in this paper not only significantly improves the convergence speed, but also can effectively reduce the network scale.
[0146] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights.
[0147] In addition, it should be understood that although this specification is described in terms of embodiments, not every embodiment contains only an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A UAV trajectory optimization and bandwidth allocation method based on hierarchical deep reinforcement learning, characterized in that: Includes steps: (1) Construct a drone-assisted IoT node data collection system model; (2) Based on the constructed system model, an optimization problem is established to maximize the amount of data collected by adjusting the position of the UAV and the bandwidth allocation ratio; (3) The optimization problem is modeled as a semi-Markov decision process, and a hierarchical deep reinforcement learning method is used to train an intelligent agent that can autonomously adjust the position of drones and the bandwidth allocation ratio.
2. The method for unmanned aerial vehicle trajectory optimization and bandwidth allocation based on hierarchical deep reinforcement learning according to claim 1 is characterized in that: The drone-assisted IoT node data collection system model in step (1) includes a drone-assisted IoT node data collection scenario model, a LoS / NLoS communication channel model, and a drone energy consumption model; The UAV-assisted IoT node data collection scenario model considers a Y*Y grid square area, which contains three types of obstacle areas: the area that only affects the communication area, the area that only affects the flight area, and the area that affects both the flight and the communication area. There are a total of I ground IoT nodes in the scene, and the position of the i-th IoT node is represented by p i ={x i ,y i ,0}, where i∈{1,2,...,I}; the scene contains J ground jammers, and they are unknown to the drone. The position of the jth jammer is represented by p j ={x j ,y j ,0}, where j∈{1,2,...,J}; the UAV starts from the take-off area and flies at a fixed altitude H. The duration of the entire mission is represented by T; the mission time T is divided into N flight periods at equal intervals, and the nth flight period is represented by δ n , where n∈{1,2,...,N}; the real-time position of the drone can be expressed as p u ={x u (δ n ),y u (δ n ),H}; Assume that the speed of the drone in each flight period is 0 or v, that is, v u ∈{0,v}; a single flight period is divided into M communication time slots at equal intervals, and the mth communication time slot under the nth flight period is denoted as δ n,m , where m∈{1,2,...,M}; when bandwidth allocation communication is performed in a flight period, the position of the drone is simplified, and the position during bandwidth allocation communication is regarded as the position of the drone in the next flight period; the amount of data of the IoT nodes in the scene will grow in real time. Assuming that the data growth of each IoT node is the same, in each communication time slot δ n,m After that, IoT nodes will increase the fixed amount of data d inc ; Of course, there is an upper limit to the amount of data stored in each IoT node, denoted as D max ; The LoS / NLoS communication channel model takes into account the shadow effect, logarithmic fading and the influence of the jammer; in the time slot δ n,m , the channel gain g between the drone and the i-th IoT node i for: where d i (δ n,m ) indicates that in time slot δ n,m The distance between the drone and IoT node i; α LoS and α NLoS Respectively represent the transmission path loss of line-of-sight and non-line-of-sight channels; η LoS and η NLoS represents the shadow effect of line-of-sight and non-line-of-sight channels, which follow normal distributions respectively and The jammer model is designed, assuming that the jammer's main lobe direction is always perpendicular to the ground; the jammer's main lobe gain G m Exists in beam width θ m and ignore the gain in other directions; the isotropic antenna gain G s and main lobe gain G m The relationship between is defined as: Among them, d m is the distance from the drone to the centerline of the main beam, P is the transmission power; the antenna gain G of the kth jammer k (d m,k ) is expressed as: Among them, θ m,k represents the beam width of the kth jammer; based on the above, we can get the UAV in time slot δ n,m The SINR between the IoT node i and i (δ n,m ) is: Where P i represents the transmission power of IoT node i, S N represents the noise power spectral density, P k Indicates the jammer transmit power; The drone and the i-th IoT node are in time slot δ n,m The transmission rate C i (δ n,m ) is expressed as: C i (d n,m )=b i (d n,m )log2(1+SINR i (d n,m )) where b i (δ n,m ) indicates that in time slot δ n,m The bandwidth size allocated to IoT node i in the following example; the sum of the bandwidth allocated to each IoT node in any communication time slot is less than or equal to the total bandwidth size B, that is: The UAV energy consumption model only considers the hovering energy consumption and propulsion energy consumption of the UAV. The UAV propulsion power P UAV for Where P0 represents the blade profile power, u tip represents the tip speed of the UAV rotor, P1 represents the induced power, v0 represents the average induced speed of the UAV rotor when it is hovering, and v u represents the flight speed of the drone, d0 represents the fuselage drag ratio, ρ represents the air density, s0 represents the rotor strength, A r Indicates the rotor disc area; when the UAV's flight speed is 0, P UAV (0) = P0 + P1; Assume that the amount of data transmitted in a communication time slot is When the remaining data volume of IoT node i Less than C i (δ n,m )δ n,m When , only the remaining data of IoT node i can be obtained. It can be expressed as: Among them, δ n,m is the mth communication time slot in the nth flight period, C i (δ n,m ) is the time interval between the drone and the i-th IoT node in time slot δ n,m The transmission rate of ξ sets a threshold for normal communication.
3. The method for unmanned aerial vehicle trajectory optimization and bandwidth allocation based on hierarchical deep reinforcement learning according to claim 1, characterized in that: The optimization problem of maximizing the amount of data collected in step (2) is specifically as follows: The first constraint is the energy constraint. At any time, the energy of the drone should be greater than 0 to ensure that the drone can reach the designated landing point normally after completing data collection; the second constraint is the speed constraint, that is, the speed of the drone in each time slot is 0 or v; the last one is the bandwidth constraint, that is, the sum of the bandwidth allocated to each IoT node in each communication time slot is always less than or equal to the total bandwidth B.
4. The method for unmanned aerial vehicle trajectory optimization and bandwidth allocation based on hierarchical deep reinforcement learning according to claim 1, characterized in that: The system scenario map of the semi-Markov decision process in step (3) is composed of 5 maps, namely, a no-fly zone map, a communication obstacle zone map, a take-off and landing map, an IoT node map and a jammer map, specifically represented as [Y, Y, 5]. After obtaining the map and drone status information each time, the map is centralized with the drone as the center, and the centralized map is processed in two ways, one way is to perform global pooling on the map, and the other way is to perform center cropping on the map. A convolutional neural network that does not participate in learning is used to extract features from the two ways. The output of the network is flattened and added to the remaining power information of the drone to obtain the input state of the algorithm.
5. The method for unmanned aerial vehicle trajectory optimization and bandwidth allocation based on hierarchical deep reinforcement learning according to claim 1, characterized in that: Step (3) Model the optimization problem as a semi-Markov decision process and <S,A,P r ,R>, increasing the option space, solving the problem of inconsistent action time slot content; where the state space S = {m r ,m g ,m k ,m i ,m j ,E u ,p u }, where the first five elements represent the no-fly zone, communication barrier zone, take-off and landing zone, IoT node and jammer map matrix, respectively. u Indicates the remaining power of the drone, p u Indicates the position of the drone; Action space A = {a f ,a b }, where a f Indicates flying action, a b Indicates bandwidth allocation action; flight action a f ∈{a n ,a e ,a s ,a w ,a h ,a l }, where a n Indicates flying one square north, a e Indicates flying one square eastward, a s Indicates flying one square south, a w Indicates flying one square westward, a h Indicates hover, a l Indicates a drop; bandwidth allocation action a b = <b1,b2,...,b i ,...,b I >, where b i It is represented as the bandwidth allocated to the i-th IoT node; P r represents the probability of state transition after executing the action; option space O = {o f }, option is executed by triple<W,π,β> Definition: W represents the initial state of option, π represents the strategy set under option, and β represents the end condition of option. f ∈{o n ,o e ,o s ,o w ,o h ,o l }, where the element subscripts in the collection correspond to the flight action a f The subscript of ; W is defined as the state space S; the strategy set is defined as π = f ,a b1 ,a b2 ,...,a bm ,...,a bM >, where a f Corresponding to the selected o f , a bM is the bandwidth allocation action in the Mth communication time slot, and the termination condition β is that all actions in π are completed in sequence; R represents the reward obtained after performing an action in a state, which consists of two parts, in the communication time slot δ n,m The lower level reward r lo (δ n,m )for: r lo (d n,m )=r collection (d n,m )+r loss (d n,m ) where r collection (δ n,m ) represents the positive data collection reward, r loss (δ n,m ) indicates the negative data loss reward; the data collection amount reward is a reward equal to the amount of data collected multiplied by a fixed ratio each time data is collected, specifically: r collection (d n,m )=ò cen ·r cen (d n,m ) Among them is cen The proportionality factor of the data collection reward, and is the amount of data collected in a single communication time slot, D i (δ n,m ) is the amount of remaining data of node i in the mth communication slot under the nth flight slot, D i (δ n,m+1 ) is the amount of remaining data of node i in the m+1th communication slot under the nth flight slot; the data loss reward is a fixed penalty value given when the data storage capacity exceeds the upper limit of the IoT node storage capacity, specifically: where r ls represents the penalty constant value; In the flight time slot δ n The upper reward r up (δ n ) includes all communication slot rewards under this flight slot and the rewards obtained by performing flight actions, specifically: r up (d n )=r f (d n )+r lo (d n,1 )+…+r lo (d n,m )+…r lo (d n,M ) where r f (δ n )=r collision (δ n )+r nland (δ n )+r return (δ n ), respectively negative collision reward, negative return reward and negative crash reward; the collision reward is a fixed penalty value given when the drone enters the no-fly zone, and the specific expression is: Among them, r csn represents the penalty constant value, p u (δ n ) represents the real-time position of the UAV during the nth flight period; The return reward is a penalty when the drone's battery power is less than the threshold. The lower the battery power and the farther away from the landing point, the greater the penalty. The specific expression is: Among them, d re (δ n ) indicates the distance from the current drone to the take-off and landing areas, re Indicates the return penalty coefficient, E(δ n ) indicates the current remaining battery power of the drone, E tsd represents the return penalty threshold; the last is the landing reward, which is a penalty given when the drone runs out of power and has not landed. The specific expression is: Among them re1 and re2 Indicates the penalty factor for not landing.
6. The method for unmanned aerial vehicle trajectory optimization and bandwidth allocation based on hierarchical deep reinforcement learning according to claim 2 is characterized in that: The hierarchical deep reinforcement learning adopts the actor and critic network architecture. The actor network formulates a strategy based on the input state s, denoted as π(s; θ π ), where θ π represents the network parameters. The critic network evaluates the value of action a selected in state s, which is recorded as Q(s,a;θ Q ); use memory bank and target network technology to stabilize the learning process. The memory bank stores the state, action, reward and next state of the agent and the environment during the interaction. Each time, a batch of samples of size batch size are randomly sampled from it for training the network; the target network is to create the same parameters π'(s; θ as the original network π' ) and Q'(s,a;θ Q' ), the network is gradually updated by soft update; the expression of soft update is: θ'=εθ+(1-ε)θ' Where ε represents the update rate, 0<ε<1, θ is the network parameter, and θ' is the target network parameter; The loss function of the actor network is: in, is the average value calculated after taking out the batch size, s n is the state at time n, θ Q is the critic network parameter; The critic network loss function is: Among them, n =(r n +γQ'(s t+1 ,π'(s t+1 θ π' );θ Q' )) represents the updated target, a n is the action chosen at time n, r n is the reward obtained at time n, and γ is the discount of the reward; The lower actor network outputs the proportional coefficient of bandwidth allocation for each IoT node, which is summed to 1 through the softmax layer. The result is recorded as: b =<λ b1 ,λ b2 ,...,λ bi ,...,λ bI >, the bandwidth allocation size of each IoT node is a b =B*λ b .
7. The method for unmanned aerial vehicle trajectory optimization and bandwidth allocation based on hierarchical deep reinforcement learning according to claim 2, characterized in that: The hierarchical deep reinforcement learning adopts the DDPG-based TBH-DDPG algorithm to initialize the network parameters, the environment map, the position and initial power of the drone; when entering the flight period, the state map is forward propagated to obtain the input state s(δ n ); simplify the position change of the UAV, and regard the position where the bandwidth allocation communication is performed as the starting position of the next flight period; the upper network is based on s(δ n ) get o f (δ n ) and a f (δ n ) after which the drone directly executes action a f (δ n ), and get the flying action reward r f (δ n ) and the next state map; enter the communication time slot, forward propagate the state map to obtain s(δ n,m ); the lower network passes s(δ n,m ) Get bandwidth allocation action a b (δ n,m ); The drone performs bandwidth allocation communication and obtains a reward r lo (δ n,m ) and the next state map; at the same time, the communication reward is added to the upper network reward r up (δ n ) in the next state map; the next state map is forward propagated to obtain s(δ n,m+1 ), and the communication time slot is ended in a loop; the upper memory bank stores <s(δ n ),a f (δ n ),r up (δ n ),s(δ n+1 )>, the lower memory bank stores <s(δ n,m ),a b (δ n,m ),r lo (δ n,m ),s(δ n,m+1 )>; When a certain value is stored, the network is updated.
Citation Information
Cited By
Multi-patrolling device software collaborative upgrading method, device, equipment and medium
CN121934865A