Unmanned aerial vehicle three-dimensional flight path planning method for wireless sensor network data collection and energy supplementation
By combining wireless power transmission and information transmission technology, dynamic cluster head selection and deep reinforcement learning algorithm are used to optimize the three-dimensional flight trajectory of the drone, solving the problem of insufficient impact of drone height planning on sensor nodes, achieving efficient energy replenishment and data acquisition, and improving the stability and reliability of the network.
Patent Information
- Application Number
- CN202510532118.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-01
AI Technical Summary
In the existing wireless rechargeable sensor network, the three-dimensional flight altitude planning of the drone does not fully consider the distance between the sensor node and the drone, resulting in low energy transmission and data acquisition efficiency, sensor nodes are prone to death due to insufficient energy, and insufficient network stability and reliability.
Combining wireless power transmission and wireless information transmission technology, a dynamic cluster head selection mechanism based on routing protocol and P-SOM neural network is adopted, combined with deep reinforcement learning algorithm DDPG, the three-dimensional flight trajectory of the drone is optimized, and the altitude and path are dynamically adjusted to realize energy supplementation and data acquisition of sensor nodes.
Effectively reduce the average flight energy consumption of drones, reduce the mortality rate of sensor nodes, improve the network data acquisition and energy replenishment efficiency, extend the network survival time, and improve system reliability and stability.
Smart Images

Figure CN120406499A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of the Internet of Things, and particularly relates to a three-dimensional flight trajectory planning method for an unmanned aerial vehicle for data collection and energy replenishment in a wireless sensor network. Background Art
[0002] A wireless rechargeable sensor network (WRSN) consists of three parts: randomly deployed sensor nodes, mobile charging devices, and base stations. Sensor nodes can monitor various environmental data such as temperature, humidity, gas composition, and soil composition in real time, and can also detect disaster omens such as earthquakes, fires, and floods in advance. At the same time, as a mobile charging device, an unmanned aerial vehicle can interact with sensor nodes for energy transfer and data collection. With its high mobility, the unmanned aerial vehicle can adapt to diverse sensor node distributions. The unmanned aerial vehicle starts from the base station, visits sensor nodes in sequence for energy replenishment and data collection, and returns to the base station for charging when the battery power of the unmanned aerial vehicle is insufficient. Nowadays, using an unmanned aerial vehicle to extend the lifespan of a wireless rechargeable sensor network has been widely applied in important fields such as military, environmental protection, agricultural monitoring, and industrial automation, and has attracted extensive attention from the industrial community, countries, and scholars at home and abroad.
[0003] The current wireless rechargeable sensor network faces many problems. The traditional wireless rechargeable sensor network uses a mobile cart to replenish energy for sensor nodes. However, in post-disaster rescue and remote areas with complex terrains, the flexibility and adaptability of the cart are limited, making it difficult to meet emergency needs. In contrast, an unmanned aerial vehicle can perform various tasks with its high mobility, wide coverage, and low cost. In addition, the wireless rechargeable sensor network also faces the problem of insufficient sensor energy. When the sensor energy is exhausted, the sensor node will die, and at the same time, the real-time monitoring of the network will be lost. Therefore, using an unmanned aerial vehicle can significantly extend the working lifespan of the network, maintain real-time monitoring and data collection of the environment, and improve the overall reliability and stability of the system.
[0004] "Energy Maximization for Ground Nodes in UAV-Enabled Wireless Power Transfer Systems" published by Min Li et al. in IEEE Internet of Things Journal in 2023 proposed a V-shaped wireless energy transfer scheme, that is, the UAV descends to the optimal hovering position to charge the ground nodes (GNs), so as to transfer more energy to the GNs. In addition, considering that the GNs far from the hovering position in the V-shaped wireless energy transfer scheme receive less energy, in order to improve the fairness of the energy received by the GNs, an inverted trapezoidal wireless energy transfer scheme was further developed, that is, when the UAV hovers or flies horizontally, it continuously charges the GNs after reducing its altitude, and two algorithms were developed to solve the optimal hovering position and altitude of the UAV, while optimizing the energy replenishment of the sensor nodes.
[0005] "UAV-Enabled Wireless Power Transfer With Base Station Charging and UAV Power Consumption" published by Hua Yan et al. in IEEE Transactions on Vehicular Technology in 2020 considered the energy consumption of the UAV during hovering and flight, the charging process from the base station to the UAV, and the conversion loss of the energy harvester in one-dimensional (1D) and two-dimensional (2D) wireless energy transfer systems. Two different charging schemes were proposed to maximize the total energy received by all sensors, so as to find the optimal strategy for UAV deployment.
[0006] From the published literature, although some studies have realized the online charging scheduling and data collection of UAVs for sensor nodes based on deep reinforcement learning, most studies still mainly focus on the energy replenishment and energy consumption problems of UAVs at a fixed flight altitude, ignoring the importance of the distance between the UAV and the sensor nodes. Especially in a three-dimensional (3D) scenario, the planning of the UAV altitude has a crucial impact on the efficiency of energy transfer and data collection. Summary of the Invention
[0007] Aiming at the above problems, the present invention proposes a three-dimensional flight trajectory planning method for UAVs for data collection and energy replenishment in wireless sensor networks, which combines wireless power transfer (WPT) and wireless information transfer (WIT) technologies to achieve the unity of energy replenishment and data collection of sensor nodes.
[0008] To achieve the above objectives, the technical solution adopted by the present invention is as follows: Step 1: Establish a wireless sensor network model: randomly deploy sensor nodes in the network. The CH nodes are represented as {ch1, ch2, ch3, …, ch m}, the CM nodes are represented as {cm1, cm2, cm3, …, cm h}, and at the same time, deploy a drone serving the sensor nodes and a base station for data storage and battery replacement of the drone. The drone combines wireless power transfer WPT and wireless information transfer WIT technologies to achieve the unification of energy replenishment and data collection for sensor nodes;
[0009] Step 2: Design a dynamic cluster head selection mechanism based on a routing protocol, and adopt a cluster head selection strategy to optimize the position of the cluster head;
[0010] Step 3: The drone inputs the information of the received cluster head nodes into the P-SOM neural network to obtain the current optimal access sequence and the cluster head node with the highest current service priority;
[0011] Step 4: Design a drone path planning scheme based on the deep reinforcement learning algorithm DDPG. Through the DDPG algorithm, the drone finds the cluster head node with the highest current service priority. When the drone reaches within the data collection range of the cluster head node with the highest current service priority, the DDPG algorithm is used to adjust the descending height of the drone, collect data from the cluster head node, and replenish energy for the nodes within the range;
[0012] Step 5: When the energy of the drone is about to run out and it cannot complete the next task, ensure that the drone can return to the base station, then replace the battery of the drone and store the data before continuing to execute the task. In addition, when the drone completes a round of access tasks for all clusters, new cluster heads will be reselected according to the cluster head selection strategy in Step 2, and the optimal sequence that the drone needs to access in the next round will be updated.
[0013] The specific steps of designing the dynamic cluster head selection mechanism based on the routing protocol in Step 2 are as follows:
[0014] Use the density peak clustering algorithm DPC to cluster the sensor nodes. In each round of competition, each cluster node elects a CH node according to its own energy state and location information to balance energy consumption. A dynamic routing protocol is adopted within each cluster. The cluster member CM nodes use the cluster head as the next-hop relay, and the cluster head collects the member data and transmits it to the drone. The weight of the CM node is obtained by the following formula:
[0015]
[0016] In the above formula, E maxRepresents the maximum energy of the sensor node, Represents the remaining energy of the CM node at time t, E ave (t) represents the average remaining energy of the CM node, α is the weighting coefficient, representing the priority and importance of distance and energy, d i,j Is expressed as cm i Node and ch j The distance between, d ave Represents the average distance between the CM node and the CH node;
[0017] The total energy consumed by the CM node to send a data packet to the CH node is
[0018]
[0019] Among them, h represents the number of CM nodes in a cluster, e t Represents the energy consumption of transmitting unit data, q i,j Represents the size of the unit data packet sent by the CM node to the CH node, and the energy consumption of the CH node Is divided into the energy consumption of receiving the data packet and the energy consumption of uploading the data to the UAV:
[0020]
[0021] Among them, m represents the number of CH nodes in the entire network, e r Is the energy consumption of the receiving unit data, Represents the energy consumption of data upload.
[0022] The specific steps of Step3 are as follows:
[0023] Introduce a penetration mechanism into the SOM neural network to form a P-SOM neural network:
[0024]
[0025] Among them, p(t) is the penetration factor, p max Is the initial maximum penetration ratio, α and β are used to control the change of the penetration ratio respectively, t represents the number of iterations. At the beginning of training, the penetration ratio is large. As the training progresses, the penetration ratio gradually decreases. The update formula is expressed as:
[0026] W n (t + 1) = p(t)·W n (t) + (1 - p(t))·G n,q (t)·η(t)·(x k - w n (t))(5)
[0027] In the above formula, Wn (t) represents the weight of the nth neuron at the t-th iteration, W n (t + 1) represents the weight of the nth neuron at the (t + 1)-th iteration, η(t) represents the learning rate, x k is the current input data sample. G n,q (t) represents the strength relationship of the influence on surrounding neurons. In P-SOM, the range that promotes the excitation of surrounding neurons is called the "winning neighborhood". Within the winning neighborhood, the relationship between the strength of influence and distance is expressed by the following formula:
[0028]
[0029] In the above formula, d n,q represents the distance between neuron n and the winning neuron q, σ(t) is the radius of the winning neighborhood, and its value gradually decreases over time. It is larger in the initial stage and smaller in the later stage for fine-tuning mode, and can be expressed by the following formula:
[0030]
[0031] In the above formula, σ0 is the initial neighborhood range, representing the neighborhood radius at the start of training, τ is a constant that controls the neighborhood decay rate, t represents the number of iterations. Input the two-dimensional coordinates of all sensor nodes into the P-SOM neural network to obtain the current optimal path, optimal access sequence, and the cluster head node with the highest current service priority.
[0032] The specific steps of Step4 are as follows:
[0033] Assume that the drag coefficient of the UAV blade cross-section is a constant. The propulsion power of a rotor UAV with a fixed altitude, speed V, and rotor thrust T is:
[0034]
[0035] Among them, represents the thrust-to-weight ratio, W is the weight of the UAV, Ω, d0, ρ, s, A, and R represent the blade angular velocity, fuselage drag ratio, air density, rotor solidity, rotor disk area, and rotor radius respectively; represents the blade profile power during hovering, represents the induced power during hovering;
[0036] When the UAV is descending vertically and ascending vertically, the rotor thrust T received by the UAV will change. The propulsion power of the UAV during descent is expressed as:
[0037]
[0038] The propulsion power during climbing is:
[0039]
[0040] Among them, represents the sum of the blade power and the induced power during hovering. Assuming that the acceleration a is less than the gravitational acceleration g = 9.8 m / s 2 , the DDPG algorithm is used to plan the descent height of the UAV. Subsequently, the UAV hovers to collect data from the CH node and replenish the energy of other nodes within the range. According to the Shannon formula, the data transmission rate between the UAV and the target node is expressed as:
[0041]
[0042] Among them, B and σ 2 represent the channel bandwidth and the noise power, P t is the node transmission power, γ0 is the channel power gain between the UAV and the target node, which varies with the distance between the UAV and the target node. The hovering time is expressed as:
[0043]
[0044] Among them, is the data buffer length of the target node. The energy consumed by the target node to upload data is:
[0045]
[0046] Assuming that compared with the data collection time of the UAV, the node battery can be fully charged in a very short time, ignoring the charging time. After the UAV completes the data collection from the target node and the energy replenishment of the sensor nodes within the range, the UAV climbs to the initial height and continues to perform path planning to find the next target node, especially in scenarios where the sensor nodes are unevenly distributed or there are nodes with high energy requirements;
[0047] In the DDPG algorithm, the UAV directly interacts with the environment. Regarding the current optimal access sequence output in step 3 and the cluster head node with the highest current service priority as the target node, combining the UAV's own position and the relative position of the target node, the UAV tries to shorten the distance to the target as much as possible. The problem is abstracted into a Markov decision process. Based on the framework of the Markov process, the DDPG algorithm is defined by a quadruple {S rl , A rl , R rl , S rl ′}. S rl is the state space, A rl is the action space, R rl is the reward function, S rl ′It is the next state that the drone enters after performing an action. The specific definition is as follows:
[0048] State space S rl
[0049]
[0050] in, Expressed as the distance between the UAV and the target node, x u (t),y u (t),z u (t) represents the lateral position, longitudinal position and height of the UAV, N f (t) represents the number of times the drone flies out of the boundary, N d (t) represents the current number of dead sensor nodes;
[0051] Action space A rl
[0052] A rl ={v u (t),θ u (t),h u}
[0053] Among them, v u (t) represents the instantaneous speed of the UAV at time t, θ u (t) represents the instantaneous heading angle of the UAV along the horizontal direction at time t, h u Indicates the height to which the drone descends after finding the target node;
[0054] Reward function R rl
[0055]
[0056] R0(t)=R connect +η1λ cover (t)+η2h u (15)
[0057]
[0058] Among them, R0(t) is the reward after the drone establishes a connection with the target node, R1(t) is the reward when the drone is looking for the target node, and R connect represents the reward after the drone establishes a connection with the target node, λ cover (t) represents the number of sensor nodes currently covered by the drone, Δd u,tar (t) represents the distance between the UAV and the target node, R dis the range of UAV data collection, and η1η2η3η4 are the reward weight factors of the optimization objectives respectively;
[0059] The DDPG network framework consists of four networks: the actor network, the critic network, the target actor network, and the target critic network. The main network and the target network share the same network structure. The actor network outputs the actions of the UAV according to the input state, and the critic network evaluates these actions to update the actor network. Use the random network parameters at time t Initialize the actor network and the critic network respectively, and copy the parameters at the same time t to initialize the target network Select the action space A according to the current policy rl , and obtain the reward R rl , and the environmental state becomes S rl ′ , store {S rl , A rl , R rl , S rl ′} into the replay pool, sample quadruples from the replay pool, and use the target actor network μ ′ and the target critic network Q ′ to calculate the target value Y as follows:
[0060]
[0061] where γ represents the discount factor, and use gradient descent to train the critic network to minimize the target loss L, as shown below:
[0062]
[0063] where M is the number of training epochs, calculate the sampled policy gradient, and use this to update the current actor network:
[0064]
[0065] where J represents the total discounted cumulative reward, represents the gradient with respect to the network parameters , represents the action output by the actor network according to the current state S rl , represents the action A generated by the current policy rl = μ(S rl ), calculate the gradient of the Q network with respect to the action, and finally use the soft update method to update, introducing the learning rate τ, as shown below:
[0066]
[0067] wherein and are the updated target network parameters.
[0068] The beneficial effects of the present invention are as follows: The present invention proposes an optimization scheme based on an improved dynamic routing protocol for cluster head selection strategy and P-SOM neural network, comprehensively considering key factors such as the flight energy consumption of drones and the mortality rate of sensor nodes. Deep reinforcement learning is used to plan the path and altitude of drones to guide drones to efficiently complete the data collection and energy replenishment tasks of sensor nodes. This method effectively reduces the average flight energy consumption of drones and at the same time reduces the mortality rate of sensor nodes, providing an efficient and reliable solution for data collection and energy replenishment in wireless sensor networks. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 is the network model diagram of the wireless sensor network;
[0070] Figure 2 is the network structure diagram of the DDPG algorithm;
[0071] Figure 3 is the flow chart of drone data collection and energy replenishment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0072] For a more detailed description of the present invention and for the convenience of those skilled in the art to understand, the present invention will be further described below in conjunction with the accompanying drawings and embodiments. The embodiments in this part are used to explain and illustrate the present invention for the purpose of understanding, and do not limit the present invention.
[0073] Embodiment 1: A three-dimensional flight trajectory planning method for drones used in data collection and energy replenishment of wireless sensor networks, comprising the following steps:
[0074] Step1: Refer to Figure 1 , establish a wireless sensor network model. Randomly deploy hundreds of sensor nodes in the wireless sensor network, set different regions, and the data update speed and energy consumption of sensor nodes in different regions are different. The CH nodes are denoted as {ch1, ch2, ch3,..., ch m}, the CH nodes are denoted as {cm1, cm2, cm3,..., cm h}, and at the same time, deploy a drone serving the sensor nodes and a base station for data storage and battery replacement of the drone. The coordinates of the drone are (x u (t), y u (t), z u(t)). The drone combines wireless power transfer (WPT) and wireless information transfer (WIT) technologies to achieve the unified process of energy replenishment and data collection for sensor nodes. To better understand the process of drone charging scheduling and data collection from a three-dimensional scenario. Figure 1 It will be divided into two planes. One is the horizontal flight plane of the drone, and the other plane is the deployment position plane of the sensors and the base station. When the drone finds the target node in the horizontal flight plane, it uses the DDPG algorithm to plan the descending height of the drone. See Step4 for details. Then it hovers for energy replenishment and data collection. Finally, the drone rises to the initial height and continues to plan the path in the horizontal flight plane to find the next target node. When the energy of the drone is about to run out and it cannot complete the next task, the drone returns to the base station. The specific process can be seen Figure 1 as shown {S1→S2→S3→S4→S5→S6→S7→S8→S9→S10→S11}.
[0075] Step2: Design a dynamic cluster head selection mechanism based on the routing protocol and adopt a cluster head selection strategy to optimize the position of the cluster head.
[0076] Compared with the static routing protocol, the dynamic routing protocol can evenly consume the energy of nodes. In the traditional leach routing protocol, each node will select a random number between 0 and 1. If it is less than the threshold, it will become the cluster head (CH) node in the current round. However, this method lacks consideration of the remaining energy and node position. Based on these two factors, a CH node selection mechanism is proposed. The present invention uses the Density Peak Clustering (DPC) algorithm to cluster the sensor nodes. Compared with the traditional clustering algorithms, the DPC algorithm does not require iterative calculation and does not depend on the artificially set number of clusters. It locates the clustering centers at one time. The main purpose of the DPC clustering algorithm is to calculate the local density and density distance of nodes and use the nodes with higher density and farther distance from other high-density points as the clustering centers. Through this method, the sensor nodes are effectively clustered, so as to ensure that more nodes can be covered when the drone reaches the clustering centers. In each round of election, each cluster node elects a CH node according to its own energy state and position information to balance the energy consumption. A dynamic routing protocol is adopted within each cluster. The cluster member (CM) nodes use the cluster head as the next-hop relay, and the cluster head collects the member data and transmits it to the drone. The weight of the CM node can be obtained by the following formula:
[0077]
[0078] In the above formula, E max represents the maximum energy of the sensor node, represents the remaining energy of the CM node at time t, Eave (t) represents the average remaining energy of the CM node, and α is the weighting coefficient, representing the priority and importance of distance and energy, d i,j is expressed as cm i the node and ch j the distance between, d ave represents the average distance between the CM node and the CH node. [[ID=!1]]
[0079] Regarding the energy consumption of the node in the sleep mode, since the energy consumption in this part is very small and can be ignored, the present invention only calculates the energy consumption of the node when receiving and sending. The total energy consumed by the CM node to send a data packet to the CH node is:
[0080]
[0081] Among them, h represents the number of CM nodes in a cluster, e t represents the energy consumption of transmitting unit data, q i,j represents the size of a unit data packet sent by the CM node to the CH node. The energy consumption of the CH node is divided into the energy consumption of receiving the data packet and the energy consumption of uploading the data to the drone:
[0082]
[0083] Among them, m represents the number of CH nodes in the entire network, e r is the energy consumption of the receiving unit data, represents the energy consumption of data uploading.
[0084] Step3: Based on the information of the cluster head node received by the drone, input it into the P-SOM neural network to obtain the current optimal access sequence and the cluster head node with the highest current service priority.
[0085] The present invention introduces a penetration factor, and the neural network will retain more historical experiences, which helps to learn the long-term patterns of the data. The penetration mechanism prevents the network from overreacting to new inputs during the training process by controlling the retention ratio of the old weights, thus retaining part of the network.
[0086] SOM is a neural network that includes an input layer and an output layer. The output layer is usually a two-dimensional or one-dimensional neuron grid. The input layer is used to receive real-world patterns. Through training, the weight vectors of the output layer neurons gradually learn and map the patterns of the input data. In the present invention, the input vector is the two-dimensional coordinates of the city, and SOM learns the spatial position relationship of the city and uses a ring neuron structure to output the optimized path.
[0087] It should be noted that there seems to be an incorrect tag "!1" in the original text which is maintained as is in the translation. You may want to double-check the accuracy of the original text for such potential errors.When the traditional Self-Organizing Maps (SOM) neural network solves optimization problems, it finds the optimal solution to the problem by mapping the input to a low-dimensional network topology. Compared with optimization algorithms such as simulated annealing and genetic algorithms, the weight update depends on the winning neuron and its neighborhood, resulting in a significant improvement in the convergence speed. However, the original SOM has strong local search ability but insufficient global search ability, and the accuracy of the algorithm is relatively low. Therefore, this invention proposes Permeable-Self-Organizing Maps (P-SOM), introducing a permeation mechanism. On the basis of retaining the advantage of the convergence speed, the P-SOM neural network significantly improves the accuracy and adaptability of the algorithm, and introduces a permeation factor:
[0088]
[0089] where p max is the initial maximum permeation ratio, α and β respectively control the change of the permeation ratio, and t represents the number of iterations. At the initial stage of training, the permeation ratio is relatively large, and the neural network will retain more historical experience. As the training progresses, the permeation ratio gradually decreases, enabling the network to gradually shift to the update of the current data and helping the network to better adapt to new inputs. The permeation mechanism prevents the network from overreacting to new inputs during training by controlling the retention ratio of the old weights, thereby retaining part of the network. The newly designed update formula enables the weights far from the winning neuron to be gradually adjusted, achieving more comprehensive optimization. The formula is expressed as:
[0090] W n (t + 1) = p(t)·W n (t) + (1 - p(t))·G n,q (t)·η(t)·(x k - w n (t))(5)
[0091] In the above formula, W n (t) represents the weight of the nth neuron at the tth iteration, W n (t + 1) represents the weight of the nth neuron at the (t + 1)th iteration, η(t) represents the learning rate, and x k is the current input data sample. G n,q (t) represents the strength relationship of the influence on the surrounding neurons. Inputting the two-dimensional coordinates of all sensor nodes into the P-SOM neural network can obtain the current optimal path, the optimal access sequence, and the cluster head node with the highest current service priority. Inputting the two-dimensional coordinates of all sensor nodes into the P-SOM neural network can obtain the current optimal path, the optimal access sequence, and the cluster head node with the highest current service priority.
[0092] In P-SOM, the range that stimulates the excitation of surrounding neurons is called the "winning neighborhood". Within the winning neighborhood, the relationship between the strength of the effect and the distance can be expressed by the following formula:
[0093]
[0094] In the above formula, d n,q represents the distance between neuron n and the winning neuron q, and σ(t) is the radius of the winning neighborhood, whose value gradually decreases over time. It has a relatively large range in the initial stage and becomes smaller in the later stage for fine-tuning. It can be expressed by the following formula:
[0095]
[0096] In the above formula, σ0 is the initial neighborhood range, representing the neighborhood radius at the start of training. τ is a constant that controls the neighborhood decay rate, and t represents the number of iterations. Inputting the two-dimensional coordinates of all sensor nodes into the P-SOM neural network can obtain the current optimal path, the optimal access sequence, and the cluster head node with the highest current service priority.
[0097] Step4: Design a UAV path planning scheme based on the deep reinforcement learning algorithm DDPG. Through the DDPG algorithm, the UAV finds the cluster head node with the highest current service priority. When the UAV reaches within the data collection range of the cluster head node with the highest current service priority, the DDPG algorithm is used to adjust the descent height of the UAV, collect data from the cluster head node, and replenish the energy of the nodes within the range.
[0098] When using the DDPG algorithm to guide the UAV to find the target node, it maintains a fixed flight height, obtains the positions of each node before taking off from the base station, and maintains communication with the base station throughout the operation. Due to the complexity of calculating the exact power consumption of the rotorcraft, the present invention assumes that the blade section drag coefficient of the UAV is a constant. The propulsion power of a rotor UAV with a fixed height, speed V, and rotor thrust T can be:
[0099]
[0100] Among them, represents the thrust-to-weight ratio, W is the weight of the UAV, and Ω, d0, ρ, s, A, and R represent the blade angular velocity, fuselage drag ratio, air density, rotor solidity, rotor disk area, and rotor radius respectively. represents the blade profile power during hovering, represents the induced power during hovering.
[0101] When the UAV is descending vertically and ascending vertically, the rotor thrust T received by the UAV will change. The descending propulsion power of the UAV is expressed as:
[0102]
[0103] The climbing propulsion power is:
[0104]
[0105] Among them, represents the sum of the blade profile power and the induced power during hovering. Assume that the acceleration a is less than the gravitational acceleration g = 9.8 m / s 2 . Use the DDPG algorithm to plan the descending height of the UAV. Subsequently, the UAV hovers to collect data from the CH node and replenish energy for other nodes within the range. According to the Shannon formula, the data transmission rate between the UAV and the target node is expressed as:
[0106]
[0107] Among them, B and σ 2 represent the channel bandwidth and the noise power, P t is the node transmission power, and γ0 is the channel power gain between the UAV and the target node, which varies with the distance between the UAV and the target node. The hovering time is expressed as:
[0108]
[0109] Among them, is the data buffer length of the target node, and the energy consumed by the target node to upload data is:
[0110]
[0111] Assume that compared with the time for the UAV to collect data, the node battery can be charged in a very short time, and the charging time is ignored in the present invention. After the UAV completes data collection from the target node and energy replenishment for the sensor nodes within the range, the UAV climbs to the initial height and continues path planning to find the next target node, especially in scenarios where the sensor nodes are unevenly distributed or there are nodes with high energy requirements. The distance between the sensor node and the UAV is significantly shortened, thereby reducing the loss during energy transmission, improving the energy transfer efficiency, enabling the UAV to complete the data collection and energy replenishment tasks faster, and saving the energy consumption during UAV hovering. In addition, the descending height of the UAV can cover more sensor nodes. Especially in scenarios where the sensor nodes are unevenly distributed, the UAV can dynamically adjust the height according to the number of covered nodes. The specific parameter settings are shown in Table 1:
[0112]
[0113] Table 1. Parameter settings
[0114] In the DDPG algorithm, the drone interacts directly with the environment. The node with the highest priority output in step 3 is used as the target node. Combining the drone's own position and the relative position of the target node, the drone is made to shorten the distance to the target as much as possible. The problem is abstracted into a Markov decision process. Based on the framework of Markov process, the DDPG algorithm uses a four-tuple {s rl ,A rl ,R rl ,S rl ′}Definition, S rl is the state space, A rl is the action space, R rl is the reward function, S rl ′ It is the next state that the drone enters after performing an action. The specific definition is as follows:
[0115] State space S rl
[0116]
[0117] in, Expressed as the distance between the UAV and the target node, x u (t),y u (t),z u (t) represents the lateral position, longitudinal position and height of the UAV, N f (t) represents the number of times the drone flies out of the boundary, N d (t) represents the current number of dead sensor nodes;
[0118] Action space A rl
[0119] A rl ={v u (t),θ u (t),h u}
[0120] Among them, v u (t) represents the instantaneous speed of the UAV at time t, θ u (t) represents the instantaneous heading angle of the UAV along the horizontal direction at time t, h u Indicates the height to which the drone descends after finding the target node;
[0121] Reward function R rl
[0122]
[0123] R0(t)=R connect +η1λcover (t) + η2h u (15)
[0124]
[0125] Among them, R0(t) is the reward after the UAV establishes a connection with the target node, R1(t) is the reward when the UAV is searching for the target node, R connect represents the reward after the UAV establishes a connection with the target node, λ cover (t) represents the number of sensor nodes currently covered by the UAV, Δd u,tar (t) represents the distance between the UAV and the target node, R d is the data collection range of the UAV, and η1, η2, η3, and η4 are the reward weight factors for the optimization objectives respectively;
[0126] The DDPG network framework consists of four networks: the actor network, the critic network, the target actor network, and the target critic network. The main network and the target network share the same network structure. The actor network outputs the actions of the UAV based on the input state, and the critic network evaluates these actions to update the actor network. Use the random network parameters at time t to initialize the actor network and the critic network respectively, and copy the parameters at the same time t to initialize the target network Select the action space A according to the current policy rl , and obtain the reward R rl , and the environmental state becomes S rl ′ , and store {S rl , A rt , R rl , S rl ′} into the replay pool, sample quadruples from the replay pool, and use the target actor network μ ′ and the target critic network Q ′ to calculate the target value Y as follows:
[0127]
[0128] Among them, γ represents the discount factor. Use gradient descent to train the critic network to minimize the target loss L, as shown below:
[0129]
[0130] Among them, M is the number of training rounds. Calculate the sampled policy gradient to update the current actor network:
[0131]
[0132] Among them, J represents the cumulative reward of the total discount, represents the gradient of the network parameters of, represents the action output by the actor network according to the current state S rl of, represents the action A generated by the current policy rl = μ(S rl ), calculate the gradient of the Q network for the action, and finally update it in a soft update manner, introducing the learning rate τ, as follows:
[0133]
[0134] Among them and are the updated target network parameters.
[0135] Step5: When the energy of the UAV is about to run out and it cannot complete the next task, ensure that the UAV can return to the base station, then replace the battery of the UAV and store the data, and then continue to execute the task. In addition, when the UAV completes a round of tasks of accessing all clusters, a new cluster head will be reselected according to the cluster head selection strategy in Step2, and the optimal sequence that the UAV needs to access in the next round will be updated.
[0136] The present invention proposes a cluster head selection strategy based on an improved dynamic routing protocol to balance the energy consumption of sensor nodes. Subsequently, the P-SOM neural network algorithm is proposed to solve the cluster head access sequence to minimize the flight energy consumption of the UAV. In addition, the present invention uses the deep reinforcement learning DDPG algorithm to design a multi-objective optimization reward function to control the instantaneous speed and heading of the UAV, guide the UAV to approach the position of the cluster head, and dynamically adjust the descent height. During this process, the UAV not only completes data collection, but also replenishes the energy of the sensor nodes within the charging range, effectively avoiding the failure of nodes due to power exhaustion. At the same time, this design significantly improves the overall operating life and reliability of the network by minimizing the mortality rate of sensor nodes and the average flight energy consumption of the UAV.
[0137] The present invention solves the problems that the flight height of the UAV is pre-fixed in the existing method for data collection and energy replenishment of the UAV-assisted wireless sensor network, and the influence of the UAV height on the relative distance between the UAV and the sensor is not deeply considered. Therefore, it significantly optimizes the distance between the UAV and the sensor for charging and data collection, improves the energy utilization efficiency, improves the energy replenishment and data collection efficiency of the network, and extends the network survival time.
[0138] The above description is only a specific idea of the present invention for the understanding of researchers. However, the present invention is not limited to the above embodiments. Those skilled in the relevant art can make improvements, modifications and substitutions based on the present invention. All improvements or variations made using the concept of the present invention are regarded as the protection scope of the present invention.
Claims
1. A three-dimensional flight trajectory planning method for an unmanned aerial vehicle for data collection and energy replenishment in a wireless sensor network, characterized in that: Including the following steps: Step1: Establish a wireless sensor network model: Randomly deploy sensor nodes in the network. CH nodes are represented as {ch1, ch2, ch3, …, ch m}, CM nodes are represented as {cm1, cm2, cm3, …, cm h}, and at the same time, deploy a drone serving the sensor nodes and a base station for data storage and battery replacement of the drone; Step2: Design a dynamic cluster head selection mechanism based on a routing protocol, and adopt a cluster head selection strategy to optimize the position of the cluster head; Step3: The drone inputs the information of the received cluster head nodes into the P-SOM neural network to obtain the current optimal access sequence and the cluster head node with the highest current service priority; Step4: Design a drone path planning scheme based on the deep reinforcement learning algorithm DDPG. Through the DDPG algorithm, the drone finds the cluster head node with the highest current service priority. When the drone reaches within the data collection range of the cluster head node with the highest current service priority, the DDPG algorithm is used to adjust the descending height of the drone, collect data from the cluster head node, and supplement the energy of the nodes within the range; Step5: When the energy of the drone is about to run out and it cannot complete the next task, ensure that the drone can return to the base station, replace the battery of the drone and store the data, and then continue to execute the task. In addition, when the drone completes a round of access tasks for all clusters, a new cluster head will be reselected according to the cluster head selection strategy in Step2, and the optimal sequence that the drone needs to access in the next round will be updated.
2. The method for three-dimensional flight trajectory planning of an unmanned aerial vehicle for data collection and energy replenishment in a wireless sensor network according to claim 1, wherein: The specific steps for designing the dynamic cluster head selection mechanism based on the routing protocol in Step2 are as follows: The density peak clustering algorithm DPC is used to cluster the sensor nodes. In each round of election, each cluster node elects a CH node according to its own energy state and location information to balance the energy consumption. A dynamic routing protocol is adopted within each cluster. The cluster member CM nodes use the cluster head as the next-hop relay, and the cluster head collects the member data and transmits it to the drone. The weight of the CM node is obtained by the following formula: In the above formula, E max represents the maximum energy of the sensor node, represents the remaining energy of the CM node at time t, and E ave (t) represents the average remaining energy of the CM node. α is a weighting coefficient, representing the priority and importance of distance and energy. d i,j is expressed in cm i between the node and ch j and the distance between them is d ave represents the average distance between the CM node and the CH node; The total energy consumed by the CM node to send a data packet to the CH node is Among them, h represents the number of CM nodes in a cluster, and e t represents the energy consumption of transmitting unit data, and q i,j represents the size of a unit data packet sent by a CM node to a CH node, and the energy consumption of the CH node is divided into the energy consumption of receiving data packets and the energy consumption of uploading data to the drone: Among them, m represents the number of CH nodes in the entire network, and e r is the energy consumption of the receiving unit for data, represents the energy consumption of data uploading.
3. The three-dimensional flight trajectory planning method of the unmanned aerial vehicle for data collection and energy replenishment in the wireless sensor network according to claim 1, wherein: The specific steps of Step3 are as follows: Introduce a penetration mechanism into the SOM neural network to form a P-SOM neural network: where p(t) is the penetration factor, p max is the initial maximum penetration ratio, α and β are used to control the change of the penetration ratio respectively, t represents the number of iterations. In the initial stage of training, the penetration ratio is large, and as the training progresses, the penetration ratio gradually decreases. The update formula is expressed as: W n (t + 1) = p(t)·W n (t) + (1 - p(t))·G n,q (t)·η(t)·(x k - w n (t)) (5) In the above formula, W n (t) represents the weight of the nth neuron at the tth iteration, and W n (t + 1) represents the weight of the nth neuron at the (t + 1)th iteration. η(t) represents the learning rate, and x k is the current input data sample. G n,q (t) represents the strength relationship of the influence on surrounding neurons. In P-SOM, the range that promotes the excitation of surrounding neurons is called the "winning neighborhood". Within the winning neighborhood, the relationship between the strength of the influence and the distance is expressed by the following formula: In the above formula, d n,q represents the distance between neuron n and the winning neuron q. σ(t) is the radius of the winning neighborhood, whose value gradually decreases over time. It has a relatively large range in the initial stage and becomes smaller in the later stage for fine-tuning mode, which can be expressed by the following formula: In the above formula, σ0 is the initial neighborhood range, representing the neighborhood radius at the beginning of training. τ is a constant that controls the neighborhood decay rate, and t represents the number of iterations. The two-dimensional coordinates of all sensor nodes are input into the P-SOM neural network to obtain the current optimal path, the optimal access sequence, and the cluster head node with the highest current service priority.
4. The three-dimensional flight trajectory planning method for the UAV used in wireless sensor network data collection and energy replenishment according to claim 1, wherein: The specific steps of Step4 are as follows: Assume that the drag coefficient of the drone blade cross-section is a constant. The propulsion power of a rotor drone with a fixed height, speed V, and rotor thrust T is: Among them, represents the thrust-to-weight ratio, W is the weight of the UAV, and Ω, d0, ρ, s, A, and R represent the blade angular velocity, fuselage drag ratio, air density, rotor solidity, rotor disk area, and rotor radius respectively; represents the blade profile power during hovering, represents the induced power during hovering; When the drone is descending vertically and ascending vertically, the rotor thrust T received by the drone will change. The descending propulsion power of the drone is expressed as: The climbing propulsion power is: Among them, represents the sum of the blade power and the induced power during hovering. Assuming that the acceleration a is less than the gravitational acceleration g = 9.8 m / s 2 , the DDPG algorithm is used to plan the descending height of the UAV. Subsequently, the UAV hovers to collect data from the CH node and supplement energy to other nodes within the range. According to the Shannon formula, the data transmission rate between the UAV and the target node is expressed as: Among them, B and σ 2 represent the channel bandwidth and the noise power, P t is the node transmission power, γ0 is the channel power gain between the UAV and the target node, which varies with the distance between the UAV and the target node, and the hovering time is expressed as: Among them, is the data buffer length of the target node, and the energy consumed by the target node to upload data is: Assume that compared with the data collection time of the drone, the node battery can complete charging in a very short time, and the charging time is ignored. After the drone completes the data collection of the target node and the energy supplement of the sensor nodes within the range, the drone climbs to the initial height and continues to perform path planning to find the next target node, especially in scenarios where the sensor nodes are unevenly distributed or there are high energy requirements; In the DDPG algorithm, the UAV directly interacts with the environment. Taking the current optimal access sequence output in step 3 and the cluster head node with the highest current service priority as the target nodes, and combining the position of the UAV itself and the relative positions of the target nodes, the UAV tries to shorten the distance to the target as much as possible. The problem is abstracted into a Markov decision process. Based on the framework of the Markov process, the DDPG algorithm is defined by a quadruple {S rl , A rl , R rl , S rl ′}. S rl is the state space, A rl is the action space, R rl is the reward function, and S rl ′ is the next state that the UAV enters after executing an action. The specific definitions are as follows: State space S rl Among them, represents the distance between the drone and the target node, x u (t), y u (t), z u (t) represent the lateral position, longitudinal position and altitude of the drone, N f (t) represents the number of times the drone flies out of the boundary, N d (t) represents the current number of sensor dead nodes; Action space A rl A rl = {v u (t), θ u (t), h u} Among them, v u (t) represents the instantaneous velocity of the UAV at time t, and θ u (t) represents the instantaneous heading angle of the UAV in the horizontal direction at time t, and h u represents the height by which the UAV descends after finding the target node; Reward function R rl R0(t) = R connect + η1λ cover (t) + η2h u (15) Among them, R0(t) is the reward after the UAV establishes a connection with the target node, R1(t) is the reward when the UAV is searching for the target node, R connect represents the reward after the UAV establishes a connection with the target node, λ cover (t) represents the number of sensor nodes currently covered by the UAV, Δd u,tar (t) represents the distance between the UAV and the target node, R d is the data collection range of the UAV, and η1, η2, η3, and η4 are the reward weight factors of the optimization objectives respectively; The DDPG network framework consists of four networks: the actor network, the critic network, the target actor network, and the target critic network. The main network and the target network share the same network structure. The actor network outputs the actions of the drone based on the input state, and the critic network evaluates these actions to update the actor network with the random network parameters at time t. Initialize the actor network and the critic network respectively, and copy the parameters at the same time t to initialize the target network respectively. Select the action space A according to the current policy. rl , obtain the reward R. rl , and the environmental state becomes S rl ′. Store {S rl , A rl , R rl , S rl ′} into the replay pool, sample quadruples from the replay pool, and calculate the target value Y using the target actor network μ′ and the target critic network Q′ as follows: Where γ represents the discount factor, and the critic network is trained using gradient descent to minimize the target loss L, as shown below: Where M is the number of training rounds, calculate the sampled policy gradient to update the current actor network: where J represents the cumulative discount reward, represents the gradient of the network parameters with respect to, represents the action output by the actor network according to the current state S rl ; represents the gradient of the Q network with respect to the action calculated at the action A rl = μ(S rl ) generated by the current policy, and finally updated using soft update, introducing the learning rate τ, as follows: Among them and are the updated target network parameters.
Citation Information
Cited By
Autonomous inspection and charging nest system based on unmanned aerial vehicle ad hoc network
CN120993933A
An autonomous inspection and charging nest system based on unmanned aerial vehicle ad hoc network
CN120993933B
Charging scheduling method for underwater wireless rechargeable sensor network
CN122411473A
An underwater wireless rechargeable sensor network charging scheduling method
CN122411473B