Multi-unmanned aerial vehicle auxiliary positioning and data collection method based on deep reinforcement learning
Through deep reinforcement learning, optimize the positioning and data collection methods of multi-UAVs, the positioning difficulties and obstacle avoidance problems of multi-UAV assisted data collection in wireless sensor networks are solved, efficient and accurate sensor node positioning and data collection are achieved, and real-time and reliability of data collection are improved.
Patent Information
- Application Number
- CN202510357066.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-04
AI Technical Summary
In wireless sensor networks, multi-UAV assisted data acquisition faces problems such as difficulty in positioning, complex obstacle avoidance and high energy consumption. Especially in the absence of prior information, traditional methods are difficult to efficiently realize the precise positioning and data collection of sensor nodes.
The collaborative positioning and data collection method of multi-UAV based on deep reinforcement learning is adopted. By building a multi-UAV positioning and data acquisition model, the geometric layout and flight trajectory of the drone are optimized, combined with improved depth deterministic strategy gradient algorithm and simulated annealing algorithm, the access sequence and obstacle avoidance strategy are optimized to achieve efficient data acquisition and obstacle avoidance.
The centimeter-level precise positioning of sensor nodes in unknown environments is achieved, the positioning error is reduced by more than 40%, the coordinated mission time of multiple drones is shortened by 30%, the collision risk is reduced by 90%, and the real-time performance of data acquisition in target areas is improved by 50%.
Smart Images

Figure CN120255535A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of communication technologies and relates to a multi-UAV assisted positioning and data collection method based on deep reinforcement learning. Background Art
[0002] The Internet of Things (IoT) technology, with its characteristics of ubiquitous interconnection, intelligent perception, real-time processing, and massive data interaction, can connect millions of devices to the Internet and is widely used in fields such as intelligent transportation, industrial IoT, and environmental monitoring. The IoT technology architecture is usually divided into four layers: the application layer, the network layer, the platform layer, and the perception layer. Among them, WSN is an important part of the IoT perception layer and consists of SNs and sink nodes. Among them, SNs sense environmental parameters, and the sink node is responsible for collecting data from SNs and transmitting it to the DC.
[0003] In recent years, WSN has been widely used in large-scale distributed monitoring scenarios due to its technical advantages such as rapid deployment, self-organization, and low power consumption. However, in scenarios where communication facilities are lacking or damaged, data collection in WSN becomes difficult. UAVs, with their characteristics of high mobility and wide coverage, are suitable as data collectors in WSN. At the same time, multiple UAVs can cooperate for data collection to improve the collection efficiency. However, UAVs consume a large amount of energy during flight and hovering, and the scheduling and trajectory planning of multiple UAVs are very complex. In addition, in scenarios where prior information is missing, UAVs need to locate SNs and avoid obstacles in the scenario, and there are still challenges in multi-UAV assisted WSN data collection. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a multi-UAV assisted positioning and data collection method based on deep reinforcement learning.
[0005] To achieve the above purpose, the present invention provides the following technical solutions:
[0006] The present invention provides a multi-UAV assisted positioning and data collection method based on deep reinforcement learning for the scenario of multi-UAV assisted WSN data collection where the accurate positions of SNs and the position information of obstacles in the target area are missing. The method includes the following steps:
[0007] S1: Construct a multi-UAV assisted sensor positioning and data collection model;
[0008] S2: Design an optimization scheme for average positioning accuracy and total task time;
[0009] S3: Optimize the geometric layout of UAVs and determine the relative angles and distances between UAVs;
[0010] S4: Optimize the access order of the head sensor distribution areas, and determine the access sequence of the UAV to the head sensor distribution areas;
[0011] S5: Optimize the UAV flight trajectory, and determine the flight trajectory that avoids obstacles in the avoidance scenario and quickly reaches the target position.
[0012] Optionally, in S1, construct a multi-UAVs assisted positioning and data collection model in the scenario where environmental prior information is missing. This model consists of WSN, DC, and UAVs. The SNs in WSN are divided into M-SNs responsible for sensing environmental data and H-SNs responsible for collecting the sensed data of M-SNs and uploading it to UAVs. M-SNs and H-SNs are scattered and deployed by the aircraft, and their exact positions are unknown. Dispatch one M-UAV and two F-UAVs. The M-UAV hovers as an anchor point above the center of the H-SNs distribution area, and the F-UAVs are deployed near the M-UAV under the limitation of the maximum communication distance of H-SNs. When not activated, H-SNs will emit positioning signals with a relatively low power. After receiving the positioning signals, the F-UAVs send the arrival time of the positioning signals to the M-UAV. Subsequently, the M-UAV calculates the positions of H-SNs through the time differences of the positioning signals arriving at the three UAVs, and sends the positions to the F-UAVs. Then, the UAVs send activation signals to activate the H-SNs, and select the UAV closest to the H-SNs to collect their data. In addition, the UAVs observe and avoid obstacles in the flight path and record their positions. After completing all positioning and data collection tasks, the UAVs return to the DC to unload the sensed data of the SNs, the exact positions of the H-SNs, and the obstacle position information. The DC is responsible for analyzing and processing the sensed data of the SNs to achieve functions such as environmental monitoring and disaster warning of the target area. Subsequently, the DC can deploy and plan the paths of the UAVs based on the complete environmental information, so that the UAVs can perform subsequent rounds of data collection more efficiently and achieve more accurate monitoring of the target area.
[0013] Optionally, in S2, design an optimization scheme for the average positioning accuracy and the total task time. First, optimize the geometric layout of the UAVs by optimizing the relative distance and angle between the F-UAVs and the M-UAV to maximize the average positioning accuracy of locating H-SNs using the TDoA method. Since the data volumes of the positioning signals sent by the H-SNs, the activation signals sent by the UAVs, and the communication signals between the UAVs are very small and the signal transmission time can be ignored, the present invention considers that the total time of multi-UAVs assisted data collection consists of the flight time of the UAVs and the data collection time. Subsequently, the present invention jointly optimizes the access order of the UAVs to the H-SNs and the flight trajectories of the UAVs, and while ensuring that the UAVs avoid obstacles in the trajectory, minimizes the total task time of multi-UAVs assisted data collection.
[0014] Optionally, in S3, a multi-UAVs assisted positioning and deployment algorithm based on improved DDPG is proposed. By optimizing the geometric layout of UAVs, the average positioning accuracy of H-SNs based on the TDoA method is maximized. The algorithm uses the Cramér-Rao Lower Bound (CRLB) that is unbiased estimation of the H-SN position to measure the positioning accuracy. The smaller the CRLB, the higher the positioning accuracy. First, according to the maximum communication distance of H-SNs, the radius of the distribution area of H-SNs, and the flight altitude of UAVs, the radius of the UAV deployment area above the H-SNs is determined. Subsequently, the problem of optimizing the geometric layout of UAVs is equivalent to the problem of determining the positions of F-UAVs, and it is modeled as an MDP. Finally, taking the F-UAVs as agents, an improved DDPG algorithm with enhanced exploration is used to train the agents. After the algorithm converges, the optimal distances and angles of the F-UAVs relative to the M-UAV are obtained, that is, the geometric layout of UAVs that maximizes the average positioning accuracy of H-SNs.
[0015] Optionally, in S4, a UAV trajectory planning algorithm based on SA and DDPG is proposed. First, the problem of determining the optimal visiting order of UAVs to H-SNs is modeled as a Traveling Salesman Problem (TSP). Subsequently, taking the straight-line flight time of the M-UAV above the center of the H-SNs area as the cost, the SA algorithm is used to solve the optimal visiting order in which the UAVs visit all areas without repetition and with the shortest flight time.
[0016] Optionally, in S5, based on the optimal visiting order determined in S4, the UAV trajectory planning problem is modeled as an MDP with a continuous action space. Subsequently, a UAV obstacle avoidance strategy is proposed: when the UAV does not observe an obstacle, it maintains a normal flight state. When the UAV observes an obstacle during flight, it makes a judgment: if the UAV will not collide with the obstacle if it continues to fly at the current heading angle, no avoidance action is taken; otherwise, if the UAV will collide with the obstacle if it continues to fly at the current heading angle, an avoidance action is taken. If the avoidance action makes the UAV move away from the obstacle, a reward is given; otherwise, a penalty is imposed. Finally, a DDPG-based algorithm is used to train the UAV. After the algorithm converges, the optimal flight trajectory that ensures the UAV avoids obstacles and quickly reaches the target position is obtained, and then the total task time of UAV-assisted data collection is calculated in combination with the total data collection time.
[0017] The beneficial effects of the present invention are as follows:
[0018] (1) Optimization of the geometric layout of unmanned aerial vehicles based on the improved DDPG algorithm realizes centimeter-level precise positioning of unknown head sensor nodes by maximizing the positioning accuracy under the Cramér-Rao lower bound constraint, with an error reduction of more than 40% compared to traditional methods;
[0019] (2) The trajectory planning technology that integrates simulated annealing and deep reinforcement learning shortens the collaborative task time of multiple unmanned aerial vehicles by 30% and reduces the collision risk by 90% through a dynamic obstacle avoidance strategy, effectively solving the problems of high energy consumption and path redundancy;
[0020] (3) Through the end-to-end continuous action space optimization model, it breaks through the local optimum bottleneck of traditional algorithms and realizes global optimum path search in unknown obstacle scenarios, improving the real-time performance of data collection in the target area by 50%, providing a highly reliable and low-latency intelligent solution for disaster warning and environmental monitoring.
[0021] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. Brief Description of the Drawings
[0022] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:
[0023] Figure 1 It is a model diagram of a multi-UAVs assisted positioning and data collection system;
[0024] Figure 2 It is a flowchart of a multi-UAVs assisted sensor positioning and data collection method based on deep reinforcement learning. Detailed Embodiments
[0025] The following illustrates the embodiments of the present invention through specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention schematically. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0026] Among them, the attached drawings are only for illustrative purposes, showing only schematic diagrams rather than physical diagrams, and should not be construed as limiting the present invention; in order to better illustrate the embodiments of the present invention, some components in the attached drawings will be omitted, enlarged or reduced, which do not represent the dimensions of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the attached drawings may be omitted.
[0027] In the attached drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the attached drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the attached drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0028] Figure 1 Fig. shows a possible structural schematic diagram of the UAVs-assisted positioning and data collection system in the scenario of missing environmental prior information involved in the embodiments of the present invention. As Figure 1 shown, the system consists of a WSN, a DC, and UAVs. The SNs in the WSN are divided into M-SNs and H-SNs responsible for sensing environmental data. Three UAVs are dispatched, divided into M-UAV and F-UAV. The M-UAV hovers as an anchor point above the center of the area where the H-SNs are distributed, and the F-UAVs are deployed near the M-UAV under the limitation of the maximum communication distance of the H-SNs. The three UAVs use a TDoA-based positioning method to locate the H-SNs when accessing the area where the H-SNs are distributed. Subsequently, the UAV closest to the H-SNs is selected to collect data. In addition, the UAVs observe and avoid obstacles in the flight path and record their positions. After completing all the positioning and data collection tasks, the UAVs return to the DC to unload the SNs sensing data, the accurate positions of the H-SNs, and the obstacle position information. The DC is responsible for analyzing and processing the SNs sensing data to achieve functions such as environmental monitoring and disaster warning of the target area.
[0029] Figure 2 The multi-UAVs-assisted positioning and data collection method based on deep reinforcement learning shown in Fig. includes the following steps:
[0030] S201: Construct a multi-UAVs-assisted positioning and data collection scenario consisting of a WSN, a DC, and UAVs;
[0031] Among them, the target area is denoted as Λ. The SNs in the target area are scattered and deployed by an aircraft and consist of M-SNs and H-SNs. Among them, the M-SNs are responsible for environmental perception, while the H-SNs are responsible for collecting data from the M-SNs and uploading it to the UAVs, and their exact positions are unknown. There are N distribution areas of H-SNs with a radius of R in the target area Λ A denoted as the set A = {A1,..., A n ,..., A N}, and the central coordinate of A n is denoted as Pos(A n ) = (x An , y An , 0). Assume that in each H-SN distribution area, an H-SN located on the horizontal plane is deployed, and its exact position is unknown. The maximum communication distance of the H-SNs is denoted as d max . The set of H-SNs in the scenario is denoted as S n and the coordinates are denoted as To reduce the energy consumption of the H-SNs, the H-SNs generally operate in a low-power state, and only emit electromagnetic wave signals with a small power while sensing environmental data to facilitate the UAVs to locate them. The UAVs fly at a fixed height H, and the set of UAVs is denoted as UAVu m and the coordinates are denoted as Pos(u m ) = (x m , y m , H). In addition, d max , R A and H need to satisfy to ensure that there exists an area such that the UAVs can receive the positioning signals emitted by the H-SNs at any position in A n when deployed at any position in this area. At this time, the radius of the UAVs deployment area
[0032] In the scenario proposed by the present invention, the UAVs consist of one M-UAV and two F-UAVs. The M-UAV and the F-UAVs depart from the DC and fly over the area where the H-SNs are distributed in the order of the shortest flight time. Among them, the M-UAV hovers directly above the center of the area as an anchor point. After the UAVs receive the positioning signals transmitted by the H-SNs in the current area, the F-UAVs send the time of arrival of the signals to the M-UAV. Subsequently, the M-UAV locates the H-SN through the time difference of arrival of the signals and sends its position to the F-UAVs. Then, the UAV closest to the current H-SN is selected to collect data. At this time, the UAV sends an activation signal to the H-SN. After being activated, the H-SN sends data with a higher power, and the other two UAVs fly to the next area where the H-SNs are distributed to wait. In addition, the UAVs will avoid the obstacles in the scenario during the flight path to the target position and record their positions and sizes. Finally, after the UAVs complete the positioning and data collection of the H-SNs in all areas where the H-SNs are distributed, they return to the DC and unload the environmental information such as the data sensed by the SNs, the positions of the H-SNs, the positions and sizes of the obstacles to the DC.
[0033] S202: Establish a multi-UAVs assisted positioning model and construct an average positioning accuracy optimization problem;
[0034] Denote the M-UAV as u1, and the two F-UAVs as u2 and u3 respectively. Taking the M-UAV as a reference, the time difference of the F-UAVs receiving the H-SN positioning signal is expressed as:
[0035]
[0036] where is the actual time when the UAV u m receives the S n positioning signal. Therefore, by solving the hyperbola equation with the positions of the two F-UAVs as the foci and the distance difference calculated by TDoA as the major axis, the coordinates of the estimated position of the H-SN can be obtained. The hyperbola equation determined by TDoA is expressed as follows:
[0037]
[0038] where c is the electromagnetic wave propagation rate. Let be the value of the distance difference caused by the arrival time measurement error, which conforms to a zero-mean Gaussian distribution. The CRLB for H-SN positioning based on the TDoA method is expressed as follows:
[0039] CRLB(S n ) = c 2 ·trace((HQ -1 HT ) -1 ) (3)
[0040] where \(Q = E\{ee T \}\) is the error covariance matrix, and \(
[0041]
[0042] is the Jacobian matrix of the hyperbola equation of the hyperbola, which is expressed as follows:
[0043]
[0044] Based on this, an average positioning accuracy optimization problem is constructed, which maximizes the average positioning accuracy of H-SNs positioning by optimizing the geometric layout of UAVs. The optimization objective function is expressed as:
[0045] S203: Establish a total task time model for multi-UAV assisted data collection and construct an optimization problem for the total task time of multi-UAV assisted data collection.
[0046] Let the data transmission power of H-SNS n be and the average path loss between UAV u m and H-SN S n be PL m,n . Then the received power of the UAV is expressed as follows:
[0047]
[0048] S n to u m The data transmission bandwidth is expressed as follows:
[0049]
[0050] where B is the uplink data transmission bandwidth and \(\sigma 2 is the additive white Gaussian noise power. Let the amount of data sensed by H-SNS n be S. Then the time for the UAV to collect S n data is expressed as follows:
[0051]
[0052] Since the data volumes of the positioning signals sent by H-SNs, the activation signals sent by UAVs, and the UAV-to-UAV communication signals are very small, the signal transmission time can be ignored. Therefore, the total task time of UAVs-assisted positioning and data collection mainly consists of two parts, namely, the flight time of UAVs from the DC to visit all the H-SNs distribution areas and then return to the DC, and the total data collection time of UAVs.
[0053] Let UAV u m fly from area A n to area A n+1 and the time be denoted as The position of the DC is A0. In area A n , the association vector between UAVs and H-SNs is denoted as ω n = [ω n,1 , ω n,2 , ω n,3 . When ω n,m = 1, it means that UAV u m is responsible for collecting the data of S n ; when ω n,m = 0, it means that UAV u m does not collect data. The association matrix of UAVs for all areas is denoted as Ω = [ω1,..., ω N T . u m' is the set of UAVs not responsible for data collection, is the set of the flight times of UAVs not responsible for data collection from one area to another area. Then the task time of multi-UAVs-assisted sensor positioning and data collection is expressed as:
[0054]
[0055] On this basis, an optimization problem of the total task time of multi-UAV-assisted data collection is constructed. The present invention minimizes the total task time of multi-UAVs-assisted data collection by jointly optimizing the access order of UAVs to the H-SNs distribution areas and the flight trajectories of UAVs among the H-SNs distribution areas. Let the access order of UAVs to the H-SNs distribution areas be expressed as the set where q0 is the position of the DC, and all elements in q1 to q N area set are arranged in a certain order. Let the trajectory matrix of three UAVs be Tr = [tr1, tr2, tr3] T , tr1 = [tr10,..., tr1 n ,..., tr1 N . Among them, tr1 n is the trajectory of the UAV from q n to qn+1 The flight trajectory, tr1 N is the flight trajectory of the UAV returning from u1 to DC. The optimization objective function is expressed as:
[0056]
[0057] The optimization objective needs to satisfy the following constraints: Each H-SN is only data-collected by one UAV; UAVs do not collide with obstacles; UAVs do not fly out of the target area.
[0058] S204: For the average positioning accuracy optimization problem, a multi-UAVs assisted positioning and deployment algorithm based on improved DDPG is proposed to train the deployment positions of F-UAVs and obtain the UAVs layout that minimizes the average CRLB;
[0059] The present invention equivalently transforms the UAVs geometric layout optimization problem into the problem of searching for the optimal deployment positions of F-UAVs and models it as an MDP in a continuous action space.
[0060] State: Since the flight altitude of UAVs is fixed, the horizontal and vertical coordinates of F-UAVu m at the τ-th step of a round of training are used as its state at the τ-th step, denoted as m ∈ {2, 3}. The horizontal and vertical coordinates of two F-UAVs are used as the joint state of the agent, denoted as
[0061] Action: The action of F-UAVu m at the τ-th step is denoted as m ∈ {2, 3}. Among them, the action value has an interval, The joint action of F-UAVs at the τ-th step is denoted as
[0062] Reward: To enhance the universality of the optimal UAVs geometric layout generated by the algorithm, the present invention samples K coordinates in the H-SN distribution area as the sampling positions of H-SNs. The reciprocal of the average CRLB of the F-UAVs at the sampling positions of H-SNs is taken and weighted as the reward term. The average CRLB is expressed as
[0063] In addition, during the training process, if the agent's position exceeds the UAVs deployment area, the agent will calculate the average CRLB based on the intersection position of the line connecting F-UAVs and M-UAVs with the boundary of the UAVs deployment area, denoted as Meanwhile, to encourage the agent to return to the deployment area, the weighted distance of the agent exceeding the boundary of the UAVs deployment area is introduced as a penalty term. Let the UAVs deployment area above the H-SNs distribution area A n be In summary, the reward function is expressed as follows:
[0064]
[0065] where θ m ∈ [0, 1], when the position of F - UAVu m exceeds then θ m = 1, otherwise θ m = 0. represents the distance that F - UAVu m exceeds the boundary. λ1 is the reward factor and λ2 is the penalty factor.
[0066] The traditional DDPG algorithm adopts an actor - critic architecture, which consists of a critic network, an actor network, and an experience replay pool. It uses a deterministic policy for action selection, which leads to a lack of sufficient randomness during the exploration process. The present invention optimizes on the basis of the traditional DDPG framework. When the number of samples ψ in the experience replay pool is less than the sampling number Ψ, the agent updates the state using random actions. This enables the agent to search more widely in the state space, thereby avoiding premature convergence to local optimal policies, helping the agent discover more potential high - reward regions, and enhancing its exploration efficiency. The present invention applies the enhanced algorithm to the problem of searching for the optimal position of F - UAVs. After the algorithm converges, the optimal deployment positions Pos(u2) * and Pos(u3) * are obtained, and then the optimal angle φ * between F - UAVs and M - UAV and the relative distance are obtained. That is, the geometric layout of UAVs that maximizes the average positioning accuracy.
[0067] S205: For the problem of optimizing the total task time of multi - UAVs assisted data collection, a UAV trajectory planning algorithm based on SA and DDPG is proposed to optimize the UAV flight trajectory. First, the algorithm models the problem of determining the optimal access order of UAVs to the distribution area of H - SNs as a TSP problem, and uses the SA algorithm to solve the optimal access sequence;
[0068] Under the condition of knowing the number N of H - SNs distribution areas and the central positions of the areas, three UAVs start from the DC and visit all H - SNs distribution areas without repetition, aiming to solve the optimal access sequence that minimizes the flight time of UAVs. The flight time of M - UAV between the centers of the UAVs deployment areas is used as the cost. To minimize the flight energy consumption of UAVs, the SA algorithm is used to determine the optimal access sequence. A random initial access sequence is formed. Subsequently, new access sequences are generated with equal probability using the swap method, the insertion method, and the reverse order method. At temperature T, the cost of the new path is less than or equal to the cost of the current path When it is, accept the new path. Otherwise, accept the new path with probability Poss(T), where Poss(T) is expressed as follows:
[0069]
[0070] where T0 is the initial temperature and μ is the temperature decay coefficient. After the SA algorithm converges, the optimal access order of the H-SNs distribution regions is obtained
[0071] S206: After determining Combine the optimal geometric layout obtained from the improved DDPG-based multi-UAVs positioning and deployment algorithm to determine the UAVs deployment positions above each H-SNs distribution region. Subsequently, model the UAVs trajectory optimization problem as an MDP with a continuous action space, and use the DDPG algorithm to train the UAVs trajectory. After the algorithm converges, the flight trajectory with the shortest flight time and ensuring that the UAVs avoid obstacles is obtained;
[0072] Each UAV is regarded as an agent, and the DDPG algorithm is used to optimize the set of trajectories for it to move between H-SNs distribution regions according to the optimal access sequence The time scale of each time step is set to 0.1 s. The MDP modeling process is as follows:
[0073] State: The heading angle φ of the UAV is set as the angle between the UAV flight direction and the positive y-axis direction. The state of UAVu m at time step τ consists of its horizontal coordinate and the current heading angle, denoted as
[0074] Action: The action of the UAV at time step τ is the change in the heading angle When the heading angle changes clockwise, Δφ m,τ has a positive sign, and vice versa.
[0075] Reward: Encourage the UAVs to quickly reach the target deployment position while avoiding obstacles. The reward function proposed in this section consists of a transition reward term, an obstacle avoidance reward term, and a time penalty term. The transition reward r at time step τ trans is expressed as:
[0076]
[0077] where λ3 and λ4 are the transition reward weighting coefficients, λ3 is a negative constant, and λ4 is a positive constant. is the straight-line distance of UAVu m from the target deployment point at time step τ. is the change in the straight-line distance of UAVu m from the target deployment point at time step τ compared to the previous time step,
[0078] The transition reward term consists of a constant transition reward term and a target arrival incentive term. The constant transition reward term is designed based on the change in the distance between the UAV and the target deployment point. When is less than 0, it indicates that the UAV is closer to the target deployment point compared to the previous time step. At this time, the value of the constant transition reward term is positive to encourage the UAV to fly towards the target deployment point. Conversely, the value of the constant transition reward term is negative to impose a penalty on the UAV to prevent it from moving further away from the target deployment point. However, the value of the constant reward term does not change significantly with the increase of time steps in a round of training, which results in it only encouraging the UAV to fly towards the target deployment point but being difficult to accurately reach the target deployment position. To make up for this deficiency, a target arrival incentive term is introduced into the transition reward term, which is the reciprocal of multiplied by the coefficient λ4. Due to the characteristics of the inverse proportional function, when the UAV gradually approaches the target deployment position, the value of the target arrival incentive term will increase significantly, thus generating a stronger incentive for the UAV to reach the target deployment point in the later stage of a round of training. At the same time, in the initial stage of a round of training, since is larger, the value of the target arrival incentive term will be relatively small. If the transition reward term only contains the target arrival incentive term, it may lead to a decrease in the convergence speed of the algorithm. Therefore, the transition reward term needs to combine the constant transition reward term and the target arrival incentive term to more strongly encourage the UAV to accurately reach the target deployment point while ensuring that the UAV flies towards the target deployment point.
[0079] The obstacle avoidance strategy of the present invention for UAVs is designed as follows: When the UAV does not observe an obstacle, it maintains a normal flight state. When the UAV detects an obstacle within the observation range, the following judgment is made: If the UAV will not collide with the obstacle when flying continuously at the current heading angle, no avoidance action is taken; conversely, if the UAV will collide with the obstacle when flying continuously at the current heading angle, an avoidance action is triggered. If the avoidance action makes the UAV move away from the obstacle, a reward is given; otherwise, a penalty is imposed. Based on the above obstacle avoidance strategy, the obstacle avoidance reward term at time step τ is expressed as follows:
[0080]
[0081] where the negative constant λ5 is the obstacle avoidance reward coefficient. is the distance between UAVu m and the center of the obstacle at time step τ. ξ is a binary variable. If the UAV will not collide with the obstacle when flying continuously at the current heading angle, then ξ = 0; otherwise, ξ = 1. is the change amount of at time step τ. When the UAV observes an obstacle and ξ = 1, if the UAV moves away from the obstacle, a reward is given; if The UAV approaches an obstacle and a penalty is imposed. r crash is a large negative constant, which is the penalty imposed on the UAV when it collides with an obstacle.
[0082] Subsequently, a time penalty term is introduced into the reward function to avoid redundant paths and encourage the UAV to quickly reach the target deployment point. The time penalty term for time step τ is expressed as follows:
[0083]
[0084] where the negative constant λ6 is the time penalty coefficient. As the time step increases, the time penalty increases, thus encouraging the UAV to avoid an increase in the mission time caused by redundant paths. To sum up, the reward function of the MDP at time step τ is expressed as follows:
[0085]
[0086] Based on the above MDP for the problem of optimizing the flight trajectories of UAVs among the distribution regions of H-SNs, the UAV trajectory planning algorithm based on SA and DDPG proposed by the present invention uses DDPG to optimize the trajectories of UAVs. After the algorithm converges, a trajectory that ensures obstacle avoidance and the shortest flight time of the UAVs is obtained. Finally, the shortest total mission time can be calculated according to the total mission time model of multi-UAV assisted data collection constructed in S203.
[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A multi-UAV assisted positioning and data collection method based on deep reinforcement learning, characterized in that: The method includes the following steps: S1: Construct a multi-UAV-assisted sensor positioning and data acquisition model; S2: Design an optimization scheme for average positioning accuracy and total mission time; S3: Optimize the geometric layout of UAVs and determine the relative angles and distances between UAVs; S4: Optimize the access order of the head sensor distribution areas and determine the access sequence of UAVs to the head sensor distribution areas; S5: Optimize the UAV flight trajectories and determine the flight trajectories that avoid obstacles in the scenario and quickly reach the target positions.
2. The multi-UAV assisted positioning and data collection method based on deep reinforcement learning according to claim 1, characterized in that: In S1, a multi-UAV (Unmanned Aerial Vehicle)-assisted sensor positioning and data acquisition model is constructed; This model consists of a wireless sensor network (WSN), a data center (DC), and UAVs; the sensor nodes (SNs) in the WSN are divided into member sensor nodes (M-SNs) responsible for sensing environmental data and head sensor nodes (H-SNs) responsible for collecting the sensed data of M-SNs and uploading it to UAVs; M-SNs and H-SNs are scattered and deployed by aircraft, and their exact positions are unknown; Dispatch three UAVs to visit all the H-SN distribution areas in the scenario, position and collect data from H-SNs; divide the UAVs into a master unmanned aerial vehicle (M-UAV) and follower unmanned aerial vehicles (F-UAVs), and the M-UAV hovers over the center of the H-SN distribution area as an anchor point; the three UAVs use a wireless positioning method based on time difference of arrival (TDoA) to position the H-SNs and then collect their data; The UAVs observe and avoid obstacles in the flight path and record their positions; After completing all the positioning and data acquisition tasks, the UAVs return to the DC to unload the sensed data of the SNs, the exact positions of the H-SNs, and the obstacle position information; The DC is responsible for analyzing and processing the sensed data of the SNs to achieve the functions of environmental monitoring and disaster warning for the target area; Based on the complete environmental information, the DC deploys and plans the paths of the UAVs so that the UAVs perform subsequent rounds of data acquisition to achieve the monitoring of the target area.
3. The multi-UAV assisted positioning and data collection method based on deep reinforcement learning according to claim 2, characterized in that: In S2, an optimization scheme for average positioning accuracy and total mission time is designed. This scheme first optimizes the geometric layout of the UAVs to maximize the average positioning accuracy of the H-SNs; subsequently, it jointly optimizes the access order of the UAVs to the H-SNs and the UAV flight trajectories to minimize the total mission time of multi-UAV-assisted data acquisition.
4. The multi-UAV assisted positioning and data collection method based on deep reinforcement learning according to claim 2, wherein: In S3, a multi-UAVs positioning and deployment algorithm based on the improved Deep Deterministic Policy Gradient (DDPG) is proposed to optimize the geometric layout of UAVs; First, the UAVs collective layout optimization problem is modeled as a Markov Decision Process (MDP) in a continuous action space; Subsequently, the F-UAVs are used as agents to train their deployment positions using the improved DDPG algorithm to determine their optimal relative positions and angles with respect to the M-UAV; Finally, the optimal UAVs geometric layout is determined to maximize the average positioning accuracy for the H-SN.
5. The multi-UAV assisted positioning and data collection method based on deep reinforcement learning according to claim 2, wherein: In S4, a UAV trajectory planning algorithm based on Simulated Annealing (SA) and DDPG is proposed; This algorithm first determines the access order with the shortest flight time for the UAVs to access the H-SN distribution area through the SA algorithm; Subsequently, based on the DDPG-based multi-UAVs positioning and deployment algorithm proposed in S3, the deployment positions of the UAVs over each H-SNs area are determined; Secondly, a UAVs obstacle avoidance scheme is proposed to ensure that the UAVs can quickly reach the deployment positions while avoiding obstacles in the scenario; Finally, the UAV trajectory planning problem is modeled as an MDP in a continuous action space, and each UAV is used as an agent. The DDPG algorithm is used for training. After the algorithm converges, the optimal flight trajectories of the three UAVs are determined, and the minimum total mission time for multi-UAVs assisted data collection is calculated.
Citation Information
Cited By
Intelligent unmanned aerial vehicle wireless energy transmission method for low-altitude economic network
CN121386885A