A multi-UAV disaster detection method and system
By decoupling the problem of disaster detection path planning for multiple drone into global track planning and local path planning, solving and motion control is performed separately, the problems of low path planning efficiency and simplification of environmental models in the existing technology are solved, and more efficient and more adaptive disaster detection is achieved.
Patent Information
- Application Number
- CN202210851483.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-20
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-07-20
AI Technical Summary
In the detection of drone disaster situations, the path planning algorithm is inefficient and the simplified environmental model has resulted in suboptimal planning results, failure to effectively consider the dynamics of the disaster situation, and it requires manual marking of points of interest, which is not efficient and lacks adaptability to different disaster situations.
By constructing the problem of multi-UAV path planning, it is decoupled into two sub-problems of global trajectory planning and local path planning, and solved separately to obtain the target solution of multi-UAV path planning, and motion control is performed according to the target solution to complete disaster detection.
It improves the efficiency of disaster detection, reduces complexity, can better adapt to changes in different disaster situations, and reduces the need to manually mark points of interest.
Smart Images

Figure CN115016540B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a multi-UAV disaster detection method and system. Background Art
[0002] Natural disasters are usually characterized by wide-area, diffusion and uncertainty. Traditional rescue methods are usually slow and inefficient in the face of large-scale natural disasters due to the lack of information about the affected area, and may even threaten the lives of rescuers. With the development of drone technology, drones have received widespread attention in the fields of disaster detection and rescue assistance due to their flexibility, low cost and low vulnerability to disasters.
[0003] By controlling drones to fly in disaster-stricken areas and taking real-time images of disaster areas, rescue workers can quickly grasp disaster area information and improve rescue efficiency and safety. However, due to the variability and wide range of disasters, how to maximize the efficiency of disaster exploration under limited energy through the dispatch of drones has become one of the key issues.
[0004] The existing technology usually uses traditional path planning algorithms and simplifies the environmental model when planning the flight trajectory of drones. Although such algorithms can ensure a certain solution efficiency, they will lead to suboptimal planning results, which will affect the acquisition of disaster information; they do not consider the highly dynamic nature of the disaster situation in the disaster-stricken area, and the fixed disaster environment model will lead to a lag in the solution results, further reducing the accuracy of trajectory planning; related methods will reduce rescue efficiency and even endanger the personal safety of rescuers in actual deployment environments. In addition, the existing technology requires manual marking of points of interest, which is inefficient and lacks adaptability to different types of disasters. Summary of the invention
[0005] In order to solve the above technical problems, the purpose of the present invention is to provide an efficient and low-complexity multi-UAV disaster detection method and system.
[0006] An aspect of an embodiment of the present invention provides a multi-UAV disaster detection method, comprising:
[0007] Construct a multi-UAV path planning problem to maximize the disaster detection effect;
[0008] Decoupling the multi-UAV path planning problem into a detection point maximization problem based on global trajectory planning and a detection effect maximization problem based on local path planning;
[0009] Solving the detection point maximization problem based on global trajectory planning and the detection effect maximization problem based on local path planning respectively, and obtaining the target solution of the multi-UAV path planning problem;
[0010] The motion of multiple UAVs is controlled according to the target solution to complete disaster detection.
[0011] Optionally, the process of solving the detection point maximization problem based on global trajectory planning includes:
[0012] Initialize the experience replay pool;
[0013] Initialize the parameters θ in the train-Q network train and the parameter θ in the target-Q network target ;
[0014] Initialize the system environment;
[0015] The state of the environment currently observed by the agent is input into the target-Q network, and the first result q{s,a|θ} is output a∈A , select action a according to the ε-greedy algorithm i ;
[0016] Configure the current environment observation state s of the agent i , the next environmental observation state s i+1 And the reward return r i ;
[0017] The target data (s i ,a i ,r i ,s i+1 ) to the experience replay pool;
[0018] When the experience replay pool is full of data, K experience values are randomly selected from it;
[0019] After multiple time steps, the θ in the target-Q network is target Update to the current moment in the train-Q network θ train , until the solution of the detection point maximization problem of global trajectory planning is completed.
[0020] Optionally, the process of solving the problem of maximizing the detection effect based on local path planning includes:
[0021] Initialize the experience replay pool;
[0022] Initialize the system environment;
[0023] Initialize the Critic network for each user and Actor Network Among them, the parameter of the Critic network is θ Qi , the parameters of the Actor network are
[0024] Initialize the targetCritic network for each user and targetActor network
[0025] In the initialization phase, an action a is randomly generated for the drone. 0 , and observe the reward r given by the environment 0 and feedback 1 ;
[0026] Then enter the outer loop stage, the agent generates the next action a according to the current strategy network and the observed state. t =μ(o t |θ μ )+N t , where N t is the added exploration noise used to encourage exploration; wherein the time parameter of the outer loop phase is t=1, 2, …, T;
[0027] The agent performs action a t , after execution, observe the next moment's state feedback o t+1 and return r t ;
[0028] The target data (o t ,a t ,r t ,o t+1 ) is stored in the experience pool and the feedback given by the environment is updated at the same time. t ←o t+1 ;
[0029] For each agent i=1,2,…,N, the following steps are performed cyclically: Randomly sample a portion of experience (o t ,a t ,r t ,o t+1 );
[0030] Use gradient descent to update the loss of the critic network;
[0031] Update the loss of the actor network using the policy gradient method;
[0032] Update the target network:
[0033] Exit the agent loop and exit the outer loop;
[0034] Complete the solution to the problem of maximizing the detection effect based on local path planning.
[0035] Optionally, the method further includes: constructing a UAV detection control model, which step includes:
[0036] Calculate the Euclidean distance between the drone and the target shooting point;
[0037] Calculate the overlap between the drone and the target shooting area;
[0038] Calculate the relative altitude of the drone;
[0039] An evaluation result of the drone photography effect is determined according to the Euclidean distance, the overlap degree and the relative height.
[0040] Optionally, the method further includes: constructing a UAV power loss model, the step comprising:
[0041] Define the total battery capacity of the drone based on the drone’s battery capacity;
[0042] Define the uplink transmission power of the drone and the LoS probability of the D2B link, and model the D2B channel as a LoS channel based on environmental feedback;
[0043] Modeling the propulsion energy consumption generated by the UAV;
[0044] Define the duration of a drone's flight from full charge to energy exhaustion;
[0045] According to the continuous flight time, the return charging time of the drone is determined according to the remaining power of the drone.
[0046] Optionally, the method further comprises: configuring restriction conditions to maximize the drone shooting effect;
[0047] The restrictions include:
[0048] Constrain the lower limit of shooting resolution;
[0049] Limit the safe length of drone flights;
[0050] Configure the drone's D2B link quality;
[0051] Limit the coverage overlap rate of multiple drones and improve the utilization efficiency of drones;
[0052] Limiting the speed of drones
[0053] Optionally, the process of solving the detection point maximization problem based on global trajectory planning also includes limiting the conditions of global planning;
[0054] The conditions of the global planning include:
[0055] No power outages are allowed for the drone;
[0056] The area detected by the drone is limited to the preset communication range;
[0057] The areas detected by different drones do not overlap.
[0058] Optionally, the process of solving the detection effect maximization problem based on local path planning also includes conditions that restrict local planning;
[0059] The conditions of the local planning include:
[0060] Limit the lower limit of shooting resolution;
[0061] Limit the flight distance of drones;
[0062] Limit the quality of the drone’s D2B link;
[0063] Limit the speed of drones.
[0064] Another aspect of the embodiment of the present invention further provides a multi-UAV disaster detection system, including:
[0065] The first module is used to construct the problem of multi-UAV path planning to maximize the disaster detection effect;
[0066] The second module is used to decouple the multi-UAV path planning problem into a detection point maximization problem based on global track planning and a detection effect maximization problem based on local path planning;
[0067] The third module is used to solve the detection point maximization problem based on global trajectory planning and the detection effect maximization problem based on local path planning, respectively, to obtain the target solution of the multi-UAV path planning problem;
[0068] The fourth module is used to control the motion of multiple UAVs according to the target solution to complete disaster detection.
[0069] Another aspect of an embodiment of the present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described above is implemented.
[0070] The beneficial effects of the present invention are as follows: the present invention constructs a multi-UAV path planning problem to maximize the disaster detection effect; decouples the multi-UAV path planning problem into a detection point maximization problem based on global track planning and a detection effect maximization problem based on local path planning; solves the detection point maximization problem based on global track planning and the detection effect maximization problem based on local path planning respectively to obtain a target solution to the multi-UAV path planning problem; controls the motion of multiple UAVs according to the target solution to complete disaster detection. The present invention improves detection efficiency and reduces complexity. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0072] Figure 1 A schematic diagram of an overall system model provided by an embodiment of the present invention;
[0073] Figure 2 An overall step flow chart is provided for an embodiment of the present invention. DETAILED DESCRIPTION
[0074] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0075] In view of the problems existing in the prior art, an embodiment of the present invention provides a multi-UAV disaster detection method, such as Figure 2 As shown, the method includes:
[0076] Construct a multi-UAV path planning problem to maximize the disaster detection effect;
[0077] Decoupling the multi-UAV path planning problem into a detection point maximization problem based on global trajectory planning and a detection effect maximization problem based on local path planning;
[0078] Solving the detection point maximization problem based on global trajectory planning and the detection effect maximization problem based on local path planning respectively, and obtaining the target solution of the multi-UAV path planning problem;
[0079] The motion of multiple UAVs is controlled according to the target solution to complete disaster detection.
[0080] Optionally, the process of solving the detection point maximization problem based on global trajectory planning includes:
[0081] Initialize the experience replay pool;
[0082] Initialize the parameters θ in the train-Q network train and the parameter θ in the target-Q network target ;
[0083] Initialize the system environment;
[0084] The state of the environment currently observed by the agent is input into the target-Q network, and the first result q{s,a|θ} is output a∈A , select action a according to the ε-greedy algorithm i ;
[0085] Configure the current environment observation state s of the agent i , the next environmental observation state s i+1 And the reward return r i ;
[0086] The target data (s i ,a i ,r i ,s i+1 ) to the experience replay pool;
[0087] When the experience replay pool is full of data, K experience values are randomly selected from it;
[0088] After multiple time steps, the θ in the target-Q network is target Update to the current moment in the train-Q network θ train , until the solution of the detection point maximization problem of global trajectory planning is completed.
[0089] Optionally, the process of solving the problem of maximizing the detection effect based on local path planning includes:
[0090] Initialize the experience replay pool;
[0091] Initialize the system environment;
[0092] Initialize the Critic network for each user and Actor Network Among them, the parameters of the Critic network are The parameters of the Actor network are
[0093] Initialize the targetCritic network for each user and targetActor network
[0094] In the initialization phase, an action a is randomly generated for the drone. 0 , and observe the reward r given by the environment 0 and feedback 1 ;
[0095] Then enter the outer loop stage, the agent generates the next action a according to the current strategy network and the observed state. t =μ(o t |θ μ )+N t , where N t is the added exploration noise used to encourage exploration; wherein the time parameter of the outer loop phase is t=1, 2, …, T;
[0096] The agent performs action a t , after execution, observe the next moment's state feedback o t+1 and return r t ;
[0097] The target data (o t ,a t ,r t ,o t+1 ) is stored in the experience pool and the feedback given by the environment is updated at the same time. t ←o t+1 ;
[0098] For each agent i=1,2,…,N, the following steps are performed cyclically: Randomly sample a portion of experience (o t ,a t ,r t ,o t+1 );
[0099] Use gradient descent to update the loss of the critic network;
[0100] Update the loss of the actor network using the policy gradient method;
[0101] Update the target network:
[0102] Exit the agent loop and exit the outer loop;
[0103] Complete the solution to the problem of maximizing the detection effect based on local path planning.
[0104] Optionally, the method further includes: constructing a UAV detection control model, which step includes:
[0105] Calculate the Euclidean distance between the drone and the target shooting point;
[0106] Calculate the overlap between the drone and the target shooting area;
[0107] Calculate the relative altitude of the drone;
[0108] An evaluation result of the drone photography effect is determined according to the Euclidean distance, the overlap degree and the relative height.
[0109] Optionally, the method further includes: constructing a UAV power loss model, the step comprising:
[0110] Define the total battery capacity of the drone based on the drone’s battery capacity;
[0111] Define the uplink transmission power of the drone and the LoS probability of the D2B link, and model the D2B channel as a LoS channel based on environmental feedback;
[0112] Modeling the propulsion energy consumption generated by the UAV;
[0113] Define the duration of a drone's flight from full charge to energy exhaustion;
[0114] According to the continuous flight time, the return charging time of the drone is determined according to the remaining power of the drone.
[0115] Optionally, the method further comprises: configuring restriction conditions to maximize the drone shooting effect;
[0116] The restrictions include:
[0117] Constrain the lower limit of shooting resolution;
[0118] Limit the safe length of drone flights;
[0119] Configure the drone's D2B link quality;
[0120] Limit the coverage overlap rate of multiple drones and improve the utilization efficiency of drones;
[0121] Limiting the speed of drones
[0122] Optionally, the process of solving the detection point maximization problem based on global trajectory planning also includes limiting the conditions of global planning;
[0123] The conditions of the global planning include:
[0124] No power outages are allowed for the drone;
[0125] The area detected by the drone is limited to the preset communication range;
[0126] The areas detected by different drones do not overlap.
[0127] Optionally, the process of solving the detection effect maximization problem based on local path planning also includes conditions that restrict local planning;
[0128] The conditions of the local planning include:
[0129] Limit the lower limit of shooting resolution;
[0130] Limit the flight distance of drones;
[0131] Limit the quality of the drone’s D2B link;
[0132] Limit the speed of drones.
[0133] Another aspect of the embodiment of the present invention further provides a multi-UAV disaster detection system, including:
[0134] The first module is used to construct the problem of multi-UAV path planning to maximize the disaster detection effect;
[0135] The second module is used to decouple the multi-UAV path planning problem into a detection point maximization problem based on global track planning and a detection effect maximization problem based on local path planning;
[0136] The third module is used to solve the detection point maximization problem based on global trajectory planning and the detection effect maximization problem based on local path planning, respectively, to obtain the target solution of the multi-UAV path planning problem;
[0137] The fourth module is used to control the motion of multiple UAVs according to the target solution to complete disaster detection.
[0138] Another aspect of an embodiment of the present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described above is implemented.
[0139] The specific implementation process of the present invention is described in detail below in conjunction with the accompanying drawings:
[0140] First, the proper nouns appearing in the embodiments of the present invention are explained:
[0141] DRL (deep reinforcement learning): is a commonly used machine learning algorithm in which the agent learns how to map states to actions to maximize long-term rewards through continuous interaction with the environment. Among them, reinforcement learning uses rewards to guide the agent to make better decisions.
[0142] DDPG (deep deterministic policy gradient) is a reinforcement learning algorithm based on policy neural network and value neural network. The optimal strategy obtained through learning can give the optimal action by using only local information when applied, and does not need to know the dynamic model of the environment and special communication requirements.
[0143] DQN (deep Q-network) Deep Q network: is a deep reinforcement learning algorithm that combines neural networks and Q-learning.
[0144] Drone cell (DC) is an unmanned aircraft controlled by radio remote control equipment and self-contained program control devices, or operated completely or intermittently autonomously by an on-board computer.
[0145] Signal Car (SC): The mobile communication car is a mobile communication tool for on-site image acquisition, transmission, and live broadcast of various conferences for flood control, drought relief, emergency disasters, etc. It provides a mobile video interactive platform for flood control dispatching and command, provides timely and intuitive on-site real-time information, and realizes remote consultation.
[0146] Figure 1 The disaster detection scenario of multiple drones (DC) is shown, in which two rotary drones are released by mobile signal vehicles in different areas to explore the disaster situation in a larger area. The present invention defines this scenario as It can accommodate multiple mobile signal vehicles (SC). Signal vehicles distributed in different areas are considered as a collection Where B is used to represent the total number of signal cars. The communication range of the signal car is limited, and the area it covers is expressed as Each signal vehicle can release a different number of drones.
[0147] The present invention defines a set of D drones as
[0148] Each The communication coverage is modeled as a circle with a limited radius R b hexagonal area.
[0149] Figure 1 The blue and yellow areas on the right represent the coverage of the two SCs, respectively. In order to simplify the environment of DC trajectory planning, the present invention evenly divides the entire scene into multiple hexagonal grids.
[0150] Therefore, the global trajectory of each DC is defined as the sequence of cells it serves. In each cell in the global trajectory of DC, DC dynamically adjusts its movement according to the scope and changes of the disaster in the cell to better analyze the disaster, thereby forming a local trajectory of DC such as Figure 1 Shown on the left.
[0151] In order to avoid possible collision and interference between DCs, the present invention defines that each unit is allowed to have at most one DC fly over it. The position is defined as l d ={x d ,y d ,z d}, where {x d ,y d ,z d} is a 3D Cartesian coordinate. The present invention defines the disaster area set of different units as in represents the total number of disaster-affected areas in unit g. In addition, the center position of each area is defined as l p ={x p ,y p}, where {x d ,y d} is a 2D Cartesian coordinate. The present invention converts time t The shooting point is defined as P d (t), where when DC is in the shooting state, |P d (t)|=1; otherwise, |P d (t)|=0. Due to the repeated and changing characteristics of disasters, the entire scene The area and location of the disaster area change over time. For the sake of clarity, the present invention considers that when the DC makes a decision, and DC vary within a small enough range.
[0152] The DC detection control model is described in detail below:
[0153] According to the characteristics of DC shooting, the present invention considers that the best shooting effect can be achieved at the disaster center, that is, the closer the distance between DC and the center is, the better the shooting effect. In addition, the relative height of DC affects the accuracy and overlap rate of disaster shooting. The higher the values of these two indicators, the more detailed the disaster analysis effect. Therefore, the present invention considers the Euclidean distance between DC and point p. And the overlap of the captured area (IoU) α dp As an indicator to measure the shooting effect, the calculation formulas are:
[0154]
[0155]
[0156] Among them, S p S represents the area of the disaster area p. dp Indicates the actual shooting area of DC in the disaster area p. In addition, the relative height of DC directly affects the resolution of the shooting, which in turn affects the accuracy of the disaster analysis. The calculation formula of the ground resolution is:
[0157]
[0158] Among them, F d Indicates the focal length of the DC carrying lens, S d Indicates the horizontal size of the sensor, HP d represents the size of horizontal pixels. Therefore, the present invention defines the shooting quality of DC at time t as:
[0159]
[0160] Among them, μ1 and μ2 represent different weight coefficients, R min is the minimum required resolution, κ pel Represents a penalty amount.
[0161] The DC power loss model is described in detail below:
[0162] In the detection activities of DC, energy loss mainly includes three aspects: computing energy consumption, data transmission energy consumption and propulsion energy consumption. Among them, computing energy is mainly used for signal processing and computing processing. According to a large number of recent studies, this part of energy consumption is much smaller than data transmission energy consumption and propulsion energy consumption. Therefore, the influence of computing energy consumption is ignored in this paper. The battery capacity of each DC is limited, and its total amount is expressed as E d The present invention assumes that all DC can fly back to its corresponding SCs to charge the battery, and the charging speed is expressed as p per time t. c For the convenience of expression, the present invention uses e d (t) represents the time t of current energy.
[0163] The present invention expresses the DC uplink transmission power as p u (t), and consider the state-of-the-art D2B model to represent the high LoS probability on the D2B link. Specifically, the D2B channel is modeled as a LoS channel with environmental feedback. The path loss of D2B is calculated as follows:
[0164]
[0165] Among them, r d(t) and h d (t) are the horizontal distance of D2B and the flight height of DC respectively. a and θ p They are respectively represented as angle offset and excess path loss cancellation. α represents the path loss coefficient, ζ represents the path loss scalar, Expressed as an angle scalar.
[0166] The propulsion power energy is used to maintain the DC's rise and adjust the movement. At time t, DC flies at a speed of v(t), and the propulsion energy consumption it generates can be modeled as:
[0167]
[0168] Among them, p b and p i are the blade profile power and induced power of DC in hovering state respectively. o and v m They represent the rotor blade tip speed and the average induced rotor speed in the hovering state, respectively. d , χ s , χ a and ρ are the fuselage drag ratio, rotor solidity, rotor disk area and air density, respectively.
[0169] The present invention defines the duration of flight of a DC from a fully charged state to energy exhaustion as T. Therefore, the present invention can obtain:
[0170]
[0171] When the flight time is greater than T, the DC may crash due to limited battery capacity. This requires the DC to determine whether to return for charging based on the remaining power. The remaining power of the DC can be expressed as:
[0172]
[0173] Where x(τ) is a binary variable. When the DC is in the detection state at time τ, x(τ) = 1; vice versa.
[0174] In addition, the present invention models the optimization conditions of the drone planning problem in a specific scenario:
[0175] The goal of the multi-DC collaborative detection problem is to detect the disaster area by formulating a suitable trajectory plan for each DC at each time t, thereby maximizing the shooting effect. Therefore, the present invention takes maximizing the shooting effect of all DCs as the main performance to formulate the following optimization problem and construct the following optimization model.
[0176] P1:
[0177] stC1:R dp (t)≥R min ,
[0178] C2:e d (t)≥0,
[0179] C3:L d (t)≤L max ,
[0180] C4:
[0181] C5:v d (t)≤v max ,
[0182] Wherein, Π represents the strategy formulated by multiple DCs to achieve cooperative detection. max Denotes the maximum allowed D2B large-scale path loss. max Expressed as DC maximum flight speed.
[0183] Five conditions must be met at the same time:
[0184] (1) Constraint (a) restricts the lower limit of the shooting resolution;
[0185] (2) Restriction (b) ensures DC flight safety and avoids accidents such as crashes;
[0186] (3) Constraint (c) ensures the quality of the D2B link;
[0187] (4) Constraint (d) means that multiple DCs do not overlap in coverage, so as to improve the utilization efficiency of DCs. ;
[0188] (5) Constraint (e) constrains the navigation speed of DC;
[0189] Considering the disaster area There are a large number of dynamically changing disaster points on the network, and the global observation values of DC are too large to be solved by traditional optimization algorithms or simple DRL algorithms. In order to solve this complexity problem, the hierarchical DRL framework is used to decouple the problem into multiple sub-problems with smaller state spaces, and then the entire problem is solved by iteratively solving all sub-problems. This paper decouples the multi-DC collaborative detection problem into two hierarchical sub-problems, namely the multi-DC global trajectory planning sub-problem and the single DC local path planning sub-problem.
[0190] (1) Multi-DC global trajectory planning sub-problem
[0191] In this section, the present invention considers represents the time gap t in the global planning g , the average sum of the points detected by DC in unit g. Therefore, the optimization goal of this part is to maximize the number of points detected in the unit.
[0192] P2:
[0193] stC1:e d (t a )≥0,
[0194]
[0195] C3:
[0196] Among them, g d (t g ) represents the global decision t g At time t, DC d detects unit g.
[0197] Three conditions must be met at the same time:
[0198] (1) Constraint (a) indicates that DC energy interruption is not allowed;
[0199] (2) Constraint (b) means that the area detected by DC is limited to the communication range of SC b;
[0200] (3) Constraint (c) means that the detection areas of different DCs do not overlap, thus avoiding waste of resources.
[0201] (2) Single DC local path planning sub-problem
[0202] According to the route planning made for each DC based on the global trajectory planning, each DC makes its own path planning in its assigned unit to maximize the shooting effect of the detection points.
[0203] P3:
[0204] stC1:R dp (t)≥R min ,
[0205] C2:e d (t)≥0,
[0206] C3:L d (t)≤L max ,
[0207] C4:v d (t)≤v max ,
[0208] Four conditions must be met at the same time:
[0209] (1) Constraint (a) restricts the lower limit of the shooting resolution;
[0210] (2) Restriction (b) ensures DC flight safety and avoids accidents such as crashes;
[0211] (3) Constraint (c) ensures the quality of the D2B link;
[0212] (4) Constraint (d) constrains the navigation speed of DC.
[0213] The solution provided by the embodiment of the present invention is:
[0214] 1. All SCs share information and build a set of reinforcement learning networks (DQN) for all DCs to formulate global trajectory planning, which includes two neural networks with the same structure, namely train Q-network and train Q-network, where the parameters of the target Q-network are copied from the train Q-network at a certain frequency.
[0215] 2. Consider adopting a central decision-making approach to formulate corresponding actions for each DC, i.e., the unit for the next stage of flight, based on the position coordinates, remaining battery capacity, and an overview of the unit points of each DC.
[0216] 3. Then, based on the flight unit decided by the global planning, a set of neural networks is constructed for DC for local path planning decisions, which includes a pair of execution networks critic and actor, and a pair of lagging target networks targetCritic and targetActor. The lagging updated network copies the parameters of the execution network at a certain ratio at regular intervals (see the specific steps).
[0217] 4. Consider the intelligent RS as an intelligent agent, take the disaster area point distribution, remaining battery capacity and current location coordinates in the unit as the current observation information, input it into the execution network Actor and output action a, that is, select the appropriate v x 、v y 、v z , so that DC can maximize the effect of shooting disaster areas without power outages.
[0218] 5. According to the competition window setting in a dynamic environment, the user obtains the return and reward r in the current window, and uses the reward and time difference method to calculate the loss of the Critic network. The Critic network is continuously updated using gradient updates to achieve an accurate estimation of the value of action a.
[0219] 6. Using the Critic’s estimate of the Actor’s value, use gradient ascent to continuously adjust the Actor network so that it selects actions with higher value with a greater probability.
[0220] Through repeated iterations, the strategy network continuously updates its own parameters and finds a set of optimal pricing strategies suitable for itself.
[0221] In summary, the present invention first proposes the problem of maximizing the disaster detection effect by multi-DC path planning. Then, this complex problem is decoupled into the problem of maximizing the detection points based on global trajectory planning and the problem of maximizing the detection effect based on local path planning. Based on this, the present invention develops a central decision algorithm based on DQN and a distributed algorithm based on DDPG for DC path planning, aiming to further maximize the shooting effect on the basis of ensuring that DC does not experience energy interruption.
[0222] (1) Global trajectory planning: For the discrete action space of the global trajectory planning subproblem, the deep Q-network (DQN) can be used to solve the subproblem with fast convergence. The deep Q-network (DQN) is composed of two neural networks with the same structure but different functions, one is the train Q-network and the other is the train Q-network. The two neural networks have different parameters θ train and θ target θ train The Q value (expected reward q(s,a|θ) used to evaluate the optimal action target )),θ target Used to select the action corresponding to the maximum Q value (through the ε greedy algorithm). These two sets of parameters separate action selection and strategy evaluation, reducing the risk of overfitting in the process of estimating Q values. The present invention uses an experience pool (relpay buffer) to store the experience generated by all agents, and uses the experience randomly sampled from the experience pool as the input of the train Q-network to update its parameters. This can not only greatly reduce the memory and computing resources required for training, but also reduce the coupling between data. After executing action a i After that, TI gets the reward signal r from the environment i And observe the next state s i+1 , then (s i ,a i ,r i ,s i+1) is stored as an experience in the experience replay pool for training the neural network. After every F time steps, the target Q-network θ target Will be updated to the current moment train Q-network θ train In one round of training, a small batch D consisting of K random experiences is extracted from the experience replay as the input of the train Q-network, and the loss function is calculated by the mean square error (MSE). a∈A is the output value of the train Q-network, indicating that in state s, the parameter is θ train The neural network outputs the expected reward for action a. Finally, the parameters in the train Q-network are updated using the gradient descent method. Every F time steps, the target Q-network is updated to the train Q-network.
[0223] (2) Local path planning: Considering the time slot t of global trajectory planning g The unit to be flown is determined for each DC in the present invention. The DDPG-LTPRA algorithm is proposed for each DC to maximize the number of disaster area points photographed within a time step t. The present invention builds four neural networks for each DC: an execution strategy network for selecting actions (the input is the state observed by the user, denoted as trainActor), an execution evaluation network for action evaluation (the input is the observed state of all users and the action selected by the user, trainCritic), a target strategy network for stabilizing training and providing actions for updating the execution value network (the input is the state observed by the user, the network is denoted as targetActor), and a target evaluation network for updating the execution evaluation network to provide the next state-action value (the input is the observation of all users and the action selected by the user, denoted as targetCritic). Among them, the network structures of trainActor and targetActor are the same, and the parameters of targetActor are copied from trainActor in a certain proportion in each round for slow update. The update process is as follows θ μ′ ←τθ μ +(1-τ)θ μ , where θ μ′ is the parameter of the targetActor network, θ μ is the parameter of the trainActor network; similarly, the network structure of trainCritic is the same as that of targetCritic, and the update process of trainCritic is as follows Q′ ←τθ Q +(1-τ)θ Q , where θ Q′is the parameter of the targetCritic network, θ Q is the parameter of the trainCritic network. The Actor network consists of an input layer, three hidden layers and an output layer. The three hidden layers use the ReLU function as their activation function, and the output layer uses the Tanh function as its activation function, outputting the action a under the current observation; the Critic network consists of an input layer, three hidden layers and an output layer. All layers use the ReLU function as the activation function to generate the state-action value Q. Each DC is regarded as an intelligent agent, and the distribution of disaster points in the unit, the remaining battery capacity, and the current location coordinates are used as the current observation information as observation input into the trainActor, and the output action a is selected. t Specifically, DC inputs the observation into the neural network and outputs the current appropriate v through the Tanh function. x 、v y 、v z . According to the reward return formula, the intelligent GS calculates the current reward r. At the same time, the DC uses the reward r to calculate the loss function of the trainCritic network, and then updates the trainCritic network by reverse gradient transfer; at the same time, the observations of all DCs and the actions of the target DC at the current moment are input into its own trainCritic network to obtain the state-action value Q, and use this value to update the trainActor network by reverse gradient transfer. In addition, the targetActor network and the targetCritic network are gradually updated by copying a certain proportion each step. By repeatedly iterating the above process, when the actions of all users no longer change, the current action is the optimal action. That is, DC will make adaptive adjustments according to changes in dynamic conditions to achieve the best effect under the current circumstances.
[0224] The specific steps of the present invention include:
[0225] 1. Global track planning:
[0226] (1) Initialize the experience replay pool;
[0227] (2) Initialize the parameters θ in the train-Q network and target-Q network train ,θ target ;
[0228] (3) Initialize the system environment;
[0229] (4) Input the state of the environment currently observed by the agent into the target Q-network and output q{s,a|θ} a∈A , select action a according to the ε-greedy algorithmi ;
[0230] (5) The agent observes the state s from the environment i+1 And the reward return r i ;
[0231] (6) i ,a i ,r i ,s i+1 ) to the experience replay pool;
[0232] (7) When the experience replay pool D is full of data, K experience values are randomly selected from it;
[0233] (8) After F time steps, the target Q-network θ target Update to the current moment train Q-network θ train .
[0234] 2. Local path planning:
[0235] 1. Initialize the experience replay pool;
[0236] 2. Initialize the system environment;
[0237] 3. Initialize the Critic network for each user and Actor Network The parameters are and
[0238] 4. Initialize the targetCritic network for each user and targetActor network
[0239] 5. In the initialization phase, a random action a is generated for DC 0 , and observe the reward r given by the environment 0 and feedback 1 ;
[0240] 6. (Enter the outer loop t = 1, 2, ..., T) The intelligent RS generates the next action a based on the current strategy network and the observed state t =μ(o t |θ μ )+N t , where N t is the exploration noise added to encourage exploration;
[0241] 7. Intelligent RS performs action a t , after execution, observe the next moment's state feedback o t+1and return r t ;
[0242] 8. t ,a t ,r t ,o t+1 ) is stored in the experience pool, and o t ←o t+1 ;
[0243] 9. (For each agent i=1,2,…,N, execute in a loop) Randomly sample a portion of experience (o t ,a t ,r t ,o t+1 );
[0244] 10. Update the loss of the critic network using gradient descent: where y b =r b +γQ′(o b+1 ,a b+1 |θ Q′ );
[0245] 11. Update the loss of the actor network using the policy gradient method, where the gradient can be obtained as follows:
[0246] 12. Update target network: θ Q′ ←τθ Q +(1-τ)θ Q ,θ μ′ ←τθ μ +(1-τ)θ μ ;
[0247] 13. Exit the intelligent body loop and exit the outer loop.
[0248] In summary, the present invention has the following advantages:
[0249] 1. This invention proposes an effective multi-DC collaborative disaster detection scheme. Among them, the global trajectory planning solves the complexity caused by multiple DCs and long-term changes in disaster distribution. The local path planning algorithm handles the real-time changes in the number and location of disaster points, and the state space is constrained by the output of the global trajectory planning algorithm. This hierarchical DRL framework converges to a suboptimal solution with a high probability.
[0250] 2. In view of the variability of disaster distribution, the present invention designs a DDPG algorithm that is independently executed by each DC to adjust the distribution of real-time DC flight control. Specifically, DDPG can realize trajectory planning in continuous space, and through mathematical analysis of D2U communication, it reduces the complexity of algorithm input and further improves convergence performance.
[0251] 3. This algorithm does not rely on existing training data, and the agent can collect experience of interacting with the environment for training.
[0252] 4. Through historical experience and continuous trial and error, this algorithm invention can eventually learn the most suitable set of pricing strategies in a dynamic environment to ensure maximum profit.
[0253] 5. The present invention adopts the deep Q-Learning technology in deep reinforcement learning, which has the characteristics of fast convergence and can achieve rapid response in complex communication environments.
[0254] In some selectable embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the present invention is provided by way of example, for the purpose of providing a more comprehensive understanding of technology. The disclosed method is not limited to the operation and logic flow presented herein. Selectable embodiments are expected, wherein the order of various operations is changed and the sub-operation of a part for which is described as a larger operation is performed independently.
[0255] In addition, although the present invention is described in the context of functional modules, it should be understood that, unless otherwise specified, one or more of the functions and / or features described may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the present invention. More specifically, in view of the properties, functions, and internal relationships of the various functional modules in the device disclosed herein, the actual implementation of the module will be understood within the conventional skills of the engineer. Therefore, those skilled in the art can implement the present invention set forth in the claims without excessive experimentation using ordinary techniques. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0256] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program codes.
[0257] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.
[0258] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.
[0259] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0260] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0261] Although the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.
[0262] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A multi-UAV disaster detection method, characterized in that: include: Construct a multi-UAV path planning problem to maximize the disaster detection effect; Decoupling the multi-UAV path planning problem into a detection point maximization problem based on global trajectory planning and a detection effect maximization problem based on local path planning; Solving the detection point maximization problem based on global trajectory planning and the detection effect maximization problem based on local path planning respectively, and obtaining the target solution of the multi-UAV path planning problem; According to the target solution, multiple UAVs are controlled to complete disaster detection; Among them, the point captured by drone d at time t is defined as P d (t), when the drone is in shooting state, |P d (t)|=1, otherwise, |P d (t)|=0; the Euclidean distance between the drone and the point and the overlap of the shooting area are used as indicators to measure the shooting effect; The process of solving the detection point maximization problem based on global trajectory planning includes: The state of the environment currently observed by the agent is input into the target-Q network, and the first result q{s i ,a|θ target } a∈A , select action a according to the ε-greedy algorithm i ; Configure the current environment observation state s of the agent i , the next environmental observation state s i+1 And the reward return r i ; The target data (s i ,a i ,r i ,s i+1 ) to the experience replay pool; When the experience replay pool is full of data, K experience values are randomly selected from it; After multiple time steps, the θ in the target-Q network is target Update to the current moment in the train-Q network θ train , until the problem of maximizing the number of detection points in global trajectory planning is solved; The process of solving the problem of maximizing the detection effect based on local path planning includes: Initialize the Critic network for each user and Actor Network Among them, the parameters of the Critic network are The parameters of the Actor network are Initialize the targetCritic network for each user and targetActor network In the initialization phase, an action a is randomly generated for the drone. 0 , and observe the reward r given by the environment 0 and feedback 1 ; Then enter the outer loop stage, the agent generates the next action a according to the current strategy network and the observed state. t =μ(o t |θ μ )+N t , where N t is the added exploration noise used to encourage exploration; wherein the time parameter of the outer loop phase is t=1, 2, …, T; The agent performs action a t , after execution, observe the next moment's state feedback o t+1 and return r t ; The target data (o t ,a t ,r t ,o t+1 ) is stored in the experience pool and the feedback given by the environment is updated at the same time. t ←o t+1 ; For each agent i=1,2,…,N, the following steps are performed cyclically: Randomly sample a portion of experience (o t ,a t ,r t ,o t+1 ); Use gradient descent to update the loss of the critic network; Update the loss of the actor network using the policy gradient method; Update the target network: Exit the agent loop and exit the outer loop; Complete the solution to the problem of maximizing the detection effect based on local path planning.
2. A multi-UAV disaster detection method according to claim 1, characterized in that: In the process of solving the detection point maximization problem based on global trajectory planning, before inputting the state of the environment currently observed by the intelligent agent into the target-Q network, it also includes: Initialize the experience replay pool; Initialize the parameters θ in the train-Q network train and the parameter θ in the target-Q network target ; Initialize the system environment.
3. The multi-UAV disaster detection method according to claim 1, characterized in that: In the process of solving the problem of maximizing the detection effect based on local path planning, before initializing the Critic network and the Actor network for each user, it also includes: Initialize the experience replay pool; Initialize the system environment.
4. The multi-UAV disaster detection method according to claim 1, characterized in that: The method further includes: constructing a UAV detection and control model, which includes: Calculate the Euclidean distance between the drone and the target shooting point; Calculate the overlap between the drone and the target shooting area; Calculate the relative altitude of the drone; An evaluation result of the drone photography effect is determined according to the Euclidean distance, the overlap degree and the relative height.
5. The multi-UAV disaster detection method according to claim 1, characterized in that: The method further includes: constructing a UAV power loss model, the step comprising: Define the total battery capacity of the drone based on the drone’s battery capacity; Define the uplink transmission power of the drone and the LoS probability of the D2B link, and model the D2B channel as a LoS channel based on environmental feedback; Modeling the propulsion energy consumption generated by the UAV; Define the duration of a drone's flight from full charge to energy exhaustion; According to the continuous flight time, the return charging time of the drone is determined according to the remaining power of the drone.
6. A multi-UAV disaster detection method according to claim 1, characterized in that: The method further includes: configuring restriction conditions to maximize the drone photography effect; The restrictions include: Constrain the lower limit of shooting resolution; Limit the safe length of drone flights; Configure the drone's D2B link quality; Limit the coverage overlap rate of multiple drones and improve the utilization efficiency of drones; Limit the speed of drones.
7. The multi-UAV disaster detection method according to claim 1, characterized in that: The process of solving the detection point maximization problem based on global trajectory planning also includes limiting the conditions of global planning; The conditions of the global planning include: No power outages are allowed for the drone; The area detected by the drone is limited to the preset communication range; The areas detected by different drones do not overlap.
8. The multi-UAV disaster detection method according to claim 1, characterized in that: The process of solving the problem of maximizing the detection effect based on local path planning also includes conditions that restrict local planning; The conditions of the local planning include: Limit the lower limit of shooting resolution; Limit the flight distance of drones; Limit the quality of the drone’s D2B link; Limit the speed of drones.
9. A multi-UAV disaster detection system, characterized in that: include: The first module is used to construct the problem of multi-UAV path planning to maximize the disaster detection effect; The second module is used to decouple the multi-UAV path planning problem into a detection point maximization problem based on global track planning and a detection effect maximization problem based on local path planning; The third module is used to solve the detection point maximization problem based on global trajectory planning and the detection effect maximization problem based on local path planning, respectively, to obtain the target solution of the multi-UAV path planning problem; The fourth module is used to control the motion of multiple drones according to the target solution to complete disaster detection; Among them, the point captured by drone d at time t is defined as P d (t), when the drone is in shooting state, |P d (t)|=1, otherwise, |P d (t)|=0; the Euclidean distance between the drone and the point and the overlap of the shooting area are used as indicators to measure the shooting effect; The process of solving the problem of maximizing the detection points based on global trajectory planning includes: The state of the environment currently observed by the agent is input into the target-Q network, and the first result q{s i ,a|θ target } a∈A , select action a according to the ε-greedy algorithm i ; Configure the current environment observation state s of the agent i , the next environmental observation state s i+1 And the reward return r i ; The target data (s i ,a i ,r i ,s i+1 ) to the experience replay pool; When the experience replay pool is full of data, K experience values are randomly selected from it; After multiple time steps, the θ in the target-Q network is target Update to the current moment in the train-Q network θ train , until the problem of maximizing the number of detection points in global trajectory planning is solved; The process of solving the problem of maximizing the detection effect based on local path planning includes: Initialize the Critic network for each user and Actor Network Among them, the parameter of the Critic network is θ Qi , the parameter of the Actor network is θ μi ; Initialize the targetCritic network for each user and targetActor network In the initialization phase, an action a is randomly generated for the drone. 0 , and observe the reward r given by the environment 0 and feedback 1 ; Then enter the outer loop stage, the agent generates the next action a according to the current strategy network and the observed state. t =μ(o t |θ μ )+N t , where N t is the added exploration noise used to encourage exploration; wherein the time parameter of the outer loop phase is t=1, 2, …, T; The agent performs action a t , after execution, observe the next moment's state feedback o t+1 and return r t ; The target data (o t ,a t ,r t ,o t+1 ) is stored in the experience pool and the feedback given by the environment is updated at the same time. t ←o t+1 ; For each agent i=1,2,…,N, the following steps are performed cyclically: Randomly sample a portion of experience (o t ,a t ,r t ,o t+1 ); Use gradient descent to update the loss of the critic network; Update the loss of the actor network using the policy gradient method; Update the target network: Exit the agent loop and exit the outer loop; Complete the solution to the problem of maximizing the detection effect based on local path planning.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Path planning method based on multi-agent enhanced learning
CN109059931A
Multi-unmanned aerial vehicle path optimization method for rapid evaluation after earthquake disasters
CN111310992A