Unmanned aerial vehicle group formation control method in forest fire scene
By constructing a drone swarm formation model and a forest fire environment model, and combining a multi-agent deep deterministic policy gradient algorithm to optimize the drone's position, speed, attitude, and monitoring area, the problem of fire and firefighter monitoring needs in drone swarm formation control was solved, improving monitoring efficiency and mission success rate.
Patent Information
- Application Number
- CN202511001255.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-11
AI Technical Summary
Existing drone swarm control methods fail to adequately consider the monitoring needs of fire and firefighters in forest fire scenarios, resulting in low monitoring efficiency and low mission success rates.
A drone swarm formation model and a forest fire environment model are constructed. A multi-agent deep deterministic policy gradient algorithm is used to optimize the drone position, speed, attitude and monitoring area. A Markov decision process is combined to control the drone swarm formation.
It improves the monitoring efficiency and mission success rate of drone swarms in forest fire scenarios, and is better able to adapt to the changing fire conditions and the changing positions of firefighters.
Smart Images

Figure CN120928845A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of mobile communication technology and relates to a method for controlling a swarm of drones in a forest fire scenario. Background Technology
[0002] In recent years, with the rapid development of technologies in the fields of communications, materials, and drones, drones have been widely used in civilian and disaster relief fields due to their small size, simple structure, convenient operation, low cost, and high maneuverability. Forest fires are a type of natural disaster that is sudden and extremely destructive, causing serious damage to the ecological environment and human life and property once they occur.
[0003] Currently, the main methods for detecting forest fires include: first, observing forest fires through lookout towers. While this method has a wide coverage area, it may have blind spots due to terrain. Second, observing through the installation of monitoring equipment in forest areas, but this method is costly and prone to damage. Third, using drones for real-time monitoring of forest fires. Drones are simple to operate, easy to deploy quickly, and can achieve real-time monitoring. However, due to the large area of the fire, multiple drones are usually needed to monitor the fire and firefighters. Therefore, rationally planning the drone's trajectory to ensure it completes its intended task within limited energy and a reasonable trajectory is a key requirement for drones in forest fire scenarios.
[0004] However, the situation of a forest fire and the work of firefighters are highly uncertain. When drones perform missions, it is necessary to rationally allocate tasks, coordinate information exchange between drones and between drones and base stations, and plan the trajectory of each drone in a comprehensive manner to avoid mission failure due to errors. Existing drone swarm control methods, such as deep reinforcement learning-based swarm navigation flight control methods and distributed self-anti-interference swarm tracking control strategies, can achieve drone swarm control to a certain extent, but most of them fail to fully consider the overall monitoring needs of the fire and firefighters, and cannot effectively adapt to the actual needs of forest fire sites.
[0005] Therefore, there is an urgent need for a drone swarm control method that can comprehensively consider the monitoring needs of fire and firefighters, in order to improve the monitoring efficiency and mission success rate of drones in forest fire scenarios. Summary of the Invention
[0006] To address the aforementioned technical problems, this invention provides a method for controlling a swarm of unmanned aerial vehicles (UAVs) in a forest fire scenario, comprising:
[0007] S1: Construct a system model, which includes a drone swarm formation model and a forest fire environment model;
[0008] S2: Based on the UAV swarm formation model and the forest fire environment model, construct an optimization problem with the objective of minimizing the position, speed, attitude and total error of the UAV's monitoring area.
[0009] S3: The optimization problem is transformed into a Markov decision process. Based on the Markov decision process, a multi-agent deep deterministic policy gradient algorithm is used to solve the optimization problem, and the UAV swarm formation control strategy is obtained.
[0010] Optionally, the drone swarm formation model includes: upper-layer drones and lower-layer drone swarms, where the position of the upper-layer drones is represented as p. top =[x top ,y top ,z top ] T Where, x top ,y top ,z top The positions of the upper-level UAVs on the x, y, and z axes in the spatial domain are given by the following: W i,k ={p i (k),v i (k),a i (k),Φ i (k)}, i∈{1,2,3,…,N}, k∈{1,2,3,…,K}, where i is the i-th drone in the lower-level drone swarm, N is the total number of lower-level drones, k is the time slot for drone movement, K is the total number of time slots for drone movement, p i (k) represents the location of drone i in time slot k, v i (k) represents the velocity of UAV i in time slot k, a i (k) represents the angular acceleration of UAV i in time slot k, Φ i (k) represents the attitude of UAV i in time slot k.
[0011] Optionally, the position of the lower-level UAV i at time slot k satisfies: p i (k)=[x i (k),y i (k),z i (k)] T The velocity of the lower-level drone i in time slot k satisfies: The angular acceleration of the lower-level UAV i in time slot k satisfies: The attitude of the lower-level UAV i at time slot k satisfies: Φ i (k)=[α i (k),β i (k),ε i (k)] T , where x i (k),yi (k),z i (k) represents the position of the lower-level UAV i on the x, y, z axes in the spatial domain at time slot k. Let x, y, and z be the velocities of the lower-level UAV i on the x, y, and z axes respectively at time slot k. Let i be the roll acceleration of the lower-level UAV i in time slot k. Let i be the pitch acceleration of the lower-level UAV i in time slot k. Let α be the yaw acceleration of the lower-level UAV i in time slot k. i (k) represents the roll angle of the lower-level UAV i in time slot k, β i (k) represents the pitch angle of the lower-level UAV i in time slot k, ε i (k) is the yaw angle of the lower-level UAV i in time slot k.
[0012] Optional, forest fire environmental models include: S k ={h(k),f(k),w(k),r(k),z(k)}, k∈{1,2,3,…,K}, where k is the time slot of the UAV's movement, K is the total number of time slots of the UAV's movement, and h(k)={p h (k),v h (k)} is the model of forest fire at time slot k, p h (k) represents the position of the forest fire at time slot k, v h (k) represents the spread rate of the forest fire at time slot k, f(k) = {p f (k),Φ f (k),v f (k)} is the model of wind in time slot k, p f (k) represents the position of the wind in time slot k, Φ f (k) represents the direction of the wind at time slot k, v f (k) is the wind speed in time slot k, w(k) = {p w (k),Φ w (k),v w (k)} is the model of smoke in time slot k, p w (k) represents the position of the smoke in time slot k, Φ w (k) represents the direction of smoke diffusion at time slot k, v w (k) represents the diffusion velocity of the smoke in time slot k, r(k) = {x r (k),y r (k),z r (k)} represents the model of the firefighter at time slot k, x r (k),y r (k),z r(k) represents the position of the firefighter at time slot k along the x, y, and z axes in the spatial domain, where z(k) = {x z (k),y z (k),z z (k)} is a model of obstacles in the fire environment at time slot k, x z (k),y z (k),z z (k) represents the position of the obstacle at the x, y, and z axes in the spatial domain at time slot k.
[0013] Optionally, based on the drone swarm formation model and the forest fire environment model, an optimization problem is constructed with the objective of minimizing the position, velocity, attitude, and total error of the drone's monitored area, including:
[0014]
[0015] stC1:D i,j (k)≥D min
[0016] C2:D i,z (k)≥D z,min
[0017] C3:H>z i (k)>h,a=0
[0018] C4:z i (k)≥H,a=1
[0019] C5:z top >z max
[0020]
[0021]
[0022]
[0023] C9:R key,i (k)≥R key,min a = 0
[0024] C10:R eff,i (k)≥R eff,min a = 1
[0025] Wherein, C1 represents the safe distance limit between drones in the lower-level drone swarm formation; C2 represents the safe distance limit between the lower-level drone swarm and obstacles; C3 represents the flight altitude limit of the lower-level drone swarm when monitoring firefighters; C4 represents the flight altitude limit of the lower-level drone swarm when monitoring forest fires; C5 represents the altitude limit of the upper-level drones; C6 represents the energy consumption limit of a single drone in the lower-level drone swarm; C7 represents the energy consumption limit of an upper-level drone; C8 represents the maximum error limit of a single drone in the lower-level drone swarm; C9 represents the overlap rate limit of the key monitoring area of a single drone in the lower-level drone swarm; and C10 represents the coverage rate limit of the efficient coverage area of a single drone in the lower-level drone swarm. all,i (k) represents the sum of errors in position, velocity, attitude, and monitored area of the i-th UAV in the lower-level UAV swarm at time slot k, where D i,j (k) represents the distance between two adjacent UAVs in the lower-level UAV swarm at time slot k, D min To maintain a safe minimum distance between adjacent drones in the lower-level drone swarm, D i,z (k) represents the distance between a drone in the lower-level drone swarm and an adjacent obstacle at time slot k, D z,min Z represents the minimum safe distance between a drone in the lower-level drone swarm and an adjacent obstacle. H represents the maximum flight altitude of a drone in the lower-level drone swarm when monitoring firefighters, and also the minimum flight altitude of a drone when monitoring forest fires. i (k) represents the flight altitude of the drones in the lower-level drone swarm, h represents the minimum safe flight altitude for the drones in the lower-level drone swarm to monitor firefighters, a = 0 indicates that the lower-level drones are monitoring firefighters, a = 1 indicates that the lower-level drones are monitoring forest fires, and z top z is the flight altitude of the upper-level drone. max E represents the flight altitude of the highest-flying drone in the lower-level drone swarm. all,i (k) represents the total energy consumption of a single UAV in the lower-level UAV swarm at time slot k, E max,i E represents the maximum energy consumption of a single drone in the lower-level drone swarm. top (k) represents the total energy consumption of the upper-level UAV in time slot k, E max,top e represents the maximum total energy consumption of the upper-level drone. max,i R represents the maximum sum of errors in the position, velocity, attitude, and monitored area of the i-th UAV in the lower-level UAV swarm. key,i (k) represents the overlap rate of the key monitoring area of the lower-level UAV in time slot k, R key,min R represents the minimum overlap rate of the key monitoring areas of the lower-level drones. eff,i (k) represents the coverage rate of the efficient coverage area of the lower-level UAV in time slot k, R eff,minThis represents the minimum coverage rate of the area where the lower-level drones can efficiently cover.
[0026] Optionally, the total error of the lower-level drone swarm satisfies:
[0027] Among them, e p,i (k) represents the deviation between the actual position and the desired position of the i-th UAV in the lower-level UAV swarm at time slot k, e v,i (k) represents the deviation between the actual speed and the expected speed of the i-th UAV in the lower-level UAV swarm at time slot k, e a,i (k) represents the deviation between the actual angular acceleration and the desired angular acceleration of the i-th UAV in the lower-level UAV swarm at time slot k, e Φ,i (k) represents the deviation between the actual attitude and the desired attitude of the i-th UAV in the lower-level UAV swarm at time slot k, e s,i (k) represents the deviation between the actual monitored area and the expected monitored area of the i-th UAV in the lower-level UAV swarm at time slot k; the deviation satisfies the following formula:
[0028]
[0029] Where, p e,i (k) represents the expected position of the i-th UAV in the lower-level UAV swarm at time slot k, p i (k) The actual position of the i-th UAV in the lower-level UAV swarm at time slot k, v e,i (k) represents the expected velocity of the i-th UAV in the lower-level UAV swarm at time slot k, v i (k) The actual velocity of the i-th UAV in the lower-level UAV swarm at time slot k, a e,i (k) represents the expected angular acceleration of the i-th UAV in the lower-level UAV swarm at time slot k, a i (k) The true angular acceleration Φ of the i-th UAV in the lower-level UAV swarm at time slot k. e,i (k) represents the desired attitude of the i-th UAV in the lower-level UAV swarm at time slot k, Φ i (k) The true attitude of the i-th UAV in the lower-level UAV swarm at time slot k, s e,i (k) represents the expected area of the i-th UAV in the lower-level UAV swarm at time slot k, s ture,i (k) The actual area of the i-th UAV in the lower-level UAV swarm at time slot k.
[0030] Optionally, the total energy consumed by the i-th drone in the lower-level drone swarm. Total energy consumed by upper-level drones Where, n i (k)∈{0,1} represents the state of the i-th UAV in the lower-level UAV swarm at time slot k, and ni (k) = 1 indicates that the i-th drone in the lower-level drone swarm is in flight state at time slot k, and n i (k) = 0 indicates that the i-th drone in the lower-level drone swarm is in a hovering state at time slot k, E fly,i (k) represents the flight energy consumption of the i-th UAV in the lower-level UAV swarm at time slot k, E stop,i (k) represents the hovering energy consumption of the i-th UAV in the lower-level UAV swarm at time slot k, E tran,i (k) represents the total energy consumption of the i-th UAV in the lower-level UAV swarm for transmitting and receiving information in time slot k, E fly For the flight energy consumption of the upper-level drone, E stop (k) represents the hovering energy consumption of the upper-level UAV in time slot k, E com (k) represents the computational energy consumption of the upper-level UAV in time slot k, E tran (k) represents the total energy consumption of the upper-level UAV in time slot k for transmitting and receiving information.
[0031] Optionally, the overlap rate of the key monitoring area of the lower-level UAV in time slot k satisfies: The coverage of the efficient coverage area of the lower-level UAV in time slot k satisfies: Among them, s imp,i (k) represents the area of the key monitoring region for the i-th UAV in the lower-level UAV swarm at time slot k, s eff,i (k) represents the area of the efficient coverage of the i-th drone in the lower-level drone swarm at time slot k, where a = 0 indicates that the lower-level drone swarm is monitoring firefighters, and a = 1 indicates that the lower-level drone swarm is monitoring forest fires. ture,i (k)=s yuan,i (k)-s z,i (k) represents the total area of the actual monitoring area of the i-th drone in the lower-level drone swarm at time slot k. Let i be the total area that the i-th drone in the lower-level drone swarm should monitor in time slot k. Let θ be the area of obstacle obstruction within the monitoring area of the i-th drone in the lower-level drone swarm at time slot k. hor,i (k) represents the horizontal viewing angle of the camera carried by the i-th UAV in the lower-level UAV swarm at time slot k, θ ver,i (k) represents the vertical viewing angle of the camera carried by the i-th UAV in the lower-level UAV swarm at time slot k.
[0032] Optionally, the optimization problem can be transformed into a Markov decision process, including:
[0033] Construct the quintuple A =<S,D,O,R,γ> Where S is the state space, D is the action space, O is the observation set; R is the reward function, γ is the discount rate, and s i(k)=(o i (k),d i (k)),s i (k)∈S,s i (k) represents the state of the lower-level UAV i in time slot k, o i (k) represents the observation of the lower-level UAV i at time slot k, o i (k)=(p i (k),E all,i (k),e all,i (k)),o i (k)∈O, p i (k) represents the position of the i-th UAV in the lower-level UAV swarm at time slot k, E all,i (k) represents the total energy consumed by the i-th drone in the lower-level drone swarm at time slot k, e all,i (k) represents the total error of the i-th UAV in the lower-level UAV swarm at time slot k, d i (k) represents the action of the lower-level UAV i in time slot k, d i (k)=(v i (k),a i (k)),d i (k)∈D,v i (k) represents the velocity of the i-th UAV in the lower-level UAV swarm at time slot k, a i (k) represents the angular acceleration of the i-th UAV in the lower-level UAV swarm at time slot k, R i (k) represents the reward value of the i-th drone in the lower-level drone swarm at time slot k.
[0034] R i (k)=R crash,i (k)+R energy,i (k)+R error,i (k),R i (k)∈R, R crash,i (k) represents the collision reward value of the lower-level drone i in time slot k, R energy,i (k) represents the energy consumption reward value of the lower-level drone i in time slot k, R error,i (k) represents the error reward value of the lower-level UAV i in time slot k.
[0035] Optionally, the optimization problem is solved using a multi-agent deep deterministic policy gradient algorithm based on Markov decision processes, including: constructing a policy network μ for each agent i. i (k) and value network η i (k); Get the current observation o i (k) will observe the current o i (k) Input the policy network to obtain action di (k), execute action d i (k) Receive reward R i (k) and the observation at the next time step o i (k+1), will (o i (k),d i (k),R i (k),o i (k+1)), i=1,2,3,…,N are stored as a set of experience samples in the experience replay buffer L; a batch of experience samples are extracted from the experience replay buffer L to train the policy network and value network, and the trained policy network and value network are obtained; the formation control strategy of the multi-UAV swarm is obtained based on the policy network and value network.
[0036] The beneficial effects of this invention are:
[0037] 1. This invention comprehensively considers the impact of wind, smoke, flame spread, firefighter positions, and the total error of the position, speed, attitude, and monitoring area of UAVs on the execution of tasks by UAV swarms in forest fire scenarios to construct an optimization problem. The objective function of the optimization problem considers the total error of the position, speed, attitude, and monitoring area of UAVs to adapt to the actual environmental conditions of forest fires, and can more effectively control the formation of UAV swarms.
[0038] 2. Since drone swarms need to face a variable environment, this invention formulates the optimization problem as a Markov decision process and adopts a multi-agent deep deterministic policy gradient algorithm to solve the optimization problem, so as to better cope with external interference and more effectively control the drone swarm formation. Attached Figure Description
[0039] Figure 1 A flowchart illustrating a method for controlling a swarm of unmanned aerial vehicles (UAVs) in a forest fire scenario, provided by an embodiment of the present invention;
[0040] Figure 2 This is a schematic diagram of the attitude change of a drone provided in an embodiment of the present invention;
[0041] Figure 3 This is a schematic diagram of a simulated drone swarm operation provided in an embodiment of the present invention;
[0042] Figure 4 The drone swarm trajectory route map provided in the embodiments of the present invention;
[0043] Figure 5 A graph showing the reward values of a drone swarm after performing a task, comparing the multi-agent deep deterministic policy gradient algorithm provided in this embodiment of the invention with other algorithms;
[0044] Figure 6The reward value of a drone swarm after avoiding different numbers of obstacles under the multi-agent deep deterministic policy gradient algorithm provided in this embodiment of the invention is shown.
[0045] Figure 7 A graph showing the total reward value of a drone swarm with different target cluster numbers, comparing the multi-agent deep deterministic policy gradient algorithm provided in this embodiment of the invention with other algorithms.
[0046] Figure 8 The data transmission revenue graph of UAVs is shown in the figure when comparing the multi-agent deep deterministic policy gradient algorithm provided in this embodiment of the invention with other algorithms, and the number of target clusters.
[0047] Figure 9 The energy consumption diagram of UAV swarms under different target cluster numbers is shown in the multi-agent deep deterministic policy gradient algorithm provided in this embodiment of the invention, compared with other algorithms. Detailed Implementation
[0048] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0049] like Figure 1 As shown, this invention employs a method for controlling a swarm of unmanned aerial vehicles (UAVs) in a forest fire scenario, comprising:
[0050] S1. Construct a system model, which includes: a drone swarm formation model and a forest fire environment model;
[0051] S2. Construct an optimization problem based on the drone swarm formation control model and the forest fire environment model;
[0052] S3. The optimization problem is transformed into a Markov decision process. Based on the Markov decision process, the multi-agent deep deterministic policy gradient algorithm is used to solve the optimization problem and obtain the UAV swarm formation strategy.
[0053] like Figure 3 As shown, the drone swarm formation is divided into two layers. The upper layer consists of individual drones with fixed positions, denoted as p. top =[x top ,y top ,z top ] T , where x top ,y top ,z topThese represent the positions of the upper-level drones along the x, y, and z axes in the spatial domain. The next layer represents the drone swarm, and the drone swarm model W is shown. i,k Includes: W i,k ={p i (k),v i (k),a i (k),Φ i (k)}, i∈{1,2,3,…,N}, k∈{1,2,3,…,K}; where i is the i-th UAV in the lower-level UAV swarm, N is the total number of lower-level UAVs, k is the time slot for UAV movement, and K is the total number of time slots for UAV movement, p i (k) represents the location of drone i in time slot k, v i (k) represents the velocity of UAV i in time slot k, a i (k) represents the angular acceleration of UAV i in time slot k, Φ i (k) represents the attitude of UAV i in time slot k.
[0054] The position p of the lower-level UAV i at time slot k i (k)=[x i (k),y i (k),z i (k)] T The speed of the lower-level drone i in time slot k angular acceleration of the lower-level UAV i in time slot k The attitude Φ of the lower-level UAV i at time slot k i (k)=[α i (k),β i (k),ε i (k)] T ;where x i (k),y i (k),z i (k) represents the position of the lower-level UAV i on the x, y, z axes in the spatial domain at time slot k. Let x, y, and z be the velocities of the lower-level UAV i on the x, y, and z axes respectively at time slot k. Let i be the roll acceleration of the lower-level UAV i in time slot k. Let i be the pitch acceleration of the lower-level UAV i in time slot k. Let i be the yaw acceleration of the lower-level UAV i in time slot k, such as Figure 2 As shown, the attitude of the UAV is related to the rotation of the fuselage around the x, y, and z axes, α i (k) represents the roll angle of the lower-level UAV i in time slot k, β i (k) represents the pitch angle of the lower-level UAV i in time slot k, ε i(k) is the yaw angle of the lower-level UAV i in time slot k.
[0055] Forest fire environmental models include: S k ={h(k),f(k),w(k),r(k),z(k)}, k∈{1,2,3,…,K}; where k is the time slot of the UAV's movement, K is the total number of time slots of the UAV's movement, and h(k)={p h (k),v h (k)} is the model of forest fire at time slot k, p h (k)=[x h (k),y h (k),z h (k)] T Let v be the position of the forest fire at time slot k. h (k)=v h (k-1)·K r ·K f ·cosλ represents the spread rate of the forest fire at time slot k, where K r The magnitude of the correction factor for surrounding combustibles is related to the type of combustibles, K f The correction factor for wind force is related to the wind force, λ is the slope at the current location, and f(k) = {p f (k),Φ f (k),v f (k)} is the model of wind in time slot k, p f (k)=[x f (k),y f (k),z f (k)] T Let Φ be the position of the wind in time slot k. f (k)=[α f (k),β f (k),ε f (k)] T The direction of the wind at time slot k. Let v be the wind speed at time slot k, where v f,average (k) represents the average wind speed at the current moment, z f,average (k) represents the average wind height at the current moment, q is the empirical coefficient related to wind force and is usually taken as [0.4, 0.6], w(k) = {p w (k),Φ w (k),v w (k)} is the model of smoke in time slot k, p w (k)=[x w (k),y w (k),z w (k)]T Let Φ be the position of the smoke in time slot k. w (k)=[α w (k),β w (k),ε w (k)] T Let K be the direction of smoke diffusion at time slot k. Let N be the diffusion velocity of the smoke in time slot k, where N is the diffusion velocity of the smoke. w (k) represents the smoke concentration at the current moment, ξ is an environmental factor whose magnitude is related to the current environment, and r(k) = {x r (k),y r (k),z r (k)} represents the model of the firefighter at time slot k, x r (k),y r (k),z r (k) represents the position of the firefighter at time slot k along the x, y, and z axes in the spatial domain, where z(k) = {x z (k),y z (k),z z (k)} is a model of obstacles in the fire environment at time slot k, x z (k),y z (k),z z (k) represents the position of the obstacle at the x, y, and z axes in the spatial domain at time slot k.
[0056] Based on the impact of various factors in forest fires on drone swarm formations, conditional models of these factors' effects are established, including:
[0057] (1) Based on the safety distance constraint model between lower-level UAVs
[0058] Failure to control the distance between drone swarms can lead to collisions and mission failure. Therefore, lower-level drones need to maintain a certain distance, which is defined as the safe distance between them, and is expressed as:
[0059] D i,j (k)≥D min
[0060]
[0061] Among them, D min This indicates the minimum safe distance between lower-level drones to avoid collisions.
[0062] (2) Model based on the safety distance constraint between the lower-level UAV and obstacles
[0063] When a drone is too close to an obstacle, it will collide with the obstacle, causing the mission to fail. Therefore, a certain distance needs to be maintained between the drone and the obstacle. This distance is defined as the safe distance between the drone and the obstacle, and is expressed as:
[0064] D i,z (k)≥D z,min
[0065]
[0066] Among them, D z,min This indicates the minimum safe distance between the drone and the obstacle to avoid collision.
[0067] (3) Conditional constraint model based on monitoring firefighters
[0068] In forest fire monitoring, drones aim to monitor the largest possible area of fire spread, often flying high. However, this can lead to insufficient clarity in their monitoring of firefighters. Therefore, it's necessary to define the primary functions of the drones. Drones primarily monitoring firefighters should focus on tracking the distance between firefighters and the flames and smoke, with a secondary task of monitoring the fire's intensity. Thus, their function can be defined as follows:
[0069] H>z i (k)>h,a=0
[0070] Where H is the maximum flight altitude of the drones in the lower-level drone swarm when monitoring firefighters, h is the minimum safe flight altitude of the drones in the lower-level drone swarm when monitoring firefighters, and a = 0 means that the lower-level drones are mainly monitoring firefighters.
[0071] (4) Conditional Constraint Model Based on Forest Fire Monitoring
[0072] In forest fire monitoring, the primary function of drones is to monitor the spread of flames and smoke, with a secondary task of monitoring the situation of firefighters. The conditions are as follows:
[0073] z i (k)≥H,a=1
[0074] Where H is the lowest altitude at which drones in the lower-level drone swarm fly when monitoring forest fires, and a = 1 indicates that the lower-level drones are monitoring forest fires.
[0075] (5) Conditional constraint model based on the altitude of two-layer UAVs
[0076] In drone swarm control, a two-layer drone swarm is used. The upper layer consists of individual drones whose tasks are to calculate and plan the tasks uploaded by the lower layer drone swarm, and then download them to the lower layer drone swarm for execution. Therefore, the upper layer drones must be above the lower layer drone swarm, hence the condition is expressed as:
[0077] z top >z max
[0078] Among them, z max This refers to the altitude at which the tallest drone in the lower-level drone swarm flies.
[0079] (6) Conditional Constraint Model Based on Energy Consumption of Lower-Level UAV Swarm
[0080] The energy consumption of a lower-level UAV can be mainly divided into three aspects: flight energy consumption, hovering energy consumption, and transmission energy consumption. Since the total energy consumption of a UAV is limited, the total energy consumption of the UAV must be less than or equal to the maximum rated energy consumption of the UAV. The condition is expressed as follows:
[0081]
[0082] Among them, E max,i This represents the maximum total energy consumed by the i-th drone in the lower-level drone swarm. Let n be the total energy consumed by the i-th drone in the lower-level drone swarm. i (k)∈{0,1} represents the state of the i-th UAV in the lower-level UAV swarm at time slot k, and n i (k) = 1 indicates that the i-th drone in the lower-level drone swarm is in flight state at time slot k, and n i (k) = 0 indicates that the i-th drone in the lower-level drone swarm is in a hovering state at time slot k, E fly,i (k)=T fly,i (k)·P fly,i (k) represents the flight energy consumption of the i-th UAV in the lower-level UAV swarm at time slot k, T fly,i (k) represents the flight time of the i-th UAV in the lower-level UAV swarm within time slot k, P fly,i (k) represents the flight power of the i-th UAV in the lower-level UAV swarm within time slot k. Let T be the hovering energy consumption of the i-th UAV in the lower-level UAV swarm at time slot k. stop,i (k) represents the hovering time of the i-th UAV in the lower-level UAV swarm within time slot k. Indicator of whether there is wind Indicates no wind. Indicates there is wind, P stop,i (k) represents the hovering power of the i-th UAV in the lower-level UAV swarm under windless conditions within time slot k, P'stop,i (k) represents the hovering power of the i-th UAV in the lower-level UAV swarm under windy conditions within time slot k, E tran,i (k)=T tran,i (k)·P tran,i (k) represents the total energy consumption of the i-th UAV in the lower-level UAV swarm for transmitting and receiving information in time slot k, T tran,i (k) represents the total time for the i-th UAV in the lower-level UAV swarm to transmit and receive information within time slot k, P tran,i (k) represents the total power of the i-th UAV in the lower-level UAV swarm for transmitting and receiving information in time slot k.
[0083] (7) Conditional constraint model based on the energy consumption of upper-layer UAVs
[0084] The energy consumption of upper-level UAVs can be mainly divided into four aspects: flight energy consumption, hovering energy consumption, computing energy consumption, and transmission energy consumption. Since the total energy consumption of a UAV is limited, the total energy consumption of the UAV must be less than or equal to the maximum rated energy consumption of the UAV. The condition is expressed as follows:
[0085]
[0086] Among them, E max,top This represents the maximum total energy consumed by the upper-level drone. E represents the total energy consumed by the upper-level drones. fly =T fly ·P fly For the flight energy consumption of the upper-level drone, T fly For the flight time of the upper-level drone, P fly For the flight power of the upper-level drone, T represents the hovering energy consumption of the upper-level UAV at time slot k. stop (k) represents the hovering time of the upper-level UAV within time slot k. A sign indicating whether there is wind. Indicates no wind. Indicates there is wind, P stop (k) represents the hovering power of the upper-level UAV in windless conditions within time slot k, P' stop (k) represents the hovering power of the upper-level UAV under windy conditions within time slot k, E com (k)=T com (k)·P com (k) represents the computational energy consumption of the upper-level UAV in time slot k, T com (k) represents the computation time of the upper-level UAV within time slot k, P com (k) represents the computational power of the upper-level UAV within time slot k, E tran (k)=T tran (k)·Ptran (k) represents the total energy consumption of the upper-level UAV in transmitting and receiving information at time slot k, T tran (k) represents the total time for the upper-level UAV to transmit and receive information within time slot k, P tran (k) represents the total power of the upper-level UAV in transmitting and receiving information within time slot k.
[0087] (8) Conditional Constraint Model Based on Total Error of Lower-Level UAV Swarm
[0088] The actual situation of a drone swarm during operation will deviate somewhat from the expected situation. This deviation is acceptable, but it cannot be too large, as excessive deviation will prevent the mission from being completed. Therefore, it is necessary to limit the maximum error to ensure the successful completion of the mission. The condition is expressed as follows:
[0089]
[0090] Among them, e max,i Let be the maximum total error of the i-th drone in the lower-level drone swarm. Let e be the total error of the i-th drone in the lower-level drone swarm. p,i (k) represents the deviation between the actual position and the desired position of the i-th UAV in the lower-level UAV swarm at time slot k, e v,i (k) represents the deviation between the actual speed and the expected speed of the i-th UAV in the lower-level UAV swarm at time slot k, e a,i (k) represents the deviation between the actual angular acceleration and the desired angular acceleration of the i-th UAV in the lower-level UAV swarm at time slot k, e Φ,i (k) represents the deviation between the actual attitude and the desired attitude of the i-th UAV in the lower-level UAV swarm at time slot k, e s,i (k) represents the deviation between the actual monitored area and the expected monitored area of the i-th UAV in the lower-level UAV swarm at time slot k. The formula for this deviation is as follows:
[0091]
[0092] In the formula, p e,i (k) represents the expected position of the i-th UAV in the lower-level UAV swarm at time slot k, p i (k) The actual position of the i-th UAV in the lower-level UAV swarm at time slot k, v e,i (k) represents the expected velocity of the i-th UAV in the lower-level UAV swarm at time slot k, v i (k) The actual velocity of the i-th UAV in the lower-level UAV swarm at time slot k, a e,i (k) represents the expected angular acceleration of the i-th UAV in the lower-level UAV swarm at time slot k, a i(k) The true angular acceleration Φ of the i-th UAV in the lower-level UAV swarm at time slot k. e,i (k) represents the desired attitude of the i-th UAV in the lower-level UAV swarm at time slot k, Φ i (k) The true attitude of the i-th UAV in the lower-level UAV swarm at time slot k, s e,i (k) represents the expected area of the i-th UAV in the lower-level UAV swarm at time slot k, s ture,i (k) The actual area of the i-th UAV in the lower-level UAV swarm at time slot k.
[0093] (9) Conditional Constraint Model Based on Key Monitoring Areas of Lower-Level UAV Swarms
[0094] When the lower-level drone swarm monitors firefighters and forest fires, to ensure the personal safety of firefighters, each firefighter is monitored by two drones. The overlapping area monitored by the two drones is the key monitoring area, which must be greater than or equal to the minimum key monitoring area to ensure the safety of the firefighters. The condition is expressed as follows:
[0095] R key,i (k)≥R key,min a = 0
[0096] Among them, R key,min This represents the minimum overlap rate in the key monitoring areas of the lower-level drones. s represents the overlap rate of the key monitoring area of the lower-level UAV in time slot k. imp,i (k) represents the area of the key monitoring region for the i-th UAV in the lower-level UAV swarm at time slot k, s ture,i (k)=s yuan,i (k)-s z,i (k) represents the total area of the actual monitoring area of the i-th drone in the lower-level drone swarm at time slot k. Let i be the total area that the i-th drone in the lower-level drone swarm should monitor in time slot k. Let θ be the area of obstacle obstruction within the monitoring area of the i-th drone in the lower-level drone swarm at time slot k. hor,i (k) represents the horizontal viewing angle of the camera carried by the i-th UAV in the lower-level UAV swarm at time slot k, θ ver,i (k) represents the vertical viewing angle of the camera carried by the i-th UAV in the lower-level UAV swarm at time slot k.
[0097] (10) Conditional Constraint Model for High-Efficiency Monitoring Area Based on Lower-Level UAV Swarm
[0098] When the lower-level drone swarm monitors firefighters and forest fires, to ensure the personal safety of the firefighters, each firefighter is monitored by two drones. The overlapping area monitored by the two drones is the key monitoring area. While ensuring the area of this key monitoring area is sufficient, the area not being monitored by the drones should be expanded as much as possible. This area is the efficient monitoring area, and it must be of a certain size to improve the efficiency of the drones' work. Therefore, its condition is expressed as:
[0099] R eff,i (k)≥R eff,min a = 1
[0100] Among them, R eff,min This represents the minimum coverage rate for the efficient coverage area of the lower-level drones. s represents the coverage rate of the area efficiently covered by the lower-level UAV in time slot k. eff,i (k) represents the area of the efficient coverage region of the i-th UAV in the lower-level UAV swarm at time slot k.
[0101] Therefore, the optimization problem is derived from all the constraints in the model:
[0102]
[0103] stC1:D i,j (k)≥D min
[0104] C2:D i,z (k)≥D z,min
[0105] C3:H>z i (k)>h,a=0
[0106] C4:z i (k)≥H,a=1
[0107] C5:z top >z max
[0108]
[0109]
[0110]
[0111] C9:R key,i (k)≥R key,min a = 0
[0112] C10:R eff,i (k)≥R eff,min a = 1
[0113] Wherein, C1 represents the safe distance limit between drones in the lower-level drone swarm formation, C2 represents the safe distance limit between the lower-level drone swarm and obstacles, C3 represents the flight altitude limit of the lower-level drone swarm when monitoring firefighters, C4 represents the flight altitude limit of the lower-level drone swarm when monitoring forest fires, C5 represents the altitude limit of the upper-level drones, C6 represents the energy consumption limit of a single drone in the lower-level drone swarm, C7 represents the energy consumption limit of the upper-level drones, C8 represents the maximum error limit of a single drone in the lower-level drone swarm, C9 represents the overlap rate limit of the key monitoring area of a single drone in the lower-level drone swarm, and C10 represents the coverage rate limit of the efficient coverage area of a single drone in the lower-level drone swarm.
[0114] Transforming an optimization problem into a Markov decision process includes:
[0115] The problem is formulated as a Markov decision process, represented by a set of five elements, A =<S,D,O,R,γ> , where S is the state space, D is the action space, O is the observation set, R is the reward function, and γ is the discount rate and γ∈(0,1).
[0116] (1) State space
[0117] The state space S is defined by the current observation of the UAV. i (k) and current action d i (k) is used to represent s i (k)=(o i (k),d i (k)),s i (k)∈S;o i (k) represents the observation information of the i-th UAV in the lower-level UAV swarm at time slot k, expressed by the UAV's current position, total energy consumption, and total error, and is defined as o i (k)=(p i (k),E all,i (k),e all,i (k)), d i (k) represents the action information of the i-th UAV in the lower-level UAV swarm at time slot k, expressed by the UAV's current velocity and angular acceleration, and defined as d. i (k)=(v i (k),a i (k)).
[0118] (2) Action space
[0119] The action space D represents the set of all actions that the drone can take. The action of the drone at time k is defined as:
[0120] d i (k)=(vi (k),a i (k))
[0121] d i (k)∈D
[0122] Among them, v i (k) represents the velocity of the i-th UAV in the lower-level UAV swarm at time slot k, a i (k) represents the angular acceleration of the i-th UAV in the lower-level UAV swarm at time slot k.
[0123] (3) Observation set
[0124] The observation set O represents the information collected from the UAV at the current time. The observation information of the UAV at time k is defined as follows:
[0125] o i (k)=(p i (k),E all,i (k),e all,i (k))
[0126] o i (k)∈O
[0127] Where, p i (k) represents the position of the i-th UAV in the lower-level UAV swarm at time slot k, E all,i (k) represents the total energy consumed by the i-th drone in the lower-level drone swarm at time slot k, e all,i (k) represents the total error of the i-th UAV in the lower-level UAV swarm at time slot k.
[0128] (4) Reward function
[0129] To enable the drone swarm to achieve its objectives, a reward value will be designed based on drone energy consumption, collision rate, and total error.
[0130] The bonus value R of drone energy consumption energy,i (k) is related to the drone's flight, hovering state, and transmission power consumption, and is set as follows:
[0131]
[0132] Where, n i (k)∈{0,1} represents the state of the i-th UAV in the lower-level UAV swarm at time slot k, and n i (k) = 1 indicates that the i-th drone in the lower-level drone swarm is in flight state at time slot k, and n i (k) = 0 indicates that the i-th drone in the lower-level drone swarm is in a hovering state at time slot k, E fly,i(k) represents the flight energy consumption of the i-th UAV in the lower-level UAV swarm at time slot k, E stop,i (k) represents the hovering energy consumption of the i-th UAV in the lower-level UAV swarm at time slot k, E tran,i (k) represents the total energy consumption of the i-th UAV in the lower-level UAV swarm for transmitting and receiving information in time slot k.
[0133] The reward value R for the collision rate between the drone and the obstacle crash,i (k) is related to the distance between the drone and obstacles, which include trees, drones, mountains, and smoke. A greater distance between the drone and obstacles results in a higher reward value, while a closer distance results in a lower reward value. This will affect R. crash,i (k) is set to the distance between the drone and the nearest obstacle, expressed as:
[0134] R crash,i (k)=||p i (k)-z(k)||
[0135] Total error reward value R of the drone error,i (k) is related to the difference between the expected value of the upper-level drones for the lower-level drone swarm and the actual flight situation of the lower-level drone swarm. A large difference results in a small reward value, and a small difference results in a large reward value. R error,i (k) is represented as:
[0136]
[0137] In summary, the total reward value is:
[0138] R i (k)=R crash,i (k)+R energy,i (k)+R error,i (k)
[0139] A multi-agent deep deterministic policy gradient algorithm is employed, where each agent i includes: a policy network μ i (k) and value network η i (k), by finding a strategy This ensures that all states and actions of the drone swarm will receive a long-term maximized reward value η. i (k)=(o1(k),o2(k),o3(k),…,o N (k),d1(k),d2(k),d3(k),…,d N (k); ψ i ),in and ψ i For the network parameters, a deep neural network is used to approximate μ. i (k) and η i The (k) function updates μ through the following steps.i (k) and η i (k) Approximate value of the function:
[0140] Step 1, Initialize the policy network μ i (k) and value network η i (k) and the target policy network μ' i (k) and value network η' i (k).
[0141] Step 2, Agent interacts with the environment: Obtain current observations. i (k) will observe the current o i (k) Input policy network μ i (k), to obtain action d i (k), execute action d i (k) then obtains the reward value R. i (k) then obtain the observation for the next time step. i (k+1).
[0142] Step 3, Experience Replay: Replay the interaction data (o1(k), o2(k), ..., o N (k),d1(k),d2(k),…,d N (k),R1(k),R2(k),…,R N (k),o1(k+1),o2(k+1),…,o N (k+1) is stored in the experience replay buffer L.
[0143] Step 4, Target Calculation: For each experience (o1(k), o2(k), ..., o N (k),d1(k),d2(k),…,d N (k),R1(k),R2(k),…,R N (k),o1(k+1),o2(k+1),…,o N (k+1)), calculate the target Y i (k):
[0144] Y i (k)=R i (k)+γη i (k+1)(o1(k+1),o2(k+1),…,o N (k+1),μ1(k+1),μ2(k+1),…,μ N (k+1))
[0145] Step 5, Calculation of the minimum mean square error function for the value network: Define the minimum mean square error function for the value network to measure the difference between the actual value and the expected value:
[0146] Q(ψ i )=Ε[(η i (k)-Y i (k)) 2 ]
[0147] Where, Q(ψ) i ) is the minimum mean square error function of the value network.
[0148] Step 6, Value Network Parameter Update: Update the value network parameters using gradient descent.
[0149]
[0150] Where g represents the learning rate.
[0151] Step 7, Calculate the expected return function of the policy network: Define the expected return function of the policy network to measure the return obtained after the policy network is executed:
[0152]
[0153] in, The function is the one that maximizes the expected return for the policy network.
[0154] Step 8, Policy Network Parameter Update: Update the policy network parameters using gradient descent.
[0155]
[0156] Where u represents the learning rate.
[0157] Step 9, Target network update:
[0158]
[0159] in, This is the soft update coefficient.
[0160] Step 10: Repeat steps 2-9 until the stopping condition is met.
[0161] In one embodiment, experimental parameters are set according to the above model, as shown in Table 1:
[0162] Table 1 Model Parameter Table
[0163] parameter numerical values Learning rate (g, u) 0.001 Number of agents 2 Discount factor 0.99 Number of drones 3 Training cycle 500 Number of obstacles 50,100,150
[0164] Figure 4This is a trajectory map of a drone swarm, where Target represents the target location, gray dots represent obstacles in the environment, and the curve represents the flight path of the drone swarm. Experiments show that the drone swarm can successfully navigate through obstacles to reach the target location.
[0165] Figure 5 The graph shows the reward values of the UAV swarm after performing the task, comparing the Multi-Agent Deep Deterministic Policy Gradient Algorithm (MADDPG) with other algorithms. As can be seen from the graph, the reward value of the MADDPG is higher than that of other algorithms, indicating that the MADDPG can effectively solve the problem studied in this invention.
[0166] Figure 6 The graph shows the reward values of a drone swarm after avoiding different numbers of obstacles under the multi-agent deep deterministic policy gradient algorithm. It can be seen from the graph that there are good reward values for different numbers of obstacles, but the fewer the number of obstacles, the higher the reward value. This is consistent with the actual situation, as fewer obstacles indicate a simpler environment and better performance.
[0167] Figure 7 The graph shows the total reward value of the UAV swarm under different target cluster numbers, comparing the multi-agent deep deterministic policy gradient algorithm with other algorithms. As can be seen from the graph, the reward value increases rapidly with the increase of the target cluster number. Furthermore, the reward value obtained by the UAV swarm under the multi-agent deep deterministic policy gradient algorithm is higher than that of other algorithms, indicating that the UAV swarm under the multi-agent deep deterministic policy gradient algorithm can complete the predetermined task and obtain higher rewards for different target cluster numbers.
[0168] Figure 8 The graph shows the data transmission gains of UAV swarms under different target cluster numbers, comparing the multi-agent deep deterministic policy gradient algorithm with other algorithms. As can be seen from the graph, the total data transmission volume of the UAV swarm increases with the increase in the number of target clusters, and thus the data transmission gains also increase. Furthermore, the data transmission gains of the UAV swarm under the multi-agent deep deterministic policy gradient algorithm are higher than those of other algorithms, indicating that the UAV swarm under the multi-agent deep deterministic policy gradient algorithm can obtain better data transmission gains under different target cluster numbers.
[0169] Figure 9 The graph shows the energy consumption of UAV swarms at different target cluster numbers, comparing the multi-agent deep deterministic policy gradient algorithm with other algorithms. As can be seen from the graph, the energy consumption of the UAV swarm increases relatively slowly as the number of target clusters increases, which can effectively solve the problem studied in this invention.
[0170] The above-described embodiments further illustrate the purpose, technical solution, and advantages of the present invention. It should be understood that the above-described embodiments are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made to the present invention within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for controlling a swarm of unmanned aerial vehicles (UAVs) in a forest fire scenario, characterized in that, The method includes: S1: Construct a system model, which includes a drone swarm formation model and a forest fire environment model; S2: Based on the UAV swarm formation model and the forest fire environment model, construct an optimization problem with the objective of minimizing the position, speed, attitude and total error of the UAV's monitoring area. S3: The optimization problem is transformed into a Markov decision process. Based on the Markov decision process, a multi-agent deep deterministic policy gradient algorithm is used to solve the optimization problem, and the UAV swarm formation control strategy is obtained.
2. The method according to claim 1, characterized in that, The drone swarm formation model includes: upper-layer drones and a lower-layer drone swarm, where the position of the upper-layer drones is represented by p. top =[x top ,y top ,z top ] T Where, x top ,y top ,z top The positions of the upper-level UAVs on the x, y, and z axes in the spatial domain are given by the following: W i,k ={p i (k),v i (k),a i (k),Φ i (k)}, i∈{1,2,3,…,N}, k∈{1,2,3,…,K}, where i is the i-th drone in the lower-level drone swarm, N is the total number of lower-level drones, k is the time slot for drone movement, K is the total number of time slots for drone movement, p i (k) represents the location of drone i in time slot k, v i (k) represents the velocity of UAV i in time slot k, a i (k) represents the angular acceleration of UAV i in time slot k, Φ i (k) represents the attitude of UAV i in time slot k.
3. The method according to claim 2, characterized in that, The position of the lower-level UAV i at time slot k satisfies: p i (k)=[x i (k),y i (k),z i (k)] T The velocity of the lower-level drone i in time slot k satisfies: The angular acceleration of the lower-level UAV i in time slot k satisfies: The attitude of the lower-level UAV i at time slot k satisfies: Φ i (k)=[α i (k),β i (k),ε i (k)] T , where x i (k),y i (k),z i (k) represents the position of the lower-level UAV i on the x, y, z axes in the spatial domain at time slot k. Let x, y, and z be the velocities of the lower-level UAV i on the x, y, and z axes, respectively, at time slot k. Let i be the roll acceleration of the lower-level UAV i in time slot k. Let i be the pitch acceleration of the lower-level UAV i in time slot k. Let α be the yaw acceleration of the lower-level UAV i in time slot k. i (k) represents the roll angle of the lower-level UAV i in time slot k, β i (k) represents the pitch angle of the lower-level UAV i in time slot k, ε i (k) is the yaw angle of the lower-level UAV i in time slot k.
4. The method according to claim 2, characterized in that, Forest fire environmental models include: S k ={h(k),f(k),w(k),r(k),z(k)}, k∈{1,2,3,…,K}, where k is the time slot of the UAV's movement, K is the total number of time slots of the UAV's movement, and h(k)={p h (k),v h (k)} is the model of forest fire at time slot k, p h (k) represents the position of the forest fire at time slot k, v h (k) represents the spread rate of the forest fire at time slot k, f(k) = {p f (k),Φ f (k),v f (k)} is the model of wind in time slot k, p f (k) represents the position of the wind in time slot k, Φ f (k) represents the direction of the wind at time slot k, v f (k) is the wind speed in time slot k, w(k) = {p w (k),Φ w (k),v w (k)} is the model of smoke in time slot k, p w (k) represents the position of the smoke in time slot k, Φ w (k) represents the direction of smoke diffusion at time slot k, v w (k) represents the diffusion velocity of the smoke in time slot k, r(k) = {x r (k),y r (k),z r (k)} represents the model of the firefighter at time slot k, x r (k),y r (k),z r (k) represents the position of the firefighter at time slot k along the x, y, and z axes in the spatial domain, where z(k) = {x z (k),y z (k),z z (k)} is a model of obstacles in the fire environment at time slot k, x z (k),y z (k),z z (k) represents the position of the obstacle at the x, y, and z axes in the spatial domain at time slot k.
5. The method according to any one of claims 1-4, characterized in that, Based on the drone swarm formation model and the forest fire environment model, an optimization problem is constructed with the objective of minimizing the total error of the drone's position, velocity, attitude, and monitored area. This problem includes: s.t.C1:D i,j (k)≥D min C2:D i,z (k)≥D z,min C3:H>z i (k)>h,a=0 C4:z i (k)≥H,a=1 C5:z top >from max C9:R key,i (k)≥R key,min ,a=0 C10:R eff,i (k)≥R eff,min ,a=1 Wherein, C1 represents the safe distance limit between drones in the lower-level drone swarm formation; C2 represents the safe distance limit between the lower-level drone swarm and obstacles; C3 represents the flight altitude limit of the lower-level drone swarm when monitoring firefighters; C4 represents the flight altitude limit of the lower-level drone swarm when monitoring forest fires; C5 represents the altitude limit of the upper-level drones; C6 represents the energy consumption limit of a single drone in the lower-level drone swarm; C7 represents the energy consumption limit of an upper-level drone; C8 represents the maximum error limit of a single drone in the lower-level drone swarm; C9 represents the overlap rate limit of the key monitoring area of a single drone in the lower-level drone swarm; and C10 represents the coverage rate limit of the efficient coverage area of a single drone in the lower-level drone swarm. all,i (k) represents the sum of errors in position, velocity, attitude, and monitored area of the i-th UAV in the lower-level UAV swarm at time slot k, where D i,j (k) represents the distance between two adjacent UAVs in the lower-level UAV swarm at time slot k, D min To maintain a safe minimum distance between adjacent drones in the lower-level drone swarm, D i,z (k) represents the distance between a drone in the lower-level drone swarm and an adjacent obstacle at time slot k, D z,min Z represents the minimum safe distance between a drone in the lower-level drone swarm and an adjacent obstacle. H represents the maximum flight altitude of a drone in the lower-level drone swarm when monitoring firefighters, and also the minimum flight altitude of a drone when monitoring forest fires. i (k) represents the flight altitude of the drones in the lower-level drone swarm, h represents the minimum safe flight altitude for the drones in the lower-level drone swarm to monitor firefighters, a = 0 indicates that the lower-level drones are monitoring firefighters, a = 1 indicates that the lower-level drones are monitoring forest fires, and z top z is the flight altitude of the upper-level drone. max E represents the flight altitude of the highest-flying drone in the lower-level drone swarm. all,i (k) represents the total energy consumption of a single UAV in the lower-level UAV swarm at time slot k, E max,i E represents the maximum energy consumption of a single drone in the lower-level drone swarm. top (k) represents the total energy consumption of the upper-level UAV in time slot k, E max,top e represents the maximum total energy consumption of the upper-level drone. max,i R represents the maximum sum of errors in the position, velocity, attitude, and monitored area of the i-th UAV in the lower-level UAV swarm. key,i (k) represents the overlap rate of the key monitoring area of the lower-level UAV in time slot k, R key,min R represents the minimum overlap rate of the key monitoring areas of the lower-level drones. eff,i (k) represents the coverage rate of the efficient coverage area of the lower-level UAV in time slot k, R eff,min This represents the minimum coverage rate of the area where the lower-level drones can efficiently cover.
6. The method according to claim 5, characterized in that, The total error of the lower-level drone swarm satisfies: Among them, e p,i (k) represents the deviation between the actual position and the desired position of the i-th UAV in the lower-level UAV swarm at time slot k, e v,i (k) represents the deviation between the actual velocity and the expected velocity of the i-th UAV in the lower-level UAV swarm at time slot k, e a,i (k) represents the deviation between the actual angular acceleration and the desired angular acceleration of the i-th UAV in the lower-level UAV swarm at time slot k, e Φ,i (k) represents the deviation between the actual attitude and the desired attitude of the i-th UAV in the lower-level UAV swarm at time slot k, e s,i (k) represents the deviation between the actual monitored area and the expected monitored area of the i-th UAV in the lower-level UAV swarm at time slot k. The deviation satisfies the following formula: Where, p e,i (k) represents the expected position of the i-th UAV in the lower-level UAV swarm at time slot k, p i (k) The actual position of the i-th UAV in the lower-level UAV swarm at time slot k, v e,i (k) represents the expected velocity of the i-th UAV in the lower-level UAV swarm at time slot k, v i (k) The actual velocity of the i-th UAV in the lower-level UAV swarm at time slot k, a e,i (k) represents the expected angular acceleration of the i-th UAV in the lower-level UAV swarm at time slot k, a i (k) The true angular acceleration Φ of the i-th UAV in the lower-level UAV swarm at time slot k. e,i (k) represents the desired attitude of the i-th UAV in the lower-level UAV swarm at time slot k, Φ i (k) The true attitude of the i-th UAV in the lower-level UAV swarm at time slot k, s e,i (k) represents the expected area of the i-th UAV in the lower-level UAV swarm at time slot k, s ture,i (k) The actual area of the i-th UAV in the lower-level UAV swarm at time slot k.
7. The method according to claim 5, characterized in that, Total energy consumed by the i-th drone in the lower-level drone swarm Total energy consumed by upper-level drones Where, n i (k)∈{0,1} represents the state of the i-th UAV in the lower-level UAV swarm at time slot k, and n i (k) = 1 indicates that the i-th drone in the lower-level drone swarm is in flight state at time slot k, and n i (k) = 0 indicates that the i-th drone in the lower-level drone swarm is in a hovering state at time slot k, E fly,i (k) represents the flight energy consumption of the i-th UAV in the lower-level UAV swarm at time slot k, E stop,i (k) represents the hovering energy consumption of the i-th UAV in the lower-level UAV swarm at time slot k, E tran,i (k) represents the total energy consumption of the i-th UAV in the lower-level UAV swarm for transmitting and receiving information in time slot k, E fly For the flight energy consumption of the upper-level drone, E stop (k) represents the hovering energy consumption of the upper-level UAV in time slot k, E com (k) represents the computational energy consumption of the upper-level UAV in time slot k, E tran (k) represents the total energy consumption of the upper-level UAV in time slot k for transmitting and receiving information.
8. The method according to claim 5, characterized in that, The overlap rate of the key monitoring area of the lower-level UAV in time slot k satisfies: The coverage of the efficient coverage area of the lower-level UAV in time slot k satisfies: Among them, s imp,i (k) represents the area of the key monitoring region for the i-th drone in the lower-level drone swarm at time slot k, s eff,i (k) represents the area of the efficient coverage of the i-th drone in the lower-level drone swarm at time slot k, where a = 0 indicates that the lower-level drone swarm is monitoring firefighters, and a = 1 indicates that the lower-level drone swarm is monitoring forest fires. ture,i (k)=s yuan,i (k)-s z,i (k) represents the total area of the actual monitoring area of the i-th drone in the lower-level drone swarm at time slot k. Let i be the total area that the i-th drone in the lower-level drone swarm should monitor in time slot k. Let θ be the area of obstacle obstruction within the monitoring area of the i-th drone in the lower-level drone swarm at time slot k. hor,i (k) represents the horizontal viewing angle of the camera carried by the i-th UAV in the lower-level UAV swarm at time slot k, θ ver,i (k) represents the vertical viewing angle of the camera carried by the i-th UAV in the lower-level UAV swarm at time slot k.
9. The method according to claim 5, characterized in that, The optimization problem is transformed into a Markov decision process, including: Construct the quintuple A =<S,D,O,R,γ> Where S is the state space, D is the action space, O is the observation set; R is the reward function, γ is the discount rate, and s i (k)=(o i (k),d i (k)),s i (k)∈S,s i (k) represents the state of the lower-level UAV i in time slot k, o i (k) represents the observation of the lower-level UAV i at time slot k, o i (k)=(p i (k),E all,i (k),e all,i (k)),o i (k)∈O, p i (k) represents the position of the i-th UAV in the lower-level UAV swarm at time slot k, E all,i (k) represents the total energy consumed by the i-th drone in the lower-level drone swarm at time slot k, e all,i (k) represents the total error of the i-th UAV in the lower-level UAV swarm at time slot k, d i (k) represents the action of the lower-level UAV i in time slot k, d i (k)=(v i (k),a i (k)),d i (k)∈D,v i (k) represents the velocity of the i-th UAV in the lower-level UAV swarm at time slot k, a i (k) represents the angular acceleration of the i-th UAV in the lower-level UAV swarm at time slot k, R i (k) represents the reward value of the i-th drone in the lower-level drone swarm at time slot k, R i (k)=R crash,i (k)+R energy,i (k)+R error,i (k),R i (k)∈R, R crash,i (k) represents the collision reward value of the lower-level drone i in time slot k, R energy,i (k) represents the energy consumption reward value of the lower-level drone i in time slot k, R error,i (k) represents the error reward value of the lower-level UAV i in time slot k.
10. The method according to claim 9, characterized in that, The optimization problem is solved using a multi-agent deep deterministic policy gradient algorithm based on Markov decision processes, including: constructing a policy network μ for each agent i. i (k) and value network η i (k); Get the current observation o i (k) will observe the current o i (k) Input the policy network to obtain action d i (k), execute action d i (k) Receive reward R i (k) and the observation at the next time step o i (k+1), will (o i (k),d i (k),R i (k),o i (k+1)), i=1,2,3,…,N are stored as a set of experience samples in the experience replay buffer L; a batch of experience samples are extracted from the experience replay buffer L to train the policy network and value network, and the trained policy network and value network are obtained; the formation control strategy of the multi-UAV swarm is obtained based on the policy network and value network.
Citation Information
Cited By
Urban fire monitoring-oriented unmanned aerial vehicle-WSN space-time collaborative optimization method
CN121578725A
Unmanned aerial vehicle-wsn spatiotemporal coordination optimization method for urban fire monitoring
CN121578725B