A method and system for forest fire detection and extinguishment

Through drone, satellite remote sensing and remote video combined with infrared thermal imaging maps and multi-source sensor data, the Markov decision-making process and MAPPO multi-agent reinforcement learning algorithm optimized drone scheduling, solving the problem of untimely detection and inaccurate response in forest fire detection and fire extinguishing, and achieving efficient and accurate forest fire recognition and destruction.

CN119838173BActive Publication Date: 2025-07-11NANJING UNIV OF INFORMATION SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510317102.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-11
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

The existing technology has problems in forest fire detection and fire extinguishing, such as untimely detection, unfast response and low intelligence, especially in complex terrain, which is difficult to detect early and effectively curb the spread of fire.

Method used

UAVs, satellite remote sensing and remote video are used for dynamic monitoring, combined with infrared thermal imaging maps and multi-source sensor data, and optimized drone scheduling through Markov decision-making process and MAPPO multi-agent reinforcement learning algorithm, realize fire point positioning and efficient throwing of fire extinguishing bombs, and build a forest fire detection and destruction system.

Benefits of technology

Accurate identification and efficient fire extinguishing of forest fires have been achieved, with wide coverage, small blind spots and high efficiency, and can respond in a timely manner and optimize resource allocation to ensure the rapid and safe fire extinguishing operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119838173B_ABST
    Figure CN119838173B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for forest fire detection and extinction, which relates to the technical fields of forest fire prevention and data processing. The method includes: carrying out fire patrol on the planned forest area through multiple inspection drones, and transmitting images and airborne sensor data back to the command center in real time; using artificial intelligence methods to identify fires, and making fire judgments and fire level classifications; the command center screens fire points according to the thermal imaging map and updates the fire point positions in real time; judging the spread trend of the fire scene based on the transmitted data, and providing a basis for the dispatching task of the fire extinguishing drones to drop fire extinguishing bombs; using the MAPPO multi-agent reinforcement learning algorithm to train the optimal dynamic fire point cooperative extinction control strategy; after extinguishing the open fire, the inspection drones continue to patrol to identify rekindling points or smoke points to prevent large-scale secondary fires. The method of the present invention can achieve efficient cooperative fire extinguishing, dynamic environment adaptation and resource optimization allocation, and maintain robustness when some agents fail.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of forest fire detection and extinction, and particularly to a forest fire detection and extinction method and system. Background Art

[0002] As an important ecological resource, once a forest fire breaks out, if the fire cannot be detected and the spread of the fire cannot be controlled in time, it will cause immeasurable ecological damage, economic losses and even casualties. However, traditional forest fire fighting is restricted by complex terrain, and it is difficult to deploy manpower, making it difficult to achieve the goals of early fire detection, rapid response and effective fire suppression. In recent years, unmanned aerial vehicle (UAV) technology has emerged and developed rapidly in the field of forest fire fighting, showing many unique advantages in the aspects of emergency warning and rapid fire extinguishing. Although there are already some related inventions based on UAVs that attempt to solve the problems of forest fire fighting, there are still many limitations.

[0003] Chinese Patent Application CN118718292A discloses a forest fire extinguishing method and system based on a UAV cluster. The initial fire is detected by the monitoring front end, and then the reconnaissance UAV conducts a secondary fire judgment. After the fire is determined, the fire extinguishing UAV is remotely controlled to carry fire extinguishing bombs to extinguish the fire. This application requires manual monitoring and remotely controls the fire extinguishing UAV to extinguish the fire manually, which is not convenient and intelligent enough.

[0004] Chinese Patent Application CN112132090A discloses a method for automatic detection and early warning of fireworks based on YOLOV3, which focuses on image processing and fire recognition, does not make a comprehensive judgment by combining UAV sensor data, has a certain possibility of misjudgment, and cannot take fire extinguishing related measures in time.

[0005] Based on this, how to conduct dynamic inspections through UAVs, fuse sensor data to accurately identify forest fires and locate the fire points, and how to use UAVs for rapid fire extinguishing response and re-ignition prevention remains to be further studied. Summary of the Invention

[0006] The problem to be solved by the present invention is: to provide a forest fire detection and extinction method and system, which conducts dynamic monitoring through UAVs, satellite remote sensing and remote videos, accurately identifies forest fires and locates the fire points, and autonomously learns the optimal strategy through interaction with the environment to perform resource allocation and task scheduling of UAVs in complex environments.

[0007] The present invention adopts the following technical solutions: A forest fire detection and extinction method, comprising the following steps:

[0008] S1. Fire situation data collection and fire situation identification: Fire situation data are collected in real time through satellite remote sensing, remote video monitoring, and cruising drones, transmitted to the command center, an infrared thermal imaging map is generated, and fire situation judgment and fire risk level classification are carried out;

[0009] S2. Fire point positioning: The command center marks potential fire points based on the infrared thermal imaging map, sorts the potential fire points according to the temperature value, obtains the geographical coordinates of the fire points from the GIS database, and updates them in real time according to the movement and change of the fire points in the continuous thermal imaging map, and marks and locates newly emerged fire points. ;

[0010] S3. Calculation of the fire spread trend: Based on the Rothermel model, fire spread simulation is carried out, the fire spread trend is calculated, and the fire spread boundary is updated;

[0011] S4. Construction of a drone scheduling model and model training: A drone scheduling model based on the Markov decision process is constructed, and the drone scheduling model is pre-trained through the MAPPO multi-agent reinforcement learning algorithm;

[0012] S5. Extinguishing of fire by the drone swarm: The fire extinguishing drones are respectively equipped with fire extinguishing bombs. The command center optimizes the scheduling through the trained drone scheduling model, determines the throwing order of each drone, takes the shortest time as the optimization goal, and cooperatively completes the investigation and extinguishing of all potential fire points;

[0013] S6. Prevention of re-ignition: After the open fire is extinguished, the fire site is continuously patrolled by the inspection drone to identify re-ignition points or smoke points.

[0014] Preferably, in step S1, satellite remote sensing is used to determine the fire site boundary and estimate the fire area; remote video monitoring is used to obtain forest site images and videos in real time; there are multiple cruising drones, which form a drone swarm, equipped with lidar, infrared thermal imagers, and temperature, smoke, and gas sensor devices. The ground surface is detected by lidar, the collected point cloud data are solved and orthophoto mosaicked, high-precision DEM is extracted and a topographic map is generated; the command center receives the fire situation images and sensor data transmitted back by satellite remote sensing, remote video monitoring, and cruising drones in real time, conducts fire situation judgment and fire risk level classification, displays them in the form of charts and heat maps, and issues review and fire extinguishing operation instructions after confirming the fire situation.

[0015] Preferably, in step S1, the command center conducts fire situation judgment and fire risk level classification, and the method is as follows:

[0016] S1.1. Use the YOLOv5 method to detect fire on image data. Perform real-time detection through a single forward propagation network. Use data enhancement, pre-trained weights, and mixed precision training to strike a balance between speed and accuracy. Determine the confidence level of the image containing flames or smoke, with a value range of 0 to 1.

[0017] S1.2. Generate an infrared thermal image of the monitored area, divide the image into several temperature areas by setting a temperature threshold, and determine the high-temperature area; obtain the current fire situation by calculating the ratio of the high-temperature area to the entire image area; and preliminarily conclude that the fire is in a state of expansion or attenuation by comparing the changes in the area of ​​the high-temperature area in the same area over a period of time.

[0018] S1.3, process the field data collected by the drone's temperature, smoke, and gas sensors; when the ambient temperature exceeds the temperature alarm threshold, activate the fire alarm command, and continue to collect temperature data and transmit it to the command center. The command center pre-processes the temperature data over a period of time and takes the average as the regional temperature;

[0019] The preprocessing adopts interpolation method to fill the missing data values, including linear interpolation and Lagrange interpolation, and estimates the value of the missing point based on the adjacent valid data. In principle, abnormal data is eliminated.

[0020] S1.4. For fire situation assessment, the command center uses a machine learning model to integrate data from different sensors with visual recognition results to comprehensively determine whether a fire has occurred. For fire level classification, the command center uses historical data to determine the fire risk of the area in the current season, and combines temperature, humidity, wind speed, and precipitation factors to set weights to obtain a fire score, and then classifies the current fire level based on the fire score.

[0021] Preferably, in step S3, the fire spread trend calculation includes the following sub-steps:

[0022] S3.1. Collect geographic information and environmental data of the monitoring area as input of the Rothermel model: obtain topographic data corresponding to the geographic coordinates of the fire point from the GIS database, including slope and aspect information; obtain wind speed, wind direction data and fuel moisture information near the fire point from meteorological data; obtain fuel characteristic parameters by scanning the monitoring area through satellite remote sensing or drones, and establish an estimation model to estimate the fuel load by analyzing the relationship between the spectral characteristics of the image and the fuel load;

[0023] The fuel characteristic parameters include: dried combustible content, dried particle density, surface area to volume ratio, combustibles, moisture content of combustibles, and extinguished moisture content of combustibles.

[0024] S3.2. Conduct fire spread simulation based on the Rothermel model.

[0025] S3.3. Calculate the fire spread speed under the conditions of initial calm wind and no slope:

[0026] According to the law of heat conduction and the input fuel characteristic parameters, under the ideal conditions of assuming calm wind and no slope, use the fire spread formula to calculate the basic fire spread speed R1 when the fire is under the initial conditions of calm wind, no slope, and the fuel moisture is the reference value (usually taken as the standard dry state).

[0027] S3.4. Comprehensively consider the influence of wind speed, slope, and fuel moisture, and calculate the actual spread speed under the combined action of wind speed and slope and the headwind spread speed:

[0028] Based on the wind speed, wind direction data collected in real time and the slope and aspect information corresponding to the ignition point location, considering the stretching and combustion-supporting effects of the wind on the flame and the fuel accumulation and heat radiation change factors caused by the slope, adjust the spread speed to obtain the corrected fire spread speed, and calculate the actual spread speed R2 of the fire under the combined action of wind speed and slope and the headwind spread speed R3 in the opposite direction of the fire head according to the correction formula.

[0029] S3.5. Update the fire spread boundary:

[0030] Based on the fire spread speeds R1, R2, and R3, combined with the preset time step , based on the kinematic principle, by the fire spread position , predict the coordinates of the next fire characteristic point spread ; Connect the spread coordinate points with a curve to obtain the new boundary of the fire spread, calculate the area of the new region, update the fire spread trend in real time and visualize it to form a fire scene development situation map.

[0031] S3.6. Update the fire spread speed and trend:

[0032] Based on the real-time updated infrared thermal imaging map, sensor data, and meteorological data, adjust through the Rothermel model and recalculate the fire spread speed and trend.

[0033] Preferably, in step S4, construct a UAV scheduling model based on the Markov decision process, and the method is as follows:

[0034] First, conduct problem description: The command center obtains the potential fire point sequence obtained by real-time fire point positioning , assume at time the set has a total of M fire points , the position is denoted as ; The UAV swarm has UAVs, each carrying a number of fire extinguishing bombs. The command center decides the throwing order of each UAV, aiming at the shortest time for optimization, and collaboratively completes the investigation and elimination of all potential fire points.

[0035] Construct a UAV scheduling model based on the Markov decision process, which is represented as a seven-tuple:

[0036] ;

[0037] Among them, is the state space; is the joint action space of is the local observation space of the agent, is the agent at the time step of the global state the local observation; is the state transition probability, indicating that given all agents in the state taking the joint action when, from the state transferring to probability; is the single-step reward after taking the joint action ; is the number of agents; is the discount factor, and the value range is [0, 1].

[0038] The random policy is a mapping from the observation space to the action space, which is represented as:

[0039] ;

[0040] Among them, represents the random policy of the agent used to give the probability distribution of a certain action under a specific observation.

[0041] Preferably, in the UAV scheduling model, the state space is specifically set as follows:

[0042] Each agent is assigned an independent state space, and the state observation of the agent, defined as the position, linear velocity, rotation matrix and ammunition number of the quadcopter UAV, is represented as:

[0043] ;

[0044] Among them, 、 、 、 respectively represent the position, linear velocity, rotation matrix and ammunition number of the agent at .

[0045] The environmental observation of the th agent is expressed as the relative positions of other agents within the communication range: 、 relative position , the relative position of the landing bay , approximated by the horizontal Euclidean distance:

[0046] ;

[0047] The environmental observation of the agent state is expressed as:

[0048] ;

[0049] Among them, is the flag bit indicating whether the th k agent has been destroyed at the moment in the fire field, is the priority of the th k agent at the moment.

[0050] t The state of the environment in the detection area at the moment is composed of the environmental observations of all agent states, expressed as:

[0051] ;

[0052] The joint action space of the agents is expressed as a discrete action space , and the action space of each agent is set the same ,

[0053] respectively representing left, right, forward, backward, up, down, throw, and return.

[0054] The throw action is only allowed to be executed when the distance to the target throwing location is less than the set threshold; the return action is only allowed to be executed when all tasks are completed, the set of potential fire points is empty and no longer updated for a period of time in the future.Preferably, a shared reward function is constructed in the UAV scheduling model , to maximize the advancement of the next one in the case of collision avoidance , the shared reward function is expressed as:

[0055] ;

[0056] where, is the reward function at each simulation time step :

[0057] ;

[0058] where, respectively represent the obstacle avoidance reward, task reward, all tasks completed reward, distance reward, and return reward, is the weight coefficient, used to adjust the importance of each part of the reward, and the sum of them is 1.

[0059] Preferably, the obstacle avoidance reward , gives different rewards according to whether the action taken causes the UAV to collide with an obstacle:

[0060] ;

[0061] where, is the distance between the centroid of the UAV and the obstacle, is the maximum wheelbase of the UAV.

[0062] Preferably, the task reward , is associated with the task priority, and gives different rewards according to whether the correct throw is completed:

[0063] ;

[0064] where, is the set completed at the th step, is the corresponding priority, corresponds to the flag bit indicating whether is destroyed.

[0065] Preferably, the all tasks completed reward :

[0066] ;

[0067] Only when all tasks are completed, the potential fire point set is empty, and no longer updated for a period of time later, give a reward once. ​

[0068] Preferably, the distance reward is related to the closest distance from each UAV to the uncompleted node. By adjusting the coefficient, the curvature of the reward function is adjusted. As long as the target point is destroyed by any UAV, all UAVs will receive a reward:

[0069] ;

[0070] For each uncompleted , the distance between the UAV closest to it is: For:

[0071] ;

[0072] Among them, is the throwing height, ([[]] ), ([[]] ) respectively represent the corresponding and the global coordinate positions of the UAVs.

[0073] Preferably, the return reward is:

[0074] ;

[0075] When all task points are completed, if the UAV successfully flies to the landing pad nearby, a large positive reward will be given, otherwise a small negative reward will be given.

[0076] Preferably, in step S4, the MAPPO multi-agent reinforcement learning algorithm adopts a centralized training - decentralized execution framework, training two independent neural networks: an actor network and a central critic network;

[0077] The actor network is represented as a policy network with parameters , and each agent has an independent actor network, making independent decisions based on their respective local information; The central critic network is represented as a state value function with parameters

[0078] , and all agents share a central critic network, evaluating the value of each agent by accessing the global state information of all agents. The input feature of the actor network is the normalized environmental information, and the output is the probability distribution of discrete actions; the input of the central critic network is the normalized centralized global state information, and the output is the state value estimate;

[0079] The input feature of the actor network is the normalized environmental information, and the output is the probability distribution of discrete actions; the input of the central critic network is the normalized centralized global state information, and the output is the state value estimate;

[0080] Both the actor network and the central critic network use a multi-layer perceptron to process the observation information. The number of neurons in the hidden layer is 256. The ReLU activation function is used, and the Adam optimizer is used to maximize the policy gradient and minimize the value function loss for network update.

[0081] Preferably, in step S4, the UAV scheduling model is pre-trained based on the MAPPO algorithm, including the following sub-steps:

[0082] S4.1. The agent selects and executes an action according to the current policy in the environment. The environment feeds back a new state and reward according to the agent's action. The agent stores the information in the experience replay buffer and records the trajectory.

[0083] S4.2. According to the collected trajectory, the Generalized Advantage Estimation (GAE) method is used to estimate the advantage function , and the estimation method is:

[0084] S4.2.1. Calculate the TD error, through reflecting the error in the estimation of the current state value function :

[0085] ;

[0086] where represents the discount factor; represents the termination flag; is the return reward after the agent takes the corresponding action, and after PopArt regularization, it is .

[0087] S4.2.2. Calculate the GAE advantage estimation by introducing the decay factor and combining the TD errors of different time steps:

[0088] ;

[0089] where represents the offset of the time step.

[0090] S4.3. For the training of each actor, calculate the following objective function:

[0091] ;

[0092] where the parameter represents the actor network parameters, is the policy entropy hyperparameter, is the policy entropy, represents the clipped expected cumulative discounted reward.

[0093] S4.4. Maximize using gradient ascent : Calculate the gradient of the policy loss , and then update the actor network parameters using the stochastic gradient ascent method , where is the learning rate.

[0094] S4.5. For the training of the critic, minimize the following value function loss using gradient descent:

[0095] ;

[0096] where, represents a more accurate estimate of the future expected return given the current state, The function represents restricting x within the range [a, b], represent the state value functions under the old and new parameters respectively, used to calculate the expected cumulative discounted reward; represents under the policy The following expected operator.

[0097] S4.6. Repeat steps S4.1 to S4.5 until the policy converges to the optimal solution or reaches the maximum number of training rounds.

[0098] The technical solution of the present invention also provides: A forest fire detection and extinguishment system for implementing any of the above methods, including: a fire situation identification module, a fire point location and fire field spread trend calculation module, a fire situation extinguishment module, and a rekindling prevention module;

[0099] The fire situation identification module collects fire situation data in real time through satellite remote sensing, remote video monitoring, and cruising drones, transmits it to the command center, generates an infrared thermal imaging map, and conducts fire situation judgment and fire level classification;

[0100] The fire point location and fire field spread trend calculation module marks potential fire points based on the infrared thermal imaging map, sorts the potential fire points according to the temperature value, accurately determines the geographical coordinates of the fire points in combination with multi-source geographic information data, and updates in real time according to the movement and change of the fire points in the continuous external thermal imaging map, marks and locates newly emerging fire points; and conducts fire spread simulation based on the Rothermel model, calculates the fire field spread trend and updates the fire spread boundary;

[0101] The fire situation extinguishment module constructs a drone scheduling model based on the Markov decision process, and pre-trains the drone scheduling model through the MAPPO multi-agent reinforcement learning algorithm; the cruising drones are equipped with fire extinguishing bombs, and through the command center based on the trained drone scheduling model, optimize the scheduling, decide the throwing order of each drone, and cooperate to complete the investigation and extinguishment of all potential fire points with the shortest time as the optimization goal;

[0102] After extinguishing the open fire, the afterburning prevention module continues to patrol the fire area through the inspection UAV to identify afterburning points or smoke points.

[0103] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:

[0104] 1. The forest fire detection and extinction method of the present invention combines satellite remote sensing, remote video monitoring, and dynamic inspection of UAVs, with a larger detection range, smaller blind spots, and higher efficiency.

[0105] 2. The forest fire detection and extinction method of the present invention conducts comprehensive early warning through fire recognition and sensor data fusion, which is more timely, accurate, and effective.

[0106] 3. In the UAV scheduling task of the method of the present invention, reinforcement learning can autonomously learn the optimal strategy through interaction with the environment, which is used to optimize the resource allocation and task scheduling of UAVs in a dynamic and complex environment, making the rescue faster, more optimized, and safer. Brief Description of the Drawings

[0107] Figure 1 It is a flow block diagram of the forest fire detection and extinction method of the present invention;

[0108] Figure 2 It is a schematic diagram of the application scenario of the MAPPO optimization strategy of the present invention;

[0109] Figure 3 It is a schematic diagram of the MAPPO algorithm framework of the present invention. Detailed Embodiments

[0110] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the application will be further elaborated in detail below with reference to the accompanying drawings. The described embodiments are only a part of the embodiments involved in the present invention. All non-innovative embodiments made by other researchers in the field based on this embodiment fall within the protection scope of the present invention. At the same time, for the step numbers in the embodiments of the present invention, they are only set for the convenience of elaboration and explanation, and no limitation is placed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0111] In an embodiment of the present invention, a forest fire detection and extinction method, as Figure 1 shown, includes the following steps:

[0112] S1. Fire situation data collection and fire situation recognition: Fire situation data is collected in real time through satellite remote sensing, remote video monitoring, and cruising UAVs, transmitted to the command center, an infrared thermal imaging map is generated, and fire situation judgment and fire level classification are carried out;

[0113] S2. Fire point positioning: The command center marks potential fire points based on the infrared thermal imaging map, sorts the potential fire points according to the temperature values, obtains the geographical coordinates of the fire points from the GIS database, and updates them in real time according to the movement and change of the fire points in the continuous thermal imaging map, and marks and locates newly emerging fire points;

[0114] S3. Calculation of the fire spread trend: Based on the Rothermel model, fire spread simulation is carried out to calculate the fire spread trend and update the fire spread boundary;

[0115] S4. Construction of the UAV scheduling model and model training: Construct a UAV scheduling model based on the Markov decision process, and pre-train the UAV scheduling model through the MAPPO multi-agent reinforcement learning algorithm;

[0116] S5. Extinguishing forest fires by the UAV swarm: The fire-extinguishing UAVs are respectively equipped with fire extinguishing bombs. The command center optimally schedules through the trained UAV scheduling model, decides the throwing order of each UAV, and takes the shortest time as the optimization goal to cooperate to complete the investigation and extinction of all potential fire points;

[0117] S6. Prevention of re-ignition: After extinguishing the open fire, continue to patrol the fire area by the inspection UAV to identify re-ignition points or smoke points.

[0118] Specifically, in step S1, the fire situation data collection and fire situation identification include the following sub-steps:

[0119] 1. Combine three forest fire detection means: satellite remote sensing, remote video monitoring, and UAV cruise to collect fire situation data.

[0120] Satellite remote sensing technology can cover a vast geographical area and has significant advantages for large-scale and cross-regional forest fire monitoring. It can relatively accurately determine the fire area boundary and estimate the fire area, providing important references for subsequent fire extinguishing and rescue work.

[0121] For some small-area fires or initial fires, satellite remote sensing technology may not be able to accurately identify and locate them, and it is easy to have missed reports or false reports. Therefore, in this embodiment, monitoring points are set in key forest areas or fire-prone areas, and the remote video monitoring system combined with forest layout is used to obtain the images and videos of the forest scene in real time and transmit them back to the command center.

[0122] Affected by factors such as terrain and tree occlusion, some areas may not be monitored, thus forming monitoring blind spots. Therefore, in this embodiment, multiple inspection UAV clusters are dispatched to conduct fire inspections on the planned forest area. The UAVs are equipped with sensing devices such as temperature sensors, smoke alarm sensors, infrared thermal imaging sensors, and gas sensors, and transmit relevant images and sensor data collected in real time to the command center for fire situation judgment and fire level classification.

[0123] The UAV cluster is also equipped with lidar for detecting the ground surface. After the point cloud data is solved and the orthophoto is mosaicked, high-precision DEM is extracted and a topographic map is generated.

[0124] 2. The command center makes a fire situation judgment and fire level classification by combining data from all parties.

[0125] The command center uses the data from all parties as the fire situation judgment criteria, sets their respective weights, calculates the weighted fire situation score, and divides the fire level according to the score. The data used includes:

[0126] 1) Image data: Use YOLOv5 to detect the fire situation in the image data, determine the confidence level that the image contains flames or smoke, and the value range is 0-1.

[0127] YOLOv5 is an efficient and flexible object detection model. In this embodiment, based on the YOLOv5 model, real-time detection is achieved through a single forward propagation network, and a good balance between speed and accuracy is achieved through techniques such as data augmentation, pre-trained weights, and mixed-precision training.

[0128] 2) Infrared thermal imaging map: By setting a temperature threshold, the image is divided into several temperature regions, and the ratio of the area of the high-temperature region to the total area of the image is calculated to obtain the current fire situation; further, by comparing the change in the high-temperature area in the same region over a period of time, it is initially concluded whether the fire is expanding or decaying.

[0129] 3) Sensor data: The data collected by sensor devices such as temperature, smoke, and gas mainly comes from on-site collection by UAVs.

[0130] Taking the temperature sensor as an example, when the ambient temperature exceeds the temperature alarm threshold, the temperature sensor activates the fire alarm, continuously collects the temperature around the area, and transmits it to the command center; the command center preprocesses the temperature collected during this period and takes the average value as the temperature of the area.

[0131] The specific method of preprocessing is as follows: If data is missing, interpolation methods such as linear interpolation and Lagrange interpolation are used to fill in the missing values, and the values of the missing points are estimated based on the adjacent valid data before and after; if data is abnormal, it is excluded according to the principle of 3 Principle for rejection.

[0132] 4) The daily forest fire danger weather level (hereinafter referred to as the fire danger level): Released by the China Meteorological Administration and the Forestry Bureau, it is mainly calculated based on factors such as temperature, humidity, wind force, and precipitation.

[0133] In this embodiment, the fire danger levels 1 to 5 are quantified as 0, 20, 60, 80, and 100. The command center uses machine learning models (such as SVM, random forest, neural network) to fuse data from different sensors with visual recognition results and comprehensively judge whether a fire has occurred.

[0134] 5) Historical data: The fire risk level of this area in the current season obtained from historical data.

[0135] 3. After confirming the fire, inform the inspection personnel to conduct a review and start the fire extinguishing operation. The fire situation data and analysis results collected by the UAV are displayed in the form of charts, heat maps, etc., which is convenient for relevant departments to monitor and make decisions.

[0136] Specifically, in step S2, the fire point positioning includes the following sub-steps:

[0137] 1. Thermal imaging map data collection and transmission

[0138] Based on the infrared thermal imaging equipment to collect real-time thermal imaging map data, the infrared thermal imaging equipment includes remote video monitoring and inspection UAVs. Adopting a combination of multiple technologies such as 5G network and satellite communication, a stable and high-speed data transmission link is established to ensure that the image data collected by the thermal imaging equipment is transmitted to the command center in real time, avoiding affecting the fire point positioning and trend judgment due to transmission delay or interruption.

[0139] 2. Fire point screening and positioning in the command center

[0140] After the command center receives the infrared thermal imaging map, it uses professional image processing software and algorithms to quickly analyze the image. By setting a temperature threshold, the area above a certain range of the normal ambient temperature is marked as a potential fire point, and the fire points are prioritized according to the temperature value. Further combined with multi-source geographic information data, the geographical coordinates of the fire points are accurately determined, and the error is controlled within a very small range for the accurate implementation of subsequent fire extinguishing operations.

[0141] 3. Real-time fire point position update

[0142] The infrared thermal imaging monitoring equipment continuously scans the monitoring area, re-collects image data at a fixed time interval (such as every 30 seconds) and transmits it to the command center. The command center automatically compares the thermal imaging maps of the previous and subsequent time periods, and combines with the geographical information system (GIS) technology to update the geographical position coordinates of the fire points in real time according to the movement and change of the fire points in the image. At the same time, newly emerged fire points are marked and located in a timely manner to ensure a comprehensive grasp of the fire scene dynamics.

[0143] Specifically, in step S3, the calculation of the fire spread trend includes the following sub-steps:

[0144] 1. Collection of geographical information and environmental data

[0145] Using high-precision geographical mapping technology, detailed topographic and geomorphic data within the monitoring area are obtained in advance, including slope and aspect information, and integrated into the GIS database.

[0146] For the location of the ignition point, after precise positioning through infrared thermal imaging combined with GIS technology, the corresponding slope and aspect information is retrieved from the GIS database in real time and fed back to the input data interface of the subsequent fire spread model, providing topographic basic data for accurate simulation of fire spread.

[0147] The command center establishes a close data sharing and cooperation mechanism with the local meteorological bureau to obtain real-time wind speed, wind direction data and fuel moisture-related information near the ignition point. These data will be used as key input parameters for subsequent fire spread calculations.

[0148] In addition, the monitoring area is regularly scanned by satellite remote sensing, and the fuel load is estimated through professional data processing algorithms, specifically including: dry combustible content, dry particle density, surface area to volume ratio, combustibles, and combustible moisture content, combustible extinction moisture content and other parameters, providing comprehensive and accurate fuel characteristic data for the fire spread model.

[0149] 2. Fire spread simulation based on the Rothermel model

[0150] In this embodiment, the Rothermel model takes into account factors such as slope and aspect, fuel bed and forest density, complies with various heat conduction laws, and has wide applications. The expression is as follows:

[0151] ;

[0152] In the formula, R is the forest fire spread speed, unit (m / min); is the reaction intensity in the flame zone, unit ; is the forest fire spread rate, is the wind speed correction coefficient, is the slope correction coefficient, is the effective heat coefficient, all dimensionless coefficients; is the density of combustibles, unit ; is the heat required to ignite a unit mass of combustibles, unit (kJ / kg).

[0153] The specific calculation steps of the Rothermel model are as follows:

[0154] 1) Input data

[0155] Parameters such as the location information of the ignition point, wind speed, wind direction, dry combustible content, dry particle density, surface area to volume ratio, combustibles, moisture content of combustibles, and moisture content at which combustibles extinguish obtained, as well as the parameters ( , ) are obtained from the empirical formula fitted from historical data, and are accurately input into the model calculation interface in sequence according to the requirements of the Rothermel model.

[0156] 2) Calculate the fire spread speed under the conditions of no wind and no slope initially

[0157] Based on the law of heat conduction and the input fuel characteristic parameters, under the ideal conditions of assuming no wind and no slope, complex operations are performed on each parameter using the built-in fire spread formula to calculate the fire spread speed under the conditions of no wind and no slope initially, and this speed is used as the basic data for considering the influence of environmental factors subsequently.

[0158] 3) Calculate the actual spread speed under the combined action of wind speed and slope and the headwind spread speed

[0159] Furthermore, the wind speed, wind direction data collected in real time and the slope and aspect information corresponding to the ignition point location are introduced. Considering comprehensively factors such as the stretching and combustion-supporting effects of wind on the flame and the fuel accumulation and heat radiation changes caused by the slope, the specific correction formula is used again to accurately adjust the spread speed, and the actual spread speed of the fire under the combined action of wind speed and slope and the headwind spread speed in the opposite direction of the fire head are calculated respectively.

[0160] The specific correction formula accurately adjusts the spread speed, and the fire spread speed after comprehensively considering the influence of wind speed, slope, and fuel moisture:

[0161] ;

[0162] Among them, U represents the wind speed, generally using the average wind speed at a certain height (such as 10 meters) above the ground; S the slope, usually represented by the tangent value of the slope angle, that is, the ratio of the vertical height to the horizontal distance; M represents the fuel moisture, expressed as a percentage; is an empirical coefficient, calibrated according to experimental data or historical fire data.

[0163] According to the above correction formula, the actual spread speed R2 of the fire under the combined action of wind speed and slope and the headwind spread speed R3 in the opposite direction of the fire head are calculated respectively.

[0164] 4) Update the fire spread boundary

[0165] Based on the calculated fire spread speeds under different conditions, combined with a pre-set time step (such as 15 minutes as one calculation step), using the principles of kinematics, by analyzing the fire spread position and the velocity vector , predict the coordinates of the next fire characteristic points (such as the coordinates of the flame front and the boundary of the high-temperature core area) :

[0166] ;

[0167] ;

[0168] Connect these spread coordinate points with curves to obtain the new boundary of the fire spread during this period, and calculate the area of the new region; meanwhile, update the fire spread trend information in real time and visualize it to provide an intuitive and dynamic map of the fire scene development. As the fire develops and evolves, the thermal imaging map, sensor data, and meteorological data are continuously updated. The system will continuously feedback the real-time new data to the Rothermel model, and the model will automatically adjust the parameters according to the new information and recalculate the fire spread speed and trend to ensure that the prediction results are always close to the actual situation of the fire scene.

[0169] Through the above fire scene positioning and fire spread trend calculation, provide accurate and timely decision-making basis for fire extinguishing operations, such as providing reliable operation guidelines for the dispatching task of fire extinguishing drones to drop fire extinguishing bombs, and ensuring the efficient allocation and utilization of fire extinguishing resources.

[0170] Specifically, in step S4, constructing a drone dispatching model and conducting model training includes the following sub-steps:

[0171] Problem description: The command center obtains a dynamically updatable sequence of potential fire points in real time through fire point positioning

[0172] , assuming that at time , the set contains a total of fire points M , and their positions are denoted as . .

[0173] In this embodiment, the existing problem is: there are drones, each carrying a quantity of fire extinguishing bombs. The system needs to make autonomous decisions on the throwing order of each drone to complete the investigation and extinguishment of all potential fire points in the shortest time as the optimization goal.

[0174] The multi-UAV scheduling problem involves the allocation of multiple complex task points, and the dynamic update of task points further increases the difficulty of the problem, making it difficult for traditional methods to respond in real time and effectively. In this embodiment, a reinforcement learning method is adopted to learn the optimal strategy through the interaction between the agent and the environment, providing a good solution to this problem.

[0175] Specifically, this embodiment uses a Markov decision process (DEC-POMDP) to describe the decision-making model for the dynamic scheduling and control of a UAV swarm, and defines the model as a seven-tuple:

[0176] ;

[0177] Among them, is the state space; is the joint action space of agents; is the local observation of agent i at time step under the global state is the state transition probability, indicating the probability of transitioning from state to when all agents take the joint action in state ; is the single-step reward after taking the joint action , which is calculated by the reward function for the reward obtained by taking the joint action in state ; is the number of agents; is the discount factor, with a value range of [0,1], which decays with the number of training steps.

[0178] The stochastic policy is a mapping from the observation space to the action space, giving the probability distribution of a certain action under a specific observation, denoted as ; among them, represents the stochastic policy of agent , which is used to give the probability distribution of a certain action under a specific observation.

[0179] 1. State observation space

[0180] Each agent is assigned an independent state space, and the state observation of the th agent is defined as the position, linear velocity, rotation matrix, and ammunition number of the quadrotor UAV, expressed as:

[0181] ;

[0182] Among them, , , , respectively represent the position, linear velocity, rotation matrix, and ammunition number of the agent at .

[0183] The th agent's environmental observation is the relative positions with other agents within the communication range, and the landing bay (with a fixed position). Among them, the relative distances between the agent and the observed and all landing bays are represented by the horizontal Euclidean distance:

[0184] ;

[0185] Agent state environmental observation , which is expressed as:

[0186] ;

[0187] Among them, is the flag bit indicating whether the th k agent has been destroyed at the th moment, and k th is the priority of the

[0188] The state of the environment within the detection area at time t is composed of all agent state environmental observations, which is expressed as:

[0189] .

[0190] 2. Joint action space

[0191] The joint action space is represented as a discrete action space , and the action space of each agent is set the same , which respectively represent left, right, forward, backward, up, down, throw, and return.

[0192] The throw action is only allowed to be executed when the distance to the target throwing location is less than a certain value, and this action is prohibited at other times; similarly, the return action is only allowed to be executed when all tasks are completed, the set of potential fire points is empty and will not be updated for some time in the future, and this action is prohibited at other times.

[0193] 3. Shared reward function

[0194] The goal is to directly maximize the advancement of the next one in the case of avoiding collisions ; For a fully collaborative problem, the reward functions of all agents are the same, and the reward function is defined at each simulation time step t as follows:

[0195] ;

[0196] Among them, is the weight coefficient, used to adjust the importance of each part of the reward, and the sum of them is 1; respectively represent the obstacle avoidance reward, task reward, all tasks completed reward, distance reward, and return reward.

[0197] Therefore, the shared reward function is:

[0198] ;

[0199] (1) Obstacle avoidance reward:

[0200] ;

[0201] Among them, is the distance between the centroid of the UAV and the obstacle, is the maximum wheelbase of the UAV.

[0202] If the action taken causes the UAV to collide with an obstacle, a large negative reward is given; if it is relatively close to the obstacle, a certain negative reward is given; otherwise, a small positive reward is given.

[0203] (2) Task reward:

[0204] ;

[0205] Among them, is the completed at the set, is the corresponding priority.

[0206] Once the throwing is completed at an uncompleted , a positive reward will be obtained, and the reward value is linked to the task priority; otherwise, an incorrect throwing will get a huge negative reward.

[0207] In the actual process, the UAV checks whether the horizontal Euclidean distance between the current position and is within the acceptable range before each throwing, and rechecks the potential fire point. After confirming it as a fire point, the fire extinguishing bomb is thrown.

[0208] (3) One-time reward for task completion:

[0209] ;

[0210] Only when the set of potential fire points is empty and there is no update for a period of time later, a reward is given once.

[0211] (4) Distance reward:

[0212] For each unfinished , the distance to the nearest drone is:

[0213] ;

[0214] where is the throwing height; ([[]] ) and ([[]] ) respectively represent the global coordinate positions of the corresponding and the drone.

[0215] Construct the following distance function, which is related to the nearest distance of each drone to the unfinished node. As long as the target point is destroyed by any drone, all drones will receive a positive reward to ensure that there is no competition among multiple drones:

[0216] ;

[0217] In addition, this term is used as a guiding term to make the reward more intensive, and the curvature of the reward function can be adjusted by adjusting the coefficient.

[0218] (5) Return reward:

[0219] ;

[0220] When all task points are completed, if the drone successfully flies to the nearest landing bay, a large positive reward is given, otherwise a small negative reward is given.

[0221] Furthermore, the drone scheduling model is pre-trained through the MAPPO multi-agent reinforcement learning algorithm. Different numbers of batches of drones are pre-trained to handle the fire extinguishing tasks in various fire levels. The command center dispatches a total of N fire extinguishing drones in sequence from the landing bay closest to the fire scene according to the fire level, and calls the pre-trained model.

[0222] In this embodiment, as Figure 2 shown, it is set that each drone can carry at most four fire extinguishing bombs. Each drone will achieve autonomous obstacle avoidance, assist in extinguishing the dynamically updated fire points, and return to the nearest bay.

[0223] ​​Specifically, PPO (Proximal Policy Optimization) is an on-policy reinforcement learning algorithm improved based on the policy gradient algorithm. It has achieved strong final rewards and sample efficiency in various cooperative multi-agent challenges and has been widely applied.

[0224] In this embodiment, the MAPPO (PPO in Multi-Agent Settings) multi-agent reinforcement learning algorithm is an extension of the PPO algorithm in the multi-agent field. As Figure 3 shown, it adopts the centralized learning - decentralized execution (CLDE) framework to train two independent neural networks: one is a policy network with parameters (also known as the actor), and the other is a state value function with parameters (also known as the critic).

[0225] All agents share a central critic network, which can more accurately evaluate the value of each agent by accessing the global state information of all agents. At the same time, each agent has an independent actor network, which can make independent decisions based on its own local information, improving the flexibility and robustness of the system.

[0226] In this embodiment, the specific training process of the UAV scheduling model is as follows:

[0227] The agent selects and executes actions according to the current policy in the environment. The environment feedbacks new states and rewards according to the agent's actions. The agent stores this information in the experience replay buffer and records the trajectory. MAPPO maximizes the cumulative reward based on the policy gradient and updates the value network of the agent using the stored experience data.

[0228] For the training of each actor, maximize the objective function:

[0229] ;

[0230] Among them, the parameter represents the actor network parameters, is the policy entropy hyperparameter.

[0231] represents the policy entropy:

[0232] ;

[0233] Among them, Represents the current policy The probability of taking action a in state s;

[0234] Represents the clipped expected cumulative discounted reward:

[0235] ;

[0236] Where, Represents the expected operator under policy And, Is the clipping coefficient, and the ratio function Represents the action probability ratio of the old and new policies at time t.

[0237] Maximize using gradient ascent , calculate the policy loss gradient , and then update the actor network parameters using the stochastic gradient ascent method :

[0238] ;

[0239] Where, Is the learning rate; Are the actor network parameters before update.

[0240] Is the advantage function estimate, estimated using the GAE (Generalized Advantage Estimation) method and regularized using the PopArt (Preserving Outputs Precisely, while Adaptively Rescaling Targets) method. The calculation steps are as follows:

[0241] 1) Calculate the TD (Temporal Difference) error: Reflects the error in the current state value function Estimation:

[0242] ;

[0243] Where, Represents the discount factor; Represents the termination flag; Is the return reward after the agent takes the corresponding action, and after PopArt regularization, it is ;

[0244] 2) Calculate the GAE advantage estimate: By introducing a decay factor To combine the TD errors of different steps:

[0245] ;

[0246] When it is, it degrades to a single-step TD estimate; when it is, it is equivalent to a Monte Carlo estimate.

[0247] In actual calculation, since a trajectory is a finite time step [0,…,T], the above formula can be approximated as:

[0248] ;

[0249] For the training of the critic, minimize the value function loss:

[0250] ;

[0251] where represents a more accurate estimate of the future expected return given the current state, the function means to limit x within a , b , represent the state value functions under the old and new parameters respectively, used to calculate the expected cumulative discounted reward, represents the clipping coefficient.

[0252] During the actual execution process, the agents call their respective optimal policies to interact with the environment, repeating the above environment interaction and learning steps until an optimal policy that can maximize the long-term cumulative reward is learned. The specific algorithm steps are as follows:

[0253] 1) Initialize the weights and and initialize them orthogonally, initialize the experience replay pool ;

[0254] 2) Initialize each agent state and start the iteration of the training round;

[0255] 3) Each agent interacts with the environment according to its respective policy function to collect trajectories, record the experience data at each time step t and store it in the experience replay pool;

[0256] 4) According to the collected trajectories, use the GAE method to estimate the advantage function ;

[0257] 5) Randomly sample a batch of experience data from the experience replay pool for the calculations in steps 6 and 7 below;

[0258] 6) Calculate the policy function loss using advantage estimation, and update the actor network parameters using the gradient ascent method ;

[0259] 7) Calculate the value function loss and update the critic network parameters using the gradient descent method ;

[0260] 8) Repeat this process for multiple iterations until the policy converges to the optimal solution or reaches the maximum number of training epochs.

[0261] In the specific training process of this embodiment, the policy network inputs the environment information after feature normalization and outputs the probability distribution of discrete actions. The value network inputs the centralized global state information after normalization and outputs the state value estimate. Both use a multi-layer perceptron (MLP) to process the observation information, with 256 neurons in the hidden layer and the ReLU activation function is used. The Adam optimizer is used to maximize the policy gradient and minimize the value function loss to achieve network update.

[0262] Furthermore, model evaluation: Use the test environment to evaluate the model. The test environment is different from the training environment in some parameter or scenario settings, but maintains the same basic structure and rules. The evaluation of the model includes the cumulative reward in an evaluation episode, calculating the average reward and its standard deviation over multiple evaluation episodes, the success rate of the agent in completing the task, etc.

[0263] Furthermore, model hyperparameter tuning: If the model performance is not ideal, it is necessary to further adjust hyperparameters such as the learning rate, discount factor, discount factor of advantage estimation, entropy coefficient, clipping range, training batch size, exploration rate, etc., or adjust network structure parameters such as the number and layers of neurons, and the type of activation function.

[0264] Specifically, in step S5, for the extinguishment of the fire in the UAV swarm, the method is as follows:

[0265] The cruise UAVs are respectively equipped with fire extinguishing bombs. The command center uses the trained UAV scheduling model for optimal scheduling, makes autonomous decisions on the throwing order of each UAV, and collaboratively completes the investigation and extinguishment of all potential fire points with the shortest time as the optimization goal.

[0266] Specifically, in step S6, for preventing re-ignition, the method is as follows:

[0267] After extinguishing the open fire, the inspection UAV patrols the fire scene to identify rekindling points or smoke points, preventing large-scale secondary fires. After receiving the rekindling point information feedback by the inspection UAV, the command center immediately activates the emergency response plan. On the one hand, it quickly dispatches the nearest fire-fighting forces according to the location of the rekindling point, such as ground personnel carrying fire extinguishers, UAVs equipped with fire bombs, etc. to go for rescue, ensuring that the fire is controlled in the initial stage of rekindling. On the other hand, it conducts secondary security around the rekindling area, evacuating the potentially affected people to prevent greater losses caused by large-scale secondary fires.

[0268] This embodiment further includes: a forest fire detection and extinguishment system for implementing the above method, including: a fire situation identification module, a fire point positioning and fire scene spread trend calculation module, a fire situation extinguishment module, and a rekindling prevention module;

[0269] The fire situation identification module collects fire situation data in real time through satellite remote sensing, remote video monitoring, and cruising UAVs, transmits it to the command center, generates an infrared thermal imaging map, and conducts fire situation judgment and fire risk level classification;

[0270] The fire point positioning and fire scene spread trend calculation module marks potential fire points based on the infrared thermal imaging map, sorts the potential fire points according to the temperature value, accurately determines the geographical coordinates of the fire points in combination with multi-source geographical information data, and updates them in real time according to the movement and change of the fire points in the continuous infrared thermal imaging map, marking and positioning newly emerging fire points; and conducts fire spread simulation based on the Rothermel model, calculates the fire scene spread trend and updates the fire spread boundary;

[0271] The fire situation extinguishment module constructs a UAV scheduling model based on the Markov decision process and pre-trains the UAV scheduling model through the MAPPO algorithm; through the cruising UAV carrying fire bombs, the command center conducts optimized scheduling based on the trained UAV scheduling model, decides the throwing order of each UAV, and takes the shortest time as the optimization goal to cooperate to complete the investigation and extinguishment of all potential fire points;

[0272] The rekindling prevention module, after extinguishing the open fire, continues to patrol the fire scene through the inspection UAV to identify rekindling points or smoke points.

[0273] In summary, for the forest fire detection and extinguishment method of the present invention, multiple patrol drones conduct fire patrols in the planned forest area, and transmit images and airborne sensor data back to the command center in real time. Artificial intelligence methods are used to identify fires and make fire judgments and fire level classifications in combination with various data. After confirming the fire, relevant personnel are informed to conduct a review and start operations. The command center screens fire points based on the thermal imaging map, updates the fire point positions in real time, judges the spread trend of the fire scene based on the transmitted data, and provides a basis for the dispatching task of the fire extinguishing drones to drop fire extinguishing bombs. The MAPPO multi-agent reinforcement learning algorithm is used to train the optimal dynamic fire point collaborative extinguishment control strategy to achieve efficient collaborative fire extinguishment, dynamic environment adaptation, and resource optimization allocation, and maintain robustness when some agents fail. At the same time, after extinguishing the open fire, the patrol drones patrol the fire scene to identify re-ignition points or smoke points to prevent large-scale secondary fires.

[0274] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for forest fire detection and extinguishment, characterized in that, It includes the following steps: S1. Fire situation data collection and fire situation identification: Fire situation data is collected in real time through satellite remote sensing, remote video monitoring, and cruising drones, and transmitted to the command center to generate an infrared thermal imaging map for fire situation judgment and fire risk level classification; S2. Fire point positioning: The command center marks potential fire points based on the infrared thermal imaging map, sorts the potential fire points according to the temperature values, obtains the geographical coordinates of the fire points from the GIS database, and updates them in real time according to the movement and change of the fire points in the continuous thermal imaging map, and marks and locates newly emerging fire points; S3. Calculation of the fire spread trend: Based on the Rothermel model, fire spread simulation is carried out to calculate the fire spread trend and update the fire spread boundary; S4. Construction of a drone scheduling model and model training: A drone scheduling model based on the Markov decision process is constructed, and the drone scheduling model is pre-trained through the MAPPO multi-agent reinforcement learning algorithm; The problem description of the UAV scheduling model based on the Markov decision process is as follows: The command center obtains the sequence of potential fire points in real time through fire point positioning , assuming at time there are M fire points , and their positions are recorded as ; the UAV swarm has UAVs, each carrying a number of fire extinguishing bombs. Through the decision of the command center on the throwing order of each UAV, with the shortest time as the optimization goal, the UAVs cooperate to complete the investigation and extinction of all potential fire points; S5. Fire extinguishing of the drone swarm: The fire extinguishing drones are respectively equipped with fire extinguishing bombs. The command center optimally schedules through the trained drone scheduling model, decides the throwing order of each drone, and takes the shortest time as the optimization goal to cooperate to complete the investigation and extinguishing of all potential fire points; Construct a shared reward function , maximize the advancement of the next while avoiding collisions : ; wherein, is the reward function at each simulation time step t: ; Among them, respectively represent obstacle avoidance reward, task reward, all tasks completed reward, distance reward, and return reward, are weight coefficients used to adjust the importance of each part of the reward; Obstacle avoidance reward , different rewards are given according to whether the action taken causes the drone to collide with an obstacle: ; wherein, is the distance between the centroid of the UAV and the obstacle, is the maximum wheelbase of the UAV; Task Reward , associated with the task priority, and different rewards are given according to whether the correct throw is completed: ; Among them, For the completed in the set, is the corresponding priority, corresponding to the flag bit indicating whether it has been destroyed; Reward for Completing All Tasks , only when all tasks are completed, the set of potential fire points is empty and no longer updated for a period of time in the future, a reward will be given once: ; Distance Reward , related to the shortest distance from each drone to the uncompleted node, by adjusting the coefficient to adjust the curvature of the reward function. As long as the target point is destroyed by any drone, all drones will receive a reward: ; For each unfinished , the distance between the unmanned aerial vehicle that is closest to is: ; Among them, is the throwing height, ( ), ( ) respectively represent the corresponding and the global coordinate positions of the UAV; Return Reward : ; When all task points are completed, the drone successfully flies to the nearest landing bay for a positive reward, otherwise a negative reward is given; S6. Re-ignition prevention: After extinguishing the open fire, continue to patrol the fire scene through cruising drones to identify re-ignition points or smoke points.

2. The forest fire detection and extinction method according to claim 1, characterized in that, In step S1, the satellite remote sensing is used to determine the fire scene boundary and estimate the fire area; the remote video monitoring is used to obtain forest scene images and videos in real time; there are multiple cruising drones, which form a drone swarm, equipped with lidar, infrared thermal imagers, and temperature, smoke, and gas sensors. The ground surface is detected by lidar, the collected point cloud data is solved and orthophoto mosaicked, and a high-precision DEM is extracted and a topographic map is generated.

3. The forest fire detection and extinguishment method according to claim 2, wherein In step S1, the command center receives the fire situation images and sensor data transmitted back by satellite remote sensing, remote video monitoring, and cruising drones in real time, and conducts fire situation judgment and fire risk level classification. The method is as follows: S1.

1. Use the YOLOv5 method to perform real-time fire situation detection on the image data through a single forward propagation network, balance speed and accuracy through data augmentation, pre-trained weights, and mixed-precision training, and determine the confidence level that the image contains flames or smoke; S1.

2. Generate an infrared thermal imaging map of the monitoring area, divide the image into several temperature regions by setting a temperature threshold to determine the high-temperature region; calculate the ratio of the area of the high-temperature region to the area of the entire image to obtain the current fire situation; compare the change in the area of the high-temperature region in the monitoring area over a period of time to preliminarily judge whether the fire is expanding or decaying; S1.

3. Continuously collect the data of the temperature, smoke, and gas sensors of the drones and transmit them to the command center. The command center preprocesses the temperature data over a period of time and takes the average value as the regional temperature. When the regional temperature exceeds the temperature alarm threshold, the fire alarm instruction is activated; S1.

4. The command center fuses different sensor data with the visual recognition results through a machine learning model to comprehensively judge whether a fire has occurred; and based on the fire risk of the monitoring area obtained from historical data in the current season, combined with factors such as temperature, humidity, wind speed, and precipitation, set weights to obtain a fire score, and divide the current fire level according to the fire score.

4. The forest fire detection and extinguishing method according to claim 2, wherein, In step S3, the calculation of the fire spread trend includes the following sub-steps: S3.

1. Collect the geographical information and environmental data of the monitoring area as the input of the Rothermel model: Obtain the topographical and morphological data corresponding to the fire point geographical coordinates from the GIS database, including: slope and aspect information; obtain the wind speed, wind direction data and fuel moisture information near the fire point from the meteorological data; obtain the fuel characteristic parameters by satellite remote sensing or drone scanning of the monitoring area, and estimate the fuel load by analyzing the relationship between the spectral characteristics of the image and the fuel load; The fuel characteristic parameters include: dry combustible content, dry particle density, surface area to volume ratio, combustibles, combustible moisture content, and combustible extinction moisture content; S3.

2. Conduct fire spread simulation based on the Rothermel model. The fire spread formula is: ; Wherein, R is the forest fire spreading speed; is the reaction intensity in the flame area; is the forest fire spread rate; is the wind speed correction coefficient; is the slope correction coefficient; is the effective heat coefficient; is the density of combustibles; is the heat required to ignite a unit mass of combustibles; S3.

3. Calculate the fire spread speed under the conditions of initial no wind and no slope: Based on the heat conduction law and the input fuel characteristic parameters, under the assumption of no wind and no slope, use the fire spread formula to calculate the basic fire spread speed R1 when the fire is in the initial no wind, no slope and the fuel moisture is the reference value; S3.

4. Calculate the actual spread speed and the headwind spread speed under the combined action of wind speed and slope: Based on the real-time wind speed, wind direction data and the slope and aspect information corresponding to the ignition point location, considering the stretching and combustion assisting effects of the wind on the flame, and the fuel accumulation and heat radiation change factors caused by the slope, correct the spread speed to obtain the corrected fire spread speed. The correction formula is: ; Among them, U represents the wind speed, which is the average wind speed at a preset height above the ground; S represents the slope, which is represented by the tangent value of the slope angle; M represents the fuel moisture, which is expressed as a percentage; is an empirical coefficient, which is calibrated according to experimental data or historical fire data; According to the correction formula, calculate the actual spread speed R2 of the fire under the combined action of wind speed and slope, and the headwind spread speed R3 in the direction opposite to the fire head direction respectively; S3.

5. Update the fire spread boundary: Based on the fire spread speeds R1, R2, and R3, combined with the time step , based on the kinematic principle, according to the fire spread position , velocity vector , predict the coordinates of the spread point of the next fire feature point : ; ; Connect the spread coordinate points with a curve to obtain the new boundary of the fire spread, calculate the area of the new area, update the fire spread trend in real time and visualize it to form a fire field development situation map; S3.

6. Update the fire spread speed and trend: Based on the real-time updated infrared thermal imaging map, sensor data and meteorological data, repeat steps S3.1 to S3.5, and adjust through the Rothermel model to update the calculation of the fire spread speed and trend.

5. The forest fire detection and extinguishment method according to claim 4, characterized in that, In step S4, build a UAV scheduling model based on the Markov decision process. The method is as follows: Build a UAV scheduling model based on the Markov decision process, which is represented as a seven-tuple: ; Among them, is the state space; is the number of agents; is the joint action space of agents; is the local observation space of an agent, is the local observation of the agent at time step under the global state is the state transition probability, indicating the probability that all given agents transfer from state to when taking the joint action to ; is the single-step reward after taking the joint action ; is the discount factor; Random policy is a mapping from the observation space to the action space, expressed as: ; Among them, represents the stochastic policy of the agent.

6. The forest fire detection and extinguishment method according to claim 5, characterized in that, In the UAV scheduling model, each agent is assigned an independent state space, and the state observation of the th agent is defined as the position, linear velocity, rotation matrix, and number of ammunitions of the quadrotor UAV, expressed as: ; Among them, , , , respectively represent the position, linear velocity, rotation matrix, and ammunition number of the agent at t ; The environmental observation of the first agent is approximately represented using the horizontal Euclidean distance as follows: ; Among them, represents the relative position of the th agent; represents the relative position of the fire point ; represents the relative position of the landing bay; Agent state environmental observation It is expressed as: ; Among them, is the flag bit indicating whether it has been destroyed at the th k time, and is the priority of the th k at the time. t The state of the environment within the moment detection area consists of all agent state environment observations and is expressed as: ; The joint action space of agents, represented as a discrete action space where the action space of each agent is set the same and represented as , which respectively represent moving left, right, forward, backward, up, down, throwing, and returning.

7. The forest fire detection and extinguishing method according to claim 5, characterized in that In step S4, the MAPPO multi-agent reinforcement learning algorithm adopts a centralized training - decentralized execution framework, and trains two independent neural networks: the actor network and the central critic network; The actor network, denoted as a policy network with parameters ; each agent has an independent actor network and makes decisions independently based on its respective local information. The input feature of the actor network is the normalized environmental information, and the output is the probability distribution of discrete actions; ​ The central critic network, denoted as a state-value function with parameters , is shared by all agents. The central critic network evaluates the value of each agent by accessing the global state information of all agents. The input of the central critic network is the centralized global state information after normalization, and the output is the state-value estimate; ​ The actor network and the central critic network process the observation information through multi-layer perceptrons respectively. The number of neurons in the hidden layer is 256. The ReLU activation function is used, and the Adam optimizer is used to maximize the policy gradient and minimize the value function loss for network update.

8. The forest fire detection and extinguishing method according to claim 7, wherein, In step S4, the UAV scheduling model is pre-trained based on the MAPPO algorithm, including the following sub-steps: S4.

1. The agent selects and executes an action according to the current policy in the environment. The environment feeds back a new state and reward according to the agent's action. The agent stores the information in the experience replay buffer and records the trajectory. S4.

2. Estimate the advantage function using the GAE method based on the collected trajectories , and the estimation method is as follows: Calculate the TD error by reflecting the current state value function estimated error: ; Among them, represents the discount factor; represents the termination flag; is the return reward after the agent takes the corresponding action, and after PopArt regularization, it is ; Calculate the GAE advantage estimate by introducing a decay factor Combine TD errors for different numbers of steps: ; Among them, represents the offset of the time step; S4.

3. Train each actor network and calculate the objective function: ; Among them, the parameter represents the actor network parameter, is the policy entropy hyperparameter, is the policy entropy, represents the clipped expected cumulative discounted reward; S4.

4. Maximize using gradient ascent , calculate the gradient of the policy loss , and then use the stochastic gradient ascent method to update the actor network parameters : ; Among them, is the learning rate; are the actor network parameters before update; S4.

5. Train the central critic network and use gradient descent to minimize the value function loss: ; Among them, represents the estimate of the future expected return under the given current state, The function represents x constrained within a , b , represent the state value functions under the old and new parameters respectively, which are used to calculate the expected cumulative discounted reward, represents the clipping coefficient; represents the expected operator under the policy ; S4.

6. Repeat steps S4.1 to S4.5 until the policy converges to the optimal solution or reaches the maximum number of training rounds.

9. A forest fire detection and extinguishing system for implementing the method according to any one of claims 1 to 8, characterized in that, Including: A fire recognition module, a fire point location and fire spread trend calculation module, a fire extinguishment module, and a re-ignition prevention module; The fire recognition module collects fire data in real time through satellite remote sensing, remote video monitoring, and cruise UAVs, transmits it to the command center, generates an infrared thermal imaging map, and conducts fire judgment and fire rating classification. The fire point location and fire spread trend calculation module marks potential fire points based on the infrared thermal imaging map, sorts the potential fire points according to the temperature value, accurately determines the geographical coordinates of the fire points in combination with multi-source geographic information data, and updates them in real time according to the movement and change of the fire points in the continuous external thermal imaging map, marks and locates newly emerging fire points; and conducts fire spread simulation based on the Rothermel model, calculates the fire spread trend, and updates the fire spread boundary. The fire extinguishment module constructs a UAV scheduling model based on the Markov decision process and pre-trains the UAV scheduling model through the MAPPO multi-agent reinforcement learning algorithm; the cruise UAVs are equipped with fire extinguishing bombs, and the command center conducts optimized scheduling based on the trained UAV scheduling model, decides the throwing order of each UAV, and takes the shortest time as the optimization goal to jointly complete the investigation and extinguishment of all potential fire points. The re-ignition prevention module continues to patrol the fire scene through cruise UAVs after extinguishing the open fire to identify re-ignition points or smoke points.

Citation Information

Patent Citations

  • YOLOV3-based smoke and fire automatic detection and early warning method

    CN112132090A

  • Forest fire extinguishing method and system based on unmanned aerial vehicle cluster

    CN118718292A

  • Unmanned aerial vehicle group collaborative fire extinguishing intelligent management system

    CN119478829A