Forest Fire Prediction Method and System Based on XGBoost-AttentionBiLSTM Model

Through the XGBoost-AttentionBiLSTM model combined with multiple data sources, the problem of inaccurate forest fire prediction in the existing technology is solved, and efficient forest fire risk warning and prediction is achieved.

CN119807891BActive Publication Date: 2025-07-04ZHEJIANG SHIZIZHIZI BIG DATA CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411792829.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-07-04
Estimated Expiration
2044-12-06

AI Technical Summary

Technical Problem

Existing shallow machine learning models are difficult to capture nonlinear relationships and deep features when processing complex and high-dimensional forest fire data, resulting in insufficient precision in forest fire prediction.

Method used

The XGBoost-AttentionBiLSTM model is adopted, combining topographic data, vegetation data, historical meteorological data and predicted meteorological data, and time series information is extracted through the BiLSTM model. The attention mechanism enhances the long-term dependency capture ability. The full connection layer and the XGBoost layer perform factor data fusion and binary classification processing to output the probability of forest fire occurrence.

Benefits of technology

More accurate forest fire predictions are achieved, high-risk areas can be identified in a timely manner and early warnings are provided to reduce fire losses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119807891B_ABST
    Figure CN119807891B_ABST
Patent Text Reader

Abstract

The present invention discloses a forest fire prediction method and system based on an XGBoost-AttentionBiLSTM model. The method includes: S1, constructing a forest fire prediction sample data set; S2, constructing an XGBoost-AttentionBiLSTM model and using the forest fire prediction sample data set for model training and learning; S3, an other factor data processing module extracts terrain factor data and vegetation factor data using terrain data and vegetation data, and the XGBoost layer outputs the wildfire occurrence probability through a sigmoid activation function; S4, collecting meteorological data, terrain data, and vegetation data containing spatio-temporal information at the prediction time and obtaining historical meteorological data before the prediction time, inputting them into the XGBoost-AttentionBiLSTM model, and obtaining the wildfire occurrence probability. The present invention can obtain the wildfire occurrence probability at each position point at the prediction time, and can also output the early warning wildfire points and the corresponding wildfire occurrence probabilities, which is convenient for timely preventive treatment of the early warning wildfire points.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of forest fire prediction, and particularly to a forest fire prediction method and system based on an XGBoost-Attention BiLSTM model. Background Art

[0002] As the main body of the terrestrial ecosystem, forests contain thousands of material or non-material resources, which have direct or indirect effects on the survival and reproduction of humans and the development of the economic society. Forests can not only conserve water sources, regulate the climate, and purify the air, but also play the roles of windbreak and sand fixation, reducing soil erosion, and protecting biodiversity. In addition, forests can not only provide social and economic resources such as timber resources, food resources, and drug resources, but also bring many conveniences to human production and life. Forest fires are one of the important natural disturbance processes in the forest ecosystem, which can help the forest ecosystem to carry out nutrient transformation and species renewal, but they are also the main natural disasters threatening forest resources. On average, hundreds of thousands of forest fires occur globally every year, which not only damage the forest ecological environment, but also cause a large number of casualties and economic losses. Since the 1990s, the forest fire area in the United States has shown an obvious upward trend. In 1996 alone, the forest area damaged by forest fires reached 2.45 million hectares. In addition, as of the beginning of September 2021, according to the statistics of the National Interagency Fire Center in the United States, more than 43,400 wildfire incidents occurred in 2021, and the burned area exceeded 2 million hectares. In February 2009, a particularly large forest fire that lasted for more than a month occurred in Australia, with a burned area of 410,000 hectares, 273 deaths, and an economic loss of up to $2 billion. It can be seen that forest fires have had a significant impact on personal safety and social economy.

[0003] Moreover, with the warming of the global climate, it not only affects some human production activities, but also has a certain impact on forest fires. Due to climate change, the amount of forest combustibles has increased. Some studies have shown that as the temperature gradually rises, the volatile oil content of Scots pine has increased by 39% compared with the normal content, and the resin content has increased by 32%, resulting in a significant increase in its flammability. In addition, due to climate change, the number of high-temperature and dry days has increased, resulting in an increase in the evaporation of surface combustibles, an increase in lightning strikes with climate change, and at the same time, climate change makes forest fire behavior more complex and changeable, greatly increasing the difficulty of forest fire fighting and posing a greater threat to social economy and people's safety. Therefore, in order to ensure the sustainable development of China's forest resources, there is an urgent need for large-scale and high-frequency forest fire warning means.

[0004] In recent years, with the rise of machine learning algorithms, more and more studies have attempted to apply machine learning algorithms to predict forest fire risks and found that machine learning algorithms have relatively superior performance in forest fire prediction. The existing machine learning models still belong to shallow machine learning models and have limited capabilities in processing highly complex and high-dimensional data, and may not be able to fully capture all potential non-linear relationships and deep features. Therefore, based on remote sensing data and artificial intelligence algorithms, establishing a forest fire risk early warning and prediction model can not only provide new ideas for forest fire prevention and control, achieve forest fire early warning, provide technical support for the relevant work of local forestry departments, reduce the losses caused by some forest fires, and unnecessary casualties. Summary of the Invention

[0005] The purpose of the present invention is to provide a forest fire prediction method and system based on the XGBoost-AttentionBiLSTM model, which constructs a forest fire prediction sample data set including terrain data, vegetation data, historical meteorological data, predicted meteorological data, and wildfire data with spatio-temporal information. The XGBoost-AttentionBiLSTM model uses the forest fire prediction sample data set for model training and learning. Among them, the BiLSTM model extracts information in two directions, forward and reverse, according to the time series of historical meteorological data, which can better understand the lag effect in meteorological data and facilitate accurate prediction; the attention mechanism module enhances the ability to capture long-term dependence relationships in time series data and can also more flexibly allocate the influence weights on the occurrence of forest fires; the fully connected layer can fuse factor data such as the output results of meteorological factors, predicted meteorological data, terrain factor data, vegetation factor data, and wildfire data, and the XGBoost layer performs binary classification processing and uses the sigmoid activation function to output the probability of wildfire occurrence.

[0006] The purpose of the present invention is achieved through the following technical solutions:

[0007] A forest fire prediction method based on the XGBoost-AttentionBiLSTM model, the method includes:

[0008] S1. Construct a forest fire prediction sample data set, and the data in the forest fire prediction sample data set is stored based on spatio-temporal coding in accordance with terrain data, vegetation data, historical meteorological data, predicted meteorological data, and wildfire data with spatio-temporal information. The predicted meteorological data is the meteorological data at the predicted time point, and the historical meteorological data is the meteorological data before the corresponding predicted time point;

[0009] S2. Construct the XGBoost-AttentionBiLSTM model and use the forest fire prediction sample dataset for model training and learning. The XGBoost-AttentionBiLSTM model includes a BiLSTM model, an attention mechanism module, an other factor data processing module, a fully connected layer, and an XGBoost layer. The BiLSTM model extracts information from historical meteorological data in both the forward and reverse directions according to the time series and obtains the output vector X. n And the query vector q of the last hidden state, and calculate the score S according to the following formula: n :

[0010] S n = q T * X n + b, where q T is the transpose of the query vector q, n represents the sequence number of meteorological data according to the time series, and b is the regularization term;

[0011] The attention mechanism module combines the score S n and the output vector X n to capture relationships, assign weights and fuse them to obtain the output result;

[0012] S3. The other factor data processing module extracts terrain factor data and vegetation factor data using terrain data and vegetation data. The fully connected layer performs fully connected training processing on the output result of step S2, predicted meteorological data, terrain factor data, vegetation factor data, and wildfire data and outputs them to the XGBoost layer. The XGBoost layer outputs the wildfire occurrence probability through the sigmoid activation function;

[0013] S4. Collect meteorological data, terrain data, and vegetation data containing spatio-temporal information at the prediction time and obtain historical meteorological data before the prediction time, and input them into the XGBoost-AttentionBiLSTM model to obtain the wildfire occurrence probability.

[0014] To better implement the present invention, the XGBoost-AttentionBiLSTM model obtains the wildfire occurrence probability according to the spatial position; sets a probability threshold, extracts the spatial positions where the wildfire occurrence probability is greater than the probability threshold as the early warning wildfire points, and then synchronously outputs the early warning wildfire points and the corresponding wildfire occurrence probabilities or visualizes them in a geographical map.

[0015] Preferably, a fire risk level classification standard is constructed, and the wildfire occurrence probabilities at the obtained spatial positions are classified according to the fire risk level classification standard, and then visualized in a geographical map.

[0016] Preferably, the spatio-temporal information includes time information and spatial information. The spatial information is longitude and latitude coordinate information, and the spatial position is longitude and latitude coordinate information.

[0017] Preferably, the points and areas repeatedly determined as early warning wildfire points and areas and / or the points and areas where wildfires have occurred repeatedly in history are set as key monitoring objects. Real-time data of the key monitoring objects are collected and input into the XGBoost-AttentionBiLSTM model, and the real-time wildfire occurrence probability of the key monitoring objects or the real-time visualization expression on the geographical map is output.

[0018] Preferably, in step S2, the BiLSTM model includes a forward LSTM processing unit and a backward LSTM processing unit. The forward LSTM processing unit fuses all meteorological data in the forward time series to extract information and inputs it into the backward LSTM processing unit, and the backward LSTM processing unit fuses all meteorological data in the backward time series to extract information.

[0019] Preferably, the processing of the attention mechanism module includes the following method: first, normalize the score S n to obtain the weight coefficient α n , and the expression is as follows:

[0020] where N represents the total number of sequence numbers of historical meteorological data and predicted meteorological data;

[0021] Then, multiply the weight coefficient α n by the output vector X n to obtain the output vector C, and the expression is as follows:

[0022] Preferably, the vegetation data is derived from two data products, MCDl2Q1.061 and MCD43A4.061, the terrain data is derived from the SRTM3 global digital elevation model dataset, and the historical meteorological data and predicted meteorological data are derived from the meteorological center.

[0023] A forest fire prediction system based on the XGBoost-AttentionBiLSTM model, including the XGBoost-AttentionBiLSTM model, a forest fire prediction sample data set, and a data acquisition module. The data in the forest fire prediction sample data set is stored associatively based on spatio-temporal coding according to terrain data, vegetation data, historical meteorological data, predicted meteorological data, and wildfire data that contain spatio-temporal information. The predicted meteorological data is the meteorological data at the predicted time point, and the historical meteorological data is the meteorological data corresponding to before the predicted time point. The XGBoost-AttentionBiLSTM model uses the forest fire prediction sample data set for model training and learning. The XGBoost-AttentionBiLSTM model includes a BiLSTM model, an attention mechanism module, an other factor data processing module, a fully connected layer, and an XGBoost layer. The BiLSTM model extracts information in two directions, forward and backward, from the historical meteorological data in time series and obtains an output vector X n and the last hidden state query vector q, and calculates the score S according to the following formula n : S n =q T *X n +b, where q T is the transpose of the query vector q, n represents the sequence number of the meteorological data in time series, and b is a regularization term. The attention mechanism module combines the score S n and the output vector X n to capture relationships, assign weights, and fuse them to obtain an output result. The other factor data processing module extracts terrain factor data and vegetation factor data using the terrain data and vegetation data. The fully connected layer performs fully connected training processing on the output result of step S2, the predicted meteorological data, the terrain factor data, the vegetation factor data, and the wildfire data and outputs them to the XGBoost layer. The XGBoost layer outputs the wildfire occurrence probability through a sigmoid activation function. The data acquisition module is used to collect meteorological data, terrain data, and vegetation data that contain spatio-temporal information at the predicted time, obtain the historical meteorological data before the predicted time, input them into the XGBoost-AttentionBiLSTM model, and obtain the wildfire occurrence probability.

[0024] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0025] (1) The present invention constructs a forest fire prediction sample data set including terrain data, vegetation data, historical meteorological data, predicted meteorological data, and wildfire data containing spatio-temporal information. The XGBoost-AttentionBiLSTM model uses the forest fire prediction sample data set for model training and learning. Among them, the BiLSTM model extracts information in both the forward and reverse directions of the historical meteorological data in time series, which can better understand the lag effect in the meteorological data and facilitate accurate prediction. The attention mechanism module enhances the ability to capture long-term dependencies in time series data and can more flexibly allocate the influence weights on the occurrence of forest fires. The fully connected layer can fuse factor data such as the output results of meteorological factors, predicted meteorological data, terrain factor data, vegetation factor data, and wildfire data. The XGBoost layer performs binary classification and uses the sigmoid activation function to output the probability of wildfire occurrence.

[0026] (2) After determining the prediction time, the present invention obtains the historical meteorological data before the prediction time and uses the meteorological data, terrain data, and vegetation data containing spatio-temporal information at the prediction time as data sources. The trained XGBoost-AttentionBiLSTM model uses the data sources to predict the probability of wildfire occurrence.

[0027] (3) The present invention can obtain the probability of wildfire occurrence at a spatial location at the prediction time according to the spatial location, and can also output the early warning wildfire points and the corresponding wildfire occurrence probabilities based on the probability threshold, which is convenient for timely preventive treatment of the early warning wildfire points. Description of the Drawings

[0028] Figure 1 is the method flow chart of the forest fire prediction method of the present invention;

[0029] Figure 2 is the schematic diagram of the principle of the XGBoost-AttentionBiLSTM model in the embodiment;

[0030] Figure 3 is the principle structure diagram of the BiLSTM model in the embodiment;

[0031] Figure 4 is the principle structure diagram of the attention mechanism module in the embodiment;

[0032] Figure 5 is the ROC curve comparison diagram of the XGBoost-AttentionBiLSTM model and other models in the embodiment;

[0033] Figure 6 is the forest fire risk level distribution diagram of a certain area in the embodiment;

[0034] Figure 7For an embodiment, an example of a wildfire point and area distribution map of key monitoring objects in a certain area;

[0035] Figure 8 For an embodiment, an example of a warning wildfire point distribution map of two certain areas;

[0036] Figure 9 For an embodiment, a curve graph of the wildfire occurrence probability of a certain wildfire point arranged in chronological order. Detailed implementation manners

[0037] The present invention will be further described in detail below in conjunction with embodiments:

[0038] Embodiment

[0039] As Figure 1 、 Figure 2 shown, a forest fire prediction method based on the XGBoost-AttentionBiLSTM model, the method includes:

[0040] S1. Construct a forest fire prediction sample dataset. The data in the forest fire prediction sample dataset is stored associatively based on spatio-temporal coding according to terrain data, vegetation data, historical meteorological data, predicted meteorological data, and wildfire data (the wildfire data in this embodiment is sourced from remote sensing products such as MODIS and VIIRS. For example, the VIIRS remote sensing product is used, which can provide information such as the location, extent, and intensity of global wildfire activities). The predicted meteorological data is the meteorological data at the predicted time point, and the historical meteorological data is the meteorological data corresponding to the time before the predicted time point. The vegetation data in this embodiment is sourced from two data products, MCD12Q1.061 and MCD43A4.061, and the land cover type and normalized difference vegetation index (NDVI) are collected respectively; download the MCD12Q1.061 and MCD43A4.061 data from NASA's LAADS DAAC, and use the HDF-EOS to GeoTIFF conversion tool (such as the HEG Tool or GDAL) to convert the HDF format data into GeoTIFF format, and use the time series analysis method to smooth the NDVI data to reduce the influence of noise and outliers; this embodiment can also fuse and adopt the American GFS numerical prediction model as the meteorological prediction factor dataset, and the historical meteorological data comes from the fifth-generation reanalysis dataset (ERA5-Land) of the European Centre for Medium-Range Weather Forecasts (ECMWF). The data features used include surface net solar radiation, 2-meter air temperature, 2-meter dew point temperature, wind speed, surface temperature, etc. The vegetation data in this embodiment at least includes air temperature (or 2-meter air temperature), dew point temperature (or 2-meter dew point temperature), soil temperature (or surface temperature), wind speed, precipitation, surface net solar radiation, surface coverage, vegetation surface temperature, vegetation self-ignition temperature, etc., where the vegetation self-ignition temperature is set separately according to the vegetation type, and the vegetation surface temperature is inversely calculated from the surface temperature, vegetation moisture content, and surface coverage. The vegetation moisture content is set separately according to the vegetation type. The surface temperature of the vegetation, the vegetation self-ignition temperature, and the potential fire source are direct indicators reflecting the likelihood of forest fires. In past forest fire prediction studies, the reanalysis dataset was mainly used as the meteorological factor. However, due to the time lag of the reanalysis data, it is difficult to synchronize with the occurrence of fires. Therefore, the present invention innovatively proposes to introduce meteorological prediction factors and combine them with meteorological lag factors (the meteorological lag factors in the present invention mainly consider the predicted meteorological data), terrain factors, and vegetation factors to achieve forest fire prediction.The terrain data is sourced from the SRTM3 global digital elevation model dataset (with a spatial resolution of 30 meters). The SRTM3 global digital elevation model dataset undergoes preprocessing operations such as reprojection, cropping, and mosaicking in ArcGIS, and the "Slope" tool and "Aspect" tool in ArcGIS are used to calculate the slope and aspect data of the terrain respectively. The terrain data includes at least elevation longitude and latitude information, slope data, and aspect data. The historical meteorological data and predicted meteorological data are sourced from the meteorological center.

[0041] S2. Construct the XGBoost-AttentionBiLSTM model and use the forest fire prediction sample dataset for model training and learning. The XGBoost-AttentionBiLSTM model includes a BiLSTM model (the Bi-LSTM is used to capture the context information before and after in time series data to help understand the impact of weather conditions in the past several days on the occurrence of fires), an attention mechanism module (the attention mechanism allows the model to focus on the most important part of the input data to determine which days' weather data is the most critical for predicting fires), an other factor data processing module, a fully connected layer (the fully connected layer integrates the information from the Bi-LSTM and the attention mechanism to understand the input data from a global perspective), and an XGBoost layer (the XGBoost layer serves as the final output layer to generate the final prediction result, leveraging its powerful classification and regression capabilities). The BiLSTM model extracts information from the historical meteorological data in both the forward and reverse directions of the time series and obtains the output vector X n and the query vector q of the last hidden state. In some embodiments, as Figure 3 shown, the BiLSTM model includes a forward LSTM processing unit (the forward LSTM processing unit includes several STM modules, Figure 3 the row below in Figure 3 contains the forward LSTM processing layer with multiple STM modules) and a reverse LSTM processing unit (the reverse LSTM processing unit includes several STM modules, Figure 3 the row above in n contains the reverse LSTM processing layer with multiple STM modules). The forward LSTM processing unit fuses all meteorological data in the forward direction of the time series to extract information and inputs it to the reverse LSTM processing unit (as Figure 3 shown, W1 to W nFor n historical meteorological data for the first n days before and after, the reverse LSTM processing unit has n STM modules. The n STM modules extract information from the n historical meteorological data one by one. The next STM module also receives the information output by the previous STM module and fuses it with the information extracted from the corresponding historical meteorological data of the next STM module and the information transmitted from the STM module corresponding to the forward LSTM processing unit and then transmits it downward. Each STM module of the reverse LSTM processing unit will obtain a hidden state (i.e., h1~h n )), and the last hidden state h n is the query vector q, and the output vector X n of the BiLSTM model will Figure 3 be the vectors h1~h n output in. The BiLSTM model can effectively process the information extraction of meteorological data. The BiLSTM model captures bidirectional information in the sequence through two LSTM module layers in two directions - one processes the natural order of the sequence (from front to back), and the other processes the reverse order (from back to front); by learning these complex spatio-temporal patterns, the BiLSTM model can better understand the lag effect in meteorological data, thereby providing more accurate results.

[0042] The score S is calculated according to the following formula n :

[0043] S n = q T * X n + b, where q T is the transpose of the query vector q (the query vector of the last hidden state of the BiLSTM model), n represents the sequence number of the meteorological data in the time series, and b is the regularization term; as Figure 4 shown, the input query vector q and the output vector X n of the BiLSTM model are calculated to obtain the score S n .

[0044] The attention mechanism module combines the score S n and the output vector X n to perform relationship capture, weight assignment and fusion to obtain the output result. As Figure 4As shown, after the BiLSTM model passes through the bidirectional long short-term memory network, the self-attention mechanism of the attention mechanism module is introduced; in the attention mechanism, through the distribution of weights, it realizes the attention to important parts and the discard of unimportant parts; it enhances the model's ability to capture long-term dependencies in time series data and can also more flexibly allocate the influence weights of different lag meteorological factors on the occurrence of forest fires. In some embodiments, the processing of the attention mechanism module includes the following methods: first, the score S n is normalized to obtain the weight coefficient α n (as Figure 4 shown, the score S n is normalized), and the expression is as follows:

[0045] where N represents the total number of sequence numbers of historical meteorological data and predicted meteorological data; then the weight coefficient α n is multiplied by the output vector X n to obtain the output vector C (as the output result of the attention mechanism module), and the expression is as follows:

[0046] S3. The other factor data processing module extracts terrain factor data and vegetation factor data using terrain data and vegetation data. The fully connected layer performs fully connected training on the output result of step S2, predicted meteorological data, terrain factor data, vegetation factor data, and wildfire data and outputs it to the XGBoost layer. The XGBoost layer outputs the wildfire occurrence probability through the sigmoid activation function. In the XGBoost layer, the classification probability can be obtained by performing a logical transformation on the output of the decision tree; for a binary classification problem such as forest fire, the model converts the decision tree output score into a probability through the sigmoid activation function; based on the XGBoost module, the final wildfire occurrence probability is output. Through the regularization term optimization in XGBoost, overfitting of the model is prevented; at the same time, the second derivative is used to make the loss more accurate and improve the prediction accuracy. The present invention inputs different types of features into different layers respectively to ensure that each feature can be properly processed; through the ensemble learning method, a powerful learner composed of multiple "weak learners" generates the final output, reducing the risk of overfitting and improving the robustness and prediction accuracy of the model. The present invention can comprehensively capture and utilize the influence of multiple factors, so as to make more accurate forest fire predictions in complex environments.

[0047] S4. Collect meteorological data, terrain data, and vegetation data containing spatio-temporal information at the prediction time and obtain historical meteorological data before the prediction time, and input them into the XGBoost-AttentionBiLSTM model to obtain the wildfire occurrence probability.

[0048] The XGBoost-AttentionBiLSTM model is a model creatively constructed to solve the technical problems of the present invention. In this embodiment, the XGBoost-AttentionBiLSTM model of the present invention is compared with Logistic Regression, Random Forest, XGBoost, etc. in terms of model performance. The evaluation results of all models in the test set are as follows in the table:

[0049]

[0050] The XGBoost-AttentionBiLSTM model of the present invention performs excellently in all indicators. The accuracy rate reaches 94.15%, the precision is 0.9229, the recall rate is 0.9586, and the F1 score is 0.9404. To more comprehensively evaluate the performance of the XGBoost-AttentionBiLSTM model of the present invention, the present invention not only introduces a variety of common machine learning models for comparative analysis, including XGBoost, Random Forest, and Logistic Regression, but also further evaluates each model through the ROC curve; evaluation indicators such as accuracy rate and recall rate are easily affected by the model segmentation threshold. Therefore, using the ROC curve can more objectively evaluate the overall performance of the model. As Figure 5 shown, the AUC value of the XGBoost-AttentionBiLSTM model of the present invention is the highest, reaching 0.9828, significantly higher than other models. The AUC values of Random Forest and XGBoost are 0.9684 and 0.9670 respectively, while the AUC value of the Logistic Regression model is the lowest, only 0.7706. The higher the AUC value of the XGBoost-AttentionBiLSTM model of the present invention, the better the comprehensive performance of the model at different thresholds, and it can better distinguish positive and negative samples, further verifying the superiority of the XGBoost-AttentionBiLSTM model in the forest fire prediction task.

[0051] In some embodiments, the XGBoost-AttentionBiLSTM model obtains the wildfire occurrence probability according to the spatial position (i.e., the longitude and latitude coordinate information). The present invention also sets a probability threshold, extracts the spatial positions where the wildfire occurrence probability is greater than the probability threshold as the early warning wildfire points, and then synchronously outputs the early warning wildfire points and the corresponding wildfire occurrence probabilities or visually expresses them on a geographical map. Preferably, the present invention sets the points and regions that are repeatedly determined as early warning wildfire points and regions or / and the points and regions where wildfires have occurred repeatedly in history as the key monitoring objects, collects real-time data of the key monitoring objects and inputs them into the XGBoost-AttentionBiLSTM model, and outputs the real-time wildfire occurrence probability of the key monitoring objects or visually expresses them in real time on a geographical map. As Figure 6 shown, in this embodiment, the Yunnan-Guizhou-Sichuan region is taken as an example in January 2023, and January 20, 2023 is taken as the prediction time. The visualization chart of the wildfire occurrence probability in the Yunnan-Guizhou-Sichuan region on January 20, 2023 is obtained according to the method of the present invention. It can also be seen from the distribution of forest fire points on that day that multiple forest fires occurred in the central and southern regions of Yunnan Province. The verification results show that: the high-risk areas predicted by the present invention are highly consistent with the positions of the actually occurring fire points, showing a high sensitivity to forest fires, and further verifying the prediction ability of the model. This advantage stems from the innovation in the model structure design, especially the combination of the efficient feature selection ability of XGBoost and the advantage of time series data processing of the Attention mechanism; in addition, the deep learning characteristics of the BiLSTM model of the present invention allow it to learn more complex information from a large amount of historical data and prediction data sets, thereby improving the prediction accuracy. The present invention not only improves the prediction performance of forest fires, but more importantly, it provides more reliable support for the forest fire early warning system, helps to take measures in advance to reduce fire losses, and ensures the safety of forest resources.

[0052] In some embodiments, the present invention also constructs a classification standard for fire risk levels, classifies the wildfire occurrence probabilities of the obtained spatial positions according to the classification standard for fire risk levels, and then visually expresses them on a geographical map. In this embodiment, the Yunnan-Guizhou-Sichuan region is taken as an example in January 2023. The classification standard for fire risk levels in the Yunnan-Guizhou-Sichuan region is as follows:

[0053]

[0054]

[0055] In this embodiment, ArcGIS Pro is used to classify the fire danger levels, and the forest fire risk levels in the Yunnan-Guizhou-Sichuan region are obtained as shown in the trademark.

[0056] In some embodiments, the present invention sets the points and regions where wildfires are determined to be warning points multiple times and / or the points and regions where wildfires have occurred multiple times in history as key monitoring objects, collects real-time data of the key monitoring objects and inputs them into the XGBoost-AttentionBiLSTM model, and outputs the real-time wildfire occurrence probability of the key monitoring objects or visualizes them in a geographical map in real time (in this embodiment, Yunnan, Guizhou, and Sichuan are taken as examples, and the visualization of the key monitoring objects in the geographical map is as shown in Figure 7 ).

[0057] A forest fire prediction system based on the XGBoost-AttentionBiLSTM model includes an XGBoost-AttentionBiLSTM model, a forest fire prediction sample data set, and a data collection module. The data in the forest fire prediction sample data set is stored associatively based on spatio-temporal coding according to terrain data, vegetation data, historical meteorological data, predicted meteorological data, and wildfire data that include spatio-temporal information. The predicted meteorological data is the meteorological data at the predicted time point, and the historical meteorological data is the meteorological data corresponding to the time before the predicted time point. The XGBoost-AttentionBiLSTM model uses the forest fire prediction sample data set for model training and learning. The XGBoost-AttentionBiLSTM model includes a BiLSTM model, an attention mechanism module, an other factor data processing module, a fully connected layer, and an XGBoost layer. The BiLSTM model extracts information in two directions, forward and backward, in time series from the historical meteorological data and obtains an output vector X n and the last hidden state query vector q, and calculates the score S according to the following formula n : S n =q T *X n +b, where q T is the transpose of the query vector q, n represents the sequence number of the meteorological data in time series, and b is a regularization term. The attention mechanism module captures the relationship, assigns weights, and fuses the score S n and the output vector X n to obtain an output result. The other factor data processing module extracts terrain factor data and vegetation factor data using the terrain data and vegetation data. The fully connected layer performs fully connected training processing on the output result of step S2, the predicted meteorological data, the terrain factor data, the vegetation factor data, and the wildfire data and outputs them to the XGBoost layer. The XGBoost layer outputs the wildfire occurrence probability through a sigmoid activation function. The data collection module is used to collect meteorological data, terrain data, and vegetation data that include spatio-temporal information at the predicted time, obtain the historical meteorological data before the predicted time, input them into the XGBoost-AttentionBiLSTM model, and obtain the wildfire occurrence probability.

[0058] In some embodiments, the XGBoost-AttentionBiLSTM model obtains the wildfire occurrence probability according to the spatial position (i.e., the longitude and latitude coordinate information). The present invention also sets a probability threshold, extracts the spatial positions where the wildfire occurrence probability is greater than the probability threshold as the early-warning wildfire points, and then synchronously outputs the early-warning wildfire points and the corresponding wildfire occurrence probabilities or visually expresses them on a geographical map. As Figure 8 shown, in this embodiment, two study areas, namely Yuxi Mountain in Jiangchuan District, Yuxi City, Yunnan Province (from April 11th to April 16th, 2023) and Guanzhuang Village in Wenquan Sub-district, Anning City, Yunnan Province (from April 9th to April 25th, 2023), are selected for model testing. The model of the present invention respectively extracts the historical meteorological data, predicted meteorological data (the prediction time is any day or several days from April 11th to April 20th, 2023), terrain data, and vegetation data from 5 days before the occurrence to the occurrence time, and obtains the results of the wildfire occurrence probability according to the method of the present invention. The probability threshold is set to 0.7. Figure 9 shows the wildfire occurrence probability at a certain location in Guanzhuang Village, Wenquan Sub-district, Anning City, Yunnan Province from April 0th to April 25th, 2023. Among them, from April 11th to April 22nd, 2023, the probability threshold of 0.7 is exceeded, thus predicting the time interval of the forest fire. As Figure 8 shown, in this embodiment, the spatial positions exceeding the probability threshold in Yuxi Mountain in Jiangchuan District, Yuxi City, Yunnan Province and Guanzhuang Village in Wenquan Sub-district, Anning City, Yunnan Province on a certain day are selected as the early-warning wildfire points and visually expressed on a geographical map.

[0059] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A forest fire prediction method based on the XGBoost-AttentionBiLSTM model, characterized in that: The method includes: S1. Construct a forest fire prediction sample data set. The data in the forest fire prediction sample data set is stored in an associated manner based on spatio-temporal coding according to terrain data, vegetation data, historical meteorological data, predicted meteorological data, and wildfire data that contain spatio-temporal information. The predicted meteorological data is the meteorological data at the predicted time point, and the historical meteorological data is the meteorological data corresponding to the time before the predicted time point. S2. Construct the XGBoost-AttentionBiLSTM model and use the forest fire prediction sample dataset for model training and learning. The XGBoost-AttentionBiLSTM model includes a BiLSTM model, an attention mechanism module, an other factor data processing module, a fully connected layer, and an XGBoost layer. The BiLSTM model extracts information from historical meteorological data in both the forward and reverse directions of the time series and obtains an output vector and the last hidden state query vector . The BiLSTM model includes a forward LSTM processing unit and a reverse LSTM processing unit. Both the forward LSTM processing unit and the reverse LSTM processing unit contain n STM modules. The forward LSTM processing unit fuses all meteorological data in the forward direction of the time series to extract information and inputs it into the reverse LSTM processing unit. The reverse LSTM processing unit fuses all meteorological data in the reverse direction of the time series to extract information. Each STM module of the reverse LSTM processing unit obtains a hidden state h1~h n as the output vector . The last hidden state h n is the query vector . Calculate the score according to the following formula : , where is the query vector is the transpose of, n represents the sequence number of meteorological data in time series, and b is the regularization term; Attention mechanism module combined score and the output vector to perform relationship capture, weight assignment and fusion to obtain the output result; the processing of the attention mechanism module includes the following methods: first, normalize the score to obtain the weight coefficient , and the expression is as follows: , where N represents the total number of sequence numbers of historical meteorological data and predicted meteorological data; Then, the weight coefficient is multiplied by the output vector to obtain the output vector , and the expression is as follows: ; S3. The other factor data processing module extracts terrain factor data and vegetation factor data using the terrain data and vegetation data. The fully connected layer performs fully connected training processing on the output result of step S2, the predicted meteorological data, the terrain factor data, the vegetation factor data, and the wildfire data and outputs it to the XGBoost layer. The XGBoost layer outputs the wildfire occurrence probability through the sigmoid activation function. S4. Collect meteorological data, terrain data, and vegetation data that contain spatio-temporal information at the predicted time and obtain the historical meteorological data before the predicted time, input them into the XGBoost-AttentionBiLSTM model, and obtain the wildfire occurrence probability.

2. The forest fire prediction method based on the XGBoost-AttentionBiLSTM model according to claim 1, characterized in that: The XGBoost-AttentionBiLSTM model obtains the wildfire occurrence probability according to the spatial position. Set a probability threshold, extract the spatial positions where the wildfire occurrence probability is greater than the probability threshold as the early warning wildfire points, and then synchronously output the early warning wildfire points and the corresponding wildfire occurrence probabilities or visually express them on the geographical map.

3. The forest fire prediction method based on the XGBoost-AttentionBiLSTM model according to claim 2, characterized in that: Construct a fire risk level classification standard, classify the wildfire occurrence probabilities at the obtained spatial positions according to the fire risk level classification standard, and then visually express them on the geographical map.

4. The forest fire prediction method based on the XGBoost-AttentionBiLSTM model according to claim 2, wherein: The spatio-temporal information includes time information and spatial information. The spatial information is longitude and latitude coordinate information, and the spatial position is longitude and latitude coordinate information.

5. The forest fire prediction method based on the XGBoost-AttentionBiLSTM model according to claim 2, wherein: Set the points and regions that are repeatedly determined as early warning wildfire points and regions or / and the points and regions where wildfires have occurred repeatedly in history as key monitoring objects, collect real-time data of the key monitoring objects and input them into the XGBoost-AttentionBiLSTM model, and output the real-time wildfire occurrence probability of the key monitoring objects or visually express it on the geographical map in real time.

6. The forest fire prediction method based on the XGBoost-AttentionBiLSTM model according to claim 1, wherein: The vegetation data is sourced from two data products, MCD12Q1.061 and MCD43A4.

061. The terrain data is sourced from the SRTM3 global digital elevation model data set. The historical meteorological data and the predicted meteorological data are sourced from the meteorological center.

7. A forest fire prediction system based on the XGBoost-Attention BiLSTM model for implementing the forest fire prediction method described in claim 1, characterized in that: It includes an XGBoost-AttentionBiLSTM model, a forest fire prediction sample dataset, and a data acquisition module. The data in the forest fire prediction sample dataset is stored in a spatio-temporally encoded association based on terrain data, vegetation data, historical meteorological data, predicted meteorological data, and wildfire data that contain spatio-temporal information. The predicted meteorological data is the meteorological data at the predicted time point, and the historical meteorological data is the meteorological data corresponding to the time before the predicted time point. The XGBoost-AttentionBiLSTM model uses the forest fire prediction sample dataset for model training and learning. The XGBoost-AttentionBiLSTM model includes a BiLSTM model, an attention mechanism module, an other factor data processing module, a fully connected layer, and an XGBoost layer. The BiLSTM model extracts information in both the forward and backward directions of the historical meteorological data in time series and obtains an output vector and the query vector of the last hidden state , and the score is calculated according to the following formula :[[]] , where is the transpose of the query vector , n represents the sequence number of the meteorological data in time series, and b is the regularization term. The attention mechanism module combines the score and the output vector to capture relationships, assign weights, and fuse them to obtain an output result; The other factor data processing module extracts terrain factor data and vegetation factor data using the terrain data and vegetation data. The fully connected layer performs fully connected training processing on the output result of step S2, the predicted meteorological data, the terrain factor data, the vegetation factor data, and the wildfire data and outputs it to the XGBoost layer. The XGBoost layer outputs the wildfire occurrence probability through the sigmoid activation function. The data collection module is used to collect meteorological data, terrain data, and vegetation data that contain spatio-temporal information at the predicted time and obtain the historical meteorological data before the predicted time, input them into the XGBoost-AttentionBiLSTM model, and obtain the wildfire occurrence probability.

Citation Information

Patent Citations

  • Shared bicycle demand quantity prediction method based on Attention-LSTM-LightGBM

    CN116797274A

  • Atmospheric temperature prediction method based on combined network model

    CN116894524A