Forest grassland fire risk assessment method based on variable screening and machine learning
By performing variable screening and machine learning on forest and grassland fire remote sensing data, a fire risk assessment model is built, which solves the accuracy and efficiency of fire risk assessment in the existing technology, and achieves more efficient fire risk assessment and protection.
Patent Information
- Application Number
- CN202510353447.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-08-15
AI Technical Summary
The prior art cannot quickly and accurately reflect the real risk situation of forest and grassland fires, resulting in the problems of high information overlapping autocorrelation, high prediction cost and low efficiency in fire risk assessment.
By obtaining historical fire remote sensing data, extracting environmental factor variables and performing correlation analysis, variable screening is performed using cable regression algorithm, principal component analysis method and stepwise regression method, multiple fire risk assessment models are constructed, and the optimal model is used for inversion to generate forest and grassland fire risk maps.
A scientific and accurate assessment of forest and grassland fire risks has been achieved, fire protection efficiency has been improved, fire prevention work can be better guided, forecasting costs have been reduced, and evaluation accuracy has been improved.
Smart Images

Figure CN120493196A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of forest and grassland fire risk assessment, and is related to but not limited to a forest and grassland fire risk assessment method based on variable screening and machine learning. Background Art
[0002] Forests and grasslands cover approximately 55% of the world's land area and play a vital role in ecological, environmental, and socioeconomic development by providing various ecosystem services, such as water conservation, soil conservation, carbon sequestration, and oxygen release. However, with continued global warming and the frequent occurrence of extreme weather, the number of forest fire days and incidents has increased dramatically each year. Due to the uneven distribution and continuity of forests and grasslands, disasters occurring in these areas are highly contagious. Among these disasters, fires are characterized by their sudden onset, rapid spread and diffusion, high destructiveness, difficulty in monitoring, widespread impact, and a latent period that is difficult to detect, making them difficult to contain once they break out. Therefore, forest and grassland fire management is imperative.
[0003] In related technologies, fire risk assessment is carried out using a system of model construction, model evaluation, and spatiotemporal analysis of fire risk. However, this method has problems with high autocorrelation of overlapping information, which leads to overfitting of the model, high prediction cost, low efficiency, and inability to accurately reflect the actual fire risk situation.
[0004] Therefore, how to quickly and accurately reflect the actual fire risk situation has become an urgent problem to be solved. Summary of the Invention
[0005] In view of this, an embodiment of the present invention provides a forest and grassland fire risk assessment method based on variable screening and machine learning, which at least solves the problem that related technologies cannot quickly and accurately reflect the actual fire risk situation.
[0006] According to a first aspect of an embodiment of the present invention, a forest and grassland fire risk assessment method based on variable screening and machine learning is provided, comprising: Obtain historical fire remote sensing data, and extract environmental factor variables from the pre-processed historical fire remote sensing data; Performing correlation analysis on the environmental factor variables to obtain correlation results, and performing variable screening on the environmental factor variables based on the correlation results, Lasso regression algorithm, principal component analysis method, and stepwise regression method to obtain screened environmental factor variables; Performing data division on the screened environmental factor variables to obtain training data and test data, and using the training data and the test data to perform model training and testing on multiple fire risk assessment models to be trained, until multiple fire risk assessment models are obtained; An optimal fire risk assessment model is determined based on the multiple fire risk assessment models and the acquired historical fire records, and the screened environmental factor variables are inverted using the optimal fire risk assessment model to obtain a fire risk map of the forest and grassland.
[0007] According to a second aspect of an embodiment of the present invention, there is provided an electronic device comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; the memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the method described in the first aspect.
[0008] According to a third aspect of an embodiment of the present invention, a computer storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method according to the first aspect is implemented.
[0009] According to the solution provided by an embodiment of the present invention, historical fire remote sensing data is obtained, and environmental factor variables are extracted from the pre-processed historical fire remote sensing data; correlation analysis is performed on the environmental factor variables to obtain correlation results, and the environmental factor variables are screened based on the correlation results, the Lasso regression algorithm, the principal component analysis method, and the stepwise regression method to obtain screened environmental factor variables; data is divided on the screened environmental factor variables to obtain training data and test data, and the training data and the test data are used to perform model training and testing on multiple fire risk assessment models to be trained until multiple fire risk assessment models are obtained; an optimal fire risk assessment model is determined based on the multiple fire risk assessment models and the obtained historical fire records, and the optimal fire risk assessment model is used to invert the screened environmental factor variables to obtain a fire risk map of forests and grasslands. In this process, correlation analysis was conducted on the collected environmental factor variables, and effective dimensionality reduction was performed on the environmental factor variables based on principal component analysis, stepwise regression, and LASSO algorithm. The screened environmental factor variables were then tested for collinearity to ensure that there was no multicollinearity between the variables used to construct the machine learning model. Multiple fire risk assessment models were constructed, and the optimal variable screening method and machine learning method were scientifically and practically selected to ensure the goodness of fit of the optimal fire risk assessment model and its applicability in the study area. Not only can the relevant indicators of forest and grassland fire risk be considered more comprehensively, but by selecting multiple environmental factor variables and combining multiple machine learning models, the forest and grassland fire risk can be better evaluated. Compared with existing forest and grassland fire risk assessments, this process can better reflect the actual fire risk situation, thereby better guiding forest and grassland fire prevention work, improving the efficiency of forest and grassland fire protection, and enhancing the effectiveness of forest and grassland fire protection. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive work, among which: Figure 1 A schematic diagram of a process for forest and grassland fire risk assessment based on variable screening and machine learning provided by an embodiment of the present invention; Figure 2 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0011] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0012] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0013] It should be pointed out that the terms "first\second\third" involved in the embodiments of the present invention are only used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present invention described here can be implemented in an order other than that illustrated or described here.
[0014] Those skilled in the art will understand that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those skilled in the art in the art to which the embodiments of the present invention pertain. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art and, unless specifically defined as herein, should not be interpreted in an idealized or overly formal sense.
[0015] Figure 1A flow chart of a forest and grassland fire risk assessment method based on variable screening and machine learning provided in an embodiment of the present invention. The forest and grassland fire risk assessment method based on variable screening and machine learning provided in an embodiment of the present invention can be executed by an electronic device, such as a computer, a server, etc.
[0016] like Figure 1 As shown in the figure, the forest and grassland fire risk assessment method based on variable screening and machine learning includes: S101. Acquire historical fire remote sensing data, and extract environmental factor variables from the pre-processed historical fire remote sensing data.
[0017] In an embodiment of the present invention, remote sensing data is information about the earth's surface collected by sensors on platforms such as satellites, drones, and airplanes. These data can be in various forms such as optical images, radar data, infrared data, etc., reflecting the different characteristics of the earth's surface. Environmental factor variables refer to various natural conditions or factors that affect biological, physical, and chemical processes. Environmental factor variables are obtained from the collected remote sensing data. Environmental factor variables include topography, climate, combustibles, and human activities. Topographic variables include altitude, slope, and aspect. Climate variables include average monthly temperature, average monthly precipitation, average monthly maximum temperature, average monthly relative humidity, average monthly net surface solar radiation, average monthly wind speed, and average monthly saturated vapor pressure difference. The average monthly saturated vapor pressure difference is calculated according to the following formula: ; In the above formula, VPD is the monthly average saturated vapor pressure difference, T is the monthly mean temperature, rH is the monthly average relative humidity.
[0018] Fuel variables include monthly average vegetation cover and annual average aboveground biomass. Human activity variables include annual average population density, annual average GDP, distance to rivers, roads, railways, and residential areas. Topographic factors have been shown to significantly influence wildfire occurrence and spread. Elevation, slope, and aspect influence the amount and structure of vegetation fuels, as well as the likelihood of human-induced ignition sources, which are closely related to wildfire occurrence. Climate is a key factor in wildfires. High temperatures, drought, strong winds, low relative humidity, and insufficient precipitation have been shown to be key and necessary factors for fire occurrence. Soil factors also play a role in fire occurrence, with seasonal influences on wildfire probability. Grassland vegetation is the primary carrier of grassland fires. As a key component of grassland vegetation, the spatial distribution and abundance of fuels have a significant impact on the occurrence and spread of grassland fires. In addition, human activities are closely related to wildfire occurrence, significantly influencing fuel distribution, fire ignition intensity, and firefighting resources. Therefore, human activity variables are widely considered to be one of the most important wildfire drivers. All the downloaded images were cropped and resampled to 1 km to ensure smooth subsequent analysis.
[0019] Furthermore, the historical fire remote sensing data were screened and filtered to obtain preprocessed historical fire remote sensing data, namely, effective fire occurrence information. Specifically, based on the MODIS Collection 6 Burned Area data product (MODIS Collection 6 Burned Area, MCD64A1), the historical fire remote sensing data images were projected and spliced. The raster to point tool in Arcmap 10.8 (ArcGIS Desktop Advanced 10.8) software was used to convert the fire area data into fire point vector maps to obtain effective fire data. Finally, environmental factor variables were extracted from the effective fire data.
[0020] Historical fire remote sensing data only includes the location and date of the fire. Using the "Buffer" tool in ArcMap 10.8, a buffer with a radius of 1 km was generated centered on the fire point. Using the "Extract Multi-Values to Points" tool in ArcMap 10.8, an equal number of non-fire points were randomly generated outside the buffer zone. Using the "Extract Multi-Values to Points" tool in ArcMap 10.8, the values of various environmental variables at the fire and non-fire points were extracted. This constituted a complete fire dataset, including the location information of the fire and non-fire points and the corresponding values of each environmental variable.
[0021] S102. Perform correlation analysis on the environmental factor variables to obtain correlation results, and screen the environmental factor variables based on the correlation results, Lasso regression algorithm, principal component analysis method, and stepwise regression method to obtain screened environmental factor variables.
[0022] In an embodiment of the present invention, a correlation analysis is performed on the environmental factor variables, specifically, Spearman correlation analysis and Pearson correlation analysis are performed on all environmental factor variables to check whether there are environmental factor variables with strong correlation, specifically, it is determined whether there are environmental factor variables with a correlation coefficient > 0.75 between the environmental factor variables to obtain correlation results, and the environmental factor variables are screened by using the correlation results, Lasso regression algorithm, principal component analysis method and stepwise regression method to obtain the screened environmental factor variables.
[0023] Principal component analysis (PCA) can effectively reduce the dimensionality of environmental factor variables. In this embodiment, environmental factor variables with an explanation rate greater than 85% were selected for subsequent modeling. PCA can be performed in SPSS 27 (Statistical Package for the Social Sciences 27), while stepwise regression and the LASSO algorithm can be performed in RStudio.
[0024] Stepwise regression: Stepwise regression is a method for selecting variables for inclusion in a linear regression model. The basic idea is to introduce variables into the model one by one. After each variable is introduced, a t-test is performed on each selected variable, and then any insignificant variables are removed. This process is repeated until no new variables can be introduced. If a correlation > 0.7 exists between environmental factor variables selected using stepwise regression, the less significant environmental factor variable is removed.
[0025] Lasso regression algorithm (LASSO): The LASSO algorithm is a biased estimator for multicollinear data. Based on linear regression, it constructs a penalty function to limit the sum of the absolute values of the model's regression coefficients to below a set threshold and minimize the sum of squared residuals. This reduces the regression coefficients of independent variables with a weak influence on the dependent variable to even zero, while retaining important variables.
[0026] S103. When the multicollinearity test results of the screened environmental factor variables meet the preset conditions, the screened environmental factor variables are divided into data to obtain training data and test data, and the training data and test data are used to train and test multiple fire risk assessment models to be trained until multiple fire risk assessment models are obtained.
[0027] In an embodiment of the present invention, multiple fire risk assessment models to be trained can be constructed, and the screened environmental factor variables include the results of environmental factor variables screened using the Lasso regression algorithm, the results of environmental factor variables screened using the principal component analysis method, and the results of environmental factor variables screened using the stepwise regression method. A multicollinearity test is performed on the screened environmental factor variables. When the multicollinearity test results of the screened environmental factor variables meet the preset conditions, the screened environmental factor variables are subjected to data division, that is, the results of environmental factor variables screened using the Lasso regression algorithm, the results of environmental factor variables screened using the principal component analysis method, and the results of environmental factor variables screened using the stepwise regression method are divided, and corresponding training data and test data are obtained respectively. The training data and test data are used to train and test multiple fire risk assessment models to be trained until multiple fire risk assessment models are obtained. When the multicollinearity test results of the screened environmental factor variables do not meet the preset conditions, the environmental factor variables are subjected to correlation analysis and variable screening again until the multicollinearity test results of the screened environmental factor variables meet the preset conditions.
[0028] S104. Determine the optimal fire risk assessment model based on multiple fire risk assessment models and the acquired historical fire records, and use the optimal fire risk assessment model to invert the screened environmental factor variables to obtain a fire risk map of the forest and grassland.
[0029] In an embodiment of the present invention, historical fire records are obtained from relevant local authorities, and an optimal fire risk assessment model is determined based on multiple trained fire risk assessment models and the historical fire records. The historical fire records include information such as the location and date of historical fires. The historical fire records can be historical fire risk maps, derived from a historical fire inventory dataset and distinct from the fire risk maps corresponding to the training and test sample data. The years corresponding to the historical fire records fall within the years corresponding to the collected remotely sensed forest and grassland data. For example, if the remotely sensed data corresponds to years between 2001 and 2022, the historical fire records can correspond to years between 2008 and 2022.
[0030] The selected environmental factor variables were then input into the optimal fire risk assessment model for inversion, resulting in a forest and grassland fire risk map. Fire risk was then categorized into five levels: high risk, medium-high risk, medium risk, medium-low risk, and low risk, ultimately generating a forest and grassland fire risk level map.
[0031] Specifically, the optimal fire risk assessment model was used to invert the screened environmental factor variables to obtain the inversion results of the forest and grassland fire risk map. The inversion results were exported in raster format and combined with land use data (LUCC). The forest and grassland images were extracted using the mask extraction tool in Arcmap, and the natural breakpoint method was used to divide the fire risk into five levels: high risk, medium-high risk, medium risk, medium-low risk and low risk. Finally, a forest and grassland fire risk level map was generated.
[0032] It can be understood that in an embodiment of the present invention, historical fire remote sensing data is obtained, and environmental factor variables are extracted from the preprocessed historical fire remote sensing data; correlation analysis is performed on the environmental factor variables to obtain correlation results, and the environmental factor variables are screened based on the correlation results, Lasso regression algorithm, principal component analysis method and stepwise regression method to obtain screened environmental factor variables; the screened environmental factor variables are divided into data to obtain training data and test data, and the training data and test data are used to train and test multiple fire risk assessment models to be trained until multiple fire risk assessment models are obtained; the optimal fire risk assessment model is determined based on multiple fire risk assessment models and the obtained historical fire records, and the optimal fire risk assessment model is used to invert the screened environmental factor variables to obtain a fire risk map of forests and grasslands. In this process, correlation analysis was conducted on the collected environmental factor variables, and effective dimensionality reduction was performed on the environmental factor variables based on principal component analysis, stepwise regression, and LASSO algorithm. The screened environmental factor variables were then tested for collinearity to ensure that there was no multicollinearity between the variables used to construct the machine learning model. Multiple fire risk assessment models were constructed, and the optimal variable screening method and machine learning method were scientifically and practically selected to ensure the goodness of fit of the optimal fire risk assessment model and its applicability in the study area. Not only can the relevant indicators of forest and grassland fire risk be considered more comprehensively, but by selecting multiple environmental factor variables and combining multiple machine learning models, the forest and grassland fire risk can be better evaluated. Compared with existing forest and grassland fire risk assessments, this process can better reflect the actual fire risk situation, thereby better guiding forest and grassland fire prevention work, improving the efficiency of forest and grassland fire protection, and enhancing the effectiveness of forest and grassland fire protection.
[0033] In some embodiments of the present invention, S101 may be implemented through S201, which is explained through the following steps.
[0034] S201. Perform a multicollinearity test on each of the first environmental factor variable, the second environmental factor variable, and the third environmental factor variable to obtain a multicollinearity test result.
[0035] In some embodiments of the present invention, the screened environmental factor variables include a first environmental factor variable, a second environmental factor variable, and a third environmental factor variable, wherein the first environmental factor variable is the variable screening result corresponding to the Lasso regression algorithm, the second environmental factor variable is the variable screening result corresponding to the principal component analysis method, and the third environmental factor variable is the variable screening result corresponding to the stepwise regression method. A multicollinearity test is performed on the first environmental factor variable, a multicollinearity test is performed on the second environmental factor variable, and a multicollinearity test is performed on the third environmental factor variable to obtain corresponding multicollinearity test results. Through the multicollinearity test, within the same set of environmental factor variables, there should not be a high correlation between any two environmental factor variables, thereby reducing the risk of overfitting.
[0036] Specifically, taking the first environmental factor variable as an example, SPSS27 software can be used to perform a collinearity test. If the tolerance is >0.1 or the variance inflation factor (VIF) is <10, it means that there is no multicollinearity between the first environmental factor variables and they can be used to construct the model.
[0037] Furthermore, when the multicollinearity test results of the screened environmental factor variables meet the preset test conditions, a data division operation is performed, that is, when the multicollinearity test results of the first environmental factor variable meet the preset test conditions, the multicollinearity test results of the second environmental factor variable meet the preset test conditions, and the multicollinearity test results of the third environmental factor variable meet the preset test conditions, the first environmental factor variable, the second environmental factor variable, and the third environmental factor variable are divided respectively to obtain subsequent operations such as corresponding training operations and test data.
[0038] In some embodiments of the present invention, data division of the screened environmental factor variables to obtain training data and test data in S103 can be achieved through S1031 to S1033, which is explained in the following steps.
[0039] S1031. Divide the first environmental factor variable into data to obtain first training data and first test data.
[0040] S1032. Divide the data of the second environmental factor variable to obtain second training data and second test data.
[0041] S1033. Perform data division on the third environmental factor variable to obtain third training data and third test data.
[0042] In some embodiments of the present invention, data is partitioned for the first environmental factor variable to obtain first training data and first test data, data is partitioned for the second environmental factor variable to obtain second training data and second test data, and data is partitioned for the third environmental factor variable to obtain third training data and third test data. The resulting training data and test data are used to train and test multiple subsequent fire risk assessment models to be trained.
[0043] Among them, in the first environmental factor variable, the second environmental factor variable and the third environmental factor variable, 80% of the data are input as the training set and 20% of the data are input as the test set.
[0044] In some embodiments of the present invention, the step of performing model training and testing on multiple fire risk assessment models to be trained using training data and test data until multiple fire risk assessment models are obtained can be implemented through S103A to S103C, which is explained in the following steps.
[0045] S103A: Perform model training on multiple fire risk assessment models to be trained using the first training data, the second training data, and the third training data to obtain multiple trained fire risk assessment models.
[0046] In some embodiments of the present invention, multiple fire risk assessment models to be trained can be constructed, and the first training data, the second training data, and the third training data can be separately input into multiple fire risk assessment models for model training. For example, the first training data can be input into model 1 for model training, and the first training data can be input into model 2 for model training, etc., and an iteration condition can be set, which can be the number of iterations. When the iteration condition is met, multiple trained fire risk assessment models are obtained.
[0047] For example, when constructing three fire risk assessment models to be trained, the first training data, the second training data and the third training data are used to train the three fire risk assessment models to be trained separately until the training conditions are met, and nine trained fire risk assessment models can be obtained.
[0048] S103B. When the performance of multiple trained fire risk assessment models meets the preset indicator conditions, the multiple trained fire risk assessment models are tested using the first test data, the second test data and the third test data to obtain their corresponding fire risk maps.
[0049] In some embodiments of the present invention, the first training data, the second training data, and the third training data are used to train multiple fire risk assessment models until the training conditions are met, and multiple trained fire risk assessment models are obtained. When the performance of the multiple trained fire risk assessment models meets the preset index conditions, the first test data, the second test data, and the third test data are used to test the multiple trained fire risk assessment models, and the multiple fire risk assessment models to be trained output their corresponding fire risk maps. For example, when the first training data is used to train model 1 and the trained model 1 is obtained, the trained model 1 is tested using the corresponding first test data to obtain the corresponding fire risk map. The preset index conditions include accuracy ( Acc ), sensitivity ( Sen ), specificity ( Spe )、 F1 value( F1 Scor e value), Cohen's Kappa Coefficient, area under the curve ( AUC ), the calculation methods are as follows: ; ; ; ; ; ; ; In the above formula, TP ( True Positive ): The prediction is positive, and the actual is also positive, TN ( True Negative ): The prediction is negative, and the actual result is also negative. FP ( False Positive ): The prediction is positive, but it is actually negative. FN ( False Negative ): The prediction is negative, but it is actually positive. po represents the observed accuracy, pe Indicates the expected accuracy.
[0050] S103C. When the corresponding fire risk maps all meet the preset indicator conditions, multiple fire risk assessment models are obtained.
[0051] In some embodiments of the present invention, the values corresponding to the preset indicators of the corresponding fire risk maps are calculated, that is, the accuracy, sensitivity and specificity of the fire risk maps are calculated, and the calculated accuracy, sensitivity and specificity are compared with the preset indicator conditions. When the preset indicators meet the preset indicator conditions (such as Acc>85% and kappa>0.75, AUC>0.8), that is, when multiple trained fire risk assessment models meet the test results, multiple fire risk assessment models are obtained.
[0052] In some embodiments of the present invention, the plurality of fire risk assessment models to be trained include a random forest model, a support vector machine model, and a BP neural network model. S103A can be implemented through S301 to S303, which is described in the following steps.
[0053] S301. Perform model training on the random forest model using the first training data, the second training data, and the third training data, respectively, until a trained first random forest model, a second random forest model, and a third random forest model are obtained, respectively.
[0054] In some embodiments of the present invention, the random forest model is trained using the first training data until a trained first random forest model is obtained, the random forest model is trained using the second training data until a trained second random forest model is obtained, and the random forest model is trained using the third training data until a trained third random forest model is obtained.
[0055] S302 : Perform model training on the support vector machine model using the first training data, the second training data, and the third training data, respectively, until a trained first support vector machine model, a second support vector machine model, and a third support vector machine model are obtained.
[0056] In some embodiments of the present invention, the support vector machine model is trained using the first training data until a trained first support vector machine model is obtained, the support vector machine model is trained using the second training data until a trained second support vector machine model is obtained, and the support vector machine model is trained using the third training data until a trained third support vector machine model is obtained.
[0057] S303 , using the first training data, the second training data, and the third training data to respectively perform model training on the BP neural network model until a trained first BP neural network model, a second BP neural network model, and a third BP neural network model are obtained.
[0058] In some embodiments of the present invention, the BP neural network model is model-trained using the first training data until a trained first BP neural network model is obtained, the BP neural network model is model-trained using the second training data until a trained second BP neural network model is obtained, and the BP neural network model is model-trained using the third training data until a trained third BP neural network model is obtained.
[0059] In some embodiments of the present invention, S104 may be implemented through S1041 to S1042, which is described in the following steps.
[0060] S1041. Overlapping the fire risk maps corresponding to the multiple fire risk assessment models with the historical fire records respectively to obtain multiple overlapping rates.
[0061] S1042. Obtain an optimal fire risk assessment model based on multiple overlap rates.
[0062] In some embodiments of the present invention, the corresponding fire risk maps are overlapped with the obtained historical fire records to obtain multiple overlap rates, that is, the historical fire points are superimposed with the corresponding high-risk and medium-high-risk areas in the corresponding fire risk maps, and multiple overlap rates are counted to obtain the optimal fire risk assessment model based on the multiple overlap rates.
[0063] Specifically, when overlaying the maps, the coordinates of historical fire points were imported into Arcmap10.8 to generate a fire point vector map, which was then overlaid with the corresponding fire risk maps, and the overlap rates of historical fire points with high-risk and medium-high-risk areas were counted.
[0064] In some embodiments of the present invention, S1042 can be implemented through S301, which is explained through the following steps.
[0065] S301 : Select the largest overlap rate from multiple overlap rates, and use the fire risk assessment model corresponding to the largest overlap rate as the optimal fire risk assessment model.
[0066] In some embodiments of the present invention, the maximum overlap rate between historical fire points and high-risk and medium-high-risk areas is selected from multiple overlap rates, and the fire risk assessment model corresponding to the maximum overlap rate is used as the optimal fire risk assessment model.
[0067] In an embodiment of the present invention, the following is a specific example 1 of a forest and grassland fire risk assessment method based on variable screening and machine learning provided in an embodiment of the present invention.
[0068] Example 1: The climate of the study area is temperate continental, with cold winters, hot summers and dry climate. The average temperature over the years is -27-16℃ and the average rainfall over the years is about 154mm. It has rich forest and grassland resources. In terms of forest and grassland fires, due to its unique terrain and climate characteristics, the fire risk is relatively high. In recent years, a comprehensive forest and grassland fire risk survey has been carried out, and the protection and management of forest and grassland resources have been continuously strengthened to reduce fire risks and reduce the occurrence of fires.
[0069] The specific implementation steps are as follows: S1. Obtain sample data: Download and process historical fire remote sensing data: Based on the MCD64A1 fire area product, perform projection conversion and splicing on the remote sensing images, and use the raster to point tool in Arcmap 10.8 software to convert the fire area data into fire point vector maps.
[0070] S2. Select environmental factor variables: In the embodiment of the present invention, 20 environmental factor variables for fire risk assessment were preliminarily selected in combination with relevant literature and characteristics of the study area. Topographic variables in environmental factor variables include altitude, slope, and aspect; climate variables include monthly average temperature, monthly average precipitation, monthly average maximum temperature, monthly average relative humidity, monthly average net surface solar radiation, monthly average wind speed, and monthly average saturated vapor pressure difference. The monthly average saturated vapor pressure difference is based on Calculations show that fuel variables include monthly average vegetation coverage and annual average aboveground biomass; human activity variables include annual average population density, annual average GDP, distance to rivers, distance to roads, distance to railways, and distance to residential areas.
[0071] Topographic factors have been shown to significantly influence the occurrence and spread of wildfires. Elevation, slope, and aspect influence the amount and structure of vegetation fuels, as well as the potential for human-induced ignition sources, which are closely related to wildfire occurrence. Climate is a key factor in fire development. High temperatures, drought, strong winds, low relative humidity, and insufficient precipitation have been shown to be key and essential factors in fire occurrence. Soil factors also play a role in fire development, with seasonal influences on wildfire probability. Grassland vegetation is the primary vehicle for grassland fires. As a key component of grassland vegetation, the spatial distribution and abundance of fuels significantly influence the occurrence and spread of grassland fires. Furthermore, human activities are closely linked to wildfire occurrence, significantly influencing fuel distribution, fire ignition intensity, and firefighting resources. Therefore, human activity variables are widely considered to be one of the most important wildfire drivers. All downloaded images were cropped and resampled to 1 km to ensure smooth subsequent analysis. Table 1 below lists the environmental factor variables.
[0072] Table 1 Environmental factor variables
[0073] S3. Variable correlation analysis: Perform Spearman correlation analysis and Pearson correlation analysis on all environmental factor variables to check whether there are environmental factor variables with strong correlation (|rho|>0.7). If there are environmental factor variable combinations with strong correlation, it is necessary to combine step S4 to screen out suitable environmental factor variables.
[0074] S4. Variable screening: Use principal component analysis, stepwise regression, and LASSO algorithm respectively, combined with step S3 to determine the final environmental factor variables used to construct the model, namely the first environmental factor variable, the second environmental factor variable, and the third environmental factor variable.
[0075] Principal component analysis: Principal component analysis can effectively reduce the dimensionality of variables. Here, variables with an explanation rate > 85% are selected for subsequent modeling; Stepwise regression: Stepwise regression is a method for selecting variables for a linear regression model. The basic idea is to introduce variables into the model one by one. After each variable is introduced, a t-test is performed on each selected variable, and then any insignificant variables are removed. This process is repeated until no new variables can be introduced. The independent variables retained in the final model are both significant and free of significant multicollinearity. A multicollinearity test is then performed on the final variables to ensure that there is no multicollinearity among the variables used to construct the model.
[0076] LASSO Regression Algorithm: The LASSO algorithm is a biased estimator for handling multicollinear data. Based on linear regression, it constructs a penalty function to limit the sum of the absolute values of the model's regression coefficients to below a set threshold and minimize the sum of squared residuals. This reduces the regression coefficients of independent variables with a weak influence on the dependent variable to even zero, while retaining important variables.
[0077] S5. Construct a machine learning model: Use the first environmental factor variable, the second environmental factor variable, and the third environmental factor variable after screening as model inputs, and construct random forest, support vector machine, and BP neural network models respectively. Determine the optimal hyperparameters of the model through ten-fold cross-validation and multiple iterations.
[0078] Random forest (RF) is a nonlinear modeling method that improves prediction accuracy by integrating a series of decision trees (Breiman, 2001). RF regression uses bootstrap sampling, where each sample is used to construct a regression tree. Training samples are continuously selected to minimize the sum of squared residuals until a complete tree is formed. After constructing multiple regression trees, the average prediction value of all individual trees is used as the final RF estimate. Among them, the support vector machine (SVM) is a supervised non-parametric machine learning algorithm. It consists of a set of hyperplanes in high-dimensional or infinite-dimensional space. It can find the optimal balance between model complexity and learning accuracy based on limited sample information to achieve the best generalization ability of the model. BP-ANN (Back Propagation Artificial Neural Network) is a multi-layer neural network model trained using the error backpropagation algorithm. It consists of an input layer, hidden layers, and an output layer, simulating the connections and propagation of biological neurons. After obtaining a predicted value through forward propagation, feedback adjustments are made based on the error between the actual value and the predicted value. The network structure's weights and thresholds are continuously updated to ensure that the network output is as close to the expected output as possible.
[0079] S6. Model Evaluation: Accuracy, sensitivity, specificity, F1 value, kappa coefficient, and area under the curve were calculated for the model training and test sets, respectively. The model fit was evaluated using these metrics. By comparison, the kappa coefficients of the Step_RF, LAS_RF, and LAS_SVM models were > 0.75, indicating high model consistency. Table 2 below shows the model evaluation results for different models on the training and test sets. Step represents the stepwise regression method, while Step_RF uses the stepwise regression method to filter variables to obtain the filtered environmental factor variables, which are then used to train and test the random forest. The interpretations of other tables are similar.
[0080] Table 2 Model evaluation results of different models on the training set and test set
[0081] S7. Fire risk inversion: The three best models obtained in S6 were used to invert the fire risk map of the study area from 2001 to 2022. The mask tool in Arcmap 10.8 was used to extract the fire risk map of the forest-steppe area. The natural breakpoint method was used to classify the fire risk values into five levels: high risk (H), medium-high risk (MH), medium risk (M), medium-low risk (ML), and low risk (L). S8. Selecting the Optimal Model: The forest and grassland fire risk zoning results were validated using forest and grassland fire point data from the study area from 2008 to 2022. The results showed that the proportion of 210 observed forest and grassland fire points falling into high- and medium-high-risk zones was as follows: Step_RF (63.33%) > LAS_RF (57.14%) > LAS_SVM (42.68%). Therefore, considering both model evaluation indicators and the overlap rate of historical fire points, the Step_RF model was the optimal model for western grassland fire risk, with the highest overlap rate of both model evaluation indicators and historical fire points. This demonstrates that the forest and grassland fire risk zoning simulated by the Step_RF model accurately identifies areas with high and high risk of forest and grassland fires, confirming the reliability of the model zoning results.
[0082] S9. Results Visualization: A grid-based graphical display makes the assessment results more intuitive. Different colors are assigned to areas with different forest and grassland fire risk levels to facilitate zoning. This research helps fire prevention professionals and other stakeholders understand the forest and grassland fire risk level of their target areas directly through the color of the grid.
[0083] The forest and grassland fire risk assessment model, constructed by combining multiple indicator screening and machine learning models, can provide intuitive, detailed, and highly accurate technical means for forest and grassland fire risk assessment and forest and grassland fire risk zoning. The present invention utilizes nine fire risk assessment models, including principal component analysis, stepwise regression, and the LASSO algorithm, combined with random forest, support vector machine, and BP-neural network models, to construct a forest and grassland fire risk assessment model. Regional-scale forest and grassland fire risk assessment and zoning research was conducted, and the results showed that: (1) Taking into account the comprehensive model evaluation indicators and the overlap rate between historical fire data and inversion results, Step_RF is the optimal model for simulating forest and grassland fire risks in the study area and is also the target fire risk assessment model.
[0084] (2) The overall spatial distribution of forest and grassland fire risk in the study area is high in the northwest and low in the southeast. High-risk and medium-high-risk areas account for 21.34% of the total study area and are primarily located in central Ili, western Bozhou, western Tacheng, and northwestern Altay. In the south, forest and grassland fire risk is relatively low, a pattern similar to the spatial distribution of burned area.
[0085] The present invention combines multiple indicator screening and machine learning models to carry out forest and grassland fire risk assessment, and then divides forest and grassland fire risk levels. Compared with existing forest fire risk assessment methods, it can better provide effective reference in guiding forest and grassland fire prevention work, curbing the occurrence of forest and grassland fires, and mitigating the environmental effects of forest and grassland fires, thereby improving the efficiency of forest and grassland fire protection and enhancing the effectiveness of forest and grassland fire protection.
[0086] Reference Figure 2 , shows a schematic structural diagram of an electronic device according to an embodiment of the present invention. The specific embodiment of the present invention does not limit the specific implementation of the electronic device.
[0087] like Figure 2 As shown, the electronic device may include: a processor (processor) 502, a communication interface (Communications Interface 504), a memory (memory) 506, and a communication bus 508.
[0088] in: The processor 502 , the communication interface 504 , and the memory 506 communicate with each other via a communication bus 508 .
[0089] The communication interface 504 is used to communicate with other electronic devices or servers.
[0090] The processor 502 is configured to execute the program 510 , and specifically may execute the relevant steps in the above method embodiment.
[0091] Specifically, the program 510 may include program codes, which include computer operation instructions.
[0092] Processor 502 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. The one or more processors included in a smart device may be of the same type, such as one or more CPUs, or different types, such as one or more CPUs and one or more ASICs.
[0093] The memory 506 is used to store the program 510. The memory 506 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk storage.
[0094] The program 510 may be specifically configured to enable the processor 502 to execute operations corresponding to the methods described in the above method embodiments.
[0095] The specific implementation of each step in program 510 can be found in the corresponding descriptions of the corresponding steps and units in the above-mentioned method embodiments, and will not be repeated here. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the above-mentioned devices and modules can refer to the corresponding process descriptions in the above-mentioned method embodiments, and will not be repeated here.
[0096] It should be pointed out that, according to the needs of implementation, the various components / steps described in the embodiments of the present invention can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present invention.
[0097] The methods according to the embodiments of the present invention described above can be implemented in hardware, firmware, or as software or computer code that can be stored on a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code originally stored on a remote recording medium or non-transitory machine-readable medium downloaded over a network and then stored on a local recording medium. Thus, the methods described herein can be processed by such software stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It will be understood that a computer, processor, microprocessor controller, or programmable hardware includes a storage component (e.g., RAM, ROM, flash memory, etc.) that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the methods described herein are implemented. Furthermore, when a general-purpose computer accesses the code for implementing the methods described herein, the execution of the code transforms the general-purpose computer into a dedicated computer for performing the methods described herein.
[0098] Those skilled in the art will appreciate that the units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present invention.
[0099] The above implementation methods are only used to illustrate the embodiments of the present invention, and are not intended to limit the embodiments of the present invention. Ordinary technicians in the relevant technical field may make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of the present invention, and the scope of patent protection of the embodiments of the present invention should be defined by the claims.
Claims
1. A forest and grassland fire risk assessment method based on variable screening and machine learning, characterized by: include: Obtain historical fire remote sensing data, and extract environmental factor variables from the pre-processed historical fire remote sensing data; Performing correlation analysis on the environmental factor variables to obtain correlation results, and performing variable screening on the environmental factor variables based on the correlation results, Lasso regression algorithm, principal component analysis method, and stepwise regression method to obtain screened environmental factor variables; When the multicollinearity test results of the screened environmental factor variables meet a preset condition, data segmentation is performed on the screened environmental factor variables to obtain training data and test data, and model training and testing are performed on multiple fire risk assessment models to be trained using the training data and the test data until multiple fire risk assessment models are obtained; An optimal fire risk assessment model is determined based on the multiple fire risk assessment models and the acquired historical fire records, and the screened environmental factor variables are inverted using the optimal fire risk assessment model to obtain a fire risk map of the forest and grassland.
2. The method according to claim 1, characterized in that The screened environmental factor variables include a first environmental factor variable, a second environmental factor variable and a third environmental factor variable; Before dividing the screened environmental factor variables into data to obtain training data and test data, the method further includes: A multicollinearity test is performed on each of the first environmental factor variable, the second environmental factor variable, and the third environmental factor variable to obtain corresponding multicollinearity test results.
3. The method according to claim 2, characterized in that The screened environmental factor variables are divided into data to obtain training data and test data, including: Performing data division on the first environmental factor variable to obtain first training data and first test data; Performing data division on the second environmental factor variable to obtain second training data and second test data; Data division is performed on the third environmental factor variable to obtain third training data and third test data.
4. The method according to claim 2, characterized in that The method of using the training data and the test data to perform model training and testing on a plurality of fire risk assessment models to be trained until a plurality of fire risk assessment models are obtained includes: Using the first training data, the second training data, and the third training data, respectively, to perform model training on the plurality of fire risk assessment models to be trained, to obtain a plurality of trained fire risk assessment models; When the performance of the multiple trained fire risk assessment models all meet the preset indicator conditions, the multiple trained fire risk assessment models are tested using the first test data, the second test data, and the third test data to obtain corresponding fire risk maps; When the corresponding fire risk maps all meet the preset indicator conditions, the multiple fire risk assessment models are obtained.
5. The method according to claim 4, characterized in that The multiple fire risk assessment models to be trained include a random forest model, a support vector machine model and a BP neural network model; The method of using the first training data, the second training data, and the third training data to respectively perform model training on the plurality of fire risk assessment models to be trained to obtain a plurality of trained fire risk assessment models includes: Performing model training on the random forest model using the first training data, the second training data, and the third training data, respectively, until a trained first random forest model, a second random forest model, and a third random forest model are obtained, respectively; Performing model training on the support vector machine model using the first training data, the second training data, and the third training data, respectively, until a trained first support vector machine model, a second support vector machine model, and a third support vector machine model are obtained; The BP neural network model is trained using the first training data, the second training data, and the third training data, respectively, until a trained first BP neural network model, a second BP neural network model, and a third BP neural network model are obtained.
6. The method according to claim 3, characterized in that The determining of the optimal fire risk assessment model based on the multiple fire risk assessment models and the acquired historical fire records includes: Overlapping the fire risk maps corresponding to the multiple fire risk assessment models with the historical fire records to obtain multiple overlapping rates; The optimal fire risk assessment model is obtained based on the multiple overlap rates.
7. The method according to claim 5, characterized in that The obtaining of the optimal fire risk assessment model based on the multiple overlap rates includes: A maximum overlap rate is selected from the multiple overlap rates, and a fire risk assessment model corresponding to the maximum overlap rate is used as the optimal fire risk assessment model.
Citation Information
Cited By
Climate mode optimization and set estimation method based on interpretable and machine learning
CN121350498A
Explainable and machine learning based climate model preference and ensemble prediction method
CN121350498B