A method for identifying the type of forest fire source considering the environmental characteristics of the fire ignition point
By combining MODIS data and multi-source environmental data, using the Jenks-DBSCAN model and classification algorithm, the shortcomings of satellite remote sensing in forest fire research were solved, and the detailed information extraction of forest fire events and accurate identification of fire source types were achieved, which improved the accuracy and data integrity of forest fire management.
Patent Information
- Application Number
- CN202311043674.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-18
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2043-08-18
AI Technical Summary
When using satellite remote sensing to conduct forest fire research, the existing technology failed to effectively identify important information such as the fire location and spreading direction of each forest fire event. The collection of fire source information relies on artificial statistics and has great uncertainty. It lacks methods to distinguish different types of fire sources. The acquisition method of multi-source remote sensing information ignores the integrity and dynamic characteristics of forest fire events.
By obtaining MODIS combustion products and land cover data, the forest fire footprints are extracted using the Jenks-DBSCAN model, combined with multi-source environmental data for in-depth information mining, and a classification model is established to identify the fire source types of forest fire events, including combustible materials, terrain, man-made activities and meteorological variables, and logistic regression, random forests and support vector machines are used for training and verification.
Accurate acquisition of the location, date and cause of fire incidents, improve the accuracy of fire source type identification, complete the forest fire database information, and support more effective fire prevention policy formulation and forest fire risk assessment.
Smart Images

Figure CN117113177B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote sensing technology, in particular to a method for identifying forest fire source types by taking into account environmental characteristics of fire starting points. Background Art
[0002] Forest fire, a natural ecological factor and one of the largest potential carbon release mechanisms in forest ecosystems, plays a key role in determining the function and structure of forest ecosystems. Fire monitoring, based on individual fire events, primarily relies on the precise location of fire points and has historically been a key focus of forest fire prevention efforts. Building a fire database containing rich attribute information on each fire, based on the extraction of individual fire events, is crucial for better understanding regional fire ecology. Therefore, timely and accurate understanding of the specific time, location, and cause of forest fires facilitates the exploration of fire causes and patterns, and is crucial for sustainable forest management supported by fire prevention management.
[0003] Effectively extracting information and managing data from historical fires, and thereby predicting and preventing the occurrence and development of forest fires, has become a key concern for natural resource managers worldwide, particularly forest managers. Currently, in my country, forest fire data management still uses the latitude and longitude of the reported fire point to represent the location of a fire. Remote sensing technology, with its large-scale observations and diverse temporal and spatial resolutions, can effectively describe surface processes at various levels, to some extent addressing the limitations of statistical data in the information and spatialization of fire management. This technology can be leveraged to improve forest fire data management. A fire footprint refers to the spatiotemporal information, including the outline and boundaries, of each forest fire event. In recent years, my country has gradually incorporated network video monitoring, wireless Internet of Things (IoT), and geographic information systems into its forest fire monitoring methods. For example, Xu Aijun et al. proposed a method for identifying forest fire events from drone videos using image segmentation and matching algorithms, based on the color characteristics of forest fire flames captured in drone videos. Chen Ke also proposed an algorithm for extracting suspected fire areas using superpixel segmentation combined with color models, enabling real-time forest fire event information. Moderate Resolution Imaging Spectroradiometer (MODIS) data, due to its high temporal resolution, moderate spatial resolution, and freely available surface observation applications, has been widely used for large-scale forest fire detection, fire risk mapping, and post-fire vegetation recovery assessment. One of the most commonly used satellite-based global burn area products is the sixth edition of the MODIS Global Burned Area product, MCD64A1. Based on the Burn Date information in this data, Luo Kaiwei et al. developed a method for extracting forest fire footprints from remote sensing data using the DBSCAN algorithm. In a 2023 paper, Su Huiyi et al. optimized a Jenks-DBSCAN algorithm for extracting forest fire footprints, and performed spatial mapping of forest fires and burned area calculations.
[0004] The factors that influence the occurrence and spread of forest fires are primarily categorized into four categories: fuels, topography, climate, and fire sources. Based on the extraction of forest fire events, various spatial environmental data can be used to represent different types of environmental factors at different spatiotemporal scales. Comprehensively analyzing the relationship between these factors and forest fires, and then identifying the dominant factors influencing forest fire occurrence and constructing various models, has become a powerful technical tool for scientific forest fire management. For example, Salame et al. used the logistic method to conclude that road density and agricultural expansion have significant impacts on forest fire occurrence. Zhang et al. used environmental factors such as topography, human activities, meteorological conditions, climate factors, and vegetation cover to estimate the probability of forest fire occurrence in Heilongjiang Province. Li Zhequan used historical MODIS hotspot data and principal component analysis to screen forest fire risk warning factors and constructed an indicator system for forest fire risk warning factors in Guangdong Province. These studies have demonstrated that the occurrence of forest fires is related to many factors. Fuel, topography, and climate create the conditions for the ignition and spread of forest fires. Fire sources, as the key factor directly causing forest fires, are also an important element to record in forest fire data management.
[0005] Forest fire sources can be categorized into two main types: natural and human. Natural fire sources primarily include volcanic eruptions, meteorite falls, spontaneous peat combustion, and lightning strikes. Human fire sources are further divided into productive fire sources (such as burning wasteland, smelting mountains, and burning ash for fertilizer) and non-productive fire sources (such as smoking outdoors and burning paper at graves). Natural fires in Australia are primarily caused by a combination of high temperatures, drought, and strong winds. Natural fires in my country are primarily caused by lightning, primarily occurring in the Greater Khingan Range region of Heilongjiang, Hulunbuir League in Inner Mongolia, and the Altai region of Xinjiang, with the Greater Khingan Range being particularly prominent. Regarding human fires, the majority of forest fires in my country are caused by human activities, particularly burning wasteland for charcoal and visiting graves for ancestral worship. Currently, human fires are the primary cause of forest fires in my country. Research indicates that human factors account for over 98% of the causes of forest fires nationwide, and many records of fires with unknown sources exist.
[0006] However, the existing technology still has certain defects:
[0007] 1) Currently, large-scale forest fire studies using satellite remote sensing rarely monitor individual fire events for regional ecological management. Furthermore, when establishing and managing forest fire databases, important information about fire succession, such as the location of a fire and its direction of spread, is missed.
[0008] 2) At present, forest fire information mining based on remote sensing can be extracted from fire scars to establish a large-scale, long-term forest fire database. However, the collection of fire source information still relies on manual statistics and is subject to many uncertainties, which weakens the richness of forest fire attribute information.
[0009] 3) Currently, although a large number of studies have combined various environmental variables to classify a region's fire risk, and some have predicted the probability of occurrence of either human-caused or natural fires, forest fires are a complex system, and accurate simulation of their causes remains a challenge. Few researchers have explored the distinction between different fire source types based on the multi-source environmental variable information provided by remote sensing data.
[0010] 4) Currently, when acquiring multi-source remote sensing information about a forest fire, environmental variables are either captured using the center point of an irregular polygon representing the fire location to represent the entire event, or the mean of the variable values across all pixels within the polygon is calculated. This approach ignores the completeness of the forest fire event and its dynamic characteristics. Summary of the Invention
[0011] In response to the problems existing in the above-mentioned background technology, the present invention provides a method for identifying the type of forest fire source taking into account the environmental characteristics of the fire starting point.
[0012] To achieve the above object, the present invention provides the following technical solution: a method for identifying the type of forest fire source taking into account the environmental characteristics of the fire point, comprising the following steps:
[0013] Step 1: Obtain the Burn Date data from the MODIS remote sensing product MCD64A1 and the land cover type data MCD12Q1, preprocess them, extract the forest fire footprint using the Jenks-DBSCAN model, and select the fire footprints that match the fire points in the ground survey data to obtain the forest fire footprints containing fire source information;
[0014] Step 2: Combine the MCD64A1 data from step 1 with the extracted forest fire footprints, and use spatial partitioning statistical tools to conduct in-depth information mining on each footprint to obtain the fire point location and fire date, center point location and main burning date of the forest fire event;
[0015] Step 3: Obtain various environmental data related to forest fires, including fuel, terrain, human activities, and weather, and pre-process them to obtain variable factors;
[0016] Step 4: Using the different spatiotemporal attributes of the fire point and center point extracted in step 2, prepare a forest fire event sample dataset for model training and divide it into a modeling dataset and a validation dataset in a ratio of 7:3;
[0017] Step 5: Use the modeling data of the forest fire event sample data set in step 4 to select the optimal variable combination for different classification models, where the classification models include logistic regression, random forest and support vector machine, and debug the corresponding model training parameters; then input the verification data into the trained model, and evaluate the accuracy of the model's fire source identification results and the fire source information provided in the investigation report to obtain the optimal fire point-model-variable combination that can identify the fire source of the forest fire event.
[0018] Preferably, the specific operation steps in step 1 are as follows:
[0019] 1) Perform format conversion, reprojection, and image stitching on the original layered HDF files in the MODIS combustion product MCD64A1 and land cover type product MCD12Q1, and extract the data layer containing the required burning date information in the MCD64A1 product;
[0020] 2) Mask the vegetation pixels in the land cover type product MCD12Q1 over the burnt pixels in MCD64A1 to exclude non-combustible pixels and reduce the uncertainty of burnt pixels;
[0021] 3) Fire point records obtained from ground surveys were screened. Record information included: latitude and longitude of the fire point, burned area, fire cause, time of fire discovery, and time of fire extinguishing. Records with a total burned area greater than 25 hectares were screened based on the 500-meter spatial resolution of MODIS data. Records with unknown or missing fire source information were excluded. Vector fire point files were generated in ArcGIS software based on the longitude and latitude provided by the remaining records. The projection method was consistent with that of the preprocessed MODIS data. Fire source types were then classified into natural and human-caused fires, and assigned as categorical variables.
[0022] 4) Input the burned pixels into the Jenks-DBSCAN model for clustering to obtain multiple clusters, each of which is tentatively defined as the footprint of a forest fire event; calculate the burned area based on the number of burned pixels and spatial resolution, and calculate the coefficient of determination To reflect the matching degree between the fire area investigated on the ground and the fire footprint area extracted by the Jenks-DBSCAN model:
[0023]
[0024] in, Represents the true value of the fire area surveyed on the ground; The representative model extracts the statistical value of the fire footprint area; The mean represents the true value; represents the number of samples, i.e. the number of forest fires;
[0025] 5) Use ground survey data to match the forest fire footprints of the clustering results to screen out forest fire event footprints with fire source information; establish a circular buffer zone with a radius of 1 km with each survey record fire point as the center, and then spatially overlay it with the forest fire footprints obtained by clustering for comparison; based on spatial association, perform a time verification to determine the footprint corresponding to the forest fire event, and verify whether the monthly fire information contained in the verification fire point and the Julian day (DOY) value of the burned pixel covered by the fire footprint can match the Gregorian calendar month; after performing the above spatial association and time verification on each ground survey fire point and the fire footprint extracted by clustering, the forest fire footprints that match one by one are used as training samples, and the dependent variable is the fire source type recorded by the ground fire point.
[0026] Preferably, the comparison method includes: if the buffer zone overlaps with a cluster of burning areas, it is determined to be the same fire event; if multiple forest fire footprints overlap with the buffer zone, the cluster with the largest overlapping pixels is defined as the overlapping or matching fire event.
[0027] Preferably, the extraction of the ignition point and the center point in step 2 is specifically as follows:
[0028] The forest fire footprint samples obtained in step 1 are clustered and the raster clusters are converted into vector irregular polygons. The zoning statistics tool of ArcGIS software is used to count the DOY values of all MCD64A1 burning pixels spatially covered by the vector polygons to obtain the minimum, maximum, median and mode DOY information of the forest fire footprint.
[0029] Preferably, the recorded minimum DOY value is used as the ignition date of the forest fire event, and the geometric center of the minimum value pixel is the ignition point; the recorded mode of DOY is used as the main burning date of the forest fire event, and the geometric center of all pixels covered by the positioning polygon is the center point of the forest fire event.
[0030] Preferably, the specific steps in step three are as follows:
[0031] 1) The fuel variable was extracted using the 1:1,000,000 vegetation map of the People's Republic of China. Based on this map, an environmental variable representing the fuel type of different forest tree species related to forest fires was extracted. The different forest plants were classified according to the local fuel type as non-flammable, relatively flammable, flammable, and extremely flammable.
[0032] 2) Extraction of terrain variables Based on digital elevation model data, terrain variables are extracted, including elevation, slope, aspect, slope length and relief;
[0033] 3) Human activity variables include two variables: spatially, local road network and residential vector layer data are collected and transformed and projected. Based on the location of the fire point or center point of the forest fire incident in step 2, the minimum distance from the fire point to the nearby road network or residential area is calculated as a variable; temporally, for festivals where burning objects pose a forest fire hazard, the fire start date and the main burning date are calculated as the closest day of the corresponding festival each year as another variable;
[0034] 4) Meteorological elements include daily and monthly meteorological variables. Unlike variables related to topography, human activities, and combustible materials, which are relatively fixed in characteristic space, meteorological variables are easily affected by local seasonal patterns and various complex components. Therefore, meteorological variables need to achieve the highest possible temporal resolution to better represent the specific conditions when forest fires occur. The daily meteorological variables used are obtained by first cleaning and completing the data from daily meteorological station data. For records with missing data on a certain day, the average of the most recent records before and after the missing date is calculated as the completion value. Then, co-kriging interpolation is performed with elevation as a covariate to obtain the meteorological element raster data corresponding to the entire region. Monthly meteorological variables include monthly average temperature and precipitation.
[0035] 5) All the above environmental variables are spatially resampled and projected to have the same spatial resolution and projection as MODIS products. Variable factors are mainly divided into numerical and categorical types. Categorical variables need to be quantified during variable selection and model building.
[0036] Preferably, the product of step 1) includes the distribution information of 11 vegetation type groups, 55 vegetation types, 960 vegetation communities and subcommunities, and more than 2,000 dominant species and major crops in China.
[0037] Preferably, the meteorological element grid data of step 4) includes average dew point, maximum temperature, minimum temperature, maximum sustained wind speed, precipitation, average sea level pressure, average temperature, average visibility and average wind speed.
[0038] Preferably, the specific steps in step five are:
[0039] 1) Input the training set of forest fire incident sample data set, use the backward feature method under ten-fold cross-validation to select the most suitable number of variables for classification modeling, and then select the optimal variable combination for three different classification models based on this number;
[0040] 2) Using the variance inflation factor combined with the chi-square test as a variable selection method for the logistic regression model; using the combined ranking of the mean decreasing precision and mean decreasing Gini index of each variable in the random forest importance analysis as a variable selection method for the random forest classification model; using the optimal variable combination obtained by the backward feature method under ten-fold cross-validation as the variable selection method for the input support vector machine;
[0041] 3) After obtaining the optimal variable combination, further debug the optimal training parameters based on different model classification principles to minimize the model error;
[0042] 4) The independent variables in the validation dataset were put into the classification model to obtain the fire source identification results. The area under the curve (AUC) calculated by the receiver operating characteristic (ROC) curve was used as the performance evaluation indicator of the model. The closer the AUC value is to 1, the higher the accuracy of the model.
[0043]
[0044] Where M is the number of positive samples, N is the number of negative samples, M + N = n represents the total number of samples, the denominator is the total number of combinations of positive and negative samples, the numerator is the number of groups in which the scores of positive samples are greater than the scores of negative samples, and Score represents the probability that each test sample belongs to a positive sample;
[0045] At the same time, the model recognition results were compared with the fire source types in the investigation records, and the overall accuracy, user accuracy and Kappa coefficient were used as evaluation indicators for the binary classification effect of different fire source discrimination models in the test data set;
[0046]
[0047]
[0048]
[0049]
[0050]
[0051] In the set of samples to be modeled containing fire source information, P represents the number of real human-caused fires, N represents the number of real natural fires, T represents accurate predictions, and F represents incorrect predictions. FP, FN, TP, and TN represent the number of samples with different judgment results for the two fire source categories in the test dataset. FP indicates that a human-caused fire recorded in the investigation record is classified as a natural fire by the model, TP indicates that a human-caused fire recorded in the investigation record is classified as a human-caused fire by the model, and so on.
[0052] 5) Comparing the fire point group and the center point group of the forest fire event sample dataset, and the performance of different classification models in fire source identification in the validation set, the optimal fire point-model-variable combination is obtained. This is used as the identification method for forest fire events in the remaining area with unknown fire source information, thereby completing the forest fire database information.
[0053] Compared with the prior art, the present invention has the following beneficial effects:
[0054] 1. The present invention combines a forest fire footprint extraction model with a binary classification model for fire source identification, and develops a technical framework for identifying fire source types based on forest fire event extraction and environmental characteristics of the fire point. The present invention can accurately obtain the fire location, fire date, and fire cause of large-scale forest fire events.
[0055] 2. This invention addresses the shortcoming of large-scale remote sensing inversion of forest fire events, which fails to fully characterize fire ignition sources. By using partitioned statistical tools to extract forest fire footprints using the Jenks-DBSCAN model, this method further explores the temporal attributes of burning pixels in MCD64A1 data, obtaining the location and date of each fire footprint. This extraction then yields a more representative combination of environmental variables, enabling multi-source environmental variables to more accurately simulate the spatiotemporal characteristics of fire occurrences, thereby effectively identifying fire source types.
[0056] 3. The proposed technical framework for identifying fire source types based on forest fire event extraction and environmental characteristics of the fire origin not only facilitates systematic assessment of the combustion mechanisms of fires in a region, but also facilitates the investigation of historical forest fires with unknown source information and provides a technical means to meet the corresponding scientific research needs for data integrity. Furthermore, it lays a solid foundation for formulating more effective fire prevention and firefighting policies for forest fires of different source types.
[0057] 4. This invention leverages remote sensing and geographic information platforms to improve forest fire databases. Building on existing fire footprint extraction and research into the relationship between forest fires and environmental factors, it provides a reliable method for identifying fire source types. This framework not only supplements forest fire records with unclear source information but also further mines fire origin information from high-temporal-resolution remote sensing combustion products. This allows for effective prevention and relief of forest fires of different source types in key fire prevention areas, providing valuable insights into regional fire ecology. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.
[0059] In the attached figure:
[0060] Figure 1 This is a technical framework flow chart of the present invention for identifying fire source types based on forest fire event extraction and combining fire point environmental characteristics;
[0061] Figure 2 This is an example map of the starting points and center points of several fire footprint patches in the southeastern part of the border area between Songling District and Huma County in Daxing'anling in 2002;
[0062] Figure 3 It is a scatter plot of the statistical area of fire footprint pixels that can match the ground survey and the area of ground fire point records over the past 20 years;
[0063] Figure 4 The spatial distribution of 68 forest fire events with a burn area greater than 100 hectares that can be matched with ground surveys over the past 20 years;
[0064] Figure 5 It is the ROC curve diagram of the training set and the validation set of the fire point dataset in the logistic regression model of the present invention;
[0065] Figure 6 It is a schematic diagram of the recognition results of modeling and verification when classifying fire sources using samples of the logistic regression model of fire point data in the present invention. DETAILED DESCRIPTION
[0066] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0067] Example: Figure 1 As shown, the present invention provides a framework for identifying fire source types based on forest fire event extraction combined with fire point environmental characteristics, based on existing fire footprint extraction methods. This framework fully considers the spatiotemporal characteristics of the fire point location during the burning process of a forest fire event, and fully utilizes the high temporal resolution of MODIS products based on the thermal sensitivity principle of detecting burning pixels. In the process of obtaining forest fire event characteristic variables using fire points, multi-source environmental data is more accurately extracted into the fire point sample data set, thereby improving the accuracy of the fire source classification model and filling the technical gap in fire source type identification in the completion of forest fire relationship information. Specifically, the following steps are included:
[0068] 1. Data Acquisition and Preprocessing
[0069] (1) The geographical location of this embodiment is the Greater Khingan Range, a region with a high incidence of forest fires in my country. The 500-meter resolution MODIS data products with product numbers h25v03 and h26v03 from 2001 to 2020 downloaded from the website of the United States Geological Survey (USGS) are used. The MODIS Reprojection Tool (MRT) provided by NASA's EOSDIS is used to convert the original hierarchical data format (HDF) format, reproject, and stitch the images. The required burn date layer, BurnDate, is extracted from the MCD64A1 product and clipped based on the administrative boundary vector elements of the Greater Khingan Range. The land cover classification scheme of the MCD12Q1 product is type 1 dataset. Based on the fire study area in which the major land categories in the product are forest land, grassland, and crops as combustibles, the pixel values are reclassified year by year in ArcGIS software. The combustible type is set to 1 and the non-combustible type is set to 0. The annual MCD64A1 data are then masked to leave only the burn pixels of the combustible category.
[0070] (2) The forest fire survey data comes from the fire point records in the region from 2001 to 2020 provided by the Daxinganling Forestry Group. The latitude and longitude positions are obtained using a manual handheld GPS recorder. Vector points are generated in ArcGIS software, and a 1-kilometer buffer zone is established with each point as the center to match the overlapping fire footprints of the model judgment results and actual fire event records.
[0071] (3) The vegetation type product used is the 1:1 million vegetation map of the People's Republic of China, which includes information on China's 11 vegetation type groups, 55 vegetation types, 960 vegetation communities and subcommunities, as well as the distribution of more than 2,000 dominant species and major crops. Based on the forest type characteristics of the Greater Khingan Range in Table 1, the different forest plants are divided into four categories according to their types, as shown in the following table:
[0072]
[0073] (4) The terrain data used was a digital elevation model covering the entire study area, derived from the SRTM (Shuttle Radar Topography Mission) 3.0 version from the Geospatial Data Cloud website, with a spatial resolution of 90 meters. The elevation data was used to further extract the slope, aspect, length, and relief of the Greater Khingan Range region in ArcGIS software.
[0074] (5) Road network and settlement data. Road network data can be obtained from the open source geographic information data OSM (OpenStreet Map) platform. This example obtains vector data of all urban roads in the Greater Khingan Range (including national roads, provincial roads, county roads, township roads, and first-, second-, third-, and fourth-level urban roads); settlement data is obtained from the relevant information released by the official websites of the Ministry of Civil Affairs and the National Bureau of Statistics (including capitals, provinces, autonomous regions, municipalities directly under the central government, prefecture-level cities, counties, autonomous counties, banners, prefecture-level city districts, townships, as well as agricultural, forestry, animal husbandry, fishery, enterprises, institutions, grazing sites, etc.), organized into formatted text data, and then geocoded to produce spatial vector points with projection information. On this basis, the neighbor analysis module in the ArcGIS analysis tool can be used to calculate the distance value from the vector fire point to the nearest road or settlement point.
[0075] (6) The meteorological station data used in this example is sourced from the official NOAA website. Stations within and around the Greater Khingan Range were selected, and daily meteorological records that were continuously operational between 2001 and 2020 were selected. Elevation was used as a covariate, and co-kriging spatial interpolation was performed to obtain a daily spatial trend map of meteorological elements within the study area. The meteorological values at specific locations (the starting point or center of each fire footprint) were extracted and used as meteorological variables in the subsequent classification model. Considering that the daily precipitation in the daily meteorological station record data is sometimes too small, and that most days of the year have a minimum value close to 0, which does not meet the requirements of the interpolation algorithm, the national monthly 1 km resolution raster data product openly accessible from the Qinghai-Tibet Plateau National Data Center is used to supplement the monthly meteorological information for the two indicators of average temperature and precipitation.
[0076] 2. Use the Jenks-DBSCAN method to extract fire footprints and locate the fire point and center point of each footprint after matching with forest fire survey records
[0077] Based on the process of the existing Jenks-DBSCAN method, all fire footprint patches in the past 20 years were obtained. The spatial partitioning statistical tool was further used to obtain information on each fire footprint, and the location and date of the fire point, the center point and the main burning date of the forest fire event were obtained. Figure 2The figure shows examples of the origin and center points of several fire footprint patches in the southeastern border area of Songling District and Huma County in the Greater Khingan Range in 2002. The image shows that each forest fire event has an irregular polygonal pattern (footprint 4). Sometimes a forest fire event is composed of multiple adjacent fire patches (footprint 1), but the time difference between the burning pixels can be used to distinguish two adjacent burning pixels into two forest fires (footprints 2 and 3). The changes in the DOY values of the burned pixels can be used to locate the fire origin and the geometric center of the fire patch, thereby reconstructing the spatiotemporal migration process of a forest fire event.
[0078] The fire footprints were then spatiotemporally matched against preprocessed ground fire survey records, ultimately yielding 68 fire event records that could be spatiotemporally matched to the clustered fire footprint patches. Sampling without replacement was used to extract these 68 fire event samples containing fire source information, yielding 70% of the modeling samples (47) for classification training and the remaining 30% for validation (21).
[0079] like Figure 3 The figure shows a scatter plot of the area of fire footprint pixels that can be matched with the ground survey and the area of ground fire records over the past 20 years. The figure shows that the burned area obtained by spatially matching fire footprint statistics is verified with the fire area provided by local reports, using the determination coefficient R between the area of fire events that can reflect the field survey and the converted area of fire footprint pixel statistics extracted by the model. 2 To evaluate the accuracy of Jenks-DBSCAN model in extracting fire footprint area, R 2 The forest fire footprint obtained by the Jenks-DBSCAN model is 0.87, which means that the forest fire events on the ground can be represented.
[0080] like Figure 4 Figure 2 shows the spatial distribution of 68 forest fires with burn areas greater than 100 hectares over a 20-year period that matched ground surveys. The image shows that lightning-induced fires were widespread in the forest fire sample, occurring in all districts and counties except Jiagedaqi District, with a primary concentration in Huzhong District and Huma County within the study area. Human-caused fires, on the other hand, were primarily concentrated in Songling District, Huma County, and Jiagedaqi District in the southeastern part of the study area.
[0081] 3. Variable Preparation and Selection
[0082] According to the spatial location and time attributes of the fire point and center point of each forest fire event sample, the following table 2 of the collection of modeling candidate variables extracted from multi-source environmental data is obtained, which shows various environmental variables related to forest fires, as shown in the following table:
[0083]
[0084] Because variables extracted from different environmental factors based on fire origins and center points differ somewhat, the data groups are named the Origin Group and the Center Group, and the corresponding variable names are suffixed with _O (Origin) and _C (Center) to distinguish them. Based on the algorithmic principles of the three classification models, 70% of the modeling samples were used to obtain the optimal variable combination in Table 3, as shown in the following table:
[0085]
[0086] As shown in Table 3, the optimal variable combinations selected by each model in different data groups, the average dew point (DEWP) on the day was selected as the optimal variable by most models, followed by the average visibility (VISIB) on the day. These two variables can be used as advantageous variables to distinguish between human-caused fires and natural fires. The reasons are: (1) The dew point temperature is a temperature value. When the water vapor in the air has reached saturation, the air temperature is the same as the dew point temperature; when the water vapor has not reached saturation, the air temperature will be higher than the dew point temperature. Therefore, the difference between the dew point temperature and the air temperature can indicate the degree of saturation of the water vapor in the air. The dew point temperature also has a direct impact on the ambient humidity and serves as a substitute for the humidity variable in the study. Humidity directly affects the moisture content of forest fuels. In contrast, in reality, the mechanism of forest fire ignition is that in an environment that is not dry enough and the moisture content of combustible materials is too high to cause combustion under natural factors, the fire sources brought by humans (smoking, burning paper, etc.) can reduce the influence of environmental humidity on fire ignition. Therefore, the average dew point of the day can be selected by most models as the dominant variable for classifying man-made fire and natural fire; (2) The average visibility of the day reflects the visibility conditions in the weather, which is mainly affected by fine particulate matter in the air. It is also an indicator of the transparency of the atmosphere. The subjectivity of the human fire source brought into the forest area is affected by the weather represented by the atmospheric visibility on that day. The average visibility of the day can be used as a reference for human activities.
[0087] 4. Train the classification model based on the modeling data and select the optimal variable combination, and use the validation data to test the accuracy of the model in identifying fire sources.
[0088] The modeled samples from the fire point group and the center point group were used to train the three classification models and evaluate the model's AUC value. At the same time, the validation samples were input into the trained classification model according to the corresponding optimal variable combination. The fire source types output by the model were compared with the fire source types recorded in the ground survey of the forest fire incident. The independent validation performance statistical parameters of each fire source classification model in Table 4 were obtained, as shown in the following table:
[0089]
[0090] The results showed that the random forest model achieved the highest AUC in the training set, but the logistic regression model achieved the highest classification accuracy in the test set. Furthermore, the classification accuracy of the fire point group data was relatively high compared to the central point group, and the accuracy of human-caused fires was relatively higher than that of natural fires. The conclusion was that the logistic regression model, based on the four environmental variables obtained from the fire point: average dew point, average visibility, average monthly temperature, and slope length, provided the highest accuracy for identifying fire sources.
[0091] like Figure 5 The figure shows the receiver operating characteristic (ROC) curves for the training and validation sets of the fire origin dataset in a logistic regression model. This combination demonstrated the best fire source identification results based on the optimal variable screening results of different models. The figure shows that the area under the curve (AUC) for the training set is 0.858, indicating that the variable combination obtained through fire origin screening has good performance in the logistic regression model for identifying fire source types. The AUC for the validation set, while slightly lower than that of the training set, also performed well.
[0092] like Figure 6 The figure shows the recognition results of a logistic regression model modeled and validated using fire source data for fire source classification. The figure shows that among the 47 samples used in the training model, three natural fires and four human-caused fires were misclassified as "other" (other), achieving an accuracy of 85.11%. Among the 21 samples in the validation dataset, only one natural fire and one human-caused fire were misclassified as "other" (other), achieving an accuracy of 90.48%. Among all fire point, model, and variable combinations, this model achieved the best fire source identification results for forest fires with unknown sources.
[0093] It can be seen that, based on the existing fire footprint extraction and the study of the relationship between forest fire and environmental factors, the present invention has developed a technical framework for identifying fire source types based on forest fire event extraction combined with the environmental characteristics of the ignition point. In addition to the extracted fire footprint morphology being an important supplement to the construction of a regional forest fire database, it further explores the forest fire event information in high-temporal-resolution remote sensing combustion products to extract the ignition location and ignition date of each forest fire event. According to the spatiotemporal characteristics of the ignition point in the context of multi-source environmental data, the modeling parameters are adaptively selected, which makes the identification of the fire source type of forest fire events more accurate, improves the accuracy and effectiveness of fire risk assessment and forest fire management, and can more richly and accurately improve forest fire event information.
[0094] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A method for identifying forest fire source types taking into account environmental characteristics of the fire point, characterized in that: The following steps are involved: Step 1: Obtain the Burn Date data from the MODIS remote sensing product MCD64A1 and the land cover type data MCD12Q1, preprocess them, extract the forest fire footprint using the Jenks-DBSCAN model, and select the fire footprints that match the fire points in the ground survey data to obtain the forest fire footprint with fire source information; Step 2: Combine the MCD64A1 data from step 1 with the extracted forest fire footprints, and use spatial partitioning statistical tools to mine information on each footprint to obtain the fire point location and fire date, center point location and main burning date of the forest fire event; Step 3: Obtain various environmental data variables related to forest fires, including combustible materials, terrain, human activities, and weather, and perform preprocessing to obtain variable factors; Step 4: Using the different spatiotemporal attributes of the fire point and center point extracted in step 2, prepare a forest fire event sample dataset for model training and divide it into a modeling dataset and a validation dataset in a ratio of 7:3; Step 5: Using the modeling data of the forest fire incident sample dataset in Step 4, select the optimal variable combination for different classification models, including logistic regression, random forest, and support vector machine, and debug the corresponding model training parameters; then input the verification data into the trained model, and perform an accuracy evaluation on the fire source identification results of the model and the fire source information provided in the investigation report to obtain the optimal fire point-model-variable combination that can identify the fire source of the forest fire incident; The specific steps in step 1 are as follows: 1) Perform format conversion, reprojection, and image stitching on the original layered HDF files in the MODIS combustion product MCD64A1 and land cover type product MCD12Q1, and extract the data layer containing the required burning date information in the MCD64A1 product; 2) Mask the vegetation pixels in the land cover type product MCD12Q1 against the combustion pixels in MCD64A1 to exclude non-combustible pixels; 3) Fire point records obtained from ground surveys were screened. Record information included: latitude and longitude of the fire point, burned area, fire cause, time of fire discovery, and time of fire extinguishing. Records with a total burned area greater than 25 hectares were screened based on the 500-meter spatial resolution of MODIS data. Records with unknown or missing fire source information were excluded. Vector fire point files were generated in ArcGIS software based on the longitude and latitude provided by the remaining records. The projection method was consistent with that of the preprocessed MODIS data. Fire source types were then classified into natural and human-caused fires, and assigned as categorical variables. 4) Input the burned pixels into the Jenks-DBSCAN model for clustering to obtain multiple clusters, each of which is tentatively defined as the footprint of a forest fire event; calculate the burned area based on the number of burned pixels and spatial resolution, and calculate the coefficient of determination To reflect the matching degree between the fire area investigated on the ground and the fire footprint area extracted by the Jenks-DBSCAN model: ; in, Represents the true value of the fire area surveyed on the ground; The representative model extracts the statistical value of the fire footprint area; The mean represents the true value; represents the number of samples, i.e. the number of forest fires; 5) Use ground survey data to match the forest fire footprints of the clustering results and screen out forest fire event footprints with fire source information; establish a circular buffer zone with a radius of 1 km with each surveyed fire point as the center, and then spatially overlay it with the forest fire footprints obtained by clustering for comparison; based on spatial association, perform a time verification to determine the footprint corresponding to the forest fire event, and verify whether the monthly fire information contained in the verification fire point matches the Julian day DOY value of the burned pixels covered by the fire footprint converted into the Gregorian calendar month; after performing the above spatial association and time verification on each ground survey fire point and the fire footprint extracted by clustering, the forest fire footprints that match one by one are used as training samples, and the dependent variable is the fire source type recorded by the ground fire point.
2. The method for identifying forest fire source types taking into account environmental characteristics of the fire starting point according to claim 1, characterized in that: The comparison method includes: if the buffer zone overlaps with a cluster of burned areas, it is determined to be the same fire event; if multiple forest fire footprints overlap with the buffer zone, the cluster with the largest overlapping pixels is defined as the overlapping or matching fire event.
3. The method for identifying forest fire source types taking into account environmental characteristics of the fire starting point according to claim 2, characterized in that: The extraction of the ignition point and the center point in step 2 is carried out in the following specific steps: The forest fire footprint samples obtained in step 1 are clustered and the raster clusters are converted into vector irregular polygons. The zoning statistics tool of ArcGIS software is used to count the DOY values of all MCD64A1 burning pixels spatially covered by the vector polygons to obtain the minimum, maximum, median and mode DOY information of the forest fire footprint.
4. The method for identifying forest fire source types taking into account environmental characteristics of the fire starting point according to claim 3, characterized in that: The recorded minimum DOY value is used as the ignition date of the forest fire event, and the geometric center of the pixel with the minimum value is the ignition point; the recorded mode of DOY is used as the main burning date of the forest fire event, and the geometric center of all pixels covered by the positioning polygon is used as the center point of the forest fire event.
5. The method for identifying forest fire source types taking into account environmental characteristics of the fire starting point according to claim 4, characterized in that: The specific steps in step three are as follows: 1) The fuel variable was extracted using the 1:1,000,000 vegetation map of the People's Republic of China. Based on this map, an environmental variable representing the fuel type of different forest tree species related to forest fires was extracted. The different forest plants were classified according to the local fuel type as non-flammable, relatively flammable, flammable, and extremely flammable. 2) Extraction of terrain variables Based on digital elevation model data, terrain variables are extracted, including elevation, slope, aspect, slope length and relief; 3) Human activity variables include two variables: spatially, local road network and residential vector layer data are collected and transformed and projected. Based on the location of the fire point or center point of the forest fire incident in step 2, the minimum distance from the fire point to the nearby road network or residential area is calculated as a variable; temporally, for festivals where burning objects pose a forest fire hazard, the fire start date and the main burning date are calculated as the closest day of the corresponding festival each year as another variable; 4) Meteorological elements include daily and monthly meteorological variables. Daily meteorological variables are obtained by cleaning and completing the data from daily meteorological stations. For missing data on a particular day, the data is completed by averaging the values of the most recent records before and after the missing date. Co-kriging interpolation is then performed using elevation as a covariate to obtain the meteorological element raster data for the entire region. Monthly meteorological variables include monthly average temperature and precipitation. 5) All the above environmental variables are spatially resampled and projected to have the same spatial resolution and projection as MODIS products. Variable factors are mainly divided into numerical and categorical types. Categorical variables need to be quantified during variable selection and model building.
6. The method for identifying forest fire source types taking into account environmental characteristics of the fire starting point according to claim 5, characterized in that: in, The product of step 1) includes the distribution information of 11 vegetation type groups, 55 vegetation types, 960 vegetation communities and subcommunities, and more than 2,000 dominant species and major crops in China.
7. The method for identifying forest fire source types taking into account environmental characteristics of the fire starting point according to claim 5, characterized in that: in, The meteorological element raster data in step 4) include average dew point, maximum temperature, minimum temperature, maximum sustained wind speed, precipitation, average sea level pressure, average temperature, average visibility and average wind speed.
8. The method for identifying forest fire source types taking into account environmental characteristics of the fire point according to claim 5, characterized in that: The specific steps in step five are: 1) Input the training set of forest fire incident sample data set, use the backward feature method under ten-fold cross-validation to select the most suitable number of variables for classification modeling, and then select the optimal variable combination for three different classification models based on this number; 2) Using the variance inflation factor combined with the chi-square test as a variable selection method for the logistic regression model; using the combined ranking of the mean decreasing precision and mean decreasing Gini index of each variable in the random forest importance analysis as a variable selection method for the random forest classification model; The optimal variable combination obtained by the backward feature method under ten-fold cross validation is used as the variable selection method for the input support vector machine; 3) After obtaining the optimal variable combination, further debug the optimal training parameters based on different model classification principles; 4) The independent variables in the validation dataset were put into the classification model to obtain the fire source identification results. The area under the curve (AUC) calculated by the receiver operating characteristic (ROC) curve was used as the performance evaluation indicator of the model. The closer the AUC value is to 1, the higher the accuracy of the model. ; Among them, M is the number of positive samples, N is the number of negative samples, M+N=n represents the total number of samples, the denominator is the total number of combinations of positive and negative samples, the numerator is the number of groups in which the scores of positive samples are greater than the scores of negative samples, and Score represents the probability that each test sample belongs to the positive sample; At the same time, the model recognition results were compared with the fire source types in the investigation records, and the overall accuracy, user accuracy and Kappa coefficient were used as evaluation indicators for the binary classification effect of different fire source discrimination models in the test data set; ; ; ; ; ; In the set of samples to be modeled containing fire source information, P represents the number of real human-caused fires, N represents the number of real natural fires, T represents accurate predictions, and F represents incorrect predictions. FP, FN, TP, and TN represent the number of samples with different judgment results for the two fire source categories in the test dataset. FP indicates that a human-caused fire recorded in the investigation record is classified as a natural fire by the model, TP indicates that a human-caused fire recorded in the investigation record is classified as a human-caused fire by the model, and so on. 5) Comparing the fire point group and the center point group of the forest fire event sample dataset, and the performance of different classification models in fire source identification in the validation set, the optimal fire point-model-variable combination is obtained. This is used as the identification method for forest fire events in the remaining area with unknown fire source information, thereby completing the forest fire database information.