Remote sensing image fire point detection method based on historical image temporal and spatial variation
Through the Vision Transformer-based fire point monitoring network model and the improved SMOTE algorithm, the sample imbalance problem in fire point detection is solved, the accuracy and real-time nature of fire point detection is improved, and the model deployment process is simplified.
Patent Information
- Application Number
- CN202411810330.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-07-18
AI Technical Summary
The existing fire point detection methods face the problem of imbalance between fire point and non-fire point samples, and the deep learning model has limitations in dealing with the spatial, spectral and spatial-temporal characteristics of fire points, resulting in low detection accuracy.
The fire point monitoring network model based on Vision Transformer is adopted, combined with the improved SMOTE algorithm for data balancing processing, background features, spectral features and spatiotemporal features are extracted, and fire point classification is performed through the full connection layer, and the trained model is used for real-time fire point monitoring.
It improves the accuracy and real-time nature of fire point detection, solves the problem of sample imbalance, enhances the efficiency and accuracy of fire point detection, and simplifies the deployment process of the model.
Smart Images

Figure CN120339665A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and in particular to a method for detecting fire points in remote sensing images based on spatio-temporal changes in historical images. Background Art
[0002] In recent years, global warming has led to frequent extreme weather events, such as droughts, high temperatures, and storms. These adverse conditions have caused a significant increase in the frequency and scale of forest fires. Forest fires not only reduce the global forest area but also damage the integrity of the ecosystem, leading to the loss of wildlife habitats and a subsequent decline in biodiversity. Widespread forest fires also release large amounts of carbon dioxide and other greenhouse gases, accelerating the process of global warming and forming a harmful feedback loop. In addition, the smoke and pollutants generated by fires pose a significant threat to human health. Therefore, it is crucial to identify and prevent the spread of fires in a timely and accurate manner.
[0003] Currently, fire point detection mainly relies on satellite remote sensing technology. Satellite remote sensing technology has advantages such as wide coverage, high resolution, real-time monitoring, cost-effectiveness, and long observation cycles, and has been widely used in fire point detection research in recent years. However, existing fire point detection methods still face many challenges.
[0004] First, the problem of imbalance in the number of fire point and non-fire point samples in fire point detection seriously affects the detection accuracy. Fire points usually appear as a small number of hot spots in the image, while most of the image area is background. This sample imbalance problem causes deep learning models to be prone to overfitting, thus reducing the detection accuracy.
[0005] Secondly, the introduction of deep learning technology provides a new method for feature extraction, and significantly improves the accuracy of fire point detection by automatically learning features in images. However, existing deep learning models still have certain limitations when dealing with the spatial, spectral, and spatio-temporal features of fire points. For example, traditional convolutional neural networks (CNNs) perform poorly in capturing the spatio-temporal change features of fire points, which limits their application in dynamic fire point detection tasks. Therefore, it is necessary to design a new deep learning method for fire point detection of Himawari-8 satellite data. Summary of the Invention
[0006] The present invention provides a method for real-time monitoring of satellite remote sensing based on spatio-temporal feature fusion, which solves technical problems such as the imbalance in the number of samples used for training fire point monitoring by existing network models and the poor effect of dynamic fire point monitoring.
[0007] The present invention can be achieved through the following technical solutions:
[0008] A method for detecting fire points in remote sensing images based on spatio-temporal changes in historical images, which obtains historical remote sensing image data corresponding to the moment to be detected, establishes a data set, and first performs data balancing processing on the historical remote sensing images in the data set, and then divides the fire point image blocks and non-fire point image blocks to obtain a training data set;
[0009] Then, the historical remote sensing images in the training data set are divided into units with a set continuous time period. Using the fire point image blocks and non-fire point image blocks corresponding to each unit as inputs and the corresponding actual fire point situation as outputs, the fire point monitoring network model is trained and learned;
[0010] Among them, the fire point monitoring network model first extracts background features, spectral features, and spatio-temporal features from the fire point image blocks and non-fire point image blocks, then performs splicing and fusion and classification in sequence, and classifies the pixel points with high scores as fire points; the pixel points with low scores are classified as non-fire points;
[0011] Finally, the trained fire point monitoring network model is used to monitor the fire points in the remote sensing image to be detected at the detection moment of the date to be detected.
[0012] Furthermore, the monitoring network model includes a feature extraction part and a classification part. The feature extraction part includes a background feature extraction module, a spectral feature extraction module, and a spatio-temporal feature extraction module.
[0013] The background feature extraction module uses the Vision Transformer encoder structure to capture the spatial environment and context information around each pixel point; the spectral feature extraction module uses the multi-layer perceptron MLP structure to extract the band information corresponding to each pixel point; the spatio-temporal feature extraction module uses the Vision Transformer encoder structure to capture the spatio-temporal change information of each pixel point.
[0014] The classification part includes a fully connected layer, which is used to comprehensively learn the features extracted by the feature extraction part and output the classification scores corresponding to each pixel point. If the score exceeds the score threshold, it is determined as a fire point, otherwise it is determined as a non-fire point.
[0015] Furthermore, the background feature extraction module uses the historical remote sensing images of the current day or the fire point image blocks and non-fire point image blocks corresponding to the remote sensing image to be detected as inputs; the spectral feature extraction module uses the numerical values of each band in each historical remote sensing image or the remote sensing image to be detected within each unit to form a column vector as an input; the spatio-temporal feature extraction module uses the fire point image blocks and non-fire point image blocks corresponding to each historical remote sensing image or the remote sensing image to be detected within each unit as inputs.
[0016] Further, an improved SMOTE algorithm is used to balance each historical remote sensing image. First, new fire point data is generated using the following rules, and then the generated new fire points are copied into cloud-free and water-free forest / grassland pixels;
[0017] The rules are set as follows
[0018]
[0019] where rand(0, 0.01) is a random factor, f1 and f2 are fire point 1 and fire point 2 respectively, and forespots represents the total number of fire points in the same image.
[0020] Further, for each historical remote sensing image within the same unit, taking the fire points including the new fire points as the center, fire point image patches are selected, and at the same time, multiple non-fire point image patches are randomly selected. These image patches are of the same size.
[0021] The beneficial technical effects of the present invention are as follows:
[0022] (1) An improved SMOTE algorithm is used to solve the problem of unbalanced sample numbers of fire points and non-fire points in fire point detection. Through the improved SMOTE algorithm, new fire points are generated from the original fire points and copied into the background information similar to the original fire points. This not only increases the number of fire points but also enriches the diversity of fire points.
[0023] (2) A deep neural network model based on the Vision Transformer (ViT) encoder structure, namely the fire point monitoring network model Fireformer, is proposed. ViT performs well in the task of extracting context-related information of fire points, and the transformer performs excellently in long-term sequence feature extraction. Applying ViT to capture the change characteristics of the spatial information of fire points over time, ViT can not only learn the characteristics of fire points changing over time but also effectively learn the spatio-temporal change characteristics of fire points by only changing the input and output without increasing the model complexity.
[0024] (3) An end-to-end fire point detection strategy is proposed, which requires no preprocessing and postprocessing. This end-to-end method simplifies the model deployment and usage process, improves the efficiency of the entire fire point detection, and ensures the real-time performance of fire point detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a schematic structural diagram of the monitoring network model of the present invention;
[0026] Figure 2 is a schematic diagram of the ViT structure of the present invention;
[0027] Figure 3Schematic diagrams of the input data for the SBT-FireNet temporal feature extraction module and the input data for the spatio-temporal feature extraction module of the present invention respectively. Among them, (a) is the input of the SBT-FireNet temporal feature extraction module, (b) is the input of the spatio-temporal change feature part of Fireformer, and t i represents the i-th day, and b i represents the i-th band;
[0028] Figure 4 Schematic diagrams of the comparison between the fire point maps and the ground truth values of Hunan, Jiangxi, and Guangxi generated by the method of the present invention and other methods respectively. Detailed implementation manners
[0029] The following will describe in detail the specific implementation manners of the present invention in conjunction with the accompanying drawings and preferred embodiments.
[0030] As Figure 1 shown, the present invention provides a method for detecting fire points in remote sensing images based on spatio-temporal changes in historical images. Obtain the historical remote sensing image data corresponding to the detection moment, establish a data set, and first perform data balancing processing on the historical remote sensing images in the data set, and then perform the division of fire point image patches and non-fire point image patches to obtain a training data set; then divide the historical remote sensing images in the training data set in units of a set continuous time period, use the corresponding fire point image patches and non-fire point image patches as inputs, and use the corresponding actual fire point situations as outputs to train and learn the fire point monitoring network model; finally, use the trained fire point monitoring network model to monitor the fire points of the remote sensing image to be detected at the detection moment of the date to be detected, where the fire point monitoring network model first extracts background features, spectral features, and spatio-temporal features from the fire point image patches and non-fire point image patches, and then performs splicing, fusion, and classification in sequence, and classifies the pixel points with high scores as fire points; the pixels with low scores are classified as non-fire points. In this way, using the fire point monitoring network model for fire point detection can automatically extract the relevant features of fire points, improve the accuracy of fire point detection, and at the same time comprehensively consider the background, spectral, and spatio-temporal features of fire points to participate in the fire point detection calculation, avoiding data loss, and providing a good data basis for the accuracy of subsequent detections, specifically as follows:
[0031] Step 1: Establish a data set
[0032] Considering the main features of fire points, the input data should contain as much information as possible about the background, spectrum, and spatio-temporal changes of pixels, that is, image elements. The input data mainly includes the following three parts of pixel points: context patch, band data, and time series data of spatio-temporal changes in the continuous time series.
[0033] The context patch of a pixel point refers to a window of size M×N centered on that pixel point. The context patch contains the environmental information of the pixel, and the context block contains the spatial information of the central pixel point. The band data of a pixel point refers to the numerical values of each band on that pixel. The spatio-temporal change data in a continuous time series refers to the sequence data of the spatial information of a pixel within a certain continuous time period. The spatial information refers to the context patch, and the continuous time refers to the same moment for D consecutive days.
[0034] Use data augmentation methods to expand the data:
[0035] Use the improved SMOTE algorithm to generate new fire points from existing fire points, and then copy the generated new fire points to the pixels where the land cover, i.e., the underlying surface, is forest or grassland. This method can increase the richness of fire points while ensuring the authenticity of the new fire points.
[0036] The improved SMOTE algorithm improves the formula based on the original SMOTE and copies the generated fire points to a similar background, i.e., forest or grassland. The specific steps are as follows:
[0037] (1) First, generate new fire point data from existing fire point data (within an image) based on the following rules:
[0038]
[0039] Among them, rand(0,0.01) is a random factor, f1 and f2 are fire point 1 and fire point 2 respectively, and forespots represents the total number of fire points within the same image.
[0040] (2) Copy the generated fire points to cloud-free and water-free forest / grassland pixels.
[0041] Therefore, according to the above method, as long as we obtain the historical remote sensing image data of geostationary satellites corresponding to the detection time, divide these historical remote sensing images in units of a set time period, such as seven days, and then take each fire point, including the new fire points, in each historical remote sensing image within each unit as the center, select fire point image blocks, such as image blocks of size 21*21, and also randomly select multiple non-fire point image blocks, also of size 21*21, to construct a training dataset. Then, use the fire point image blocks and non-fire point image blocks within each unit as inputs, and the actual fire point situation corresponding to each historical remote sensing image as the output, and we can train and learn the monitoring network model. During training, 40% of the training set is not involved in training and is used as the validation set to observe the training situation of the network in real time.
[0042] Step 2: Construct the fire point monitoring network model Fireformer
[0043] The model proposed in this invention is called Fireformer, and its structure is as Figure 1 shown. This model mainly includes the following three parts: a spatio-temporal feature extraction module, a spectral feature extraction module, and a background feature module. Among them, the spatio-temporal feature extraction module is used to extract the changing information of the spatio-temporal information of the fire point over time, the spectral feature extraction module is used to extract sensitive band information, and the background feature extraction module is used to analyze the difference information between the background and the fire point of the fire point. The three parts of the extracted features are spliced and combined, and then the features extracted in the feature extraction process are comprehensively considered. The pixel points with high fire point classification scores are classified as fire points; the pixels with low fire point classification scores are classified as non-fire points. Among them, the spatio-temporal feature extraction module of the FireFormer model is composed of ViT (Vision Transformer), the spectral feature extraction module is composed of MLP, and the background feature extraction module is also composed of ViT (Vision Transformer). The specific structure of the FireFormer model is as Figure 1 shown.
[0044] 1. Background feature extraction module
[0045] The background feature extraction module is used to capture and analyze the spatial environment and context information around the fire point. This module is composed of Vision Transformer (ViT). Vision Transformer (ViT) was proposed by Google Research in 2020 and is the first attempt to successfully apply the Transformer model to computer vision tasks. Before this, the transformer was designed as a model for processing sequence data, and deep learning in the field of image processing mainly relied on convolutional neural networks (CNNs). Although CNNs have achieved great success in tasks such as image recognition and object detection, they tend to focus on local features when processing images and may sometimes ignore the global context information in the images. The Transformer model, especially its self-attention mechanism, provides an effective method for capturing long-range dependencies, which has shown great potential in the field of natural language processing (NLP). The internal structure of ViT follows the core design of the original Transformer but makes specific modifications suitable for image processing.
[0046] ViT first divides the input image into patches of a fixed size and transforms each image patch into a vector of a fixed dimension through a linear transformation. To enable the model to understand the relative or absolute positions of the image patches, positional encoding is added to the embedding vectors of each image patch. The encoder of the Transformer consists of multiple Transformer layers, and each layer is composed of a self-attention mechanism and a feed-forward neural network (FFN). In the ordinary attention mechanism (Scaled Dot-Product Attention), the input vector is first transformed into three different vectors: the query vector d q the key vector d k and the value vector d v . Then, the vectors derived from different inputs are packed into three different matrices, namely Q, K, and V. We calculate the output matrix of the attention mechanism as:
[0047]
[0048] The attention mechanism used in ViT is multi-head attention. Multi-head attention is a mechanism that can be used to improve the performance of ordinary self-attention layers. Single-head attention limits our ability to focus on one or more specific positions without affecting the ability to simultaneously pay attention to other equally important positions. Multi-head attention is achieved by endowing the attention layer with different representation subspaces. Specifically, different heads use different query matrices, key matrices, and value matrices. These matrices, due to random initialization, can project the input vector into different representation subspaces after training. The multi-head attention process is as follows:
[0049] MultiHead(Q, K, V) = Concat(head1,…,head h )W O
[0050] where head i = Attention((QW i Q , KW i K , VW i V )
[0051] where h is the number of heads.
[0052] At the end of the model, the fully connected layer is used to convert the output of the Transformer encoder into the final prediction result. Compared with traditional Convolutional Neural Networks (CNNs), the Vision Transformer (ViT) can directly capture the global information of the entire image through the self-attention mechanism, while traditional CNNs obtain global information by gradually expanding the local receptive field when processing images. This enables ViT to better understand the relationships between different regions in the image and thus more effectively utilize the global background information.
[0053] 2. Spectral Feature Extraction Module
[0054] The spectral feature extraction module is composed of an MLP. We input the band information of the fire point itself (a column vector with a shape of (1, 6)), where the size of the column vector is 6 because we have selected six bands. The MLP includes an input layer, two hidden layers, and an output layer. Each layer consists of a series of neurons. Among them, the number of neuron nodes in the two hidden layers is 1024 and 512 respectively, and the number of neuron nodes in the output layer is 128. The neurons are connected by learnable weights. Assuming that the input value z of each neuron is the weighted sum of the output a of its previous layer and the weight w, plus the bias b, the forward propagation process of the MLP can be expressed by the following formula:
[0055] z l = w l a l-1 + b l
[0056] where l represents the current layer, and a 0 is the input layer.
[0057] Backpropagation is used to calculate the gradients of the loss function with respect to each weight and bias. Let the loss function be L, then the gradient of the weight and the gradient of the bias can be calculated by the chain rule:
[0058]
[0059] where, is the partial derivative of the loss function with respect to the output of the current layer, which needs to be recursively calculated according to the reciprocal of the activation function and the gradient of the next layer.
[0060] By applying a non-linear activation function to the neurons, the MLP can learn non-linear relationships. It makes predictions or classifications by learning the complex mapping relationship between the input data and the output results.
[0061] 3. Spatiotemporal Feature Extraction Module
[0062] In the existing monitoring model SBT-FireNet, the temporal feature extraction module is used to extract the temporal features of fire points. The input is as follows Figure 3 (a), which represents the band data of the continuous time series of the fire point itself. The model used is LSTM. LSTM (Long Short-Term Memory) is a special recurrent neural network (RNN), mainly used to solve the problems of vanishing gradients and exploding gradients in standard RNNs when dealing with long sequences. LSTM effectively controls the flow of information by designing a memory cell and gate mechanisms, including forget gate, input gate, and output gate. The forget gate determines which information in the memory cell needs to be forgotten, the input gate determines which new information is written into the memory cell, and the output gate determines which information is output from the memory cell. Through this mechanism, LSTM can maintain and utilize information over long time intervals, thus solving the deficiencies of standard RNNs in processing long sequences. LSTM performs excellently in various sequence tasks, including natural language processing, time series prediction, etc.
[0063] If we use the temporal feature extraction module of the monitoring model SBT-FireNet, we can only capture the time-varying features of the fire point itself. As our research on the background information of the fire point deepens, we find that the background information of the fire point also changes with the development of the fire point. If we only rely on LSTM to process fire point data, we cannot comprehensively capture these important background change information. In order to capture not only the temporal changes of the fire point but also the spatio-temporal changes of the fire point, we find that LSTM is no longer applicable. Therefore, we change the input to the data of the fire point and its background information of the continuous time series, as shown in Figure 3 (b), and replace LSTM with Vision Transformer (ViT).
[0064] We utilize the characteristics of ViT to capture long-range dependencies and global information to extract the spatio-temporal change features of fire points. ViT has strong global learning ability and can adaptively extract and learn the features and their associations contained in context patches. The transformer can also effectively capture long-range dependencies. As a Transformer-based model, ViT can process time and space information simultaneously through the self-attention mechanism, thus better understanding the spatio-temporal changes of fire points and their background information. Different from LSTM, ViT can capture long-range dependency relationships and maintain a high degree of parallelism during processing, improving the efficiency and performance of the model. Compared with LSTM, ViT can not only learn the change features of the fire point itself but also learn the temporal changes of the fire point background information, while RNN and LSTM can only learn the temporal features of the fire point itself.
[0065] Step 3: Real-time monitoring
[0066] Obtain the remotely sensed images to be detected for seven consecutive days counting backward from the detection moment including the date to be detected. Obtain the corresponding fire point image patches and non-fire point image patches according to the method described above, and input them into the trained fire point monitoring network model to perform real-time fire point monitoring.
[0067] To verify the feasibility of the real-time monitoring method of the present invention, remotely sensed image data collected by the Himawari-8 satellite is obtained. The study area includes 8 provinces in southern China, namely Fujian, Guangdong, Guangxi, Hunan, Jiangxi, Hainan, Yunnan, and Guizhou. Its latitude and longitude range is: 18°10′N - 30°08′N, 97°31′E - 120°30E. The climate in the study area is subtropical or tropical monsoon climate, and the long-term annual temperature difference is 12 - 29°C.
[0068] As the successor of the MTSAT series of geostationary meteorological satellites, the Himawari-8 satellite was launched on July 7, 2015, and is equipped with Advanced Himawari Imagers (AHIs). The AHIs have a total of 16 observation bands: 10 infrared bands, 3 near-infrared bands, and 3 visible light bands. The spatial resolutions of the infrared, near-infrared, and visible light bands are 0.5 km - 1 km, 1 km - 2 km, and 1 km - 2 km respectively, and the full-disk observation interval is 10 minutes.
[0069] The bands used in this method are listed in Table 1.
[0070] Table 1
[0071]
[0072] All the codes of this method are written in Python 3.6 version. In the part of building the deep learning model, the deep learning framework used is PyTorch, version 1.2. All experiments are carried out on an Intel Core I i9-10900K CPU@3.70GHz, 128GB RAM, and NVIDIA GeForce GTX 3080. During the training process of each deep learning model, the Adam optimizer is used as the parameter optimizer, the loss function is the cross-entropy loss function, the epoch is set to 500, the batch size is set to 100, and the learning rate is set to 10-6.
[0073] In addition, accuracy, precision, recall, F1-score (F1), miss detection rate (MD), and error detection rate (ED) are used as the key performance indicators of the model, and the calculation formulas are as follows:
[0074] Accuracy = (TP + TN) / (TP + TN + FP + FN)
[0075] Precision = TP / (TP + FP)
[0076] Recall = TP / (TP + FN)
[0077] F1 = (2 · Precision · Recall) / (Precision + Recall)
[0078] MD = 1 - R
[0079] ED = 1 - P
[0080] Where TP is true positive; TN is true negative; FP is false positive; FN is false negative.
[0081] Fireformer is compared with SBT-FireNet, STS-RNN, SMCNN, and the commonly used classification network EfficientNet. EfficientNet is a convolutional neural network (CNN) architecture proposed by Google in 2019, and it performs well in image classification tasks. The design goal of EfficientNet is to improve the efficiency and performance of the model by optimizing the depth, width, and resolution of the network. The core idea of EfficientNet is to use a simple and effective compound coefficient to uniformly scale the depth, width, and resolution of the network, so as to find the balance among the three, and achieve better performance with less computational cost while increasing the network depth and capacity.
[0082] The SMCNN is a single-module CNN proposed by Yoojin Kang et al. In this paper, the brightness temperature data of Himawari-8 is used as the input of the key variable, and the solar zenith angle, satellite zenith angle, relative humidity, etc. are used as the input of the auxiliary variables. This method first judges potential fire points according to the key variable, then the convolutional neural network is used for fire point detection, and finally the final result is obtained through post-processing of false alarm elimination. The convolutional neural network considers both the single-module CNN (SMCNN) and the dual-module CNN (DMCNN). Similarly, both the key variables related to fires and the auxiliary variables are considered for active fire detection. The DMCNN extracts features from the key variables using two-dimensional convolution and extracts features from the auxiliary variables using one-dimensional convolution and then combines the features. While the SMCNN extracts features from both the key variables and the auxiliary variables using two-dimensional convolution. This paper discusses the effects of the two different convolutional neural networks in fire point detection.
[0083] The Spatiotemporal Spectral Recurrent Neural Network (STS-RNN) is a framework for near-real-time and early wildfire detection using 10-minute data from the Himawari-8 satellite. After judging and preprocessing early or continuous wildfires, this method generates time-series spatial variance, time difference, and spectral difference through the time-series BT07 and BT14 information. The time-series spatial variance, time difference, and spectral difference curves are respectively imported into the STS-RNN prediction model. The STS-RNN can adaptively learn the time-series curves and predict the next value. The main body of its model is the RNN, and the RNN can effectively utilize the consistency and variability of time-series vectors. The RNN can remember the information of the previous input and use this information to affect the processing of the current input. Therefore, it is very effective in processing tasks such as time-series data, natural language processing (NLP), and speech recognition.
[0084] SBT-FireNet is a customized network designed according to the characteristics of fires by Zhonghua Hong et al. to detect fire points in Himawari-8 satellite images. To solve the problem of the imbalance in the number of fire point and non-fire point samples, this method proposes a method of replicating fire points to increase the number of fire points, thereby generating a dataset with spatiotemporal robustness. The fire feature extraction part of this deep learning network includes the following three modules: the spatial feature extraction module (SFE), the band feature extraction module (BFE), and the time-series feature extraction module (TFE), which are composed of ViT, MLP, and LSTM respectively.
[0085] As can be seen from Table 2, the Fireformer model (Precision: 0.805; Recall: 0.806; F1-score: 0.804) is higher than that of SBT-FireNet (Precision: 0.781; Recall: 0.747; F1-score: 0.759), with the best comprehensive performance. Specifically, Fireformer comprehensively utilizes the spatial, band, and spatio-temporal change features of fire points to obtain a precision 2.4% higher and a Recall value 5.9% higher than SBT-FireNet, indicating that the spatio-temporal change features of fire points play an important role in fire point detection. STS-RNN (Precision: 0.651; Recall: 0.725) also performs well, indicating that in fire point detection, the temporal change features of fire points are an important feature and judgment index for determining whether it is a fire point. The precision and Recall values are lower than those of SBT-FireNet, indicating that relying solely on the temporal change features of fire points is not enough, and the spatial and band features of fire points are also needed. SMCNN (Precision: 0.22; Recall: 0.151) performs poorly, which exactly shows that it is far from enough to only learn the spatial features of fire points, and it is difficult for a simple convolutional neural network to capture all the features of fire points. It is necessary to comprehensively utilize the combination of various networks to comprehensively learn the spatial, temporal, and spatio-temporal change features of fire points. Although EfficientNet is a network used more in image classification, it is not ultimately designed for the fire point detection task based on geostationary satellites, so it performs poorly in fire point detection.
[0086] Comparison of Experimental Results in Table 2
[0087]
[0088] Appendix Figure 4 Fire point maps of different methods are given, where "×" represents fire points. As can be seen from the figure, Fireformer and SBT-FireNet perform equivalently, and STS-RNN also performs well. However, when interfered by thin clouds, Fireformer performs relatively better. For example, from the visualization results in Jiangxi, it can be seen that SBT-FireNet is interfered by clouds and misidentifies non-fire points as fire points, and STS-RNN is more severely interfered by clouds, resulting in more false detections.
[0089] Although the specific implementation manners of the present invention are described above, those skilled in the art should understand that these are only examples. Without departing from the principles and essence of the present invention, various changes or modifications can be made to these implementation manners. Therefore, the protection scope of the present invention is defined by the appended claims.
Claims
1. A method for detecting fire points in remote sensing images based on spatio-temporal changes in historical images, characterized in that: Obtain the historical remote sensing image data corresponding to the detection time, establish a dataset, and first perform data balancing processing on the historical remote sensing images in the dataset, and then perform the division of fire point image patches and non-fire point image patches to obtain a training dataset; Then divide the historical remote sensing images in the training dataset by a set continuous time period, use the fire point image patches and non-fire point image patches corresponding to each unit as inputs, and the corresponding actual fire point situation as outputs to train and learn the fire point monitoring network model; Among them, the fire point monitoring network model first extracts background features, spectral features and spatio-temporal features from fire point image patches and non-fire point image patches, then performs splicing and fusion and classification in sequence, and classifies the pixel points with high scores as fire points; the pixels with low scores are classified as non-fire points; Finally, use the trained fire point monitoring network model to monitor the fire points of the to-be-detected remote sensing image at the detection time of the date to be detected.
2. The remote sensing image hotspot detection method based on the spatio-temporal variation of historical images according to claim 1, characterized in that: The monitoring network model includes a feature extraction part and a classification part. The feature extraction part includes a background feature extraction module, a spectral feature extraction module and a spatio-temporal feature extraction module. The background feature extraction module uses the VisionTransformer encoder structure to capture the spatial environment and context information around each pixel point; the spectral feature extraction module uses the multi-layer perceptron MLP structure to extract the band information corresponding to each pixel point; the spatio-temporal feature extraction module uses the Vision Transformer encoder structure to capture the spatio-temporal change information of each pixel point; The classification part includes a fully connected layer, which is used to comprehensively learn the features extracted by the feature extraction part and output the classification score corresponding to each pixel point. If it exceeds the score threshold, it is determined as a fire point, otherwise it is determined as a non-fire point.
3. The remote sensing image fire point detection method based on the spatio-temporal change of historical images according to claim 2, characterized in that: The background feature extraction module takes the historical remote sensing image of the current day or the fire point image patches and non-fire point image patches corresponding to the to-be-detected remote sensing image as inputs; the spectral feature extraction module takes the values of each band in each historical remote sensing image or the to-be-detected remote sensing image within each unit as a column vector as an input; the spatio-temporal feature extraction module takes the fire point image patches and non-fire point image patches corresponding to each historical remote sensing image or the to-be-detected remote sensing image within each unit as inputs.
4. The remote sensing image fire point detection method based on the spatio-temporal change of historical images according to claim 1, wherein: Use the improved SMOTE algorithm to balance each historical remote sensing image. First, generate new fire point data according to the following rules, and then copy the generated new fire points to cloud-free and water-free forest / grassland pixels; The rules are set as follows Among them, rand(0,0.01) is a random factor, f1 and f2 are fire point 1 and fire point 2 respectively, and forespots represents the total number of fire points in the same image.
5. The method for detecting hot spots in remote sensing images based on spatio-temporal changes of historical images according to claim 4, wherein: For each historical remote sensing image within the same unit, take the fire point (including the new fire point) as the center, select the fire point image patch, and at the same time randomly select multiple non-fire point image patches. These image patches have the same size.