A forest fire occurrence real-time prediction method and system

By constructing a forest fire prediction model based on convolution, graph neural networks, and attention mechanisms, the shortcomings of multi-source data fusion and spatiotemporal feature modeling are addressed, enabling precise capture and real-time early warning of forest fire risks, and improving the accuracy and real-time performance of predictions.

CN121189155BActive Publication Date: 2026-07-21INST OF FOREST ECOLOGY ENVIRONMENT & PROTECTION CHINESE ACAD OF FORESTRY +2
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INST OF FOREST ECOLOGY ENVIRONMENT & PROTECTION CHINESE ACAD OF FORESTRY
Filing Date
2025-09-15
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing forest fire prediction technologies are inadequate in terms of multi-source data fusion, spatiotemporal feature modeling, real-time performance, and interpretability, making it difficult to meet the needs of precise prevention and control.

Method used

A forest fire prediction model is constructed using convolutional neural networks, graph neural networks, and attention mechanisms. By collecting multi-source data in real time and performing spatiotemporal alignment processing, and by using graph structures and attention mechanisms for feature extraction and fusion, the model can accurately capture forest fire risks.

Benefits of technology

It improves the accuracy and real-time performance of forest fire prediction, can identify the spatial distribution characteristics and temporal evolution patterns of fire risks, provides a more comprehensive analytical perspective, and supports early warning information at the minute or hour level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121189155B_ABST
    Figure CN121189155B_ABST
Patent Text Reader

Abstract

The application discloses a forest fire occurrence real-time prediction method and system, and the method comprises the following steps: collecting multi-source data monitored in a forest area in real time; performing spatio-temporal alignment processing on the multi-source data by using sampling and terrain correction; constructing a spatio-temporal collaborative perception network model based on convolution, a graph neural network and an attention mechanism, training a sample data set by using the multi-source data, and obtaining a forest fire prediction model; and inputting the multi-source data into the forest fire prediction model to obtain a forest fire prediction result. The technical scheme provided by the application can process multi-source data in a forest area, can extract multi-level spatial features through convolution, can dynamically weight key spatio-temporal dimensions through the attention mechanism, can model the spatial relationship of complex terrain and vegetation distribution through the graph neural network, can not only capture the spatial distribution characteristics of fire risk, but also identify the law of the evolution of the fire risk with time, can effectively overcome the shortcomings of traditional methods in spatio-temporal modeling, and greatly improves the prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of forest fire prediction, and in particular to a method and system for real-time prediction of forest fire occurrence. Background Technology

[0002] In recent years, global climate change has led to frequent forest fires, necessitating forest monitoring and timely fire prediction. Currently, mainstream forest fire prediction technologies mainly include traditional machine learning methods based on statistics, remote sensing monitoring technology, and deep learning time-series prediction models. While these methods can predict fire risk to some extent, they still have many limitations in practical applications and are insufficient to meet the needs of precise fire prevention and control.

[0003] Traditional machine learning methods, such as random forests and support vector machines, primarily rely on manually extracted static features like weather and topography for modeling. While these methods are computationally efficient, they struggle to capture the spatiotemporal dynamics of forest fires. Because forest fires are influenced by a combination of factors, including weather conditions, vegetation cover, and human activities, static features cannot reflect real-time changes in fire risk. Furthermore, these methods are highly dependent on data quality and feature engineering; when faced with complex and ever-changing natural environments, prediction accuracy is often difficult to guarantee.

[0004] Remote sensing technology can monitor fires over large areas using satellite data, but it has significant shortcomings in terms of real-time performance and accuracy. Optical remote sensing data is easily affected by cloud cover, and while microwave remote sensing data has all-weather observation capabilities, its spatial resolution is low. Existing remote sensing fire detection algorithms are mostly based on threshold methods, which are easily affected by interference from clouds, solar flares, etc., leading to a high false alarm rate. More importantly, remote sensing monitoring is essentially a post-event monitoring method, making it difficult to provide proactive early warning information for fire prevention.

[0005] In recent years, deep learning time-series prediction models, such as LSTM and GRU recurrent neural networks, have shown potential in the field of fire prediction. These models can process time-series data and capture the temporal dependence of fire occurrence. However, simple time-series models struggle to effectively integrate spatial information and fail to reflect the spatial heterogeneity of fire risk. Furthermore, existing deep learning models often use a single data source for training, failing to fully utilize the complementary advantages of multi-source heterogeneous data. For example, meteorological data, remote sensing data, topographic data, and socioeconomic data have different spatiotemporal resolutions and data formats. This diversity poses significant challenges to data alignment and feature fusion, making the effective fusion of these heterogeneous data sources a key technical challenge.

[0006] Existing fire risk assessment systems suffer from insufficient model generalization ability. Due to significant differences in vegetation types, climate characteristics, and human activity intensity across regions, models trained in one area are often difficult to directly apply to others. Furthermore, with global warming, fire occurrence patterns are changing, and traditional statistical models based on historical data struggle to adapt to these dynamic changes. Particularly against the backdrop of frequent extreme weather events, fire risk exhibits new spatiotemporal distribution characteristics, placing higher demands on predictive models.

[0007] In terms of real-time performance, existing forecasting systems mostly employ offline training and periodic updates, making it difficult to provide minute-level or hour-level early warning information for fire prevention decisions. Forest fire prevention and control demands extremely high timeliness; delays in early warning information can lead to missed opportunities for optimal prevention and control. Furthermore, existing systems generally lack interpretable analysis of forecast results, hindering the provision of scientific basis for disaster prevention and mitigation decisions. Decision-makers not only need to know that fires may occur, but also need to understand the key factors contributing to high risk in order to take targeted measures.

[0008] In summary, existing forest fire prediction technologies have significant shortcomings in areas such as multi-source data fusion, spatiotemporal feature modeling, real-time performance, and interpretability. Therefore, there is an urgent need to develop a novel forest fire prediction method that can integrate multi-source data, accurately capture spatiotemporal features, and possess real-time early warning capabilities. Summary of the Invention

[0009] In view of the above-mentioned shortcomings of the current technology, the present invention provides a real-time forest fire prediction method that can process multi-source data in forest areas and construct a forest fire prediction model based on convolution, graph neural networks and attention mechanisms. This hybrid architecture can not only capture the spatial distribution characteristics of fire risk, but also identify its evolution over time. It can effectively overcome the shortcomings of traditional methods in spatiotemporal modeling and greatly improve the accuracy of prediction.

[0010] To achieve the above objectives, the embodiments of the present invention adopt the following technical solutions:

[0011] A method for real-time prediction of forest fire occurrence includes the following steps:

[0012] Real-time collection of multi-source data from forest area monitoring;

[0013] Spatiotemporal alignment of multi-source data is performed using sampling and terrain correction.

[0014] A spatiotemporal collaborative perception network model was constructed based on convolutional, graph neural networks and attention mechanisms. A sample dataset was built using multi-source data for training to obtain a forest fire prediction model.

[0015] Multi-source data is input into the wildfire prediction model to obtain wildfire prediction results.

[0016] According to one aspect of the present invention, the multi-source data includes: lightning data, atmospheric electric field data, and meteorological data.

[0017] According to one aspect of the present invention, the real-time forest fire prediction method further includes:

[0018] Preprocess the lightning data.

[0019] According to one aspect of the present invention, the preprocessing of lightning data includes:

[0020] Spherical convolutional layers are used to process lightning data, which includes the three-dimensional spatial distribution characteristics of longitude, latitude, and cloud lightning height.

[0021] By utilizing a geographic weighted attention mechanism, the weights of lightning characteristics are dynamically adjusted based on the distribution density of combustibles in forest areas;

[0022] Design a risk decay time window function to adjust the risk assessment weights of lightning characteristics.

[0023] According to one aspect of the present invention, the spatiotemporal alignment processing of multi-source data using sampling and terrain correction includes:

[0024] An interpolation algorithm is used to upsample low-frequency data, and a moving average is used to suppress noise in high-frequency data.

[0025] A gridded terrain correction field is established based on the digital elevation model, and data from different sources are uniformly mapped to a standard grid.

[0026] According to one aspect of the present invention, the spatiotemporal cooperative sensing network model constructed based on convolution, graph neural networks, and attention mechanisms includes:

[0027] A multimodal feature projection layer is constructed, which maps the input multi-source data to a unified feature space through a fully connected layer and adjusts it using a graph structure; the mapped features are then batch normalized to obtain the original features.

[0028] A spatiotemporal feature extraction layer is constructed, which uses convolutional branches, Transformer branches and graph neural network branches to extract features from the original features. The feature interaction between the convolutional branches, Transformer branches and graph neural network branches is realized through a cross-modal attention gating mechanism to obtain preliminary features.

[0029] Construct a feature enhancement and aggregation layer, use the SE attention mechanism to enhance the initial features, and perform global average pooling to obtain aggregated features;

[0030] A feature fusion and prediction layer is constructed, which uses a fully connected layer to perform feature fusion on aggregated features and uses a prediction head to predict wildfires.

[0031] According to one aspect of the present invention, the spatiotemporal cooperative sensing network model constructed based on convolution, graph neural networks, and attention mechanisms includes:

[0032] A multimodal feature projection layer is constructed, which maps the input multi-source data to a unified feature space through a fully connected layer, and positional encoding technology and graph structure are introduced for adjustment; the mapped features are then batch normalized to obtain the original features.

[0033] A spatiotemporal feature extraction layer is constructed, which uses convolutional branches, Transformer branches and graph neural network branches to extract features from the original features. The feature interaction between the convolutional branches, Transformer branches and graph neural network branches is realized through a cross-modal attention gating mechanism to obtain preliminary features.

[0034] A feature enhancement layer is constructed, which uses the SE attention mechanism and graph structure to perform dual-path feature enhancement on the initial features, and adaptive weighting is performed through gating units to obtain enhanced features;

[0035] A feature aggregation layer is constructed, which performs multi-scale pooling on the enhanced features in the time dimension and graph pooling in the spatial dimension. The aggregated features are obtained by matching and filtering through a spatiotemporal cross-attention mechanism.

[0036] A feature fusion and prediction layer is constructed, and a cross-modal attention mechanism is used to fuse aggregated features. Then, a prediction head is used to predict wildfires.

[0037] According to one aspect of the present invention, the multimodal feature projection layer further includes:

[0038] For input multi-source data, a dynamic modal weight allocation strategy for environmental perception is established.

[0039] According to one aspect of the present invention, the feature fusion of aggregated features using a cross-modal attention mechanism includes:

[0040] The local features extracted by convolution, the global temporal patterns encoded by Transformer, and the spatial relationships inferred by graph neural network are 3D aligned. A three-level fusion strategy is adopted: primary fusion is to concatenate the local features extracted by convolution, the global temporal patterns encoded by Transformer, and the spatial relationships inferred by graph neural network; intermediate fusion is to use a gating mechanism to select features; and advanced fusion is to adaptively weight the features.

[0041] According to one aspect of the present invention, the feature fusion and prediction layer further includes: constructing a multi-task prediction head for forest fire prediction and risk level assessment.

[0042] According to one aspect of the present invention, the multi-task prediction head introduces an uncertainty weighting algorithm to automatically balance the weights of the loss functions for different tasks.

[0043] According to one aspect of the invention, the multi-task prediction head introduces a triple loss function to supervise and align multiple tasks:

[0044] L_align = L_ce + λ1L_triplet + λ2L_consistency;

[0045] Where L_align is the triple loss, L_ce is the cross-entropy loss, L_triplet is the triplet loss, L_consistency is the inter-modal prediction consistency loss, and λ1 and λ2 are weight parameters.

[0046] According to one aspect of the present invention, the real-time forest fire prediction method further includes:

[0047] A geographic interpretability analysis module is set up to visualize the contribution of each monitoring station to the prediction results, providing spatial basis for decision-making.

[0048] According to one aspect of the present invention, the real-time forest fire prediction method further includes:

[0049] Configure the edge layer and the cloud;

[0050] The edge layer detects the electric field intensity of atmospheric electric field data and immediately issues an alarm when a sudden change in electric field intensity is detected.

[0051] The cloud uses forest fire prediction models to predict forest fire risks and issue alerts.

[0052] A real-time forest fire prediction system, based on the real-time forest fire prediction method described above, includes:

[0053] The data acquisition module is used to collect multi-source data for forest area monitoring in real time.

[0054] The alignment module is used to perform spatiotemporal alignment processing on multi-source data using sampling and terrain correction.

[0055] The model building module is used to build a spatiotemporal collaborative perception network model based on convolution, graph neural networks and attention mechanisms, and to use multi-source data to build a sample dataset for training to obtain a forest fire prediction model.

[0056] The prediction module is used to input multi-source data into the forest fire prediction model and obtain forest fire prediction results.

[0057] Advantages of implementing this invention:

[0058] This invention provides a real-time forest fire prediction method that can process multi-source data from forest areas. The forest fire prediction model is based on convolution, graph neural networks, and attention mechanisms. Convolution can automatically extract multi-level spatial features, while the attention mechanism can dynamically weight key spatiotemporal dimensions. Graph neural networks can enhance the modeling of spatial relationships between complex terrain and vegetation distribution. This hybrid architecture can not only capture the spatial distribution characteristics of fire risk, but also identify its evolution over time. It can effectively overcome the shortcomings of traditional methods in spatiotemporal modeling and greatly improve the accuracy of prediction. Attached Figure Description

[0059] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0060] Figure 1 This is a flowchart of a real-time forest fire prediction method according to the present invention. Detailed Implementation

[0061] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0062] Example 1

[0063] like Figure 1 As shown, a method for real-time prediction of forest fire occurrence includes the following steps:

[0064] S1: Real-time collection of multi-source data for forest area monitoring.

[0065] In practical applications, various detection devices can be deployed in forest areas, including lightning detectors, temperature and humidity sensors, wind speed sensors, video monitoring equipment, multispectral cameras, atmospheric electric field meters, and infrared detectors. Remote sensing satellites and meteorological satellites can also be used for dynamic monitoring of forest areas. Using these detection devices, monitoring equipment, and satellites, multi-source heterogeneous data such as lightning data, temperature and humidity data, wind speed data, vegetation data, atmospheric electric field data, and meteorological data can be collected from forest areas.

[0066] Preferably, the method further includes:

[0067] Preprocess the lightning data.

[0068] In practical applications, this method uses a full-waveform three-dimensional lightning locator to detect lightning data in forest areas. For example, when applying this method to forest fire monitoring in the Greater Khingan Mountains, several full-waveform three-dimensional lightning locators are deployed in the Greater Khingan Mountains to cover the entire area. They can receive lightning location data in real time, including the latitude and longitude, time, intensity, polarity, type (cloud lightning / ground lightning), cloud lightning height, and other information for each lightning strike.

[0069] This method addresses the unique characteristics of lightning data acquired by a full-waveform 3D lightning locator by preprocessing it and designing a dedicated feature extraction and fusion architecture, specifically including:

[0070] (1) Use spherical convolutional layers to process the three-dimensional spatial distribution characteristics of lightning data, including longitude, latitude and cloud lightning height.

[0071] Traditional methods typically simplify lightning data into two-dimensional planar coordinates for processing, while this method innovatively introduces spherical convolutional layers, which can better extract the features of lightning data.

[0072] (2) Using the geographically weighted attention mechanism, the weight of lightning characteristics is dynamically adjusted according to the distribution density of combustibles in the forest area.

[0073] The impact of lightning on forest fires exhibits significant geographical variability. By utilizing a geographically weighted attention mechanism, the weights of lightning characteristics can be dynamically adjusted based on the distribution density of combustible materials in the forest area (such as coniferous forest cover). For example, in areas with high larch density, lightning of the same intensity will receive a higher risk score.

[0074] (3) Design a risk decay time window function and adjust the risk assessment weight of lightning characteristics.

[0075] Lightning data has a certain time sensitivity, so a risk attenuation coefficient is defined based on the time difference: Risk attenuation coefficient = 1 / (1 + α·Δt); where α is a parameter calibrated according to the aridity of the forest area (for example, α=0.05 for the Greater Khingan Mountains), and Δt is the difference between the current time and the time of the lightning strike.

[0076] Adjusting the risk assessment weights of lightning features using a risk attenuation coefficient is significantly superior to the conventional sliding time window method. Experiments show that this method can improve the accuracy of identifying lightning-related fire hazards by 10%.

[0077] S2: Spatiotemporal alignment of multi-source data is performed using sampling and terrain correction.

[0078] In practical applications, the various heterogeneous data collected in forest areas exhibit significant spatiotemporal gaps, thus requiring spatiotemporal alignment processing. For example, lightning data is typically at a 1-minute resolution, atmospheric electric field meters are typically sampled every 10 seconds, and meteorological satellite data is at the hourly level.

[0079] Specifically, step S2 includes:

[0080] (1) The low-frequency data is upsampled using an interpolation algorithm, and the noise of the high-frequency data is suppressed by using a moving average.

[0081] For example, a cubic spline interpolation algorithm can be used to upsample low-frequency meteorological data, while moving average can be used to suppress noise in high-frequency lightning data.

[0082] (2) Establish a gridded terrain correction field based on the digital elevation model (DEM) and map data from different sources to a standard grid.

[0083] For example, the size of the grid cell can be set to 30m.

[0084] S3: Construct a spatiotemporal collaborative perception network model based on convolution, graph neural networks and attention mechanisms, and use multi-source data to build a sample dataset for training to obtain a forest fire prediction model.

[0085] Specifically, the spatiotemporal collaborative perception network model constructed in this embodiment includes, in sequence, a multimodal feature projection layer, a spatiotemporal feature extraction layer, a feature enhancement and aggregation layer, and a feature fusion and prediction layer. The output of the previous layer is the input of the next layer, and the final prediction layer outputs the required forest fire prediction result.

[0086] (1) Multimodal feature projection layer: The forest area multi-source data after alignment processing in step S2 is used as the input of the multimodal feature projection layer. The input multi-source data is mapped to a unified feature space through a fully connected layer and adjusted using a graph structure; the mapped features are batch normalized to obtain the original features.

[0087] The multimodal feature projection layer can map the input multi-source heterogeneous features (including forest lightning data, temperature and humidity data, wind speed data, vegetation data, atmospheric electric field data, meteorological data, etc.) to a unified latent feature space, effectively solving the problem of inconsistent feature scales.

[0088] Specifically, fully connected layers linearly map input features from the original dimensions to the hidden dimensions. Using a graph structure, an initial adjacency matrix is ​​constructed based on the latitude and longitude coordinates of the monitoring stations, and dynamically adjusted by fusing the cosine similarity of meteorological features from each station. Finally, a parallel embedding strategy is employed to simultaneously generate sequential features suitable for convolution, block embedding features for Transformer, and node features for graph neural networks. This design maintains the original model's ability to handle differences in feature scale while providing a rich representational foundation for subsequent spatiotemporal analysis.

[0089] Batch normalization of the mapped features not only accelerates model training convergence but also improves training stability. Finally, the ReLU activation function is used to perform a nonlinear transformation on the normalized features to further enhance the feature representation capability.

[0090] (2) Spatiotemporal feature extraction layer: The original features are extracted using convolutional branches, Transformer branches and graph neural network branches, and the feature interaction of convolutional branches, Transformer branches and graph neural network branches is realized through cross-modal attention gating mechanism to obtain preliminary features.

[0091] The spatiotemporal feature extraction layer employs a hybrid architecture with three feature extraction paths: a convolutional branch, a Transformer branch, and a graph neural network branch. The convolutional branch primarily extracts local features from the raw features, treating them as one-dimensional time-series data to capture short-term variation patterns in meteorological data. This convolutional branch is implemented using the ConvBNAct module, which sequentially includes convolutional layers, batch normalization layers, and activation functions. The Transformer branch divides the feature sequence into overlapping time-series segments and captures long-range dependencies across time steps through a multi-head self-attention mechanism. The graph neural network branch uses a graph attention network architecture, dynamically aggregating neighboring site information for node features and fusing spatial distance and environmental gradient factors for edge weights. The three branches can interact via a cross-modal attention gating mechanism, with the output of the convolutional branch serving as the query, and the outputs of the Transformer and graph neural network branches serving as the key and value, respectively.

[0092] (3) Feature enhancement and aggregation layer: The initial features are enhanced using the SE attention mechanism and global average pooling is performed to obtain aggregated features.

[0093] The core function of the Squeeze-and-Excitation (SE) attention module is to adaptively weight the initial features output from the spatiotemporal feature extraction layer through a channel attention mechanism, thereby enhancing the network's ability to focus on key meteorological and environmental features. This module employs a three-stage processing flow: first, global average pooling is used to spatially compress the initial features, obtaining global feature representations for each channel; then, two fully connected layers (containing ReLU and Sigmoid activation functions in between) generate channel attention weights; finally, the learned channel weights are multiplied by the initial features to obtain the enhanced feature representation.

[0094] The global average pooling layer, as a feature aggregation module, primarily compresses the global information of the enhanced features to obtain the global statistical features of each channel. This design ensures that the model can comprehensively utilize multi-scale feature information for accurate prediction. Specifically, it performs average pooling along the width dimension of the convolutional layer output, ultimately outputting a 16-dimensional feature vector containing global feature information for each channel.

[0095] (4) Feature fusion and prediction layer: The aggregated features are fused using a fully connected layer, and wildfire prediction is performed using a prediction head.

[0096] This layer maps the aggregated features obtained from the feature enhancement and aggregation layers to the output category space (such as "fire" or "non-fire"), generating wildfire prediction results.

[0097] Furthermore, this layer can also concatenate the original features obtained from the multimodal feature projection layer, the feature enhancement, and the aggregated features obtained from the aggregation layer to form a comprehensive feature vector; then, the concatenated features are mapped to the output category space (such as "fire" or "non-fire") through a fully connected layer, ultimately generating the forest fire prediction result. This design ensures that the model fully utilizes the complementary information of global and local features.

[0098] Preferably, the feature fusion and prediction layer can also construct a multi-task prediction head, simultaneously performing auxiliary tasks such as forest fire prediction and risk level assessment. The multi-task prediction head introduces an uncertainty-weighted algorithm to automatically balance the weights of the loss functions for different tasks.

[0099] Sample data were selected from multi-source heterogeneous data in the forest area, labeled, and a sample dataset was constructed. The spatiotemporal collaborative perception network model was trained using methods such as cross-validation to obtain a trained forest fire prediction model for actual prediction.

[0100] S4: Input multi-source data into the forest fire prediction model to obtain forest fire prediction results.

[0101] Real-time multi-source data collected in forest areas is input into a trained forest fire prediction model to obtain results such as forest fire prediction and risk level assessment. Alarms are then issued based on the prediction results to facilitate timely handling and minimize fire risks.

[0102] Preferably, the method further includes: setting up a geographic interpretability analysis module, which can visualize the contribution of each monitoring station to the prediction results and provide spatial basis for decision-making.

[0103] The beneficial effects of this embodiment are as follows: The forest fire prediction model provided by this method, by fusing multi-source environmental data and a spatiotemporal attention mechanism, utilizes the spatial feature extraction capability of convolution, the long-term dependency modeling advantage of Transformer, and graph neural networks to model the spatial relationships of complex terrain and vegetation distribution, constructing a spatiotemporal hybrid model specifically for forest fire prediction, thus enhancing the accuracy and real-time performance of forest fire occurrence prediction; it can achieve fine-grained, seasonal fire occurrence mapping of forest areas, helping to identify fire distribution characteristics and autocorrelation, and providing a more comprehensive analytical perspective; compared with traditional machine learning models, it has higher accuracy, precision, and recall. This method can provide more effective prevention strategies and response solutions for the forest fire challenges brought about by climate change.

[0104] Example 2

[0105] like Figure 1 As shown, a method for real-time prediction of forest fire occurrence includes the following steps:

[0106] S1: Real-time collection of multi-source data for forest area monitoring.

[0107] In practical applications, various detection devices can be deployed in forest areas, including lightning detectors, temperature and humidity sensors, wind speed sensors, video monitoring equipment, multispectral cameras, atmospheric electric field meters, and infrared detectors. Remote sensing satellites and meteorological satellites can also be used for dynamic monitoring of forest areas. Using these detection devices, monitoring equipment, and satellites, multi-source heterogeneous data such as lightning data, temperature and humidity data, wind speed data, vegetation data, atmospheric electric field data, and meteorological data can be collected from forest areas.

[0108] Preferably, the method further includes:

[0109] Preprocess the lightning data.

[0110] In practical applications, this method uses a full-waveform three-dimensional lightning locator to detect lightning data in forest areas. For example, when applying this method to forest fire monitoring in the Greater Khingan Mountains, several full-waveform three-dimensional lightning locators are deployed in the Greater Khingan Mountains to cover the entire area. They can receive lightning location data in real time, including the latitude and longitude, time, intensity, polarity, type (cloud lightning / ground lightning), cloud lightning height, and other information for each lightning strike.

[0111] This method addresses the unique characteristics of lightning data acquired by a full-waveform 3D lightning locator by preprocessing it and designing a dedicated feature extraction and fusion architecture, specifically including:

[0112] (1) Use spherical convolutional layers to process the three-dimensional spatial distribution characteristics of lightning data, including longitude, latitude and cloud lightning height.

[0113] Traditional methods typically simplify lightning data into two-dimensional planar coordinates for processing, while this method innovatively introduces spherical convolutional layers, which can better extract the features of lightning data.

[0114] (2) Using the geographically weighted attention mechanism, the weight of lightning characteristics is dynamically adjusted according to the distribution density of combustibles in the forest area.

[0115] The impact of lightning on forest fires exhibits significant geographical variability. By utilizing a geographically weighted attention mechanism, the weights of lightning characteristics can be dynamically adjusted based on the distribution density of combustible materials in the forest area (such as coniferous forest cover). For example, in areas with high larch density, lightning of the same intensity will receive a higher risk score.

[0116] (3) Design a risk decay time window function and adjust the risk assessment weight of lightning characteristics.

[0117] Lightning data has a certain time sensitivity, so a risk attenuation coefficient is defined based on the time difference: Risk attenuation coefficient = 1 / (1 + α·Δt); where α is a parameter calibrated according to the aridity of the forest area (for example, α=0.05 for the Greater Khingan Mountains), and Δt is the difference between the current time and the time of the lightning strike.

[0118] Adjusting the risk assessment weights of lightning features using a risk attenuation coefficient is significantly superior to the conventional sliding time window method. Experiments show that this method can improve the accuracy of identifying lightning-related fire hazards by 10%.

[0119] S2: Spatiotemporal alignment of multi-source data is performed using sampling and terrain correction.

[0120] In practical applications, the various heterogeneous data collected in forest areas exhibit significant spatiotemporal gaps, thus requiring spatiotemporal alignment processing. For example, lightning data is typically at a 1-minute resolution, atmospheric electric field meters are typically sampled every 10 seconds, and meteorological satellite data is at the hourly level.

[0121] Specifically, step S2 includes:

[0122] (1) The low-frequency data is upsampled using an interpolation algorithm, and the noise of the high-frequency data is suppressed by using a moving average.

[0123] For example, a cubic spline interpolation algorithm can be used to upsample low-frequency meteorological data, while moving average can be used to suppress noise in high-frequency lightning data.

[0124] (2) Establish a gridded terrain correction field based on the digital elevation model (DEM) and map data from different sources to a standard grid.

[0125] For example, the size of the grid cell can be set to 30m.

[0126] S3: Construct a spatiotemporal collaborative perception network model based on convolution, graph neural networks and attention mechanisms, and use multi-source data to build a sample dataset for training to obtain a forest fire prediction model.

[0127] Specifically, the spatiotemporal collaborative perception network model constructed in this embodiment includes, in sequence, a multimodal feature projection layer, a spatiotemporal feature extraction layer, a feature enhancement and aggregation layer, and a feature fusion and prediction layer. The output of the previous layer is the input of the next layer, and the final prediction layer outputs the required forest fire prediction result.

[0128] (1) Multimodal feature projection layer: The forest area multi-source data after alignment processing in step S2 is used as the input of the multimodal feature projection layer. The input multi-source data is mapped to a unified feature space through a fully connected layer, and positional encoding technology and graph structure are introduced for adjustment; the mapped features are batch normalized to obtain the original features.

[0129] The multimodal feature projection layer can map the input multi-source heterogeneous features (including forest lightning data, temperature and humidity data, wind speed data, vegetation data, atmospheric electric field data, meteorological data, etc.) to a unified latent feature space, effectively solving the problem of inconsistent feature scales.

[0130] Preferably, for the multi-source heterogeneous features of the input, a dynamic modality weight allocation strategy for environment awareness can be established to break through the traditional fixed dominant modality limitation. For example:

[0131] ①Clear / High visibility: Optical data is the dominant mode (weight 0.4-0.5);

[0132] ② Night / Cloudy: Infrared data is the dominant mode (weight 0.5-0.6);

[0133] ③ Cloudy and foggy weather: SAR is the dominant mode (weight 0.4-0.5);

[0134] ④ Strong winds during the dry season: Meteorological data is the main mode (weight 0.3-0.4).

[0135] The formula for calculating dynamic weights is:

[0136] w_i = exp(α·q_i) / Σexp(α·q_j)

[0137] Where w_i represents the weight value of the i-th mode; q_i represents the quality score of the i-th mode; and α is the environmental coefficient.

[0138] Specifically, a fully connected layer linearly maps the input features from the original dimension to the hidden dimension. A learnable location encoding technique is introduced to handle the irregular sampling problem of meteorological data through conditional location encoding. The conditional location encoding is: PE(pos)=[sin(pos / 10000^(2i / d_model)); cos(pos / 10000^(2i / d_model))]; where pos represents the standardized coordinates of the monitoring station, d_model=64 is the encoding dimension, and i represents the i-th data point.

[0139] Using a graph structure, an initial adjacency matrix is ​​constructed based on the latitude and longitude coordinates of the monitoring stations, and dynamically adjusted by incorporating the cosine similarity of meteorological features of each station. The elements of the adjacency matrix A∈R^(N×N) are calculated as follows: A_ij = softmax(MLP([x_i‖x_j‖Δg_ij])); where MLP([x_i‖x_j‖Δg_ij]) represents inputting the concatenated feature vector of x_i, x_j, and Δg_ij into a multilayer perceptron (MLP) for processing; x_i and x_j are the meteorological feature vectors of stations i and j, respectively; Δg_ij represents the meteorological gradient (temperature difference, humidity difference, etc.) between stations i and j; and N is the number of monitoring stations.

[0140] Finally, a parallel embedding strategy is employed to simultaneously generate serialized features suitable for convolution, block embedding features for Transformers, and node features for graph neural networks. This design maintains the original model's ability to handle feature scale differences while providing a rich representational foundation for subsequent spatiotemporal analysis.

[0141] Batch normalization of the mapped features not only accelerates model training convergence but also improves training stability. Finally, the ReLU activation function is used to perform a nonlinear transformation on the normalized features to further enhance the feature representation capability.

[0142] (2) Spatiotemporal feature extraction layer: The original features are extracted using convolutional branches, Transformer branches and graph neural network branches, and the feature interaction of convolutional branches, Transformer branches and graph neural network branches is realized through cross-modal attention gating mechanism to obtain preliminary features.

[0143] The spatiotemporal feature extraction layer employs a hybrid architecture with three feature extraction paths: a convolutional branch, a Transformer branch, and a graph neural network branch. The convolutional branch primarily extracts local features from the raw data, treating it as one-dimensional time-series data to capture short-term variation patterns in meteorological data. Implemented using the ConvBNAct module, the convolutional branch sequentially includes convolutional layers, batch normalization layers, and activation functions, and introduces dilated convolutions to expand the receptive field. The Transformer branch divides the feature sequence into overlapping temporal segments and captures long-range dependencies across time steps through a multi-head self-attention mechanism. The graph neural network branch uses a graph attention network architecture, dynamically aggregating neighboring site information for node features, and fusing spatial distance and environmental gradient factors for edge weights.

[0144] The three branches can achieve feature interaction through a cross-modal attention gating mechanism, where the output of the convolutional branch serves as the Query, and the outputs of the Transformer and Graph Neural Network branches serve as the Key and Value, respectively.

[0145] (3) Feature enhancement layer: The initial features are enhanced by dual-path feature enhancement using SE attention mechanism and graph structure, and the enhanced features are obtained by adaptive weighting through gating unit.

[0146] The feature enhancement layer utilizes the Squeeze-and-Excitation (SE) attention module and graph structure to form a dual-path feature enhancement channel. The SE attention is the channel attention path, and the graph structure is the graph attention path.

[0147] The core function of the SE attention module is to adaptively weight the initial features output from the spatiotemporal feature extraction layer through a channel attention mechanism, thereby enhancing the network's ability to focus on key meteorological and environmental features. This module employs a three-stage processing flow: first, global average pooling is used to spatially compress the initial features, obtaining global feature representations for each channel; then, two fully connected layers (containing ReLU and Sigmoid activation functions in between) generate channel attention weights; finally, the learned channel weights are multiplied by the initial features to obtain the enhanced feature representation. The channel attention path utilizes the compression-activation structure of the SE module to highlight important feature dimensions.

[0148] The graph attention path constructs a dynamically learnable spatial relationship graph, and achieves cross-site feature enhancement through multi-head graph attention.

[0149] The outputs of the channel attention path and the graph attention path are adaptively weighted through a gated fusion unit, with the fusion weights dynamically generated by the feature importance prediction sub-network. The gated fusion output is: y_out = σ(W_g[f_c‖f_g])⊙f_c + (1-σ(W_g[f_c‖f_g]))⊙f_g; where f_c is the output of the channel attention path, f_g is the output of the graph attention path, ‖ denotes vector concatenation, W_g is the weight, and σ is the coefficient.

[0150] This design allows the model to focus on key meteorological indicators while effectively utilizing complementary information from spatially related stations.

[0151] (4) Feature aggregation layer: Multi-scale pooling is performed on the enhanced features in the time dimension and graph pooling is performed in the spatial dimension. The features are matched and filtered through the spatiotemporal cross attention mechanism to obtain aggregated features.

[0152] The feature aggregation layer comprises a three-level information compression process: First, multi-scale pooling is performed in the temporal dimension, combining average pooling and max pooling to capture temporal patterns at different granularities; second, graph pooling is performed in the spatial dimension, progressively downsampling the monitoring station network based on node importance scores; finally, a spatiotemporal cross-attention mechanism is used to match the compressed features with learnable task prototype vectors, selecting the most discriminative feature combinations. This hierarchical aggregation strategy significantly improves the model's efficiency in processing long-sequence meteorological data, ensuring that the model can comprehensively utilize multi-scale feature information for accurate prediction.

[0153] (5) Feature fusion and prediction layer: The cross-modal attention mechanism is used to fuse aggregated features and the prediction head is used to predict forest fires.

[0154] The feature fusion stage employs a cross-modal attention mechanism to align the local features extracted by convolution, the global temporal patterns encoded by Transformer, and the spatial relationships inferred by the graph neural network in three dimensions. Specifically, a three-level fusion strategy can be adopted: primary fusion involves concatenating the local features extracted by convolution, the global temporal patterns encoded by Transformer, and the spatial relationships inferred by the graph neural network; intermediate fusion uses a gating mechanism for feature selection; and advanced fusion adaptively weights the features.

[0155] Finally, the fused features are mapped to the output category space (such as "fire" or "non-fire") using the prediction head to generate wildfire prediction results.

[0156] Preferably, the feature fusion and prediction layer can also construct a multi-task prediction head to simultaneously perform auxiliary tasks such as forest fire prediction and risk level assessment.

[0157] Multi-task prediction heads can incorporate uncertainty-weighted algorithms to automatically balance the weights of the loss functions for different tasks. For example, when the model includes two task prediction heads, the uncertainty-weighted loss is: L_total = 1 / (2σ_1^2 ) L_1 +1 / (2σ_2^2 ) L_2 + logσ_1^2 + logσ_2^2; where σ_1 and σ_2 are the task-related uncertainty parameters; and L_1 and L_2 are the losses of the two task prediction heads.

[0158] In addition, a triple loss function can be introduced for supervised alignment across multiple tasks: L_align = L_ce + λ1L_triplet + λ2L_consistency; where L_align is the triple loss, L_ce is the cross-entropy loss, L_triplet is the triplet loss, L_consistency is the inter-modal prediction consistency loss, and λ1 and λ2 are weight parameters. The triplet loss is a widely used loss function in metric learning. Its core objective is to optimize model parameters so that similar samples (positive sample pairs) are closer together in the feature space, while dissimilar samples (negative sample pairs) are farther apart, thereby enhancing intra-class clustering and expanding inter-class differences. The inter-modal prediction consistency loss is a loss function used in multi-modal learning. Its goal is to ensure consistency across different modalities (such as images, text, and audio) in common representation or prediction tasks. By minimizing the differences in prediction results across different modalities for the same task, the model can learn shared semantic information across modalities, improving the effectiveness of multi-modal fusion.

[0159] Sample data were selected from multi-source heterogeneous data in the forest area, labeled, and a sample dataset was constructed. The spatiotemporal collaborative perception network model was trained using methods such as cross-validation to obtain a trained forest fire prediction model for actual prediction.

[0160] S4: Input multi-source data into the forest fire prediction model to obtain forest fire prediction results.

[0161] Real-time multi-source data collected in forest areas is input into a trained forest fire prediction model to obtain results such as forest fire prediction and risk level assessment. Alarms are then issued based on the prediction results to facilitate timely handling and minimize fire risks.

[0162] Preferably, the method further includes: setting up a geographic interpretability analysis module, which can visualize the contribution of each monitoring station to the prediction results and provide spatial basis for decision-making.

[0163] The beneficial effects of this embodiment are as follows: The forest fire prediction model provided by this method achieves comprehensive modeling of the spatiotemporal characteristics of meteorological data through the deep synergy of three neural network architectures: the convolutional architecture excels at capturing abrupt change patterns in local meteorological indicators; the Transformer network effectively models the long-term evolution patterns of meteorological elements; and the graph neural network explicitly encodes the spatial correlation between monitoring stations. This fusion architecture achieves a dual improvement in accuracy and robustness in forest fire prediction tasks, especially demonstrating excellent generalization ability for out-of-distribution samples (such as meteorological data from unseen areas). The model's spatial attention map can also automatically identify high-risk associated areas, providing scientific guidance for joint prevention and control of forest fires.

[0164] Example 3

[0165] A method for real-time prediction of forest fire occurrence includes:

[0166] By setting up an edge layer and a cloud, an edge-cloud collaborative computing architecture is constructed, and a layered processing strategy is adopted to better meet real-time requirements.

[0167] The forest fire prediction models used in Examples 1 and 2 typically need to be deployed in data centers or the cloud. They can only be processed using computing resources after data collected from forest areas is gathered, thus incurring some latency and resource dependence. While this centralized processing mode can rely on the powerful computing capabilities of the cloud to complete complex model training and global data analysis, it is limited by data transmission bandwidth, network stability, and scheduling priorities. From the capture of fire signals by forest area monitoring equipment to the final output of early warning information, a certain time delay is often required. Especially in remote forest areas or in scenarios where communication is disrupted due to disasters, obstructed data transmission can directly lead to the failure of early warnings. Furthermore, the continuously increasing scale of monitoring equipment and data volume places higher demands on cloud storage and computing resources, highlighting long-term operating costs and energy consumption pressures, and relying on a single cloud architecture poses a single point of failure risk.

[0168] Therefore, by building an edge-cloud collaboration, a lightweight prediction model can be deployed at the edge to adapt it to embedded devices, enabling localized real-time analysis and alarms.

[0169] (1) Edge layer

[0170] Atmospheric electric field meters, infrared fire detectors, and other equipment can be deployed in forest areas, with sampling frequencies set to short intervals, such as once every 10 seconds. The collected data is then analyzed and judged. If a forest fire is directly identified, an alarm is immediately triggered, meeting the real-time requirements for forest fire prediction.

[0171] For example, the edge layer can detect the electric field strength of atmospheric electric field data and immediately issue an alarm when a sudden change in electric field strength is detected; or if the edge layer detects a fire in the forest area, it can immediately issue an alarm, so that forest rangers can take timely countermeasures to eliminate the risk of fire.

[0172] (2) Cloud

[0173] The cloud receives various signals collected by edge devices, integrates multi-source data, and uses a trained forest fire prediction model, such as the forest fire prediction model described in Example 1 or Example 2, or other convolutional neural network, multimodal neural network models, to predict forest fire risk, generate prediction results, and issue corresponding alarms based on the prediction results.

[0174] The beneficial effect of this embodiment is that by constructing collaborative computing and hierarchical processing between the edge layer and the cloud, the real-time requirements of forest fire prediction can be better met.

[0175] Example 4

[0176] A real-time forest fire prediction system, based on the real-time forest fire prediction method as described in Embodiments 1, 2, or 3, includes:

[0177] The data acquisition module is used to collect multi-source data for forest area monitoring in real time.

[0178] The alignment module is used to perform spatiotemporal alignment processing on multi-source data using sampling and terrain correction.

[0179] The model building module is used to build a spatiotemporal collaborative perception network model based on convolution, graph neural networks and attention mechanisms, and to use multi-source data to build a sample dataset for training to obtain a forest fire prediction model.

[0180] The prediction module is used to input multi-source data into the forest fire prediction model and obtain forest fire prediction results.

[0181] Preferably, the system further includes setting up an edge layer and a cloud layer.

[0182] The edge layer detects the electric field intensity of atmospheric electric field data and immediately issues an alarm when a sudden change in electric field intensity is detected.

[0183] The cloud uses forest fire prediction models to predict forest fire risks and issue alerts.

[0184] Preferably, the system further includes:

[0185] The geographic interpretability analysis module can visualize the contribution of each monitoring station to the prediction results, providing spatial basis for decision-making.

[0186] Example 5

[0187] A computer program product comprising a computer program that, when executed, implements the steps of the real-time forest fire occurrence prediction method as described in Embodiments 1, 2, or 3.

[0188] Example 6

[0189] A readable storage medium storing a computer program as described in Embodiment 5, wherein when executed, the computer program implements the steps of the real-time forest fire occurrence prediction method as described in Embodiments 1, 2, or 3.

[0190] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for real-time prediction of forest fire occurrence, characterized in that, Includes the following steps: Real-time collection of multi-source data from forest area monitoring; Spatiotemporal alignment of multi-source data is performed using sampling and terrain correction. A spatiotemporal collaborative perception network model was constructed based on convolutional, graph neural networks and attention mechanisms. A sample dataset was built using multi-source data for training to obtain a forest fire prediction model. Input multi-source data into the forest fire prediction model to obtain forest fire prediction results; The spatiotemporal collaborative perception network model constructed based on convolution, graph neural networks, and attention mechanisms includes: A multimodal feature projection layer is constructed, which maps the input multi-source data to a unified feature space through a fully connected layer and adjusts it using a graph structure; the mapped features are then batch normalized to obtain the original features. A spatiotemporal feature extraction layer is constructed, which uses convolutional branches, Transformer branches and graph neural network branches to extract features from the original features. The feature interaction between the convolutional branches, Transformer branches and graph neural network branches is realized through a cross-modal attention gating mechanism to obtain preliminary features. Construct a feature enhancement and aggregation layer, use the SE attention mechanism to enhance the initial features, and perform global average pooling to obtain aggregated features; A feature fusion and prediction layer is constructed, which uses a fully connected layer to perform feature fusion on aggregated features and uses a prediction head to predict wildfires.

2. A method for real-time prediction of forest fire occurrence, characterized in that, Includes the following steps: Real-time collection of multi-source data from forest area monitoring; Spatiotemporal alignment of multi-source data is performed using sampling and terrain correction. A spatiotemporal collaborative perception network model was constructed based on convolutional, graph neural networks and attention mechanisms. A sample dataset was built using multi-source data for training to obtain a forest fire prediction model. Input multi-source data into the forest fire prediction model to obtain forest fire prediction results; The spatiotemporal collaborative perception network model constructed based on convolution, graph neural networks, and attention mechanisms includes: A spatiotemporal collaborative perception network model was constructed based on convolutional, graph neural networks and attention mechanisms. A sample dataset was built using multi-source data for training to obtain a forest fire prediction model. Input multi-source data into the forest fire prediction model to obtain forest fire prediction results; A multimodal feature projection layer is constructed, which maps the input multi-source data to a unified feature space through a fully connected layer, and positional encoding technology and graph structure are introduced for adjustment; the mapped features are then batch normalized to obtain the original features. A spatiotemporal feature extraction layer is constructed, which uses convolutional branches, Transformer branches and graph neural network branches to extract features from the original features. The feature interaction between the convolutional branches, Transformer branches and graph neural network branches is realized through a cross-modal attention gating mechanism to obtain preliminary features. A feature enhancement layer is constructed, which uses the SE attention mechanism and graph structure to perform dual-path feature enhancement on the initial features, and adaptive weighting is performed through gating units to obtain enhanced features; A feature aggregation layer is constructed, which performs multi-scale pooling on the enhanced features in the time dimension and graph pooling in the spatial dimension. The aggregated features are obtained by matching and filtering through a spatiotemporal cross-attention mechanism. A feature fusion and prediction layer is constructed, and a cross-modal attention mechanism is used to fuse aggregated features. Then, a prediction head is used to predict wildfires.

3. The real-time forest fire prediction method according to claim 1 or 2, characterized in that, The multi-source data includes: lightning data, atmospheric electric field data, and meteorological data.

4. The real-time forest fire prediction method according to claim 3, characterized in that, The real-time forest fire prediction method also includes: Preprocess the lightning data.

5. The real-time forest fire prediction method according to claim 4, characterized in that, The preprocessing of the lightning data includes: Spherical convolutional layers are used to process lightning data, which includes the three-dimensional spatial distribution characteristics of longitude, latitude, and cloud lightning height. By utilizing a geographic weighted attention mechanism, the weights of lightning characteristics are dynamically adjusted based on the distribution density of combustibles in forest areas; Design a risk decay time window function to adjust the risk assessment weights of lightning characteristics.

6. The real-time forest fire prediction method according to claim 1 or 2, characterized in that, The spatiotemporal alignment processing of multi-source data using sampling and terrain correction includes: An interpolation algorithm is used to upsample low-frequency data, and a moving average is used to suppress noise in high-frequency data. A gridded terrain correction field is established based on the digital elevation model, and data from different sources are uniformly mapped to a standard grid.

7. The real-time forest fire prediction method according to claim 2, characterized in that, The multimodal feature projection layer also includes: For input multi-source data, a dynamic modal weight allocation strategy for environmental perception is established.

8. The real-time forest fire prediction method according to claim 2, characterized in that, The feature fusion of aggregated features using the cross-modal attention mechanism includes: The local features extracted by convolution, the global temporal patterns encoded by Transformer, and the spatial relationships inferred by graph neural network are 3D aligned. A three-level fusion strategy is adopted: primary fusion is to concatenate the local features extracted by convolution, the global temporal patterns encoded by Transformer, and the spatial relationships inferred by graph neural network; intermediate fusion is to use a gating mechanism to select features; and advanced fusion is to adaptively weight the features.

9. The real-time forest fire prediction method according to claim 2, characterized in that, The feature fusion and prediction layer also includes: constructing a multi-task prediction head for forest fire prediction and risk level assessment.

10. The real-time forest fire prediction method according to claim 9, characterized in that, The multi-task prediction head introduces an uncertainty weighting algorithm to automatically balance the weights of the loss functions for different tasks.

11. The real-time forest fire prediction method according to claim 9, characterized in that, The multi-task prediction head introduces a triple loss function to supervise and align multiple tasks: L_align = L_ce + λ1L_triplet + λ2L_consistency; Where L_align is the triple loss, L_ce is the cross-entropy loss, L_triplet is the triplet loss, L_consistency is the inter-modal prediction consistency loss, and λ1 and λ2 are weight parameters.

12. The real-time forest fire prediction method according to claim 1 or 2, characterized in that, The real-time forest fire prediction method also includes: A geographic interpretability analysis module is set up to visualize the contribution of each monitoring station to the prediction results, providing spatial basis for decision-making.

13. The real-time forest fire prediction method according to claim 3, characterized in that, The real-time forest fire prediction method also includes: Configure the edge layer and the cloud; The edge layer detects the electric field intensity of atmospheric electric field data and immediately issues an alarm when a sudden change in electric field intensity is detected. The cloud uses forest fire prediction models to predict forest fire risks and issue alerts.

14. A real-time forest fire occurrence prediction system, characterized in that, The real-time forest fire prediction method according to any one of claims 1 to 13 includes: The data acquisition module is used to collect multi-source data for forest area monitoring in real time. The alignment module is used to perform spatiotemporal alignment processing on multi-source data using sampling and terrain correction. The model building module is used to build a spatiotemporal collaborative perception network model based on convolution, graph neural networks and attention mechanisms, and to use multi-source data to build a sample dataset for training to obtain a forest fire prediction model. The prediction module is used to input multi-source data into the forest fire prediction model and obtain forest fire prediction results.