A method and system for monitoring forest fires based on multi-source heterogeneous data fusion
By spatially registering and temporally aligning multi-source heterogeneous data, a spatiotemporal heterogeneous map is constructed, and a mean heatmap and uncertainty distribution map are generated. Combined with a semi-physical fire behavior model, the timeliness and integration problems of fire monitoring data in existing technologies are solved, and the comprehensiveness and accuracy of forest fire risk assessment are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- KANGSHUO (CHONGQING) INTELLIGENT MANUFACTURING SYSTEM TECHNOLOGY RESEARCH INSTITUTE CO LTD
- Filing Date
- 2026-05-08
- Publication Date
- 2026-06-02
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Among existing forest fire monitoring technologies, satellite monitoring systems are prone to false alarms due to atmospheric interference and cloud cover, and data timeliness is insufficient; drone inspection systems have limited coverage and are prone to missing smoke or weak fire points; various solutions have failed to effectively integrate satellite, drone and ground sensor data, resulting in blurred fire area boundaries, inaccurate prediction of spread trends and deviations in the estimation of smoke spread range.
By spatially registering and temporally aligning satellite hotspot monitoring data, UAV inspection video stream data, and time-series data from ground-based IoT sensors, multimodal feature vectors are generated. Cross-modal fire risk inference is then performed using a fire inference network to construct a spatiotemporal heterogeneous map. Combined with a graph variational autoencoder network, mean heatmaps and uncertainty distribution maps are generated. Finally, a semi-physical fire behavior model is used to generate the geometric envelope of smoke diffusion, thereby improving the comprehensiveness and accuracy of fire risk assessment.
It solves the feature fusion bias caused by spatiotemporal inconsistencies in multi-source data, quickly filters out fire areas to be confirmed, realizes probabilistic estimation of fire status, and improves the comprehensiveness and accuracy of fire risk assessment.
Smart Images

Figure CN122135480A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of forest fire monitoring technology, and in particular to a forest fire monitoring method and system based on multi-source heterogeneous data fusion. Background Technology
[0002] Current forest fire monitoring practices are generally limited to single data source applications or surface data overlay models. Satellite-based monitoring systems rely solely on preset brightness and temperature thresholds to screen hotspots, but are susceptible to atmospheric interference, surface reflection, and cloud cover, leading to increased false alarm rates. Furthermore, the limited satellite revisit cycle results in insufficient data timeliness, making it difficult to capture early fire dynamics. While satellite-based solutions combining ground-based IoT sensors introduce temperature comparison mechanisms, the sparse spatial distribution of sensors and asynchronous timestamps cause spatiotemporal misalignments in data comparison, failing to effectively correct measurement biases caused by terrain obstruction or local microclimates. Unmanned aerial vehicle (UAV) inspection systems rely excessively on manual video analysis, limiting coverage due to endurance and weather conditions. Operators are prone to missing smoke or weak fire points due to visual fatigue, and subsequent analysis simply correlates ground sensor temperature and humidity data, failing to delve into the physical relationship between smoke flow characteristics and fire thermal radiation within the video stream. More notably, none of the proposed solutions explored the intrinsic connections between satellite hotspot data, drone video streams, and ground sensor time-series data. This superficial integration approach in existing technologies leads to blurred boundaries of fire zones, inaccurate predictions of spread trends, and biased estimates of smoke diffusion ranges, making it impossible to support refined fire situation awareness and firefighting resource allocation decisions.
[0003] Therefore, there is an urgent need for a forest fire monitoring method and system based on the fusion of multi-source heterogeneous data. Summary of the Invention
[0004] To address the aforementioned technical issues, this application provides a forest fire monitoring method and system based on multi-source heterogeneous data fusion.
[0005] A first aspect of this application provides a forest fire monitoring method based on multi-source heterogeneous data fusion, comprising: Spatial registration and temporal alignment are performed on multi-source heterogeneous data of the target forest area to generate multimodal feature vectors; the multi-source heterogeneous data includes satellite hotspot monitoring data, UAV inspection video stream data and time series data of ground IoT sensors; The multimodal feature vectors are input into a pre-defined fire inference network to perform cross-modal fire risk inference, generate a first fire confidence map, and extract the fire area to be confirmed. In response to the maximum value in the first fire confidence map satisfying the preset threshold condition, the area to be confirmed as a fire is rasterized. The resulting raster is used as a node, the spatial adjacency relationship between adjacent raster is used as the first type of edge, and the connection between the cosine similarity between the multimodal feature vectors corresponding to each raster node is greater than the preset similarity threshold and the Euclidean distance between the two raster nodes in geographic space is less than the preset neighborhood expansion radius is used as the second type of edge. A spatiotemporal heterogeneous graph is constructed based on the first type of edge and the second type of edge. The spatiotemporal heterogeneous map and the multimodal feature vectors corresponding to each grid node are input into a preset graph variational autoencoder network to obtain the mean heat map and uncertainty distribution map of the fire area to be confirmed. Based on the fire intensity parameters of each grid calculated from the mean heatmap, local meteorological grid data, and semi-physical fire behavior spread model, a geometric envelope of fire smoke diffusion is generated. The mean heat map and the uncertainty distribution map are overlaid on the geographic information system base map, and the spatial overlap area between the area with the heat value greater than the first threshold in the mean heat map and the geometric envelope of the fire smoke diffusion is calculated as the composite high-risk area.
[0006] A second aspect of this application provides a forest fire monitoring system based on multi-source heterogeneous data fusion, comprising: The data registration and feature generation module is used to perform spatial registration and temporal alignment of multi-source heterogeneous data of the target forest area to generate multimodal feature vectors; the multi-source heterogeneous data includes satellite hotspot monitoring data, UAV inspection video stream data and ground IoT sensor time series data; The fire inference and region extraction module is used to input multimodal feature vectors into a preset fire inference network to perform cross-modal fire risk inference, generate a first fire confidence map, and extract the fire area to be confirmed. The spatiotemporal heterogeneous graph construction module is used to perform rasterization processing on the fire area to be confirmed in response to the maximum value in the first fire confidence map meeting the preset threshold condition. The resulting raster is used as a node, the spatial adjacency relationship between adjacent raster is used as the first type of edge, and the connection between the cosine similarity between the multimodal feature vectors corresponding to each raster node is greater than the preset similarity threshold and the Euclidean distance between the two raster nodes in geographic space is less than the preset neighborhood expansion radius is used as the second type of edge. The spatiotemporal heterogeneous graph is constructed based on the first type of edge and the second type of edge. The graph variational autoencoder inference module is used to input the spatiotemporal heterogeneous graph and the multimodal feature vectors corresponding to each grid node into a preset graph variational autoencoder network to obtain the mean heat map and uncertainty distribution map of the fire area to be confirmed. The smoke diffusion envelope generation module is used to calculate the fire line intensity parameters of each grid based on the mean heat map, and generate the geometric envelope of the fire smoke diffusion by combining local meteorological grid data and semi-physical fire behavior spread model. The composite high-risk area module is used to overlay the mean heat map and the uncertainty distribution map onto the geographic information system base map, and calculate the spatial overlap area between the area with the heat value greater than the first threshold in the mean heat map and the geometric envelope of the fire smoke diffusion, as the composite high-risk area.
[0007] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the above-described forest fire monitoring method based on multi-source heterogeneous data fusion.
[0008] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described forest fire monitoring method based on multi-source heterogeneous data fusion.
[0009] The beneficial effects of the forest fire monitoring method and system based on multi-source heterogeneous data fusion provided in this application are as follows: This application spatially registers and temporally aligns three types of heterogeneous data—satellite, UAV, and terrestrial IoT—solving the feature fusion bias problem caused by spatiotemporal inconsistency of multi-source data. The generated multimodal feature vector can comprehensively characterize the fire features of different dimensions in the forest area. Secondly, the fire inference network quickly filters out the fire areas to be confirmed, significantly narrowing the scope of subsequent processing. Then, a spatiotemporal heterogeneous graph including spatial adjacency edges and feature similarity edges is constructed, which preserves the continuity of geographic space and captures the fire correlation of non-adjacent areas. Furthermore, a graph variational autoencoder simultaneously outputs a mean heat map and an uncertainty distribution map, realizing a probabilistic estimation of the fire status rather than a single deterministic judgment. Finally, a semi-physical fire behavior model is combined to generate a smoke flow diffusion envelope, which is superimposed with high-heat areas to obtain a composite high-risk area, improving the comprehensiveness and accuracy of fire risk assessment. Attached Figure Description
[0010] Figure 1 A flowchart illustrating a forest fire monitoring method based on multi-source heterogeneous data fusion provided in an embodiment of this application; Figure 2 A structural block diagram of a forest fire monitoring system based on multi-source heterogeneous data fusion provided in an embodiment of this application; Figure 3 This is a schematic block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0011] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0012] To make the purpose, technical solution, and advantages of this application clearer, the following will be described in conjunction with the appendix. Figure 1 - Appendix Figure 3 The following is an explanation using specific examples.
[0013] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating a forest fire monitoring method based on multi-source heterogeneous data fusion provided in an embodiment of this application. The method includes: S101: Spatial registration and temporal alignment of multi-source heterogeneous data in the target forest area to generate multimodal feature vectors; multi-source heterogeneous data includes satellite hotspot monitoring data, UAV inspection video stream data, and time-series data from ground IoT sensors.
[0014] In this embodiment, multi-source heterogeneous data refers to a data set originating from different sensors or platforms, and whose data formats, structures, or semantics differ. In this embodiment, multi-source heterogeneous data includes satellite hotspot monitoring data, UAV inspection video stream data, and time-series data from ground-based IoT sensors. These data characterize the state of the target forest area from different dimensions.
[0015] Specifically, satellite data provides large-scale thermal anomaly information, UAV video stream data provides high-resolution visual information, and ground sensor data provides local environmental parameters such as temperature and humidity. To ensure the collaborative use of these data from different sources and in different formats, spatial registration is required. This embodiment uses geographic coordinate transformation and image registration techniques to unify all multi-source heterogeneous data into the same geographic coordinate system; simultaneously, timestamp synchronization ensures time alignment of different data. Subsequently, a pre-defined multi-source feature encoding network is used to extract and fuse features from the spatially registered and time-aligned multi-source heterogeneous data, generating a unified multimodal feature vector. This multimodal feature vector can characterize the multi-dimensional information of the target forest area.
[0016] The multi-source feature encoding network is independent of the fire inference network and is used to map raw multi-source heterogeneous data to a unified multimodal feature vector space. This multi-source feature encoding network includes: a satellite data encoder, a UAV video stream encoder, a ground sensor temporal encoder, and a multi-source feature fusion layer; among them, the satellite data encoder adopts the convolutional part of ResNet-18 (removing the last fully connected layer), the input is multi-channel raster data of a single time-series satellite, and the output is a feature map with dimensions [H,W,512], which is then compressed to 64 dimensions by a 1×1 convolutional layer.
[0017] The UAV video stream encoder employs the spatial stream branch of a pre-trained SlowFast network. The input is a short video clip (8 frames), and the output is a 256-dimensional vector after spatiotemporal global pooling, which is then mapped to a 64-dimensional vector through a fully connected layer. For each spatial grid, video features of the UAV flight trajectory coverage area are spatially interpolated and assigned to the corresponding grid.
[0018] The ground sensor time-series encoder employs a two-layer bidirectional LSTM network with 64 hidden units per layer. The input is the time sequence of each sensor node, and the output is the concatenated hidden state (128-dimensional) at the last time step, which is then compressed to 64-dimensional through a fully connected layer. Sensor features are extended to all grids using inverse distance weighted interpolation.
[0019] The multi-source feature fusion layer concatenates the three types of 64-dimensional features along the channel dimension to obtain 192-dimensional intermediate features. Then, through a two-layer fully connected network and layer normalization, the final 64-dimensional multimodal feature vector is output.
[0020] The training objective of this multi-source feature encoding network is to maximize the mutual information of different modal features in the representation space while preserving reconstruction capability. It employs a multi-task loss function for joint training, where... The contrast loss is calculated for three pairs of modalities: satellite features and UAV features, UAV features and sensor features, and satellite features and sensor features at the same spatiotemporal location. This encourages similar representations of similar samples and dissimilar representations of dissimilar samples.
[0021] The reconstruction loss is achieved by connecting each modal encoder to a lightweight decoder (transposed convolution or inverse LSTM), which requires reconstructing the original data from the encoded features and uses mean squared error loss.
[0022] The consistency loss applies an L2 distance constraint to the feature vectors of the same grid before and after the feature update of the target area, to prevent the feature space from drifting drastically due to the supplementary data.
[0023] Total loss = α·S1 + β·S2 + γ·S3; where S1 is the contrastive loss, S2 is the reconstruction loss, and S3 is the consistency loss, α = 1.0, β = 0.5, and γ = 0.1. The optimizer used is Adam, with an initial learning rate of 5e. -4 , train 100 times.
[0024] The output (64-dimensional vector) of this multi-source feature encoding network is directly used as the input to the fire inference network and the graph variational autoencoder network. After the supplementary data is processed through the same encoding process to obtain feature vectors, the original feature vectors of the corresponding grids are replaced with the updated feature vectors to ensure that the subsequent spatiotemporal heterogeneous graph construction and graph variational inference are based on the latest and lowest uncertainty multimodal representation.
[0025] In this embodiment, the specific definition of the same spatiotemporal location is as follows: First, satellite hotspot monitoring data, UAV inspection video stream data, and ground IoT sensor time-series data are uniformly spatially registered to a grid with a resolution of 30m × 30m. For UAV video stream features, a nearest neighbor interpolation method based on grid center coordinates is used to assign feature vectors in the video frame that are less than 15m away from the center point of each grid to that grid. For ground sensor time-series features, an inverse distance weighting method is used for spatial expansion, with a distance attenuation coefficient of 2 in the weight calculation, and only grids within a 500m radius around the sensor are included in the interpolation range. After the above processing, the features of the three modalities all have a unified grid index, forming a feature map with dimensions [H, W, 64]. At this point, the same grid location is considered the same spatiotemporal location.
[0026] The specific implementation of the reconstruction loss is as follows: The satellite data decoder consists of three transposed convolutional layers, each with a kernel size of 4×4 and strides of 2, 2, and 1 respectively. The number of output channels is consistent with the original satellite multi-channel data, and the reconstruction loss is the sum of the mean square errors of each channel. The UAV video stream decoder is based on an R(2+1)D deconvolutional structure, which progressively upsamples the 64-dimensional features and restores the spatiotemporal sequence. The loss is a combination of the optical flow consistency error and pixel reconstruction error between the output and input video frames, with a combined weight ratio of 1:0.5. The ground sensor temporal decoder consists of a bidirectional LSTM followed by a fully connected layer, outputting a reconstructed temporal sequence. The loss is the mean square error between the reconstructed sequence and the original sequence. The above decoders are trained independently and are only used in the pre-training stage of the multi-source feature encoding network.
[0027] S102: Input the multimodal feature vector into the preset fire inference network to perform cross-modal fire risk inference, generate the first fire confidence map, and extract the fire area to be confirmed.
[0028] In this embodiment, a multimodal feature vector refers to a vector representation that characterizes the comprehensive features of a target region by extracting and fusing features from multi-source heterogeneous data that has undergone spatial registration and temporal alignment. This multimodal feature vector can capture the correlation between different modalities of data.
[0029] A fire inference network is a pre-defined machine learning model built using a deep neural network structure. This network is trained to receive multimodal feature vectors as input and, by learning the complex nonlinear relationships between different modalities of data and fire conditions, outputs a confidence level of fire occurrence. Specifically, the fire inference network can employ a multi-stream fusion architecture based on a multi-head attention mechanism, such as a branch consisting of satellite hotspot feature encoding, a branch for extracting spatiotemporal features from UAV video streams, and a branch for extracting temporal features from ground sensors. By using cross-modal attention layers to mine the correlation patterns between different modalities of data in spatial and temporal dimensions, it can jointly infer the fire risk status of each grid cell in the target forest area.
[0030] The first fire confidence map is a two-dimensional raster map output by the fire inference network, representing the confidence level of fire occurrence at various locations in the target forest area. The value of each raster in the two-dimensional raster map represents the probability or intensity of a fire at that location; the higher the value, the higher the confidence level of the fire.
[0031] The unconfirmed fire area refers to the area with a high fire confidence level that needs further confirmation, selected based on the first fire confidence map and by setting a threshold. These areas are considered potential fire locations.
[0032] Specifically, the fire inference network consists of three parallel modal feature encoding branches and a cross-modal fusion inference module. The first branch is the satellite hotspot feature encoding branch. Its input is a rasterized feature tensor of the satellite hotspot monitoring data after spatial registration and temporal alignment, with dimensions [H, W, Cs], where H and W are the number of raster rows and columns of the target forest area in space, and Cs is the number of satellite data channels (including brightness temperature, background brightness temperature deviation, static terrain occlusion index, etc.). This branch uses a stacked 3-layer 2D convolutional neural network for spatial feature extraction. Each convolutional kernel is 3×3 with a stride of 1, and the activation function is LeakyReLU. The output feature map maintains dimensions [H, W, D_s], where Ds = 64.
[0033] The second branch is the spatiotemporal feature extraction branch for UAV video streams. The input is a spatiotemporal sequence tensor formed by optical flow extraction and frame sampling of UAV inspection video stream data, with dimensions [T, H, W, Cu], where T is the number of sampled frames (16 frames in this embodiment), and Cu is the number of image channels and optical flow amplitude channels per frame. This branch adopts an R(2+1)D convolutional structure: first, spatial features are extracted for each frame through a 2D convolutional layer, and then motion information is aggregated between frames through a 1D temporal convolutional layer. Specifically, there are 2 spatial convolutional layers, each with 64 3×3 convolutional kernels; there is 1 temporal convolutional layer with 64 one-dimensional convolutional kernels of length 3; the final output is a temporal feature vector at each spatial grid after global average pooling, with dimensions [H, W, Du], where Du = 128.
[0034] The third branch is the temporal feature branch for terrestrial IoT sensors. The input is the temporal data of each sensor node, including parameters such as temperature, humidity, wind speed, and wind direction. The sensor nodes are spatially interpolated to form a temporal feature matrix corresponding to the grid, with dimensions [H, W, L], where L is the number of sampling steps within the time window (24 steps in this embodiment, corresponding to 2 hours of historical data). This branch uses a gated cyclic unit (GRU)-based sequence encoder with a GRU hidden layer dimension of 64. The output is the hidden state vector of each grid node at the last moment, with dimensions [H, W, Dt], where Dt = 64.
[0035] The specific construction method of the multi-head cross-attention layer in the cross-modal fusion inference module of this embodiment includes: The feature maps output from the three modal branches are concatenated along the channel dimension to obtain the feature tensor before fusion, with dimensions [H,W,Ds+Du+Dt]=[H,W,256].
[0036] For each spatial location, its corresponding 256-dimensional feature vector is taken as the joint modality representation of that location. The multi-head cross-attention mechanism does not use the three modality features as different sequence inputs. Instead, it modifies the weight matrix of the linear projection so that the query vector is primarily generated by satellite hotspot features, while the key and value vectors are jointly generated by UAV video and sensor temporal features (specifically, the UAV and sensor features are concatenated along the channel dimension and projected to obtain the K and V matrices). This allows the attention mechanism to selectively utilize video and sensor temporal information to enhance the discriminative power of satellite hotspot features. The output of each attention head is fed into a feedforward network consisting of two fully connected layers after residual connections and layer normalization. Finally, the fire confidence value is output through a linear layer and a sigmoid function, forming the first fire confidence map.
[0037] The specific hyperparameters used in this embodiment are as follows: 8 attention heads, each with a dimension of 32; 512 dimensions for the intermediate layers of the feedforward network; a batch size of 16 during training; AdamW optimizer; an initial learning rate of 0.0001; and a cosine annealing strategy to decay the learning rate to 1% of the initial value. In the loss function, the weight of positive samples is 10, and the weight of negative samples is 1. This setting was determined through validation set testing.
[0038] The training samples consist of multi-source heterogeneous data from historical real forest fire events and randomly selected fire-free periods. Each sample corresponds to a fixed spatial range (e.g., 10km × 10km) and a fixed time window (e.g., 2 hours). Positive samples are labeled as follows: if any grid within the sample's time window is confirmed to be experiencing a fire, the label for that grid is 1, and the labels for the remaining grids are 0; negative samples have all grids labeled 0. Data augmentation techniques include random horizontal flipping, random pruning, and time-series sliding sampling.
[0039] The total loss L = λ1 × L1 + λ2 × L2 is calculated as the weighted sum of the weighted binary cross-entropy loss function and the Dice loss function; where λ1 = 0.6 and λ2 = 0.4; the weight of positive samples in L1 is set to 10 times the weight of negative samples. The optimizer used is AdamW, with an initial learning rate of 1e... -4 A cosine annealing learning rate decay strategy was adopted. The number of training epochs was set to 200, and the batch size for each iteration was 16. An early stopping strategy was adopted during training, and training was terminated when the validation set loss did not decrease for 10 consecutive epochs.
[0040] S103: In response to the maximum value in the first fire confidence map satisfying the preset threshold condition, the area to be confirmed as a fire is rasterized. The resulting raster is used as a node, the spatial adjacency relationship between adjacent raster is used as the first type of edge, and the connection between the cosine similarity between the multimodal feature vectors corresponding to each raster node is greater than the preset similarity threshold and the Euclidean distance between the two raster nodes in geographic space is less than the preset neighborhood expansion radius is used as the second type of edge. A spatiotemporal heterogeneous graph is constructed based on the first type of edge and the second type of edge.
[0041] In this embodiment, the method for determining the preset threshold condition includes: obtaining a set of historical confidence maps of the target forest area output by the fire inference network in the past M fire seasons, extracting the global maximum value in each historical confidence map to form an extreme value sample sequence; fitting a generalized extreme value distribution to the extreme value sample sequence, and taking the upper 95th percentile of the generalized extreme value distribution as the benchmark threshold; obtaining the regional average combustible moisture content and the number of consecutive rainless days of the target forest area at the current time, constructing an environmental severity index, and linearly scaling the benchmark threshold with the environmental severity index to generate the preset threshold condition.
[0042] A spatiotemporal heterogeneous graph is a graph structure where nodes represent graticules within a fire zone to be confirmed, and edges are of two types: the first type represents spatial adjacency between graticules, and the second type represents connectivity between graticules based on feature similarity and geographic distance. This graph structure can simultaneously capture the spatial continuity and feature correlation of fires. Specifically, spatial adjacency between adjacent graticules is considered the first type of edge. For example, in the two-dimensional graticule index space of a digital elevation model, graticule pairs with eight-neighbor connectivity can be considered adjacent graticules, where eight-neighbor connectivity includes sharing graticule edges or sharing graticule vertices. Furthermore, connections where the cosine similarity between the multimodal feature vectors corresponding to each graticule node is greater than a preset similarity threshold and the Euclidean distance between the two graticule nodes in geographic space is less than a preset neighborhood expansion radius are considered the second type of edge. The second type of edge is an undirected weighted edge added between two graticule nodes that meet the conditions, with its weight being the cosine similarity value. For example, when two grid cells have high similarity in feature vectors and are geographically close, they can be considered to have a potential fire correlation, thus establishing a second type of edge. Therefore, by combining spatial adjacency and feature similarity relationships, a spatiotemporal heterogeneous graph that can characterize the spatial and feature-dimensional correlations of fires can be constructed.
[0043] The method for determining the preset similarity threshold includes: calculating the cosine similarity between each pair of multimodal feature vectors of all grid nodes in the fire area to be confirmed, and plotting the cumulative similarity distribution curve; taking the similarity value corresponding to the point of maximum curvature of the cumulative distribution curve as the initial similarity threshold; calculating the local spatial autocorrelation Moran index of the temperature sequence in the time series data of ground IoT sensors in the fire area to be confirmed, constructing a regulation factor with the Moran index, and lowering the initial similarity threshold to encourage the connection of feature similar edges when the Moran index indicates strong spatial positive correlation; otherwise, raising the initial similarity threshold to suppress noise correlation, and finally generating the preset similarity threshold.
[0044] The method for determining the preset neighborhood expansion radius includes: obtaining the Rosenmeiers fire spread rate reference value of the target forest area under current meteorological conditions; multiplying the reference value by the upper bound of the time alignment error of multi-source heterogeneous data; and using the product as the base radius; obtaining the digital elevation model of the target forest area; calculating the average topographic relief of the area to be confirmed as a fire; and using the average topographic relief to nonlinearly compensate the base radius to generate the preset neighborhood expansion radius. Specifically, the Rosenmeiers fire spread rate reference value is an empirical rate value obtained by combining the initial spread index in the Canadian Fire Risk Rating Scale (CFFDRS) with the types of common combustible materials in the local area, and the unit is meters per minute. In this embodiment, the Rosenmeiers fire spread rate reference value is generated based on real-time meteorological data (wind speed, relative humidity) and combustible material types (divided into four categories: grassland, shrubland, coniferous forest, and broadleaf forest). Specifically: for every 5 km / h increase in wind speed, the base spread rate increases by a factor of 1.3; when the relative humidity is below 30%, the factor increases by 1.5; the maximum spread rate limits for different types of combustibles are: grassland 60 m / min, shrubs 25 m / min, coniferous forests 10 m / min, and broadleaf forests 5 m / min.
[0045] S104: Input the spatiotemporal heterogeneous map and the multimodal feature vectors corresponding to each grid node into the preset graph variational autoencoder network to obtain the mean heat map and uncertainty distribution map of the fire area to be confirmed.
[0046] In this embodiment, the graph variational autoencoder network refers to a deep learning model built based on the principles of graph neural networks and variational autoencoders. This graph variational autoencoder network can process graph-structured data. It maps graph node features to the latent space through the encoder, and then reconstructs the original features from the latent space through the decoder. In the process, it generates a mean heatmap and an uncertainty distribution map to quantify the fire status and the uncertainty of its estimation.
[0047] A mean heatmap is a two-dimensional grid map output by a graph variational autoencoder network, representing the distribution of fire intensity or calorific value in each grid cell within a fire area to be confirmed. This mean heatmap provides spatial average intensity information about the fire. An uncertainty distribution map is a two-dimensional grid map output by a graph variational autoencoder network, representing the uncertainty of fire estimation in each grid cell within a fire area to be confirmed. This uncertainty distribution map quantifies the reliability of the fire estimation; higher uncertainty indicates a less reliable estimation result.
[0048] The graph variational autoencoder network module consists of a graph encoder and a graph decoder; the graph encoder employs a two-layer graph convolutional network (GCN). The input to the first GCN layer is the node feature matrix. (N is the number of grid nodes, feature dimension 64) and adjacency matrix (Weighted adjacency matrix of first-class and second-class edges). The first layer output dimension is 32, and the activation function is ReLU; the second layer GCN outputs two N×16 matrices in parallel, representing the latent variable mean μ and log-variance, respectively. Reparameterization techniques are used to transform Gaussian distributions. Let Z be the sampled latent variable, where I represents the identity matrix, and the covariance matrix representing the posterior distribution of the latent variable is an isotropic diagonal matrix, i.e., it is assumed that each dimension of the latent variable space is independent and the variance is σ. 2 .
[0049] The graph decoder employs an inner product decoder, which reconstructs the edge probabilities of the adjacency matrix by calculating the inner product of every pair of nodes in the latent variable Z and applying it to the Sigmoid function. Simultaneously, it reconstructs the node feature matrix from Z using a two-layer fully connected network (16-32-64). The decoder output includes the reconstructed adjacency probability matrix and the reconstructed feature matrix.
[0050] Starting from the latent variable Z, two independent linear layers are connected in parallel to output the fire heat value (which is activated by Softplus to ensure non-negativity) and the prediction uncertainty (which is activated by Sigmoid and mapped to the [0,1] interval) respectively, forming a mean heat map and an uncertainty distribution map.
[0051] The loss function includes the evidence lower bound loss, the prior constraint domain loss, and the uncertainty regularization loss.
[0052] Among them, the loss of the lower boundary of evidence .
[0053] in, The adjacency matrix reconstruction loss is a binary cross-entropy, used to measure the difference between the reconstructed adjacency probability matrix output by the graph decoder and the true adjacency matrix of the input spatiotemporal heterogeneous graph. The node feature reconstruction loss measures the difference between the node feature matrix reconstructed by the graph decoder from the latent variable Z and the input multimodal feature vector matrix. This is the Kolbec-Leibler divergence term, used to measure the approximate posterior distribution. With prior distribution The difference between them serves to regularize the latent variable space, prevent overfitting, and encourage latent variables to follow the desired prior distribution; The posterior distribution of the latent variables of each node output by the graph encoder is a diagonal Gaussian distribution parameterized by the mean and log-variance output of the second layer GCN of the graph encoder. The prior distribution is a standard normal distribution N(0,I), which assumes that the latent variables of each node are independent and follow a standard Gaussian distribution.
[0054] Prior constraint domain loss The calculation is based on the prior constraint domain generated by the semi-physical fire behavior spread model in the following embodiments. For nodes within the prior constraint domain, if the fire heat value of their decoded output is lower than the first threshold, a penalty term is applied; for nodes outside the prior constraint domain, if the fire heat value is higher than the first threshold, a penalty term is applied.
[0055] Uncertainty Regularization Loss Encourage the uncertainty distribution map to align with the region of reconstruction error, and use negative log-likelihood as a guide.
[0056] Total loss is , where, η=0.2, μ=0.05.
[0057] The training data consists of spatiotemporal heterogeneous maps of historical cases of unconfirmed fire areas selected by the fire inference network. Each training sample corresponds to a complete evolution process of an unconfirmed fire area. The Adam optimizer is used with a learning rate of 1e^(-1 / 2). -3 Weight decay 5e -4 The training run consisted of 300 epochs with a batch size of 32 graphs. All GCN layers in the graph encoder were initialized using Xavier. After training, the model's forward propagation simultaneously outputs both the mean heatmap and the uncertainty distribution map.
[0058] The spatiotemporal heterogeneous graph node features of the input graph variational autoencoder network are directly taken from the 64-dimensional multimodal feature vector output by the multi-source feature encoding network, and the edge weights are calculated based on the cosine similarity of the node features. Each grid value of the output mean heatmap corresponds to the expected fire thermal radiation intensity inferred by Bayesian posterior, and each grid value of the uncertainty distribution map corresponds to the variance normalization measure of the posterior distribution.
[0059] S105: Generate the geometric envelope of fire smoke diffusion based on the fire intensity parameters of each grid calculated from the mean heat map, local meteorological grid data, and semi-physical fire behavior spread model.
[0060] In this embodiment, the geometric envelope of fire smoke diffusion is calculated based on the fire spread model and the smoke diffusion model, and is used to describe the geometric boundary of the spatial diffusion range of fire smoke. This geometric envelope of fire smoke diffusion is used to predict the area affected by the smoke.
[0061] Specifically, based on the heat value distribution in the mean heat map, the fire line intensity parameters of each grid are estimated using empirical formulas or physical models. Simultaneously, local meteorological grid data within the target forest area, such as wind speed, wind direction, temperature, and humidity, are acquired. These parameters are then combined with a pre-defined semi-physical fire behavior spread model to predict the direction and speed of fire propagation. Based on the predicted fire spread results and combined with meteorological data, a Gaussian plume diffusion model is further used to simulate the spatial diffusion process of the smoke stream, ultimately generating the geometric envelope of the fire smoke stream diffusion.
[0062] Specifically, the pre-defined semi-physical fire behavior spread model uses the modified Rothermel surface fire spread model as its core mechanism and couples physical constraints of terrain slope and wind field. The specific input-output mapping relationship of this model is as follows: The input feature vector (for each grid cell) includes: fire line intensity parameter (kW / m), derived from the radiative heat flux derived from the mean heat map and converted using the Byram fire line intensity equation; local surface combustible moisture content (%), derived from grid data based on satellite spectral index inversion or ground sensor interpolation; terrain slope (°) and aspect (°), calculated from the spatial first derivative of the digital elevation model; local wind speed (m / s) and wind direction (°), derived from high-resolution numerical weather prediction grid fields or UAV-measured wind field interpolation; and combustible material type code, coniferous forest, broadleaf forest, shrubland, or grassland, used to index the parameter lookup table.
[0063] The output parameter vector includes: the fire spread rate vector (m / min) and the initial fire spread prediction range; wherein, the fire spread rate vector (m / min) is a scalar of spread rate and direction angle that includes the direction of the maximum gradient along the ground surface.
[0064] S106: Overlay the mean heat map and the uncertainty distribution map onto the geographic information system base map, and calculate the spatial overlap area between the area with a heat value greater than the first threshold in the mean heat map and the geometric envelope of the fire smoke diffusion, as a composite high-risk area.
[0065] In this embodiment, the composite high-risk area refers to the area where the heat value is greater than a first threshold and the geometric envelope of the fire smoke diffusion are spatially overlapped after the mean heat map and the uncertainty distribution map are overlaid on the geographic information system base map. This composite high-risk area integrates the thermal intensity and smoke diffusion effects of the fire and is used to indicate the area with the highest risk under the current fire situation.
[0066] Specifically, the geographic information system (GIS) base map provides background information such as topography, vegetation, and roads in the forest area. Overlaying the mean heat map and uncertainty distribution map as transparent layers on top of this provides a visual representation of the fire's heat intensity and estimation uncertainty. Areas with heat values exceeding a first threshold are considered active fire areas. Then, the spatial intersection of these active fire areas with the geometric envelope of the fire smoke diffusion is calculated. This intersection area is then identified as a composite high-risk area.
[0067] The method for determining the first threshold includes: obtaining the set of mean heatmaps output by the graph variational autoencoder network during the training phase when performing leave-one-out validation on historical fire cases; extracting the heat values of the grids corresponding to real fire points as positive samples and extracting the heat values of the grids in fire-free areas as negative samples; determining the optimal heat value segmentation point on the receiver operating characteristic curve with the goal of maximizing the Youden index, as the benchmark first threshold; obtaining the average uncertainty of the fire-to-be-confirmed area in the current time uncertainty distribution map, and correcting the benchmark first threshold with the average uncertainty; when the average uncertainty of the area is high, reducing the first threshold to reduce the missed fire detection caused by model uncertainty, thus generating the first threshold.
[0068] As can be seen from the above, this application spatially registers and temporally aligns three types of heterogeneous data—satellite, UAV, and terrestrial IoT—solving the feature fusion bias problem caused by spatiotemporal inconsistencies in multi-source data. The generated multimodal feature vectors can comprehensively characterize the fire features of different dimensions in forest areas. Secondly, a fire inference network quickly filters out fire areas to be confirmed, significantly narrowing the scope of subsequent processing. Then, a spatiotemporal heterogeneous graph including spatial adjacency edges and feature similarity edges is constructed, preserving the continuity of geographic space while capturing the fire correlation between non-adjacent areas. Furthermore, a graph variational autoencoder simultaneously outputs a mean heatmap and an uncertainty distribution map, achieving probabilistic estimation of fire status rather than a single deterministic judgment. Finally, a semi-physical fire behavior model is combined to generate a smoke flow diffusion envelope, which is superimposed with high-heat areas to obtain a composite high-risk area, improving the comprehensiveness and accuracy of fire risk assessment.
[0069] In one embodiment of this application, a forest fire monitoring method based on multi-source heterogeneous data fusion further includes: Spatial cluster analysis is performed on the uncertainty of each grid in the uncertainty distribution map, and connected regions with an average regional uncertainty greater than a preset upper limit threshold are taken as target regions for supplementary sampling; the average regional uncertainty is the area-weighted average of the uncertainties of each grid in the connected region; Based on the geographical coordinates, topographic elevation data, and geometric envelope of the fire smoke diffusion in the target area for remediation mining, a remediation reconnaissance path planning instruction is generated. Control at least one mobile reconnaissance platform to collect supplementary data based on supplementary reconnaissance path planning instructions; The pre-set multi-source feature coding network is used to perform feature-level fusion coding on the supplementary data and the original multi-source heterogeneous data in the supplementary target area, so as to update the multi-modal feature vectors corresponding to each grid node in the supplementary target area. The original multimodal feature vectors of each grid node in the target area of the supplementary sampling are replaced with the updated multimodal feature vectors, while keeping the feature vectors of other grid nodes unchanged; Based on the updated multimodal feature vectors, the connection relationships and weights of the second type of edges are recalculated, and the spatiotemporal heterogeneous graph is reconstructed. The updated multimodal feature vectors and the reconstructed spatiotemporal heterogeneous graph are re-input into the preset graph variational autoencoder network to iteratively update the mean heat map, uncertainty distribution map, fire smoke diffusion geometric envelope, and composite high-risk area markers. Repeat the supplementary sampling reconnaissance and iterative update steps until the area-weighted average of the uncertainties of each grid in the supplementary sampling target area in the uncertainty distribution map is less than the preset convergence threshold, or the preset maximum number of supplementary sampling iterations is reached.
[0070] In this embodiment, spatial clustering analysis aims to identify continuous regions with high uncertainty for targeted data supplementation. This spatial clustering analysis employs a density-based clustering algorithm, identifying high-density uncertainty regions by setting a neighborhood radius and a minimum sample size. The method for determining the preset upper limit threshold includes: obtaining the uncertainty distribution map output by the graph variational autoencoder network on the validation set; calculating the upper quartile of the uncertainty distribution of the actual fire line boundary grid in historical fire cases; using this upper quartile as the initial value of the preset upper limit threshold; obtaining the median of the global uncertainty distribution map from the previous iteration of the current supplementary data collection; if this median is higher than the historical normal level, lowering the initial value to expand the supplementary data collection range, thus generating the preset upper limit threshold.
[0071] The calculation of regional average uncertainty is a comprehensive assessment of the uncertainty level of the entire connected region. For example, a weighted average based on the grid area size can be used to more accurately characterize the overall uncertainty of the region.
[0072] This embodiment provides a safe, efficient, and effective reconnaissance route for mobile reconnaissance platforms to cover target areas. The path planning algorithm uses the A* search algorithm, incorporating terrain elevation data as part of the cost function. Simultaneously, it uses the geometric envelope of fire smoke diffusion as a dynamic obstacle to plan reconnaissance routes that avoid areas with high smoke concentrations.
[0073] The mobile reconnaissance platform in this embodiment includes drones, ground robots, or manned aircraft. For example, for drones, preset waypoints and flight altitude commands, combined with GPS positioning and inertial navigation systems, enable them to automatically fly along a planned reconnaissance route and activate onboard sensors to collect data. These onboard sensors include visible light cameras, infrared thermal imagers, and gas sensors. For ground robots, LiDAR or visual SLAM technology is used to autonomously avoid obstacles in complex terrain and move along the planned reconnaissance route while simultaneously collecting data from ground sensors.
[0074] In this embodiment, the update operation can directly replace the old feature vector with the feature vector obtained by encoding the newly collected data, or use a weighted average method to fuse the new and old features for a smooth transition.
[0075] Since the supplementary data updates the multimodal feature vectors of grid nodes within the target area, changes in these feature vectors affect the similarity relationships between nodes, necessitating a reassessment of the connections and weights of second-type edges. For example, the cosine similarity between the updated multimodal feature vectors can be recalculated, and new connections and weights can be determined based on a preset similarity threshold and neighborhood expansion radius. Alternatively, a time decay factor can be incorporated during recalculation, making the impact of new data greater on connections while the influence of old data gradually weakens, thus characterizing the dynamic changes in the fire situation.
[0076] This embodiment inputs the latest data representation and graph structure into a graph variational autoencoder network, enabling the network to relearn and infer fire conditions. For example, the graph variational autoencoder network utilizes updated node features and edge relationships, and through its encoder and decoder structure, generates new mean heatmaps and uncertainty distribution maps. Based on these updated fire conditions, fire line intensity parameters are further recalculated, and combined with local meteorological grid data and a semi-physical fire behavior spread model, updated geometric envelopes for fire smoke diffusion are generated, ultimately updating the composite high-risk area markers.
[0077] The process of re-collecting reconnaissance and iteratively updating is repeated until the area-weighted average of the uncertainties of each grid within the target area in the uncertainty distribution map is less than a preset convergence threshold, or the preset maximum number of re-collection iterations is reached. This termination condition ensures the effectiveness and efficiency of the iteration process. For example, when the uncertainty of the target area is reduced to a sufficiently low level (less than the convergence threshold), it indicates that the fire information in that area is sufficiently clear, and no further re-collection is needed, thus saving resources. On the other hand, setting a maximum number of re-collection iterations can prevent infinite loops, ensuring that the system provides the final result within a finite time. Even if the uncertainty in some areas does not fully converge, the real-time response capability of the system can still be guaranteed. The method for determining the preset convergence threshold includes: analyzing the decay curve of the grid uncertainty within the target area as a function of the number of re-collection iterations, fitting an exponential decay model, and taking 1.2 times the uncertainty value corresponding to the stable period of the decay model as the preset convergence threshold. The preset maximum number of re-collection iterations is dynamically determined by the ratio of the remaining endurance energy constraint of the mobile reconnaissance platform to the average time consumed per re-collection reconnaissance.
[0078] Through the above technical solution, this application effectively solves the problem that in forest fire monitoring methods based on multi-source heterogeneous data fusion, certain areas of the uncertainty distribution map have high uncertainty, leading to insufficient accuracy and reduced reliability of monitoring results. This application can improve the accuracy and reliability of forest fire monitoring, providing a more solid data foundation and decision support for accurate assessment of the fire situation and firefighting decisions, thereby effectively enhancing the ability to prevent and control forest fires.
[0079] In one embodiment of this application, a reconnaissance path planning instruction for re-mining is generated based on the geographical coordinate range, topographic elevation data, and geometric envelope of the fire smoke diffusion of the target area, including: The vector boundary of the target area to be supplemented is buffered outward by a preset distance to generate an initial security reconnaissance envelope; Identify terrain feature lines within the target area for re-mining and within the initial security reconnaissance envelope; terrain feature lines include ridgelines and canyon lines; In response to the fact that the angle between the main diffusion axis of the fire smoke diffusion geometric envelope obtained in the most recent iteration update before the supplementary reconnaissance and the direction of the terrain feature line is less than the preset flow direction coincidence threshold, the reconnaissance path segment planned along the terrain feature line is regarded as a high smoke exposure risk segment. Spatial overlay analysis is performed on the path nodes corresponding to high smoke flow exposure risk sections and the uncertainty maximum points in the uncertainty distribution map to generate a detour blind spot observation point sequence; The sequence of detour-filling observation points indicates the lateral penetration reconnaissance position of the mobile reconnaissance platform in the smoke-covered area at the bottom of the canyon from a high terrain sheltered location on the canyon flank.
[0080] In this embodiment, when generating the supplementary reconnaissance path planning instruction, the vector boundary of the supplementary reconnaissance target area is first buffered outward by a preset distance to generate an initial safe reconnaissance envelope. This provides a safe operating area for the mobile reconnaissance platform, preventing it from getting too close to the core or edge of the fire area, thereby reducing potential dangers.
[0081] The preset distance is dynamically determined based on the ratio of the mobile reconnaissance platform's flight speed to the average horizontal smoke transport velocity near the ground along the geometric envelope of the fire smoke diffusion. This ensures the reconnaissance platform has at least twice the safe evacuation reaction time when encountering a sudden smoke surge. For example, the preset distance is set to 50 to 100 meters for small drones and 200 to 500 meters for large drones or manned aircraft. Alternatively, this distance can be adaptively adjusted based on the real-time predicted range of the fire smoke diffusion geometric envelope, ensuring the reconnaissance platform remains in an area with minimal smoke impact.
[0082] The terrain feature lines in this embodiment include ridgelines and canyon lines. Terrain feature lines are an important component of the landform and influence the movement of airflow and smoke. Terrain feature lines are obtained by performing terrain analysis on digital elevation model data; for example, ridgelines and canyon lines are extracted using algorithms based on terrain factors such as slope, aspect, and curvature.
[0083] The main diffusion axis is obtained through principal component analysis or directional statistics of the geometric envelope of the smoke diffusion in the fire area. The method for determining the preset flow direction coincidence threshold includes: obtaining the atmospheric boundary layer turbulence intensity parameters of the target forest area under current meteorological conditions; using the turbulence intensity parameters as input to query a preset lookup table function; the lookup table function is obtained by parametrically modeling the smoke deflection angle of typical canyon terrain under different turbulence intensities using computational fluid dynamics simulation; the output is the terrain-induced lateral diffusion half-angle of the smoke; this lateral diffusion half-angle is used as the preset flow direction coincidence threshold. For example, when the angle is less than 15 degrees or 30 degrees, a high coincidence risk is considered to exist. When the smoke diffusion direction is highly consistent with the terrain feature line, the area corresponding to that terrain feature line will become a natural channel for the smoke, and reconnaissance platforms flying in this area will face a high risk of smoke exposure.
[0084] Spatial overlay analysis is achieved through the spatial analysis function of a geographic information system. For example, geometric data of high smoke flow exposure risk sections are overlaid with point data of uncertainty maximum points to identify points that are both located in high-risk areas and correspond to high uncertainty.
[0085] This embodiment performs spatial overlay analysis on the path nodes corresponding to high smoke flow exposure risk sections and the uncertainty maxima points in the uncertainty distribution map to identify geographical locations that are both near the path nodes of high smoke flow exposure risk sections and correspond to the uncertainty maxima. Using these locations as centers, candidate observation points that can cover the smoke flow area at the bottom of the canyon are searched in the canyon flank area where the terrain obscuration is higher than a preset safety threshold, thereby generating a detour-filling observation point sequence. The detour-filling observation point sequence is used to indicate the lateral penetration reconnaissance position of the mobile reconnaissance platform on the smoke flow coverage area at the bottom of the canyon in the high terrain obscuration area on the canyon flank.
[0086] Through the above technical solution, this application effectively addresses the problems of high smoke exposure risk and low data collection efficiency faced by reconnaissance platforms when conducting forest fire reconnaissance in complex terrain. By dynamically assessing the interaction between smoke diffusion and terrain features, this application can intelligently identify high-risk areas and plan reconnaissance routes that both avoid risks and effectively cover areas with high uncertainty. Utilizing the high terrain shielding on the canyon flanks for lateral penetration reconnaissance allows the mobile reconnaissance platform to acquire data from the smoke-covered area at the bottom of the canyon while ensuring its own safety, thereby improving the quality and efficiency of data collection.
[0087] In one embodiment of this application, multimodal feature vectors are input into a preset fire inference network to perform cross-modal fire risk inference, generate a first fire confidence map, and extract the fire area to be confirmed, including: Based on the preset multidimensional fire risk benchmark field model, cross-modal correlation verification is performed on multi-source heterogeneous data to generate a multidimensional fire risk prior feature map. After fusing the multidimensional fire risk prior feature map with the multimodal feature vector, the feature map is input into the fire inference network to obtain the first fire confidence map. Based on the connected regions in the first fire confidence map where the confidence level is greater than the second threshold, the fire areas to be confirmed are extracted.
[0088] In this embodiment, the pre-defined multidimensional fire risk benchmark model provides an objective and comprehensive fire risk assessment benchmark for comparison and verification with actual monitoring data to identify anomalies or inconsistencies. Cross-modal correlation verification refers to the mutual verification and comparison of data from different data sources (modalities) to ensure their consistency in spatial, temporal, or physical attributes. Its function is to discover and correct inconsistencies, noise, or errors in multi-source heterogeneous data, improving the reliability and accuracy of the data. This verification process can be implemented through statistical methods, such as calculating the correlation and differences of different modal data in the same area or time period and setting thresholds for judgment; or through physical consistency verification, such as comparing the brightness temperature value in satellite hotspot monitoring data with the background brightness temperature value predicted by the benchmark model, or comparing the smoke flow direction and local wind vector in UAV inspection video stream data.
[0089] A multidimensional fire hazard prior feature map refers to a spatial distribution map that, after cross-modal correlation verification, integrates information from multiple fire hazard factors and characterizes the probability or intensity of fire occurrence. This multidimensional fire hazard prior feature map contains more reliable fire hazard prior knowledge than a single data source. Its role is to provide validated and enhanced prior information for subsequent fire inference networks, guiding them to more accurately identify fires. This multidimensional fire hazard prior feature map is multi-channel raster data, with each channel representing a fire hazard factor (temperature anomaly, smoke concentration, combustible material humidity, etc.), and has undergone normalization processing; alternatively, the multidimensional fire hazard prior feature map is a fire hazard index map, represented by a weighted sum or fire hazard probability value output by a machine learning model.
[0090] Feature fusion combines validated prior knowledge (multidimensional fire risk prior feature maps) with features from raw observation data (multimodal feature vectors) to provide richer and more reliable input to the fire inference network. This fusion process can be achieved by simply stacking the two feature maps along the channel dimension.
[0091] The second threshold is used to filter areas with a high probability of fire in the first fire confidence map. Its function is to filter out noisy or uncertain areas with low confidence, focusing on high-confidence fire candidate areas. The method for determining the second threshold includes: performing Gaussian mixture model decomposition on the first fire confidence map, treating the confidence distribution as a superposition of background noise distribution and potential fire signal distribution; estimating the parameters of the two Gaussian components using the expectation-maximization algorithm, and taking the confidence value corresponding to the intersection of the probability density functions of the background noise distribution component and the potential fire signal distribution component as the initial second threshold; obtaining the spatial density of brightness-temperature anomaly pixels in the current satellite hotspot monitoring data, and fine-tuning the initial second threshold with the spatial density; the higher the density of brightness-temperature anomaly pixels, the higher the second threshold is to suppress over-segmentation.
[0092] Connected component extraction refers to combining all spatially adjacent (e.g., four-neighbor or eight-neighbor) grid points with a confidence level greater than a second threshold into a single independent region in a raster image. Its purpose is to aggregate discrete, high-confidence grid points into meaningful fire candidate regions, facilitating subsequent processing and analysis. This extraction process employs a connected component labeling algorithm from image processing to process the binarized confidence map.
[0093] A fire area to be confirmed refers to a candidate fire area that requires further verification and confirmation after threshold filtering and connected component extraction. This fire area can be represented as a set of grid points, or a polygonal vector data indicating its geographical extent; alternatively, it can be a binary mask where the pixel value of the fire area to be confirmed is 1, and other areas are 0.
[0094] Through the above technical solution, before inputting the multimodal feature vectors into the fire inference network, this application first performs cross-modal correlation verification on multi-source heterogeneous data based on a preset multi-dimensional fire hazard benchmark field model, generating a multi-dimensional fire hazard prior feature map. This verification process can effectively identify and correct inconsistencies or noise that may exist in the multi-source heterogeneous data, thereby generating more reliable and accurate fire hazard prior information. Subsequently, the verified multi-dimensional fire hazard prior feature map is fused with the original multimodal feature vectors, enabling the fire inference network to receive a more comprehensive and discriminative input that integrates prior knowledge and original observation data. This fusion mechanism enhances the completeness and synergy of the input information, thereby enabling the fire inference network to generate a more accurate first fire confidence map, effectively reducing the risk of fire misjudgment, and further improving the reliability and accuracy of overall fire monitoring.
[0095] In one embodiment of this application, cross-modal correlation verification is performed on multi-source heterogeneous data based on a preset multi-dimensional fire hazard reference field model to generate a multi-dimensional fire hazard prior feature map, including: The satellite hotspot monitoring data and UAV inspection video stream data from the multi-source heterogeneous data were physically verified against the multi-dimensional fire risk reference field model to obtain the verified grid. The joint verification confirmation area is determined based on the spatial distribution of the verified graticule; The initial fire risk value output by the multidimensional fire risk benchmark model is weighted and corrected based on the joint verification confirmation area to generate a multidimensional fire risk prior feature map.
[0096] In this embodiment, by physically verifying the consistency between satellite hotspot monitoring data and UAV inspection video stream data in multi-source heterogeneous data and the multi-dimensional fire risk reference field model, invalid data affected by environmental noise, sensor failure, or non-fire heat sources can be effectively filtered out, so that the original data on which subsequent processing depends has high authenticity and reliability.
[0097] Based on this, a joint verification confirmation area is determined according to the spatial distribution of the verified graticules. This step aims to further enhance the confidence of fire assessment through cross-verification of multi-source data. One approach is to perform spatial overlay analysis on all graticules that have passed physical consistency verification, and determine the overlapping areas that have been verified by both satellite hotspot monitoring data and UAV inspection video stream data as the joint verification confirmation area. Another approach is to use spatial proximity analysis; if the graticules verified by satellite and those verified by UAV are geographically adjacent (e.g., within a preset distance threshold) and have similar fire indication intensities, these adjacent areas are merged into a joint verification confirmation area.
[0098] Furthermore, the initial fire hazard value output by the multidimensional fire hazard benchmark model is weighted and corrected based on the joint verification confirmation region to generate a multidimensional fire hazard prior feature map. This step aims to incorporate high-confidence information rigorously validated by multi-source data into the fire hazard assessment to optimize the initial fire hazard prediction. Specifically, a Bayesian update mechanism is used, employing the initial fire hazard value output by the multidimensional fire hazard benchmark model as the prior probability, and the information from the joint verification confirmation region as observational evidence. The posterior fire hazard probability is then updated using the Bayesian formula, thereby generating the multidimensional fire hazard prior feature map.
[0099] Through the aforementioned technical solution, this application effectively eliminates false alarms or missed alarms that may exist from a single data source by independently verifying the physical consistency of satellite hotspot monitoring data and UAV inspection video stream data from multi-source heterogeneous data, thereby improving the reliability of the original data. Subsequently, based on the spatial distribution of these verified graticules, a joint verification confirmation area is determined, realizing cross-modal data mutual verification, which greatly enhances the accuracy and robustness of fire risk assessment. Finally, this high-confidence joint verification confirmation area is used to weight and correct the initial fire risk value output by the multi-dimensional fire risk benchmark model, enabling the generated fire risk prior feature map to more accurately represent the actual fire situation and fully utilize the inherent correlation of multi-source data.
[0100] In one embodiment of this application, the physical consistency verification of satellite hotspot monitoring data and UAV inspection video stream data with a multidimensional fire risk reference field model is performed, including: Calculate the brightness temperature difference between the real-time brightness temperature value of each thermal anomaly pixel in the satellite hotspot monitoring data and the background brightness temperature value of the corresponding grid in the multidimensional fire risk reference field model; In response to a brightness temperature deviation exceeding a preset dynamic threshold and a static terrain occlusion index corresponding to a thermal anomaly pixel satisfying a preset visibility condition, the geographic raster corresponding to the thermal anomaly pixel is designated as a first-class verification-passing raster; the static terrain occlusion index is the target point visibility coverage calculated based on the digital elevation model. Optical flow field analysis was performed on the video stream data of UAV inspection to extract the angle between the main direction of the smoke flow motion vector field and the local wind vector. In response to the direction angle being less than a preset angle threshold and the static terrain occlusion index corresponding to the optical flow region being less than a preset vortex determination threshold, the geographic grid corresponding to the optical flow region is regarded as a second type of verification grid. The spatial overlap area or the spatial neighboring area (meeting the preset proximity distance threshold) between the first-type verified grid and the second-type verified grid is determined as the joint verification confirmation area.
[0101] In this embodiment, the difference between the real-time brightness temperature value and the background brightness temperature value is calculated, and the difference is divided by the background brightness temperature value to obtain the brightness temperature deviation. This brightness temperature deviation is a normalized brightness temperature deviation, which is used to adapt to the comparison of thermal anomaly intensity under different background brightness temperature levels.
[0102] A preset dynamic threshold is used to determine whether the brightness-temperature deviation reaches the standard for fire assessment. Its dynamic adjustment mechanism can adapt to different environmental conditions. Specifically, the method for determining the preset dynamic threshold includes: constructing a two-dimensional lookup table, where each cell stores the mean and standard deviation of the brightness-temperature deviation for this type of land surface under historical clear sky conditions within the corresponding irradiance range; at the current time, based on the land cover type and solar irradiance of the thermal anomaly pixel, the corresponding cell is queried, and the mean plus three times the standard deviation is used as the preset dynamic threshold.
[0103] The Static Terrain Obscuration Index is a target point visibility coverage calculated based on a digital elevation model (DEM) and used to assess the visibility of locations with thermal anomaly pixels. This index determines the percentage of visible area from the thermal anomaly pixel towards a satellite sensor or potential fire source through viewpoint analysis of the DEM. Preset visibility conditions define the pass / fail criteria for this Static Terrain Obscuration Index.
[0104] Optical flow field analysis was performed on the video stream data from UAV inspections to extract the motion characteristics of the smoke flow. The optical flow field analysis employed a deep learning-based optical flow network to more robustly estimate the smoke flow motion vector. The principal direction of the smoke flow motion vector field was determined by averaging, mode statistics, or principal component analysis of all smoke flow motion vectors within a local area. Local wind vectors were obtained from data from nearby ground meteorological stations, meteorological model interpolation data, or directly measured by sensors (anemometers) onboard the UAV, with necessary corrections applied.
[0105] A preset angle threshold is used to determine the consistency between the smoke flow direction and the local wind vector direction, in order to exclude smoke or interference caused by non-fire conditions. The method for determining the preset angle threshold includes: obtaining the angular variance of the smoke flow motion vector in multiple consecutive frames of UAV inspection video stream data, constructing a fuzzy membership function based on the angular variance, and outputting a dynamic angle threshold between 15° and 45°. The more stable the smoke flow direction, the stricter the threshold.
[0106] The preset vortex detection threshold is used to identify vortex areas that may be caused by the terrain, avoiding misjudging smoke streams caused by terrain vortices as fires. This preset vortex detection threshold is determined by the curl amplitude distribution of the topographic potential gradient vector of the target forest area. Specifically, it is determined by normalizing the variance of the terrain-induced vertical velocity output by the numerical weather prediction model of the area to be verified in the past hour, and taking the upper 80th percentile as the preset vortex detection threshold.
[0107] Through the above technical solution, this application effectively solves the problem of misjudgment and omission caused by ignoring terrain shading and dynamic environmental factors in the physical consistency verification process.
[0108] In one embodiment of this application, a spatiotemporal heterogeneous graph is input into a graph variational autoencoder network to obtain a mean heatmap and an uncertainty distribution map of the fire area to be confirmed, including: The graph variational inference model in the graph variational autoencoder network is used to perform feature fusion on spatiotemporal heterogeneous graphs. The graph variational inference model takes the multimodal feature vectors of each grid node as input. In the encoder stage, the adjacency information represented by the first and second types of edges is aggregated through a multi-layer graph convolutional network to generate the latent variable mean and latent variable variance of each grid node. Then, the posterior probability distribution representing the fire status of each grid is calculated through Bayesian variational inference. In the variational inference process, the spatial constraint information generated based on the physical mechanism of forest fire spread is used to adjust the distribution of latent variables, and the mean heat map and uncertainty distribution map are obtained based on the posterior probability distribution.
[0109] In this embodiment, the graph variational autoencoder network maps graph data to a latent space through an encoder, and then reconstructs the graph data from the latent space through a decoder. Simultaneously, it learns the probability distribution of the data, thereby learning a latent representation of fire status from complex spatiotemporally heterogeneous graphs and quantifying its uncertainty. Specifically, the graph variational autoencoder network employs an encoder based on a graph convolutional network and a decoder based on a multilayer perceptron. The encoder outputs the mean and variance vector of each node, while the decoder reconstructs node features or graph structure from the sampled latent variables.
[0110] The graph variational inference model is a component of the graph variational autoencoder network, responsible for performing variational inference on graph-structured data. It infers latent variables through approximate posterior distributions, thereby learning the data generation process. This is used to extract deep features of fire status from spatiotemporally heterogeneous graphs and represent the uncertainty of fire status in the form of a probability distribution. This graph variational inference model employs mean-field variational inference, assuming that latent variables are independent to simplify the calculation of the posterior distribution.
[0111] Specifically, the graph variational inference model uses the latent variable mean and log-variance output by the graph encoder as a basis to generate the probability distribution of the fire status of each grid node through approximate posterior inference. The detailed structure and inference process of this model are as follows: The graph variational inference model comprises a latent variable sampling module, a posterior probability decoding module, and a physical constraint adjustment module. The latent variable sampling module employs a reparameterization technique to sample the latent variables of each grid node from a diagonal Gaussian distribution parameterized by the mean and log-variance of the latent variables output by the graph encoder. Specifically, it first randomly samples a noise vector with the same dimension as the latent variable from a standard normal distribution. This noise vector is then multiplied element-wise by the standard deviation vector calculated from the log-variance. Finally, the product is added element-wise to the mean vector, and the result is the latent variable. This reparameterization process makes the random sampling operation differentiable with respect to the model parameters, thus supporting end-to-end backpropagation training.
[0112] The posterior probability decoding module inputs the sampled latent variables from each node into a two-layer fully connected decoding network. The first layer maps the latent variables from the original dimension to an intermediate feature space with twice the original dimension, using a modified linear unit as the activation function. The second layer maps the intermediate features to a two-dimensional output, where one dimension, after being processed by a soft addition activation function, serves as the mean heatmap value of the grid, representing the posterior expectation of the fire radiation intensity; the other dimension, after being activated by a sigmoid function, serves as the uncertainty of the grid, representing the variance normalization measure of the posterior estimate. Arranging the mean heatmap values of all grids according to their spatial location forms the mean heatmap, and arranging the uncertainties of all grids according to their spatial location forms the uncertainty distribution map.
[0113] The physical constraint adjustment module introduces spatial constraint information based on the physical mechanism of forest fire spread during variational inference. Specifically, it obtains the initial fireline spread prediction range generated by the semi-physical fire behavior spread model based on the prior motion vector of fire head spread, and uses the set of grid nodes covered by this prediction range as the prior constraint domain. For grid nodes within the prior constraint domain, it is expected that their mean thermal value is not lower than a preset first threshold; for grid nodes outside the prior constraint domain, it is expected that their mean thermal value is not higher than a preset second threshold. A corresponding penalty term is constructed in the loss function, increasing the loss value when the above expected conditions are not met, thereby forcing the posterior fire thermal value output by the model to be consistent with the physical spread prior in spatial distribution, suppressing posterior estimates that conflict with the spread prior.
[0114] The inference objective of the graph variational inference model is to approximate the true posterior distribution of fire conditions, which is difficult to calculate directly. The model assumes that the latent variables of each grid node are independent and that the latent variables of each node follow a Gaussian distribution parameterized by the mean and variance of the encoder output. Under this assumption, the model parameters are optimized by maximizing the lower bound of evidence.
[0115] The lower bound of evidence consists of three parts: the first part is the adjacency matrix reconstruction loss, which measures the difference between the reconstructed adjacency matrix output by the decoder and the true adjacency matrix of the input spatiotemporal heterogeneous graph, using binary cross-entropy loss; the second part is the node feature reconstruction loss, which measures the difference between the node feature vector reconstructed by the decoder from the latent variables and the input multimodal feature vector, using mean squared error loss; the third part is the Kolb-Leibler divergence term, which measures the difference between the approximate posterior distribution and the standard normal prior distribution, serving to regularize the latent variable space. Based on the above lower bound of evidence loss, a penalty term generated by the aforementioned physical constraint adjustment module is superimposed to form the total loss function for model training.
[0116] After model training is complete, the forward propagation process of the graph variational inference model is as follows: The spatiotemporal heterogeneous graph (containing multimodal feature vectors and adjacency matrices of each grid node) constructed within the fire area to be confirmed is input into the graph encoder. The graph encoder aggregates the adjacency information represented by the first and second types of edges through a two-layer graph convolutional network, outputting the latent variable mean and log-variance of each grid node. The latent variable sampling module generates latent variables for each node based on the mean and log-variance. The posterior probability decoding module maps the latent variables of each node to the corresponding mean heatmap value and uncertainty. The mean heatmap values of each grid are arranged according to their spatial grid positions to generate a mean heatmap. The uncertainties of each grid are arranged according to their spatial grid positions to generate an uncertainty distribution map.
[0117] In a variational autoencoder, the latent variable mean and variance are two vectors output by the encoder, representing the mean and variance of the probability distribution of each node (or data point) in the latent space, respectively. The mean describes the central location of the latent representation, while the variance describes its uncertainty or dispersion. They are used to define the approximate posterior distribution of the fire status of each grid node in the latent space and are the foundation of Bayesian variational inference. The encoder outputs two independent linear layers, generating the mean vector and the log-variance vector, respectively.
[0118] Spatial constraint information generated from the physical mechanisms of forest fire spread is derived from the physical laws governing forest fire spread (such as the influence of fuel type, terrain, and meteorological conditions on fire propagation) and is used to guide or limit the model learning process. The role of spatial constraint information is to adjust the distribution of latent variables during variational inference, making it more consistent with the actual laws of forest fire spread and improving the physical rationality and accuracy of fire prediction.
[0119] Through the above technical solution, this application utilizes a graph variational inference model in a graph variational autoencoder network to perform feature fusion on spatiotemporally heterogeneous graphs. This effectively processes multi-source heterogeneous data, captures complex spatial and feature relationships, and uses the multimodal feature vectors of each grid node as input. This allows the graph variational inference model to acquire rich features from multiple sources, including satellite hotspot monitoring data, UAV inspection video stream data, and time-series data from ground-based IoT sensors. In the encoder stage, a multi-layer graph convolutional network aggregates the adjacency information represented by the first and second types of edges, comprehensively integrating spatial adjacency relationships and feature similarity connections, thereby enhancing the model's ability to recognize local and global fire patterns. By generating the latent variable mean and variance of each grid node, and then using Bayesian variational inference to calculate the posterior probability distribution representing the fire state of each grid, a probabilistic estimate of the fire state is achieved, effectively quantifying the uncertainty of the fire. In the variational inference process, spatial constraint information generated based on the physical mechanism of forest fire spread is introduced to adjust the distribution of latent variables. This makes the potential fire state distribution learned by the model no longer just data-driven, but integrates prior knowledge of the physical world, thereby improving the physical rationality and accuracy of fire prediction and avoiding deviation of the prediction results from the actual spread pattern.
[0120] In one embodiment of this application, adjusting the distribution of latent variables based on spatial constraint information generated from the physical mechanism of forest fire spread during variational inference includes: In response to the first fire confidence map meeting the preset threshold condition, the prior motion vector of fire spread is obtained; Input the prior motion vector of the fire spread into the preset semi-physical fire behavior spread model to generate the initial fire spread prediction range; The set of grid nodes corresponding to the initial fire spread prediction range is used as the prior constraint domain; In the decoding process of the graph variational inference model, the posterior estimation is based on the conflict between prior constraint domain suppression and spreading prior distribution.
[0121] In this embodiment, by setting a preset threshold condition, potential fires with low confidence can be filtered out, focusing on more certain fire areas. For example, by analyzing the local maximum value or the average confidence of connected regions in the first fire confidence map. When the maximum confidence value of a certain area in the first fire confidence map is greater than the preset threshold, the area can be considered to meet the preset threshold condition. At this time, starting from the geometric center or the point with the highest confidence in the area, combined with historical fire data, terrain, wind direction, and other information, the prior motion vector of fire spread is initially estimated. Alternatively, in addition to the first fire confidence map, cross-validation can be performed using original multi-source heterogeneous data. For example, when the first fire confidence map meets the preset threshold condition, further checks are made on whether there are high-brightness and temperature anomalies in satellite hotspot monitoring data, or whether there are obvious smoke and fire features in UAV inspection video stream data. Only when multi-source data jointly indicate the existence of a fire is the prior motion vector of fire spread obtained. This prior motion vector of fire spread is initially estimated by analyzing the morphological change trend of the fire area, combined with real-time wind direction and speed data, and terrain slope and other factors.
[0122] When inputting the prior motion vector of fire spread into a pre-defined semi-physical fire behavior spread model to generate the initial fire line spread prediction range, this step utilizes a mature semi-physical fire behavior spread model to transform the abstract prior motion vector of fire spread into the initial fire line spread prediction range. This provides prior constraints based on real fire behavior patterns for subsequent variational inference, enhancing the reliability and physical consistency of the prediction. For example, the semi-physical fire behavior spread model uses the Rothermel model. This Rothermel model, based on physical parameters such as the prior motion vector of fire spread (direction and velocity), combustible material type, combustible material load, moisture content, terrain slope, wind speed and direction, simulates the spread process of the fire line under different environments through a series of empirical formulas and physical laws, thereby generating the initial fire line spread prediction range for a future period. Specifically, when the first fire confidence map triggers the threshold, the coordinates of the fire point center are extracted, and the prior motion vector of fire spread is estimated based on initial meteorological and terrain data. This vector is used as the initial condition input into the semi-physical model. The internal processing flow of the semi-physical model includes: traversing the neighboring grids with the fire point center grid as the source point. For each grid cell to be calculated, the corresponding terrain slope and wind field projection are read. The Rothermel equation is executed to calculate the maximum spread rate and maximum spread direction at each grid cell; the maximum spread direction is determined by a weighted average of the effective wind speed vector (the vector sum of local wind speed and slope-induced wind speed) (effective wind speed weighting coefficient is 0.6, slope-induced weighting is 0.4). A vector field-based fast travel method is used to extrapolate the fire front location, and the connected set of continuous grid cells reachable by the maximum spread rate within a specified time is marked as the initial fire spread prediction range. The set of grid nodes corresponding to this initial fire spread prediction range serves as the prior constraint domain in subsequent graph variational inference.
[0123] In the decoding process of graph variational inference models, a crucial step—integrating prior knowledge of the physical spread mechanism into the decoding stage of the graph variational inference model when suppressing posterior estimations that conflict with the prior distribution of the spread based on the prior constraint domain—forces the posterior estimates to align with the physical spread laws by modifying them. For example, a penalty term is introduced into the loss function of the decoder in the graph variational inference model. When the posterior estimate output by the decoder (e.g., the probability that a grid node is predicted to be in a fire state) conflicts with the spread prior distribution indicated by the prior constraint domain (e.g., grids outside the prior constraint domain should not have a high probability of fire), this penalty term increases the loss, thereby guiding the model to adjust its parameters so that the posterior estimate better conforms to the prior constraints. For example, for grids outside the prior constraint domain, if their posterior fire probability is higher than a certain threshold, a larger penalty is applied. Alternatively, a post-processing step can be performed after the decoder generates the initial posterior estimate. For graticules within the prior constraint domain, increase their fire confidence; for graticules outside the prior constraint domain, if their fire confidence is too high, reduce their confidence or set it to zero to force them to conform to the prior distribution of physical spread.
[0124] Through the above technical solution, this application cleverly incorporates spatial constraint information generated based on the physical mechanism of forest fire spread into the decoding process of the graph variational inference model, thereby effectively solving the problem of how to specifically generate and apply spatial constraint information in variational inference to ensure consistency with the physical spread mechanism and avoid posterior estimation conflicts.
[0125] In one embodiment of this application, a geometric envelope of fire smoke diffusion is generated based on the fire intensity parameters of each grid calculated from the mean heatmap, local meteorological grid data, and a semi-physical fire behavior spread model, including: The estimated radiative heat flux of each grid in the mean heat map is converted into fire line intensity parameters; Input the fire intensity parameters, local surface combustible moisture content grid data and terrain slope data into the preset semi-physical fire behavior spread model to calculate the fire spread rate vector field. A Gaussian plume diffusion model is constructed based on the fire spread rate vector field and the three-dimensional wind vector field in the local meteorological grid data. The spatiotemporal distribution of near-ground smoke concentration was calculated using a Gaussian plume diffusion model, and the isosurface corresponding to a preset visibility hazard concentration threshold was used as the geometric envelope of the fire smoke diffusion.
[0126] In this embodiment, the estimated radiative heat flux refers to the total amount of heat energy radiated outward per unit area per unit time within a fire zone. Its concept lies in quantifying the intensity of the fire source and is a key physical quantity for assessing the activity level of a fire. This estimated radiative heat flux is obtained through remote sensing methods, such as using brightness temperature data detected by an infrared sensor. The estimated radiative heat flux is calculated by the difference between the infrared radiative brightness temperature and the background radiative brightness temperature using empirical formulas or physical models. For example, it can be calculated using the Stefan-Boltzmann law combined with surface emissivity, or by inversion using a pre-established radiative transfer model.
[0127] The fire line intensity parameter is an indicator describing the intensity of fire line combustion and is related to the heat release rate per unit length of the fire line. It serves as a crucial input to semi-physical fire behavior spread models, directly influencing the model's prediction of the fire spread rate. The fire line intensity parameter is calculated by combining the estimated radiative heat flux with the fire line width. For example, the Byram fire line intensity formula can be used to combine radiative heat flux with factors such as fire line width and combustion efficiency.
[0128] Infrared brightness temperature refers to the blackbody temperature corresponding to the energy radiated by an object's surface in the infrared band. In forest fire monitoring, it characterizes the surface temperature of the fire point or fire area. This infrared brightness temperature is collected by infrared sensors mounted on satellites or drones, such as thermal infrared channel data from geostationary meteorological satellites or polar-orbiting satellites, and data from infrared thermal imagers carried by drones. Background brightness temperature refers to the infrared brightness temperature of the normal ground surface surrounding the fire area that is not affected by the fire. Its function is to serve as a benchmark, used to separate the additional thermal radiation caused by the fire from the observed infrared brightness temperature. Background brightness temperature is obtained by statistically averaging the infrared brightness temperatures of non-fire grids within a certain range around the fire area, or by using historical infrared brightness temperature data from the same period in the same area under fire-free conditions as a reference.
[0129] Local surface combustible moisture content gridded data refers to the moisture content information of surface combustibles (such as dead branches and leaves, shrubs, and herbaceous plants) corresponding to each geographic grid within the target forest area. Its function is to directly affect the combustion performance of combustibles and the speed of fire spread; the higher the moisture content, the more difficult the combustibles are to ignite, and the slower the fire spreads. Local surface combustible moisture content gridded data is obtained through inversion from satellite remote sensing data, for example, by estimating using the empirical relationship between vegetation index and surface moisture content. Topographic slope data refers to the slope magnitude and direction of each geographic grid within the target forest area. Its function is to affect the speed and direction of fire spread; fire spreads faster uphill. This topographic slope data is calculated based on a digital elevation model (DEM). DEMs are generated through aerial photogrammetry, lidar scanning, or topographic map digitization. Slope calculation involves spatial analysis based on a digital elevation model, such as using slope analysis tools in geographic information system software like ArcGIS or QGIS, to calculate the rate of elevation change between each grid point and its neighboring grid points.
[0130] The fire spread rate vector field refers to the speed and direction of fire spread on each geographic grid within the target forest area. Its function is to quantify the dynamic spread trend of the fire, providing dynamic information about the fire source for subsequent smoke spread models. This vector field is calculated by a semi-physical fire behavior spread model, outputting a two-dimensional vector including velocity magnitude and direction, centered on each grid. The semi-physical model at this stage employs a high-resolution bilinear interpolation input strategy: the model accepts a 3×3 grid sliding window input; the central grid calculates fire intensity, and the surrounding grids provide terrain curvature features to correct for local wind direction disturbances. Radiant heat flux (kW / m²) is converted to fire line intensity; the three-dimensional wind field vector in the meteorological grid is projected onto the surface slope tangent plane to calculate the effective mid-surface wind speed. The logarithmic wind profile attenuation formula is used by default to convert the 10-meter high wind speed output by the meteorological model to the surface combustible bed height (the default wind speed estimate at a height of 2 meters is 60% of the 10-meter wind speed). When outputting the fire spread rate vector field, a confidence weight for each grid cell is also output (composed of water content inversion error and terrain shading error). This weight will be used for prior calibration of the uncertainty distribution map.
[0131] The three-dimensional wind vector field in local meteorological gridded data refers to the wind speed and direction information at different altitudes for each geographic grid within the target forest area. Its function is to serve as the main driving force for smoke diffusion models, determining the transport direction and diffusion range of the smoke stream. This local meteorological gridded data originates from high-resolution numerical weather prediction or is derived through spatial interpolation and vertical profile extrapolation from observation data from multiple meteorological stations deployed within the forest area.
[0132] The Gaussian plume diffusion model is a mathematical model used to simulate the diffusion of atmospheric pollutants. Its basic assumption is that the concentration distribution of pollutants in both the vertical and horizontal directions follows a Gaussian (normal) distribution. Its function is to calculate the spatial concentration distribution of smoke plumes based on ignition source parameters, meteorological conditions, and topographic information. This Gaussian plume diffusion model predicts the downwind concentration of smoke plumes using parameters such as smoke source intensity, wind speed, atmospheric stability, and mixing layer height.
[0133] The spatiotemporal distribution of near-ground plume concentration refers to the variation of particulate matter or harmful gas concentrations in the near-ground layer (e.g., 0-10 meters above ground) of a target forest area at different geographical locations within a preset time period. Its purpose is to visually demonstrate the scope and extent of the potential impact of plumes on the environment and people. This spatiotemporal distribution is calculated using a Gaussian plume diffusion model, outputting concentration values at different times and spatial locations.
[0134] The visibility hazard concentration threshold refers to the concentration limit at which particulate matter in a smoke stream will affect visibility and may even pose a hazard to human health. Its function is to serve as a basis for judging the scope and severity of the smoke stream's impact, transforming an abstract concentration value into a meaningful boundary. This visibility hazard concentration threshold is set according to national or local ambient air quality standards, visibility standards, or health risk assessment guidelines, such as specific concentration values for PM2.5 or PM10.
[0135] The geometric envelope of smoke diffusion in a fire scene refers to the geometric boundary formed by the area where the smoke concentration reaches or exceeds a preset visibility hazard concentration threshold at a specific time point. Its function is to clearly define the actual range of impact that smoke may cause, providing an intuitive spatial reference for fire monitoring, personnel evacuation, and firefighting decisions.
[0136] Through the above technical solutions, this application can fully integrate multi-source data, including gridded data of surface combustible moisture content obtained from satellite hotspot monitoring data and terrain slope data calculated based on digital elevation models. These data provide more refined and realistic input parameters for the semi-physical fire behavior spread model. In particular, the introduction of terrain slope data allows the calculation of the fire spread rate vector field to fully consider the acceleration or deceleration effect of terrain on fire spread, thereby improving the accuracy of fire spread prediction. Furthermore, combining the more accurate fire spread rate vector field with the three-dimensional wind vector field to construct a Gaussian plume diffusion model can more accurately simulate the source strength and diffusion path of the smoke stream. Finally, defining the geometric envelope of smoke stream diffusion through a visibility hazard concentration threshold ensures that the generated envelope not only has physical meaning but also directly characterizes the range of smoke stream influence on visibility and potential hazards. Compared to smoke stream diffusion models that do not fully integrate terrain details, the fire smoke stream diffusion geometric envelope generated in this application can more accurately characterize terrain-induced fire spread and smoke stream diffusion patterns, effectively solving the problem of insufficient prediction accuracy.
[0137] In one embodiment of this application, a Gaussian plume diffusion model is constructed based on the fire spread rate vector field and the three-dimensional wind vector field in the local meteorological grid data, and further includes: Based on the digital elevation model data of the target forest area, the topographic potential gradient vector and topographic curvature tensor at each grid point are calculated. Based on the dot product operation of the topographic potential gradient vector and the three-dimensional wind vector field, the parameters of the vertical motion tendency induced by the terrain are generated. The terrain-induced vertical motion tendency parameter and the terrain curvature tensor are added to the turbulent diffusion coefficient correction term of the Gaussian plume diffusion model to construct a terrain-sensitive plume settling correction factor. When calculating the spatiotemporal distribution of near-ground plume concentration, the vertical diffusion parameters of the Gaussian plume diffusion model are adjusted using a topographically sensitive plume settling correction factor.
[0138] In this embodiment, the topographic potential gradient vector represents the rate of change of terrain height and its direction, indicating the slope direction and steepness of the land surface. It is calculated by performing first-order partial derivative operations on digital elevation model data, for example, using the finite difference method.
[0139] The topographic curvature tensor describes the degree and direction of curvature of the topographic surface, and can characterize the microscopic features of the terrain, such as convexity / concavity, ridgelines, and canyon lines. It can be calculated by performing second-order partial derivative operations on digital elevation model data, for example, using the Gaussian curvature calculation method.
[0140] Based on the dot product operation of the topographic potential gradient vector and the three-dimensional wind vector field, parameters for topographically induced vertical motion tendency are generated. The dot product operation can quantify the relative relationship between the wind field and the topographic slope direction, thereby determining whether the airflow is rising or falling along the slope. For example, when the angle between the wind direction and the slope aspect is small, the dot product result is large, indicating that the wind field and the topographic slope direction tend to be consistent, which will induce vertical airflow.
[0141] The topography-induced vertical motion tendency parameter characterizes the effect of topography on the vertical motion of airflow, i.e., the intensity of topographic lifting or descending airflow. This parameter can be obtained directly from the dot product calculation result, or it can be used after normalization or thresholding.
[0142] Furthermore, topography-induced vertical motion tendency parameters and topography curvature tensors are incorporated into the turbulent diffusion coefficient correction term of the Gaussian plume diffusion model to construct a topography-sensitive plume settling correction factor. The turbulent diffusion coefficient in the Gaussian plume model describes the vertical and horizontal diffusion velocity of the plume and is influenced by factors such as atmospheric stability. By introducing topography-induced vertical motion tendency parameters and topography curvature tensors, the influence of terrain on airflow turbulence characteristics under complex terrain can be more accurately characterized. For example, these topography parameters can be used as inputs, and correction factors can be calculated using empirical formulas, then multiplied or added to the original turbulent diffusion coefficient. The topography-sensitive plume settling correction factor is a correction coefficient that comprehensively considers the influence of terrain, used to adjust the vertical diffusion and settling of the plume, enabling the Gaussian plume model to simulate the lifting, sinking, or vortex effects of the plume caused by terrain. The topographic curvature tensor includes profile curvature and planar curvature. Profile curvature represents the topographic undulation along the slope direction and is used to correct the topographic-induced uplift or subsidence effect in the vertical diffusion parameter. Planar curvature represents the topographic curvature along the contour line direction and is used to correct the local convergence or divergence effect in the horizontal diffusion parameter. This allows the topographically sensitive smoke settling correction factor to more accurately reflect the non-uniform influence of complex topography on the turbulent diffusion of smoke.
[0143] Finally, when calculating the spatiotemporal distribution of near-surface plume concentration, the vertical diffusion parameters of the Gaussian plume diffusion model were adjusted using a topographically sensitive plume settling correction factor. Vertical diffusion parameters directly affect the vertical distribution of the plume, thus influencing the near-surface plume concentration. By applying the topographically sensitive plume settling correction factor, the vertical diffusion calculation was directly optimized; for example, the correction factor was directly multiplied into or added to the original vertical diffusion parameters to convert them into new vertical diffusion parameters.
[0144] Through the above technical solution, this application calculates the topographic potential gradient vector and topographic curvature tensor at each grid point based on the digital elevation model data of the target forest area, quantifying the changes in topographic height and curvature characteristics. Based on the dot product operation of the topographic potential gradient vector and the three-dimensional wind vector field, topographically induced vertical motion tendency parameters are generated. Combined with the dynamic calculation of topographic and wind field trends, the vertical airflow trend is characterized, representing the direct influence of topography on smoke rise or settling, avoiding the neglect of topographically induced vertical motion. The topographically induced vertical motion tendency parameters and the topographic curvature tensor are added to the turbulent diffusion coefficient correction term of the Gaussian plume diffusion model to construct a topographically sensitive smoke settling correction factor. By correcting the turbulence coefficient to consider topographically induced turbulent changes, the Gaussian plume diffusion model adapts to topographic undulations, improving the accuracy of smoke settling prediction. When calculating the spatiotemporal distribution of near-surface smoke concentration, the topographically sensitive smoke settling correction factor is used to adjust the vertical diffusion parameters of the Gaussian plume diffusion model. The correction factor is directly applied to optimize the vertical diffusion calculation, making the smoke concentration distribution more accurately represent the diffusion behavior under complex terrain, thereby improving the generation quality of the geometric envelope of fire smoke diffusion.
[0145] In one embodiment of this application, following the composite high-risk area, it further includes: Based on the geometric envelope of the smoke flow diffusion in the fire field and the preset smoke flow concentration attenuation profile function, the high concentration coverage range of near-ground smoke flow within a preset future time window is calculated. Calculate the geometric intersection area between the fire line spread coverage area obtained by extrapolating the fire spread rate vector field based on the fire spread rate vector field and the near-ground smoke flow high concentration coverage area within a preset future time window; Calculate the wind direction stability index of each grid within the geometric intersection area. The wind direction stability index is the reciprocal of the variance of the wind direction azimuth change of the geometric intersection area within a preset historical time period. If the area of the geometric intersection area is greater than a preset area threshold and the wind direction stability index is less than a preset threshold, or if the area change rate of the geometric intersection area within a preset future time window is greater than a preset trend threshold, then the geometric intersection area is determined to be a high-risk interaction zone for fire and smoke coupling. For high-risk interaction zones involving smoke and fire, key monitoring marker information and smoke-fire coupling risk level assessment results are generated.
[0146] In this embodiment, when calculating the high-concentration coverage area of near-ground smoke streams, a preset smoke stream concentration decay profile function is used to describe the variation of pollutant concentration with distance or time as the smoke stream diffuses in space. This smoke stream concentration decay profile function can quantify the concentration of the smoke stream at different locations, thereby identifying high-concentration areas that pose a threat to visibility, air quality, or personnel safety. For example, this smoke stream concentration decay profile function can use an empirical exponential decay model, where the concentration... ,in, For source concentration, Where is the distance, and k is the attenuation coefficient, which is determined based on atmospheric stability.
[0147] When calculating the geometric intersection of the fire spread coverage area and the near-surface smoke high-concentration coverage area, the fire spread coverage area refers to all geographical areas that the fire line may reach within a preset future time window, as predicted by the fire spread rate vector field. This fire spread coverage area visually demonstrates the potential spread trend of the fire over a future period and serves as a basis for assessing the fire's development. For example, by applying the fire spread rate vector field to rasterized geographic information system data, the fire line position is updated at each time step based on the spread rate and direction of the raster, and all rasters covered by the fire line are accumulated to generate a dynamic spread area.
[0148] The geometric intersection area refers to the geographically overlapping portion of the fire line spread coverage area and the near-surface high-concentration smoke flow coverage area. This geometric intersection area identifies the region where fire spread and high-concentration smoke flow coexist, and is a key focus in smoke-fire coupling risk assessment. In a geographic information system environment, spatial analysis tools (intersection or overlay analysis functions) are used to perform Boolean operations on the vector or raster data of the two coverage areas to extract their overlapping portion.
[0149] When calculating the wind direction stability index for each grid within the geometric intersection area, the wind direction stability index is used to quantify the smoothness of wind direction changes within the geometric intersection area over a preset historical time period. It is defined as the reciprocal of the variance of the wind direction azimuth change. Wind direction stability is a meteorological factor affecting smoke diffusion and fire behavior. Unstable wind directions make smoke diffusion paths difficult to predict and fire behavior complex and variable, thus increasing the risk of smoke-fire coupling. The lower the wind direction stability index, the more drastic the wind direction azimuth change and the greater the uncertainty in the smoke diffusion direction. In this case, the spatial overlap between the fire front spread range and the near-surface high-concentration smoke coverage area will drift rapidly, increasing the difficulty of firefighting force deployment and personnel evacuation decisions. Therefore, a wind direction stability index below a preset threshold is used as one of the auxiliary conditions for determining high-risk interaction zones for smoke-fire coupling. This wind direction stability index can objectively assess this uncertainty. For example, within a preset historical time period, collect wind direction and azimuth sequence data from multiple weather stations or weather grid points within the geometric intersection area, calculate the statistical variance of these azimuths, and then take their reciprocal as the wind direction stability index.
[0150] When determining high-risk interaction zones for smoke-fire coupling, preset area thresholds, preset thresholds (wind direction stability), and preset trend thresholds (area change rate) are quantitative standards used to identify such zones. The method for determining the preset area threshold includes: calculating the equivalent protection area of the smallest combat unit of the firefighting force; multiplying this equivalent protection area by an amplification factor based on the ratio of fire line spread rate to smoke flow diffusion rate in smoke-fire coupling events, obtained from historical fire data statistics, to generate the preset area threshold.
[0151] The method for determining the preset threshold (the threshold corresponding to the wind direction stability index) includes: estimating the probability density of the variance of wind direction and azimuth changes at the same time of day in the geometric intersection area over the past year, and taking the 15th percentile of the cumulative probability distribution as the preset threshold.
[0152] The method for determining the preset trend threshold includes: multiplying the average rate of the fire spread rate vector field within the geometric intersection area by the perimeter of the geometric intersection area to obtain a theoretical reference value for the area growth rate, and taking 50% of the theoretical reference value as the preset trend threshold.
[0153] In this embodiment, the key monitoring and marking information is manifested by highlighting or flashing high-risk areas on the geographic information system interface, or by generating automatic early warning messages and sending them to relevant personnel. The smoke and fire coupling risk level assessment result can be a graded system from low to high, or a continuous risk score. This risk score is calculated by weighting the area of the geometric intersection region, the wind direction stability index, the area change rate, and other relevant factors.
[0154] Through the above technical solution, this application can effectively solve the problem that the existing solution ignores the interaction between the fire line and the smoke flow, resulting in incomplete risk assessment and inability to effectively identify and predict high-risk areas of smoke and fire interaction.
[0155] Through the above technical solution, this application effectively solves the problem that fixed heat sources may be misjudged as fires when identifying fire areas by generating mean heat maps based on multimodal feature vectors, thereby improving the accuracy of forest fire monitoring.
[0156] Corresponding to the forest fire monitoring method based on multi-source heterogeneous data fusion in the above embodiment, Figure 2 This is a structural block diagram of a forest fire monitoring system based on multi-source heterogeneous data fusion, provided as an embodiment of this application. For ease of explanation, only the parts relevant to the embodiment of this application are shown. References Figure 2 The forest fire monitoring system 20 based on multi-source heterogeneous data fusion includes: a data registration and feature generation module 21, a fire inference and area extraction module 22, a spatiotemporal heterogeneous graph construction module 23, a graph variational autoencoder reasoning module 24, a smoke flow diffusion envelope generation module 25, and a composite high-risk area module 26.
[0157] Among them, the data registration and feature generation module 21 is used to perform spatial registration and temporal alignment of multi-source heterogeneous data of the target forest area to generate multimodal feature vectors; the multi-source heterogeneous data includes satellite hotspot monitoring data, UAV inspection video stream data and ground IoT sensor time series data; The fire inference and region extraction module 22 is used to input multimodal feature vectors into a preset fire inference network to perform cross-modal fire risk inference, generate a first fire confidence map, and extract the fire area to be confirmed. The spatiotemporal heterogeneous graph construction module 23 is used to perform rasterization processing on the area to be confirmed fire in response to the maximum value in the first fire confidence map meeting the preset threshold condition. The obtained raster is used as a node, the spatial adjacency relationship between adjacent raster is used as the first type of edge, and the connection between the cosine similarity between the multimodal feature vectors corresponding to each raster node is greater than the preset similarity threshold and the Euclidean distance between the two raster nodes in geographic space is less than the preset neighborhood expansion radius is used as the second type of edge. The spatiotemporal heterogeneous graph is constructed based on the first type of edge and the second type of edge. The graph variational autoencoder inference module 24 is used to input the spatiotemporal heterogeneous graph and the multimodal feature vectors corresponding to each grid node into a preset graph variational autoencoder network to obtain the mean heat map and uncertainty distribution map of the fire area to be confirmed. The smoke diffusion envelope generation module 25 is used to calculate the fire line intensity parameters of each grid based on the mean heat map, and generate the geometric envelope of the fire smoke diffusion by combining local meteorological grid data and semi-physical fire behavior spread model. The composite high-risk area module 26 is used to overlay the mean heat map and the uncertainty distribution map onto the geographic information system base map, and calculate the spatial overlap area between the area with the heat value greater than the first threshold in the mean heat map and the geometric envelope of the fire smoke diffusion, as the composite high-risk area.
[0158] See Figure 3 , Figure 3 This is a schematic block diagram of an electronic device provided according to an embodiment of this application. Figure 3 The electronic device 300 in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304. Specifically, the processors 301 are configured to invoke the program instructions to perform the functions of the modules in the aforementioned device embodiments, for example... Figure 2The functions of the data registration and feature generation module 21, fire inference and area extraction module 22, spatiotemporal heterogeneous graph construction module 23, graph variational autoencoder reasoning module 24, smoke flow diffusion envelope generation module 25, and composite high-risk area module 26 are shown.
[0159] It should be understood that, in the embodiments of this application, the processor 301 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0160] Input device 302 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 303 may include a display (LCD, etc.), a speaker, etc.
[0161] The memory 304 may include read-only memory and random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include non-volatile random access memory. For example, the memory 304 may also store device type information.
[0162] In specific implementations, the processor 301, input device 302, and output device 303 described in the embodiments of this application can execute the implementation methods described in any embodiment of the forest fire monitoring method based on multi-source heterogeneous data fusion provided in the embodiments of this application, or they can execute the implementation methods of the electronic devices described in the embodiments of this application, which will not be repeated here.
[0163] In another embodiment of this application, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor, the program instructions implement all or part of the processes in the methods described above. Alternatively, the computer program can instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0164] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the electronic device. Furthermore, the computer-readable storage medium can include both internal and external storage units of the electronic device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0165] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0166] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the electronic devices and units described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0167] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces or units, or it may be an electrical, mechanical, or other form of connection.
[0168] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.
[0169] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0170] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A forest fire monitoring method based on multi-source heterogeneous data fusion, characterized in that, include: Spatial registration and temporal alignment are performed on multi-source heterogeneous data of the target forest area to generate multimodal feature vectors; the multi-source heterogeneous data includes satellite hotspot monitoring data, UAV inspection video stream data and time series data of ground IoT sensors; The multimodal feature vectors are input into a preset fire inference network to perform cross-modal fire risk inference, generate a first fire confidence map, and extract the fire area to be confirmed. In response to the maximum value in the first fire confidence map satisfying a preset threshold condition, the fire area to be confirmed is rasterized. The resulting raster is used as a node, the spatial adjacency relationship between adjacent raster is used as a first type of edge, and the connection between the cosine similarity between the multimodal feature vectors corresponding to each raster node is greater than a preset similarity threshold and the Euclidean distance between the two raster nodes in geographic space is less than a preset neighborhood expansion radius is used as a second type of edge. A spatiotemporal heterogeneous graph is constructed based on the first type of edge and the second type of edge. The spatiotemporal heterogeneous map and the multimodal feature vectors corresponding to each grid node are input into a preset graph variational autoencoder network to obtain the mean heat map and uncertainty distribution map of the fire area to be confirmed. Based on the fire intensity parameters of each grid calculated from the mean heat map, local meteorological grid data, and semi-physical fire behavior spread model, a geometric envelope of fire smoke diffusion is generated. The mean heat map and the uncertainty distribution map are overlaid on the geographic information system base map, and the spatial overlap area between the area with a heat value greater than the first threshold in the mean heat map and the geometric envelope of the fire smoke diffusion is calculated as a composite high-risk area.
2. The forest fire monitoring method based on multi-source heterogeneous data fusion according to claim 1, characterized in that, Also includes: Spatial clustering analysis is performed on the uncertainty of each grid in the uncertainty distribution map, and connected regions with an average regional uncertainty greater than a preset upper limit threshold are taken as target regions for supplementary sampling. Based on the geographical coordinate range, topographic elevation data, and geometric envelope of the fire smoke diffusion in the target area for replenishment, a replenishment reconnaissance path planning instruction is generated. Control at least one mobile reconnaissance platform to collect supplementary data based on the supplementary reconnaissance path planning command; The supplementary data and the original multi-source heterogeneous data in the supplementary target area are subjected to feature-level fusion encoding processing through a preset multi-source feature encoding network to update the multimodal feature vectors corresponding to each grid node in the supplementary target area. Based on the updated multimodal feature vectors, the connection relationships and weights of the second type of edges are recalculated, and the spatiotemporal heterogeneous graph is reconstructed. The updated multimodal feature vectors and the reconstructed spatiotemporal heterogeneous graph are re-input into the preset graph variational autoencoder network to iteratively update the mean heat map, the uncertainty distribution map, the geometric envelope of the fire smoke diffusion, and the composite high-risk area markers. Repeat the supplementary sampling reconnaissance and iterative update steps until the area-weighted average of the uncertainties of each grid in the supplementary sampling target area in the uncertainty distribution map is less than the preset convergence threshold, or the preset maximum number of supplementary sampling iterations is reached.
3. The forest fire monitoring method based on multi-source heterogeneous data fusion according to claim 2, characterized in that, The step of generating a reconnaissance path planning instruction based on the geographical coordinates, topographic elevation data, and geometric envelope of the fire smoke diffusion of the target area includes: The vector boundary of the target area to be supplemented is buffered outward by a preset distance to generate an initial security reconnaissance envelope; Identify terrain feature lines within the target area for re-sampling and within the initial security reconnaissance envelope; In response to the fact that the angle between the main diffusion axis of the fire smoke diffusion geometric envelope obtained in the most recent iteration update before the supplementary reconnaissance and the direction of the terrain feature line is less than a preset flow direction coincidence threshold, the reconnaissance path segment planned along the terrain feature line is designated as a high smoke exposure risk segment. Spatial overlay analysis is performed on the path nodes corresponding to the high smoke flow exposure risk section and the uncertainty maximum point in the uncertainty distribution map to generate a detour blind spot observation point sequence; Based on the sequence of detour-filling observation points, the supplementary mining reconnaissance path planning instruction is generated.
4. The forest fire monitoring method based on multi-source heterogeneous data fusion according to claim 1, characterized in that, The step of inputting the spatiotemporal heterogeneous graph into a graph variational autoencoder network to obtain the mean heat map and uncertainty distribution map of the fire area to be confirmed includes: The graph variational inference model in the graph variational autoencoder network is used to perform feature fusion on the spatiotemporal heterogeneous graph. The graph variational inference model takes the multimodal feature vector of each grid node as input. In the encoder stage, it aggregates the adjacency information represented by the first type of edge and the second type of edge through a multi-layer graph convolutional network to generate the latent variable mean and latent variable variance of each grid node. Then, the posterior probability distribution representing the fire status of each grid is calculated through Bayesian variational inference. In the variational inference process, the spatial constraint information generated based on the physical mechanism of forest fire spread is used to adjust the distribution of latent variables, and the mean heat map and uncertainty distribution map are obtained according to the posterior probability distribution.
5. A forest fire monitoring method based on multi-source heterogeneous data fusion according to claim 4, characterized in that, The adjustment of latent variable distribution based on spatial constraint information generated from the physical mechanism of forest fire spread during variational inference includes: In response to the first fire confidence map satisfying the preset threshold condition, the prior motion vector of fire spread is obtained; The a priori motion vector of the fire spread is input into a preset semi-physical fire behavior spread model to generate an initial fire spread prediction range; The set of grid nodes corresponding to the initial fire spread prediction range is used as the prior constraint domain; In the decoding process of the graph variational inference model, the posterior estimation is based on the conflict between the prior constraint domain suppression and the spread prior distribution.
6. A forest fire monitoring method based on multi-source heterogeneous data fusion according to claim 1, characterized in that, The process of generating the geometric envelope of fire smoke diffusion based on the fire line intensity parameters of each grid calculated from the mean heat map, local meteorological grid data, and semi-physical fire behavior spread model includes: The estimated radiative heat flux of each grid in the mean heat map is converted into fire line intensity parameters; The fire intensity parameters, local surface combustible moisture content grid data, and terrain slope data are input into a preset semi-physical fire behavior spread model to calculate the fire spread rate vector field. Based on the fire spread rate vector field and the three-dimensional wind vector field in the local meteorological grid data, a Gaussian plume diffusion model is constructed. Based on the digital elevation model data of the target forest area, the topographic potential gradient vector and topographic curvature tensor at each grid point are calculated to generate the topographically induced vertical motion tendency parameter. The terrain-induced vertical motion tendency parameter and the terrain curvature tensor are added to the turbulent diffusion coefficient correction term of the Gaussian plume diffusion model to construct a terrain-sensitive plume settling correction factor. When calculating the spatiotemporal distribution of near-ground plume concentration, the vertical diffusion parameters of the Gaussian plume diffusion model are adjusted using the topography-sensitive plume settling correction factor. The spatiotemporal distribution of near-ground smoke concentration was calculated using the adjusted Gaussian plume diffusion model, and the isosurface corresponding to the preset visibility hazard concentration threshold was used as the geometric envelope of the fire smoke diffusion.
7. A forest fire monitoring method based on multi-source heterogeneous data fusion according to claim 6, characterized in that, Following the designation as a composite high-risk area, it also includes: Based on the geometric envelope of the fire smoke diffusion and the preset smoke concentration attenuation profile function, the high concentration coverage range of near-ground smoke within a preset future time window is calculated. Calculate the geometric intersection area between the fire line spread coverage area obtained by extrapolating the fire spread rate vector field based on the fire spread rate vector field and the near-ground smoke high concentration coverage area within a preset future time window; Calculate the wind direction stability index of each grid within the geometric intersection area. The wind direction stability index is the reciprocal of the variance of the wind direction azimuth change in the geometric intersection area within a preset historical time period. If the area of the geometric intersection region is greater than a preset area threshold and the wind direction stability index is less than a preset threshold, or the area change rate of the geometric intersection region within the preset future time window is greater than a preset trend threshold, then the geometric intersection region is determined to be a high-risk interaction zone for fire and smoke coupling. For the high-risk interaction zone of smoke and fire coupling, key monitoring marking information and smoke and fire coupling risk level assessment results are generated.
8. A forest fire monitoring system based on multi-source heterogeneous data fusion, characterized in that, include: The data registration and feature generation module is used to perform spatial registration and temporal alignment on multi-source heterogeneous data of the target forest area to generate multimodal feature vectors; the multi-source heterogeneous data includes satellite hotspot monitoring data, UAV inspection video stream data and ground IoT sensor time series data; The fire inference and region extraction module is used to input the multimodal feature vector into a preset fire inference network to perform cross-modal fire risk inference, generate a first fire confidence map, and extract the fire area to be confirmed. The spatiotemporal heterogeneous graph construction module is used to perform rasterization processing on the fire area to be confirmed in response to the maximum value in the first fire confidence graph meeting the preset threshold condition. The resulting raster is used as a node, the spatial adjacency relationship between adjacent raster is used as a first type of edge, and the connection between the cosine similarity between the multimodal feature vectors corresponding to each raster node is greater than a preset similarity threshold and the Euclidean distance between the two raster nodes in geographic space is less than a preset neighborhood expansion radius is used as a second type of edge. The spatiotemporal heterogeneous graph is constructed based on the first type of edge and the second type of edge. The graph variational autoencoder inference module is used to input the spatiotemporal heterogeneous graph and the multimodal feature vectors corresponding to each grid node into a preset graph variational autoencoder network to obtain the mean heat map and uncertainty distribution map of the fire area to be confirmed. The smoke diffusion envelope generation module is used to calculate the fire line intensity parameters of each grid based on the mean heat map, and generate the geometric envelope of the fire smoke diffusion by combining local meteorological grid data and semi-physical fire behavior spread model. The composite high-risk area module is used to overlay the mean heat map and the uncertainty distribution map onto the geographic information system base map, and to calculate the spatial overlap area between the area with a heat value greater than a first threshold in the mean heat map and the geometric envelope of the fire smoke diffusion, as the composite high-risk area.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.