Short-term regional lightning prediction method based on multi-modal fusion and bidirectional state space
Patent Information
- Application Number
- CN202610655099.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-13
- Publication Date
- 2026-09-08
AI Technical Summary
模态单一:部分方法仅使用统计特征(如历史落雷次数),虽然能捕捉时间趋势,但缺乏对雷暴云团空间形态(如云团大小、边缘纹理)的感知,导致对强对流天气的识别率低,虚警率高
通过同时提取视觉模态特征(多通道时空热力图张量)和统计模态特征(多维度特征向量),并通过设置包括并行视觉特征编码器和统计特征编码器的双流特征编码器,使得雷电预测模型能够同时感知雷电活动的空间形态分布(由视觉模态特征反映)和微观物理属性(由统计模态特征反映),从而有效克服了现有技术中因仅依赖单一数据源(仅统计特征或仅雷达回波图像)而导致的识别率低、虚警率高以及难以对雷电强度进行精准分级的技术问题。
Smart Images

Figure CN122710218A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of meteorological disaster early warning technology, and in particular to a short-term regional lightning prediction method based on multimodal fusion and bidirectional state space. Background Technology
[0002] Lightning is a violent electrical discharge phenomenon that accompanies severe convective weather events such as thunderstorms, heavy rainfall, and hail, posing a serious threat to social production and the safety of people's lives and property. Achieving high-precision short-term (0-30 minute) regional lightning forecasts is of great significance for timely issuance of warnings and implementation of protective measures.
[0003] Previous technologies were mainly divided into two categories: one is the numerical model-based method, which relies on solving complex meteorological equations, has a large computational load, and is difficult to meet the needs of minute-level real-time early warning; the other is the data-driven method, which relies on multi-source heterogeneous data such as Doppler radar and satellites, which has high equipment costs and is difficult to fuse.
[0004] In recent years, deep learning-based methods have gradually emerged. Current solutions mainly suffer from the following limitations: Limited modality: Some methods rely solely on statistical features (such as the number of historical lightning strikes). While these methods can capture temporal trends, they lack an understanding of the spatial morphology of thunderstorm clouds (such as cloud size and edge texture), resulting in low recognition rates and high false alarm rates for severe convective weather. Other methods use only radar echo images, lacking quantitative references based on historical statistical data, making it difficult to accurately classify lightning intensity.
[0005] Insufficient long-range spatiotemporal sequence modeling capability: Traditional recurrent neural networks (such as LSTM) suffer from gradient vanishing and slow inference speed when processing long-range spatiotemporal sequences, making it difficult to effectively capture the long-range temporal and spatial dependencies of lightning activity. While the Transformer model can solve the long-term dependency problem in long-range spatiotemporal sequences through a self-attention mechanism, the computational complexity of the self-attention mechanism increases quadratically with the length of the long sequence, resulting in high computational resource consumption.
[0006] Therefore, there is an urgent need for a high-precision short-term regional lightning prediction method that can integrate visual spatial information and statistical quantitative information and efficiently process long-range spatiotemporal sequences. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to provide a short-term regional lightning prediction method based on multimodal fusion and bidirectional state space, which can integrate visual spatial information and statistical quantitative information, capture spatiotemporal dependence bidirectionally, and achieve refined hierarchical early warning of short-term regional lightning prediction through multi-task parallel output.
[0008] The technical solution adopted by the present invention to solve the above-mentioned technical problems is: a short-time regional lightning prediction method based on multimodal fusion and bidirectional state space, including a model training stage and a real-time prediction stage; The model training phase includes: S11. Obtain historical lightning data for the target geographical area and perform data preprocessing to obtain a lightning event sequence arranged in chronological order; S12. Based on the temporal distribution characteristics of lightning events in the lightning event sequence, multiple simulated current moments for feature extraction are generated through a dynamic time window generation strategy. S13. For each sub-region of the target geographic region, select all lightning events that occurred in the sub-region from the lightning event sequence, and then determine a historical lightning activity area in the sub-region; perform high-resolution and low-resolution gridding processing on the historical lightning activity area corresponding to each sub-region to obtain the corresponding high-resolution grid array and low-resolution grid array, and map each lightning event that occurred in the sub-region to the corresponding spatial grid in the high-resolution grid array and low-resolution grid array respectively; S14. For each simulation time, for each sub-region corresponding to the high-resolution grid array, extract the corresponding visual modal features that reflect the spatial morphological distribution of lightning. The visual modal features are multi-channel spatiotemporal heat map tensors generated based on the high-resolution grid array. For each sub-region corresponding to the low-resolution grid array, extract the corresponding statistical modal features that reflect the microscopic physical properties of lightning. The statistical modal features are multi-dimensional statistical feature vectors. S15. Construct a lightning prediction model, which includes: The dual-stream feature encoder includes a parallel visual feature encoder and a statistical feature encoder; the visual feature encoder is used to encode spatial features of visual modal features to generate a visual feature sequence; the statistical feature encoder is used to perform high-dimensional semantic mapping of statistical modal features to generate a statistical feature sequence. The multimodal fusion module is used to fuse visual feature sequences with statistical feature sequences and embed spatial location information to generate multimodal fused feature sequences; The bidirectional state space core module consists of at least two stacked bidirectional state space modules with unified parameter configurations. The bidirectional state space module includes a forward scan branch and a backward scan branch, which are used to perform bidirectional spatiotemporal sequence modeling on multimodal fusion feature sequences and output high-level spatiotemporal feature sequences. A multi-task prediction head is used to output the probability of lightning presence, the probability distribution of lightning frequency level, and the probability distribution of lightning intensity level in parallel based on high-level spatiotemporal feature sequences. S16. Using the visual modal features and statistical modal features corresponding to each sub-region at each simulated current time as training samples, and using the multi-task labels of lightning events mapped to each spatial grid in the low-resolution grid array corresponding to the sub-region within a future time window after the simulated current time as multi-task supervision signals, supervise the training of the lightning prediction model to obtain the trained lightning prediction model; wherein, the multi-task labels include lightning presence labels, lightning frequency level labels, and lightning intensity level labels; The real-time prediction phase involves using a trained lightning prediction model to perform short-term regional lightning predictions for one or more target sub-regions within the target geographical region.
[0009] The data preprocessing includes data cleaning and time-series arrangement; The data cleaning process is used to filter out lightning event data from the historical lightning data that are complete and valid in the five key fields of occurrence time, steepness, intensity, longitude and latitude, as valid data entries. The time sequence arrangement is used to arrange the lightning events corresponding to all the valid data entries in chronological order of their occurrence to form the lightning event sequence.
[0010] Step S12 includes: The occurrence time distribution of lightning events in the lightning event sequence is detected to identify dense periods and sparse events on the time axis corresponding to the lightning event sequence. A dense period refers to a time interval containing at least two lightning events, and the time interval between any two consecutive lightning events is no greater than a first threshold. The dense period begins at the occurrence time of the first lightning event it contains and ends at the occurrence time of the last lightning event it contains. A sparse event refers to an isolated lightning event, where the time interval between this lightning event and both its preceding and following lightning events is greater than the first threshold. Based on the dense time periods and the sparse events, execute the dynamic time window generation strategy: For each dense period, starting from its start time, a series of time windows are generated by sliding at a fixed time step, with the start time of each time window serving as a simulation of the current moment; For each sparse event, its occurrence time is used as a simulation of the current moment, and a corresponding time window is generated.
[0011] Here, by identifying time intervals between consecutive lightning events with an interval no greater than a first threshold as dense periods, and employing a fixed-time-step sliding window generation strategy, sufficient training samples can be generated during periods of frequent lightning activity, fully capturing the evolutionary patterns of lightning. By identifying isolated lightning events as sparse events and directly using their occurrence time as the simulated current moment, a small number of key samples can be effectively preserved without being missed during periods of sparse lightning activity. Through this adaptive processing strategy, the method of this invention can effectively cope with the uneven distribution of lightning event occurrence times (e.g., dozens of occurrences per minute during dense periods, and intervals of several days or even months during sparse periods), avoiding both the low training efficiency caused by sample redundancy during dense periods and the model failure caused by a lack of training samples during sparse periods.
[0012] The process for determining the historical lightning activity area is as follows: based on the minimum and maximum longitude, minimum and maximum latitude of all lightning events occurring within the sub-region, a geographical boundary box is defined for the sub-region as the historical lightning activity area.
[0013] The high-resolution meshing process is as follows: the historical lightning activity area is divided into M1×N1 spatial meshes to form the high-resolution mesh array; the low-resolution meshing process is as follows: the historical lightning activity area is divided into M2×N2 spatial meshes to form the low-resolution mesh array; wherein, M1 is a positive integer multiple of M2 and N1 is a positive integer multiple of N2; The method for mapping each lightning event to the corresponding spatial grid is as follows: the grid coordinates to which the lightning event belongs are determined based on the longitude and latitude of the lightning event.
[0014] Here, by setting M1 to be a positive integer multiple of M2 and N1 to be a positive integer multiple of N2, a precise integer multiple correspondence is formed between the high-resolution grid array and the low-resolution grid array in physical space. This lays the spatial alignment foundation for the subsequent downsampling of high-resolution feature maps to the low-resolution spatial dimension through adaptive average pooling. By determining the grid coordinates to which a lightning event belongs based on its longitude and latitude, each lightning event can be accurately mapped to the corresponding spatial grid in both the high-resolution and low-resolution grid arrays, ensuring the consistency of spatial data. The dual-resolution gridding design enables the method of this invention to avoid the problem of loss of the spatial topology of lightning activity caused by simply unfolding and stitching multimodal data in traditional schemes, ensuring a one-to-one correspondence between visual modal features and statistical modal features in physical space.
[0015] In step S14, a historical backtracking time window is defined based on the current simulated time. The historical backtracking time window is divided into a recent observation sub-window and a long-term reference sub-window. The multi-channel spatiotemporal heatmap tensor is constructed by splicing together spatiotemporal heatmaps from five channels along the channel dimension, specifically as follows: The recent frequency density map is generated by: counting the number of lightning events mapped to each spatial grid in the high-resolution grid array within the recent observation sub-window, generating a recent spatial count distribution matrix composed of these numbers, performing two-dimensional Gaussian smoothing on the recent spatial count distribution matrix, and then normalizing the value range to the [0,1] interval to obtain the recent frequency density map. The maximum intensity distribution map is generated by extracting the absolute values of the intensity of all lightning events mapped to each spatial grid in the high-resolution grid array within the recent observation sub-window, taking the maximum value as the maximum intensity value of the spatial grid, generating a maximum intensity spatial distribution matrix composed of these maximum intensity values, and then normalizing the value range to the interval [0, 1] to obtain the maximum intensity distribution map. The maximum steepness distribution map is generated as follows: extract the absolute steepness values of all lightning events mapped to each spatial grid in the high-resolution grid array within the recent observation sub-window, take the maximum value as the maximum steepness value of the spatial grid, generate the maximum steepness spatial distribution matrix composed of these maximum steepness values, and then normalize the value range to the interval [0, 1] to obtain the maximum steepness distribution map. The time decay density map is generated as follows: the occurrence time of all lightning events mapped to each spatial grid in the high-resolution grid array within the recent observation sub-window is recorded; the time difference between the occurrence time of each lightning event and the current simulation time is calculated; and the time decay weighted value of each lightning event is obtained based on the exponential decay weighting mechanism of the time difference. The time decay weighted values of all lightning events in each spatial grid are accumulated to obtain the time decay value of that spatial grid, and a time decay spatial distribution matrix composed of these time decay values is generated. The value range is then normalized to the [0,1] interval to obtain the time decay density map. The density difference trend map is generated as follows: the number of lightning events mapped to each spatial grid in the high-resolution grid array within the long-term reference sub-window is counted, a long-term spatial count distribution matrix composed of these numbers is generated, and the long-term spatial count distribution matrix is subjected to two-dimensional Gaussian smoothing to obtain a long-term background density map; the long-term background density map is subtracted from the recent frequency density map, and the value range is normalized to the interval [-1, 1] to obtain the density difference trend map.
[0016] In step S14, any one of the spatial grids in the low-resolution grid array is defined as the target grid. The multi-dimensional statistical feature vector includes: Global time series features, used to reflect time series patterns, include hourly cycle codes and date cycle codes. The hourly cycle code refers to the feature pairs obtained by performing sine and cosine transformations on the hour value of the simulated current moment in a day. The date cycle code refers to the feature pairs obtained by performing sine and cosine transformations on the yearly day value of the simulated current moment in a year. Local counting features are used to reflect the frequency of lightning activity within the target grid. The local counting features include long-term activity counts and recent activity counts. The long-term activity count refers to the number of lightning events mapped to the target grid within the historical backtracking time window. The recent activity count refers to the number of lightning events mapped to the target grid within the recent observation sub-window. Local physical features, used to reflect the electromagnetic properties of lightning discharges within the target grid, include maximum intensity and average steepness. The maximum intensity is determined as follows: when the long-term activity count is 0, the maximum intensity is 0; when the long-term activity count is greater than 0, the maximum intensity refers to the maximum absolute value of the intensity of all lightning events mapped to the target grid within the historical backtracking time window. The average steepness is determined as follows: when the long-term activity count is 0, the average steepness is 0; when the long-term activity count is greater than 0, the average steepness refers to the average absolute value of the steepness of all lightning events mapped to the target grid within the historical backtracking time window. Local environmental features, used to reflect the spatial context of the target grid, include neighborhood counts and dynamic ratios. The neighborhood counts refer to the number of lightning events mapped to all spatial grids in the low-resolution grid array other than the target grid within the historical backtracking time window. The dynamic ratio is determined as follows: when the long-term activity count is 0, the dynamic ratio is determined to be 0; when the long-term activity count is greater than 0, the dynamic ratio is the ratio of the recent activity count to the long-term activity count.
[0017] The visual feature encoder includes a first convolutional block, a first max pooling layer, a second convolutional block, a second max pooling layer, a third convolutional block, an adaptive average pooling layer, and a first flattening layer connected in sequence. The input of the first convolutional block is the visual modality feature, and the output of the first flattening layer is the visual feature sequence. The first convolutional block, the second convolutional block, and the third convolutional block have the same structure, each consisting of a convolutional layer, a batch normalization layer, and a ReLU activation function connected in sequence. The statistical feature encoder includes a first linear layer, a first normalization layer, a first ReLU activation function, a dropout layer, a second linear layer, a second normalization layer, a second ReLU activation function, and a second flattening layer connected in sequence. The input of the first linear layer is the statistical modal feature, and the output of the second flattening layer is the statistical feature sequence. The multimodal fusion module includes: The feature concatenation layer is used to concatenate the visual feature sequence and the statistical feature sequence in the channel dimension to form a multimodal hybrid feature sequence; The feature fusion layer is used to perform feature transformation and dimensionality compression on the multimodal hybrid feature sequence through a multilayer perceptron, so as to realize deep interaction of cross-modal information and form an interactive fusion feature sequence; The location encoding embedding layer is used to initialize a set of learnable spatial location encoding parameters, and to use a broadcast mechanism to superimpose the learnable spatial location encoding parameters onto the interactive fusion feature sequence to form the multimodal fusion feature sequence; The multi-task prediction head adopts a parallel decoupled design, including three independent prediction branches, specifically: The lightning presence prediction branch includes a first output linear layer and a Sigmoid activation function. The input of the first output linear layer is the high-level spatiotemporal feature sequence, and the output of the Sigmoid activation function is the lightning presence probability. The lightning frequency level prediction branch includes a second output linear layer and a first Softmax activation function. The input of the second output linear layer is the high-level spatiotemporal feature sequence, and the output of the first Softmax activation function is the lightning frequency level probability distribution. The lightning intensity level prediction branch includes a third output linear layer and a second Softmax activation function. The input of the third output linear layer is the high-level spatiotemporal feature sequence, and the output of the second Softmax activation function is the lightning intensity level probability distribution.
[0018] Here, by setting an adaptive average pooling layer at the end of the visual feature encoder, the high-resolution feature map is downsampled to the same spatial dimension as the low-resolution grid array, so that the high-resolution visual features and low-resolution statistical features are accurately aligned in spatial dimension, laying a spatial foundation for subsequent cross-modal fusion. By initializing a set of learnable spatial location encoding parameters and using a broadcast mechanism to superimpose them onto the fused feature sequence, the disordered grid sequence is endowed with clear spatial topological semantics. The lightning prediction model can automatically optimize the location encoding as the network is trained, and explicitly distinguish the positional differences of different grids such as the upper left corner and the center in the lightning propagation path, which has stronger adaptability compared with fixed sine and cosine encoding. By setting three independent prediction branches, using the Sigmoid activation function to output the probability of lightning presence, and using the Softmax activation function to output the probability distribution of lightning frequency level and lightning intensity level, the lightning prediction model can simultaneously output three types of prediction results: presence, frequency level, and intensity level. This achieves refined graded early warning from "whether it will happen" to "how much and how strong it will happen", providing meteorological departments with classified and graded decision support.
[0019] The bidirectional state space module further includes a third normalization layer. The forward scanning branch includes a forward processing unit, and the reverse scanning branch includes a sequence flipping operation, a reverse processing unit, and a sequence flipping restoration operation. The forward processing unit and the reverse processing unit have the same structure but independent parameters. The input of the third normalization layer is the multimodal fusion feature sequence. One output of the third normalization layer passes through the forward processing unit to obtain the forward scanning feature sequence, and the other output of the third normalization layer passes through the sequence flipping operation, the reverse processing unit, and the sequence flipping restoration operation to obtain the reverse scanning feature sequence. The forward scanning feature sequence and the reverse scanning feature sequence are added element-wise, and the addition result is fused with the multimodal fusion feature sequence through residual connection to obtain the high-level spatiotemporal feature sequence. Here, by setting a forward scanning branch to maintain the original order of the input sequence for forward scanning, the forward evolution dependency of lightning activity with spatial location can be captured. By setting a reverse scanning branch to flip the input sequence before performing a reverse scan, and then flipping the output sequence back to its original state, the global context information of the subsequent spatial grid on the preceding spatial grid can be captured, making up for the information blind spot of unidirectional scanning. By adding the forward scanning feature sequence and the reverse scanning feature sequence element by element, the output features of bidirectional scanning are fused with the spatial dependencies of both directions, constructing a global spatial perception capability without blind spots. By fusing the result of element-wise addition with the original multimodal fusion feature sequence through residual connections, the training stability of the deep network is ensured, while the original input information is preserved, avoiding information loss.
[0020] The implementation process of the forward processing unit and the reverse processing unit is as follows: the input sequence of the forward processing unit and the reverse processing unit passes through the first projection layer and the first SiLu activation function in one path, and the input sequence passes through the second projection layer, the one-dimensional convolutional layer, the second SiLu activation function, and the selective state space model in another path. The output of the first SiLu activation function is used as a gating signal and multiplied element-wise with the output of the selective state space model to realize the gating mechanism. The output of the gating mechanism passes through the third projection layer to obtain the output sequence of the forward processing unit and the reverse processing unit. Here, the input sequence generates a gating signal on one path and a transformation feature on the other, enabling the one-dimensional convolutional layer to extract local contextual features from the sequence and the selective state-space model to capture long-range dependencies. The combination of these two approaches allows the lightning prediction model to focus on both local relationships between adjacent spatial grids and global correlations between distant spatial grids. By using the output of the first SiLu activation function as a gating signal and multiplying it element-wise with the output of the selective state-space model, the lightning prediction model can dynamically adjust the information flow according to the current input, selectively retaining or suppressing specific features, thus enhancing its representational capabilities. By employing the selective state-space model as the core operator, the lightning prediction model can dynamically adjust the state transition parameters according to the current input while maintaining a linear computational complexity of O(N), achieving efficient capture of long-range dependencies. This demonstrates a significant computational efficiency advantage compared to the O(N²) complexity of the Transformer model's self-attention mechanism.
[0021] Compared with the prior art, the advantages of the present invention are as follows: By simultaneously extracting visual modal features (multi-channel spatiotemporal heatmap tensors) and statistical modal features (multi-dimensional feature vectors), and by setting up a dual-stream feature encoder that includes a parallel visual feature encoder and a statistical feature encoder, the lightning prediction model can simultaneously perceive the spatial morphological distribution of lightning activity (reflected by visual modal features) and its microscopic physical properties (reflected by statistical modal features). This effectively overcomes the technical problems of low recognition rate, high false alarm rate, and difficulty in accurately classifying lightning intensity caused by relying on only a single data source (statistical features only or radar echo images only) in existing technologies.
[0022] By constructing a bidirectional state-space core module, which consists of at least two stacked bidirectional state-space modules with unified parameter configurations, and each bidirectional state-space module containing a forward scan branch and a backward scan branch, the lightning prediction model can perform bidirectional spatiotemporal sequence modeling on multimodal fusion feature sequences. The forward scan branch captures the positive evolution dependency of lightning activity with spatial location, while the backward scan branch captures the negative spatial correlation; the two combine to construct a global spatiotemporal dependency. By employing the bidirectional state-space module as the core operator, the lightning prediction model can efficiently process long-range spatiotemporal sequences while maintaining linear computational complexity. This effectively overcomes the gradient vanishing and slow inference speed problems of traditional recurrent neural networks (such as LSTM) in existing technologies, as well as the high resource consumption and difficulty in deploying on edge devices caused by the quadratic increase in computational complexity of Transformer models with the sequence length.
[0023] By setting up a multi-task prediction head, the lightning prediction model can output the probability of lightning presence, the probability distribution of lightning frequency level, and the probability distribution of lightning intensity level in parallel based on high-level spatiotemporal feature sequences. By using multi-task labels containing lightning presence labels, lightning frequency level labels, and lightning intensity level labels as supervision signals for supervised training, the lightning prediction model can simultaneously output three types of prediction results: presence, frequency level, and intensity level. This achieves refined hierarchical early warning from "whether lightning will occur" to "how often it will occur and how strong it will be," thus effectively overcoming the technical problems of insufficient prediction capability and high false negative rate of existing technologies for sparse and high-intensity extreme weather disasters.
[0024] By performing high-resolution and low-resolution meshing on the historical lightning activity areas of each sub-region, corresponding high-resolution and low-resolution mesh arrays are obtained. Lightning events are then mapped to corresponding spatial grids in both types of arrays. Visual modal features are extracted from the high-resolution array, and statistical modal features are extracted from the low-resolution array, ensuring a clear correspondence between the visual and statistical modal features in physical space. A multimodal fusion module is then used to fuse the visual and statistical feature sequences and embed spatial location information. This ensures that the fused feature sequence retains the physical neighborhood relationships and spatial location information of lightning activity between grids, effectively overcoming the technical problem in existing technologies where simply unfolding and splicing multimodal data leads to the loss of the spatial topology of lightning activity.
[0025] The method of this invention defines a complete process from historical lightning data acquisition, preprocessing, dynamic time window generation, dual-resolution gridding, dual-modal feature extraction, model construction and training, to the real-time prediction stage using the trained model for prediction. This method enables end-to-end short-term regional lightning prediction without manual intervention or reliance on expensive external multi-source heterogeneous data (such as Doppler radar and satellite data), reducing equipment costs and data fusion difficulties, and meeting the need for minute-level real-time early warning. Attached Figure Description
[0026] Figure 1 This is a flowchart illustrating the overall implementation process of the method of the present invention; Figure 2 A schematic diagram of the composition structure of the lightning prediction model constructed by the method of the present invention; Figure 3 This is a schematic diagram of the composition of the forward processing unit and the reverse processing unit in the bidirectional state space module of the core module of the lightning prediction model. Figure 4 This is to visualize the visual modal features during the reasoning process of a real lightning sample (i.e., a test sample) using a trained lightning prediction model. Detailed Implementation
[0027] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0028] In a first aspect, embodiments of the present invention provide a short-time regional lightning prediction method based on multimodal fusion and bidirectional state space, such as... Figure 1 As shown, it includes a model training phase and a real-time prediction phase.
[0029] The model training phase includes: S11. Obtain historical lightning data for the target geographical area and perform data preprocessing to obtain a lightning event sequence arranged in chronological order.
[0030] In some embodiments, data preprocessing includes data cleaning and temporal arrangement. Data cleaning is used to filter out lightning event data from historical lightning data that are complete and valid in all five key fields: occurrence time, steepness, intensity, longitude, and latitude, as valid data entries. Temporal arrangement is used to arrange the lightning events corresponding to all valid data entries in chronological order of their occurrence time, forming a lightning event sequence.
[0031] In this embodiment, the target geographical area is a city, and each urban district in the city is a sub-region.
[0032] As an example, historical lightning data from different districts of a city over the past ten years were collected. The historical lightning data contains data from several lightning events. Since some lightning event data are missing one or more of the five key fields of occurrence time, steepness, intensity, longitude and latitude, data cleaning is required to ensure the integrity of the basic data.
[0033] S12. Based on the temporal distribution characteristics of lightning events in the lightning event sequence, multiple simulated current moments for feature extraction are generated through a dynamic time window generation strategy.
[0034] In some embodiments, step S12 includes: S121. Detect the time distribution of lightning events in a lightning event sequence, and identify dense periods and sparse events on the time axis corresponding to the lightning event sequence; wherein, a dense period refers to a time interval containing at least two lightning events, and the time interval between any two consecutive lightning events is not greater than a first threshold; a dense period begins at the occurrence time of the first lightning event it contains and ends at the occurrence time of the last lightning event it contains; a sparse event refers to an isolated lightning event, the time interval between the occurrence of the lightning event and the lightning events before and after it is greater than the first threshold, wherein, the lightning event before the isolated lightning event is the lightning event that occurs before the isolated lightning event and is adjacent to the isolated lightning event, and the lightning event after the isolated lightning event is the lightning event that occurs after the isolated lightning event and is adjacent to the isolated lightning event.
[0035] S122. Addressing the uneven timing of lightning events (which may occur dozens of times within a minute, or with intervals of several days or months between events), this invention proposes a dynamic time window generation strategy. Based on densely occurring periods and sparse events, the dynamic time window generation strategy is executed as follows: For each dense period, starting from its start time, a series of time windows are generated by sliding at a fixed time step, with the start time of each time window serving as a simulation of the current moment.
[0036] For each sparse event, its occurrence time is used as a simulation of the current moment, and a corresponding time window is generated.
[0037] As an example, the first threshold is preset to 30 minutes; the fixed time step is preset to 1 minute; the length of the time window is preset. The purpose of this step is to determine the current moment in the simulation. The length of the time window does not affect the determination of the current moment in the simulation.
[0038] As an example, suppose there are 7 lightning events in a sequence. If the intervals between lightning events 1, 2, 3, and 4 are all less than 30 minutes, the interval between lightning event 4 and lightning event 5 is more than 30 minutes, and the intervals between lightning events 5, 6, and 7 are all less than 30 minutes, then the time from the occurrence of lightning event 1 to the occurrence of lightning event 4 constitutes one dense period, and the time from the occurrence of lightning event 5 to the occurrence of lightning event 7 constitutes another dense period. If the intervals between lightning events 1, 2, 3, and 4 are all less than 30 minutes, the interval between lightning event 4 and lightning event 5 is more than 30 minutes, the interval between lightning event 5 and lightning event 6 is also more than 30 minutes, and the interval between lightning event 6 and lightning event 7 is less than 30 minutes, then the time from the occurrence of lightning event 1 to the occurrence of lightning event 4 constitutes a dense period, lightning event 5 is a sparse event, and the time from the occurrence of lightning event 6 to the occurrence of lightning event 7 constitutes another dense period.
[0039] S13. For each sub-region of the target geographic region, select all lightning events occurring within that sub-region from the lightning event sequence, and then determine a historical lightning activity area within that sub-region; perform high-resolution and low-resolution meshing processing on the historical lightning activity area corresponding to each sub-region to obtain corresponding high-resolution mesh arrays and low-resolution mesh arrays, and map each lightning event occurring within that sub-region to the corresponding spatial mesh in the high-resolution mesh array and low-resolution mesh array respectively; here, in order to capture the spatial morphology and texture details of lightning activity in detail, each sub-region is divided into a 30×30 high-resolution mesh array.
[0040] In some embodiments, the process of determining the historical lightning activity area is as follows: based on the minimum and maximum longitude, minimum and maximum latitude of all lightning events occurring in the sub-region, a geographical boundary box is defined for the sub-region as the historical lightning activity area. Since the geographical size of each sub-region is different, the geographical size of the historical lightning activity area of each sub-region is also different.
[0041] In some embodiments, high-resolution meshing is performed by dividing the historical lightning activity area into M1×N1 spatial grids to form a high-resolution mesh array; low-resolution meshing is performed by dividing the historical lightning activity area into M2×N2 spatial grids to form a low-resolution mesh array; wherein M1 is a positive integer multiple of M2 and N1 is a positive integer multiple of N2.
[0042] As an example, historical lightning activity areas are divided into 30×30 spatial grids to form a high-resolution grid array; and further divided into 3×3 spatial grids to form a low-resolution grid array. That is, M1=N1=30, M2=N2=3, and one spatial grid in the low-resolution array contains 100 spatial grids from the high-resolution array. To address the issue of inconsistent geographical sizes across different sub-regions, this invention does not divide all sub-regions' historical lightning activity areas into low-resolution and high-resolution spatial grids of equal size. Instead, it divides all sub-regions' historical lightning activity areas into 3×3 low-resolution spatial grids and 30×30 high-resolution spatial grids. This spatial gridding not only reduces the complexity of spatial analysis but also provides basic spatial units for subsequent feature extraction, facilitating the capture of lightning activity characteristics.
[0043] In some embodiments, each lightning event is mapped to a corresponding spatial grid by determining the grid coordinates to which the lightning event belongs based on its longitude and latitude.
[0044] S14. For each simulation time, for each sub-region corresponding to the high-resolution grid array, extract the corresponding visual modal features that reflect the spatial morphological distribution of lightning. The visual modal features are multi-channel spatiotemporal heat map tensors generated based on the high-resolution grid array. For each sub-region corresponding to the low-resolution grid array, extract the corresponding statistical modal features that reflect the microscopic physical properties of lightning. The statistical modal features are multi-dimensional statistical feature vectors.
[0045] In some embodiments, in step S14, a historical backtracking time window based on the current simulated time is defined, and the historical backtracking time window is equally divided into a recent observation sub-window and a long-term reference sub-window; the multi-channel spatiotemporal heatmap tensor is composed of spatiotemporal heatmaps from five channels stitched together along the channel dimension, specifically: The recent frequency density map is generated by counting the number of lightning events mapped to each spatial grid in the high-resolution grid array within the recent observation sub-window, generating a recent spatial count distribution matrix composed of these numbers, performing two-dimensional Gaussian smoothing on the recent spatial count distribution matrix, and then normalizing the value range to the [0, 1] interval to obtain the recent frequency density map.
[0046] The maximum intensity distribution map is generated by extracting the absolute intensity values of all lightning events mapped to each spatial grid in the high-resolution grid array within the recent observation sub-window, taking the maximum value as the maximum intensity value of the spatial grid, generating a maximum intensity spatial distribution matrix composed of these maximum intensity values, and then normalizing the value range to the [0,1] interval to obtain the maximum intensity distribution map.
[0047] The maximum steepness distribution map is generated by extracting the absolute steepness values of all lightning events mapped to each spatial grid in the high-resolution grid array within the recent observation sub-window, taking the maximum value as the maximum steepness value of the spatial grid, generating a maximum steepness spatial distribution matrix composed of these maximum steepness values, and then normalizing the value range to the [0,1] interval to obtain the maximum steepness distribution map.
[0048] The time decay density map is generated as follows: The occurrence times of all lightning events mapped to each spatial grid in the high-resolution grid array within the recent observation sub-window are recorded. The time difference between the occurrence time of each lightning event and the current simulation time is calculated. Based on the exponential decay weighting mechanism of the time difference, the time decay weighting value of each lightning event is obtained. The time decay weighting values of all lightning events in each spatial grid are accumulated to obtain the time decay value of that spatial grid. A time decay spatial distribution matrix composed of these time decay values is generated. The value range is then normalized to the interval [0, 1] to obtain the time decay density map. The formula for calculating the time decay weighting value of each lightning event is as follows: W represents the time decay weighted value of a single lightning event, e is the natural constant, e=2.71…, Δt represents the time difference between the occurrence time of a single lightning event and the current simulation time, in minutes, and τ represents the time decay constant (set to 10 minutes).
[0049] The density difference trend map is generated as follows: the number of lightning events mapped to each spatial grid in the high-resolution grid array within the long-term reference sub-window is counted, a long-term spatial count distribution matrix composed of these numbers is generated, and the long-term spatial count distribution matrix is subjected to two-dimensional Gaussian smoothing to obtain the long-term background density map; the recent frequency density map is subtracted from the long-term background density map, and the value range is normalized to the interval [-1, 1] to obtain the density difference trend map.
[0050] As an example, the length of the historical backtracking time window is a preset value, such as the 20 minutes immediately preceding the current simulated moment. The historical backtracking time window is divided into a recent observation sub-window and a long-term reference sub-window, both with a length of 10 minutes. The recent observation sub-window is the 10 minutes immediately preceding the current simulated moment. Based on this sub-window division mechanism, a spatiotemporal heatmap with five channels is generated that can comprehensively depict the spatial topology and physical evolution of thunderstorm clouds.
[0051] In this embodiment, the multi-channel spatiotemporal heatmap tensor has 5 channels and a spatial resolution of 30×30.
[0052] In this embodiment, the generation of the recent frequency density map not only eliminates the harsh boundary effects caused by artificially dividing the spatial grid, but also simulates the meteorological physical law of the continuous diffusion of strong convective energy from the eruption center to the surrounding areas. This transforms the originally sparse and isolated lightning events into a continuous heat map of concentration, enabling better identification of the overall spatial morphology and contiguous areas of thunderstorm clouds. In meteorological physics, extremely strong lightning currents often directly correspond to violent updrafts and complex and active microphysical ice phase processes. Therefore, the maximum intensity distribution map explicitly quantifies the upper limit of the energy destructive power of the convective system at the current stage, providing the most direct physical attribute basis for accurately identifying and warning of potential extreme severe convective weather disasters. Steepness is a key instantaneous indicator characterizing the dynamic stability of a convective system. Extremely high steepness values mean that the charge release within a local space is extremely concentrated and violent. Introducing the maximum steepness distribution map can provide important identification signals for thunderstorm cells in the rapid generation and violent eruption stages, greatly improving the sensitivity to capturing sudden convective enhancement processes. The time decay density map incorporates dynamic "time attention," enabling it to move beyond treating all lightning events within a 10-minute period indiscriminately. Instead, it allows for the keen perception of sudden incremental signals in the immediate minutes, thus more accurately judging the real-time evolution trend of lightning activity shifting from sparse to dense. The innovative differential operation of subtracting the long-term background density map from the recent frequency density map extracts the spatial momentum characteristics of thunderstorm clouds: positive values in the density difference trend map reflect the enhancement and influx of convective activity, while negative values indicate decay and outflow. This alternating positive and negative trend map allows for a global understanding of the lightning system's movement direction and generation and dissipation dynamics, compensating for the information lag inherent in single static observations.
[0053] In some embodiments, in step S14, any spatial grid in the low-resolution grid array is defined as the target grid; the multi-dimensional statistical feature vector includes: Global time series features are used to reflect time series patterns. Global time series features include hourly cycle codes and date cycle codes. Hourly cycle codes refer to feature pairs obtained by performing sine and cosine transformations on the hour value of the current simulated moment in a day, respectively, and are used to characterize the diurnal variation periodicity of lightning activity. Date cycle codes refer to feature pairs obtained by performing sine and cosine transformations on the annual day value of the current simulated moment in a year, respectively, and are used to characterize the seasonal variation periodicity of lightning activity.
[0054] Local count features are used to reflect the frequency of lightning activity within the target grid. Local count features include long-term activity counts and recent activity counts. Long-term activity counts refer to the number of lightning events mapped to the target grid within the historical backtracking time window, which is used to reflect the baseline level and historical accumulation of lightning activity in the target grid over a longer time scale. Recent activity counts refer to the number of lightning events mapped to the target grid within the recent observation sub-window, which is used to reflect the sudden intensity and instantaneous burst state of lightning activity in the target grid in a short period of time.
[0055] Local physical characteristics, used to reflect the electromagnetic properties of lightning discharges within the target grid, include maximum intensity and average steepness. Maximum intensity is determined as follows: when the long-term activity count is 0, the maximum intensity is set to 0; when the long-term activity count is greater than 0, the maximum intensity refers to the maximum absolute value of the intensity of all lightning events mapped to the target grid within the historical backtracking time window, reflecting the upper limit of the energy destructive power of lightning events within the target grid and identifying potential extreme lightning strike risks. Average steepness is determined as follows: when the long-term activity count is 0, the average steepness is set to 0; when the long-term activity count is greater than 0, the average steepness refers to the average absolute value of the steepness of all lightning events mapped to the target grid within the historical backtracking time window, reflecting the average intensity of lightning activity and the severity of convection development within the target grid. The calculation formula is as follows: , among which, S avg S represents the average of the absolute values of the steepness of all lightning events mapped to the target grid. i This represents the steepness of the i-th lightning event mapped to the target grid, where abs() is the absolute value, and n represents the number of lightning events mapped to the target grid.
[0056] Local environmental features are used to reflect the spatial context of the target grid. These features include neighborhood counts and dynamic ratios. Neighborhood counts refer to the number of lightning events mapped to all spatial grids in the low-resolution grid array other than the target grid within the historical backtracking time window. This reflects the activity level of the spatial environment in which the target grid is located and captures the clustering effect and spatial propagation potential of lightning activity. The dynamic ratio is determined as follows: when the long-term activity count is 0, the dynamic ratio is set to 0; when the long-term activity count is greater than 0, the dynamic ratio is the ratio of the recent activity count to the long-term activity count. This reflects the distribution skewness and evolution trend of lightning activity in the time dimension and distinguishes between the sudden enhancement stage and the continuous dissipation stage.
[0057] In this embodiment, the global temporal features of the low-resolution grid array are 4-dimensional, and the local counting features, local physical features, and local environmental features of each spatial grid in the low-resolution grid array are all 2-dimensional. Since the low-resolution grid array contains 9 spatial grids, the multi-dimensional feature vector is 58-dimensional.
[0058] In this embodiment, the core significance of long-term activity counting lies in providing an instantaneous and comparable quantitative benchmark for lightning activity levels in the target grid within the historical backtracking time window. By comparing the long-term activity counts of all spatial grids at the same simulation current moment, it is possible to effectively identify long-term relatively active centers and quiet areas. The core significance of recent activity counting lies in its high sensitivity to sudden and emerging small-scale convective activity, serving as a key instantaneous indicator for capturing rapidly changing phenomena such as the initiation of thunderstorms and short-duration strong convective outbursts. Maximum intensity records the maximum physical intensity (e.g., peak radiation energy) of the lightning signal detected by the target grid within the historical backtracking time window. The core significance of maximum intensity lies in the fact that strong lightning events are often associated with stronger convective updrafts and more severe weather phenomena, and maximum intensity helps identify potential high-impact weather events. Average steepness describes the average rate of change in lightning event intensity. Its core significance lies in reflecting the average severity of convective activity development. Combined with maximum steepness, it can distinguish between single pulse events and a general increase in overall activity. The core significance of neighborhood counting lies in identifying the overall instability and spatial propagation potential of the lightning convection system by comparing the local activity of the target grid with the background activity of the surrounding environment. Even if the target grid is currently in a calm state, if its neighborhood count remains high, it indicates that it is surrounded by an active, strong convective environment and is affected by systemic energy spillover or neighboring transmission, significantly increasing the risk of future lightning events. When the long-term activity count is greater than 0, the dynamic ratio is defined as the ratio of the lightning event count within the recent observation sub-window to the lightning event count within the historical backtracking time window. The core significance of the dynamic ratio is that by quantifying the proportion of "recent activity" in the "overall history," it can effectively distinguish the development stages of thunderstorms. When the dynamic ratio approaches 1.0, it is a strong warning signal, indicating that the lightning activity within the target grid is mainly concentrated near the current simulation time and belongs to the sudden generation stage. Conversely, when the dynamic ratio approaches 0.0, it indicates that the lightning activity mainly occurred in the past and is currently in the dissipation stage. This relative index allows the model to eliminate the interference of "historical remnant" signals and keenly capture the real-time strengthening trend of thunderstorms.
[0059] S15. Construct a lightning prediction model, such as Figure 2 As shown, the lightning prediction model includes: The dual-stream feature encoder includes a parallel visual feature encoder and a statistical feature encoder; the visual feature encoder is used to encode spatial features of visual modal features to generate a visual feature sequence; the statistical feature encoder is used to perform high-dimensional semantic mapping of statistical modal features to generate a statistical feature sequence.
[0060] The multimodal fusion module is used to fuse visual feature sequences with statistical feature sequences and embed spatial location information to generate multimodal fused feature sequences.
[0061] The bidirectional state space core module consists of at least two bidirectional state space modules (Bi-Mamba modules) with unified parameter configuration stacked together. The bidirectional state space module contains a forward scan branch and a backward scan branch, which are used to perform bidirectional spatiotemporal sequence modeling on multimodal fusion feature sequences and output high-level spatiotemporal feature sequences.
[0062] A multi-task prediction head is used to output the probability of lightning presence, the probability distribution of lightning frequency level, and the probability distribution of lightning intensity level in parallel based on high-level spatiotemporal feature sequences.
[0063] In some embodiments, specifically, such as Figure 2 As shown, the visual feature encoder includes a first convolutional block, a first max-pooling layer (MaxPool), a second convolutional block, a second max-pooling layer, a third convolutional block, an adaptive average pooling layer (AdaptiveAvgPool), and a first flattening layer, all connected in sequence. The input to the first convolutional block is visual modal features, and the output of the first flattening layer is a sequence of visual features. The first, second, and third convolutional blocks have the same structure, each consisting of a convolutional layer, a batch normalization layer, and a ReLU activation function connected in sequence. The first convolutional block has 32 output channels, and the second convolutional block has [missing information - likely a number of output channels]. The number of output channels in the third convolutional block is 128, and the kernel size of the adaptive average pooling layer is 3×3. Spatial features are extracted step by step through the first, second, and third convolutional blocks. The high-resolution feature map output by the third convolutional block is forced downsampled to the same spatial dimension as the low-resolution grid array by the adaptive average pooling layer. Then, a visual feature sequence is generated by the first flattening layer. The core significance of the adaptive average pooling layer is that it creatively solves the problem of aligning high-resolution image data and low-resolution grid data in spatial dimension, ensuring a one-to-one correspondence between visual features and statistical features in physical space. The first flattening layer in the visual feature encoder adopts a row-first scanning flattening strategy. Specifically, the M2×N2 (3×3 in this embodiment) two-dimensional high-resolution feature map output by the adaptive average pooling layer is unfolded row by row in strict order from top to bottom and from left to right, and converted into a one-dimensional visual feature sequence of length M2×N2.
[0064] In some embodiments, specifically, such as Figure 2 As shown, the statistical feature encoder includes a first linear layer, a first normalized layer, a first ReLU activation function, a dropout layer, a second linear layer, a second normalized layer, a second ReLU activation function, and a second flattening layer connected in sequence. The input of the first linear layer is statistical modal features, and the output of the second flattening layer is a statistical feature sequence. The number of output channels for both the first and second linear layers is 64.
[0065] In some embodiments, specifically, such as Figure 2 As shown, the multimodal fusion module includes: a feature concatenation layer, used to concatenate visual feature sequences and statistical feature sequences along the channel dimension to form a multimodal hybrid feature sequence; a feature fusion layer, used to perform feature transformation and dimensionality compression on the multimodal hybrid feature sequence through a multilayer perceptron, realizing deep interaction of cross-modal information and forming an interactive fusion feature sequence; and a positional encoding embedding layer, used to initialize a set of learnable spatial positional encoding parameters and use a broadcast mechanism to superimpose the learnable spatial positional encoding parameters onto the interactive fusion feature sequence to form a multimodal fusion feature sequence. Unlike fixed sine and cosine encoding, learnable spatial positional encoding can be automatically optimized during network training. Its core significance lies in giving the disordered grid sequence a clear spatial topological semantics, enabling the model to explicitly distinguish the positional differences of different spatial grids such as the top left corner and the center in the lightning propagation path.
[0066] In some embodiments, specifically, such as Figure 3As shown, the bidirectional state space module also includes a third normalization layer. The forward scan branch includes a forward processing unit, and the reverse scan branch includes a sequence flipping operation, a reverse processing unit, and a sequence flipping restoration operation. The forward processing unit and the reverse processing unit have the same structure but independent parameters. The input of the third normalization layer is the multimodal fusion feature sequence. One output of the third normalization layer passes through the forward processing unit to obtain the forward scan feature sequence, and the other output of the third normalization layer passes through the sequence flipping operation, the reverse processing unit, and the sequence flipping restoration operation to obtain the reverse scan feature sequence. The forward scan feature sequence and the reverse scan feature sequence are added element by element, and the addition result is fused with the multimodal fusion feature sequence through residual connection to obtain the high-level spatiotemporal feature sequence. Here, the forward and backward scanning branches are designed to address the limitation of traditional sequence models that can only capture unidirectional dependencies. The forward scanning branch is used to maintain the original order of the input multimodal fusion feature sequence during forward scanning, capturing the positive evolution dependency of lightning activity with spatial location through the forward processing unit. The backward scanning branch is used to first perform a sequence flipping operation on the input multimodal fusion feature sequence to reverse the grid index, and then input it to the backward processing unit. After the backward processing unit completes its processing, it performs a sequence flipping and restoration operation on the output sequence again. The core significance of the backward scanning branch is that it uses the backward scanning mechanism to capture the global context information of the future relative to the past or the subsequent grid relative to the preceding grid, making up for the information blind spot of unidirectional scanning.
[0067] In some embodiments, specifically, the implementation process of the forward processing unit and the reverse processing unit is as follows: Figure 3 As shown, the input sequences of the forward processing unit and the backward processing unit pass through the first projection layer and the first SiLu activation function in one path, and the input sequence passes through the second projection layer, the one-dimensional convolutional layer, the second SiLu activation function, and the selective state space model in another path. The output of the first SiLu activation function is used as a gating signal and multiplied element-wise with the output of the selective state space model to realize the gating mechanism. The output of the gating mechanism passes through the third projection layer to obtain the output sequences of the forward processing unit and the backward processing unit.
[0068] In some embodiments, specifically, the multi-task prediction head employs a parallel decoupled design, comprising three independent prediction branches, specifically: The lightning presence prediction branch (binary classification) includes a first output linear layer and a Sigmoid activation function. The input of this branch is a high-level spatiotemporal feature sequence, that is, the input of the first output linear layer is a high-level spatiotemporal feature sequence. After transformation by the first output linear layer, the lightning presence probability of each spatial network in the low-resolution grid array is output by the Sigmoid activation function as a primary screening signal to determine whether a lightning event will occur within a future time window (i.e., the number of lightning strikes N>0).
[0069] The lightning frequency level prediction branch (multi-classification) includes a second output linear layer and a first Softmax activation function. The input to this branch is a high-level spatiotemporal feature sequence; that is, the input to the second output linear layer is also a high-level spatiotemporal feature sequence. After transformation by the second output linear layer, the first Softmax activation function outputs the probability distribution of lightning frequency levels for each spatial network in the low-resolution grid array. The lightning frequency level prediction branch aims to quantify the intensity of lightning activity, classifying the prediction results into three categories: no lightning (corresponding to 0 lightning strikes, safe), sparse lightning (corresponding to sporadic activity with 1 to 5 lightning strikes), and dense lightning (corresponding to areas of strong convection activity with more than 5 lightning strikes), thereby effectively identifying strong convection centers in their outbreak phase.
[0070] The lightning intensity level prediction branch (multi-classification) includes a third output linear layer and a second Softmax activation function. The input to this branch is a high-level spatiotemporal feature sequence; that is, the input to the third output linear layer is a high-level spatiotemporal feature sequence. After transformation by the third output linear layer, the second Softmax activation function outputs the probability distribution of lightning intensity levels for each spatial network in a low-resolution grid array. The lightning intensity level prediction branch focuses on assessing the physical destructive power of lightning currents, classifying the prediction results into three categories: no lightning (corresponding to an intensity of 0 kA, safe), general lightning (corresponding to a common discharge phenomenon with an intensity greater than 0 kA and less than or equal to 48 kA, conventional risk), and extremely destructive lightning (corresponding to an intensity greater than 48 kA, extremely destructive power, extremely high risk). This effectively identifies extreme lightning events that may cause physical damage to critical infrastructure such as power grids and communication base stations.
[0071] S16. Using the visual modal features and statistical modal features corresponding to each sub-region at each simulated current moment as training samples, and using the multi-task labels of lightning events mapped to each spatial grid in the low-resolution grid array corresponding to the sub-region within a future time window after the simulated current moment as multi-task supervision signals, supervise the training of the lightning prediction model to obtain the trained lightning prediction model; wherein, the multi-task labels include lightning presence labels, lightning frequency level labels, and lightning intensity level labels.
[0072] In some embodiments, during the training of the lightning prediction model, a multi-task joint loss function is used to calculate the model loss. Specifically, for the lightning presence prediction task, a labeled smoothing focal loss function is used. This loss function effectively addresses the extreme class imbalance problem of infrequent lightning events by adjusting the weights of positive and negative samples, enabling the lightning prediction model to maintain high sensitivity even under sparse positive sample conditions. For the lightning frequency level prediction task and the lightning intensity level prediction task, a weighted cross-entropy loss function is used. This loss function automatically calculates the loss weight based on the number of samples for each level, assigning higher loss weights to categories with fewer samples, thereby addressing prediction bias caused by long-tail distribution and improving the lightning prediction model's ability to identify a few categories such as sparse lightning and highly destructive lightning. The model loss is obtained by weighted summing the losses of the three tasks.
[0073] As an example, the length of the future time window can be preset to 10 minutes.
[0074] Here, the multi-task label for each training sample within a future time window is a multi-dimensional label group containing a lightning presence label, a lightning frequency level label, and a lightning intensity level label. Specifically, the lightning presence label is a binary vector indicating whether at least one lightning event occurs within the spatial grid within the future time window; the lightning frequency level label is a multi-valued vector used to distinguish between no lightning, sparse lightning events, and dense lightning periods; and the lightning intensity level label is a multi-valued vector used to distinguish between no lightning, general lightning, and destructive strong lightning.
[0075] It should be noted that, to balance lightweight deployment and high-precision prediction of the lightning prediction model, the bidirectional state-space module constructed in this invention fully leverages the complementary advantages of the dual-stream architecture for heterogeneous modal data, as well as the deep perception capability of the bidirectional sequence scanning mechanism for global spatial topology. Compared to the traditional Transformer model based on the self-attention mechanism, the lightning prediction model of this invention effectively solves the computational resource bottleneck problem in long-sequence spatiotemporal modeling while maintaining linear computational complexity, significantly improving the sensitivity and accuracy of capturing complex lightning propagation paths and extreme severe convective weather.
[0076] The real-time prediction phase involves using a trained lightning prediction model to perform short-term regional lightning predictions for one or more target sub-regions within the target geographical region.
[0077] In some embodiments, during the real-time prediction phase, when a short-term regional lightning prediction task is triggered for a target sub-region within a target geographic region, the following steps are performed: S21. For the target sub-region, acquire real-time monitoring lightning data of all lightning events occurring within the target sub-region within the historical backtracking time window based on the current real time, and perform the same data preprocessing as in step S11 to obtain a lightning event sequence for prediction arranged in time sequence.
[0078] S22. Based on the predicted lightning event sequence, determine the historical lightning activity area of the target sub-region, and perform high-resolution and low-resolution meshing processing on the historical lightning activity area of the target sub-region using the same method as in step S13. Then, map each lightning event occurring in the target sub-region to a spatial mesh. Using the same method as in step S14, for the current real time, extract visual modal features for the high-resolution mesh array corresponding to the target sub-region; and extract statistical modal features for the low-resolution mesh array corresponding to the target sub-region.
[0079] S23. Use the visual modal features and statistical modal features corresponding to the target sub-region at the current real time as test samples and input them into the trained lightning prediction model.
[0080] S24. The lightning prediction model outputs in parallel the probability distribution of lightning presence, the probability distribution of lightning frequency level, and the probability distribution of lightning intensity level in each spatial grid of the low-resolution grid array corresponding to the target sub-region within a future time window after the current real time, providing meteorological departments with graded and classified refined early warning support.
[0081] Figure 4 This embodiment demonstrates the visualization results of visual modal features during the reasoning process of a real lightning sample (i.e., the test sample) using the trained lightning prediction model, highlighting the core representational role of visual modal features in complex lightning prediction. Figure 4 As shown on the left, the "Visual Modal Input" presents a recent frequency density map generated over the past 20 minutes based on a high-resolution grid array. This recent frequency density map visually reconstructs the spatial distribution structure and local clustering patterns of lightning activity through color gradients. Figure 4The bright yellow highlighted area on the left clearly outlines the high-density center of historical lightning strikes, while the cyan box represents one of the spatial grids selected from the low-resolution grid array, i.e., the target area (G4). This visual feature extraction method can effectively capture the spatial neighborhood texture and evolution trend that is difficult to describe with traditional statistical data. After fusing statistical modal features, the trained lightning prediction model outputs the lightning strike prediction for each spatial grid in the low-resolution grid array within the next 10 minutes. The trained lightning prediction model not only accurately calculates the lightning existence probability of lightning events in each spatial grid in the low-resolution grid array (the lightning existence probability, i.e., the lightning strike probability, of the selected target area reaches 0.82), but also further performs high-confidence classification of the physical attributes of high-risk areas, accurately predicting that the selected target area will be accompanied by "sparse" frequency and have "strong lightning" destructive power, fully verifying the effectiveness of feature learning based on visual images in improving the accuracy of lightning disaster early warning.
[0082] In a second aspect, embodiments of the present invention provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the short-time regional lightning prediction method.
[0083] Thirdly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the short-time regional lightning prediction method.
[0084] To further demonstrate the feasibility and effectiveness of the method of the present invention, experimental verification was conducted.
[0085] In the experiment, the lightning prediction model constructed in this invention is a lightweight deep learning model, and the training and optimization strategy is as follows: Evaluation metrics include: AUC, accuracy, precision, recall (hit rate), F1 score, false alarm rate (FAR), success index (CSI), and Heidke skill score (HSS). Accuracy = (TP+TN) / (TP+TN+FP+FN); Precision = TP / (TP+FP); Recall (also known as hit rate POD) = TP / (TP+FN); F1 score F1 = 2(Precision×Recall) / (Precision+Recall); False alarm rate FAR = FP / (TP+FP); Success index CSI = TP / (TP+FP+FN); Heidke skill score HSS = 2(TP×TN-FP×FN) / ((TP+FN)(FN+TN)+(TP+FP)(FP+TN)); AUC is the area under the ROC curve. Where TP (True Positive) represents true positives, FN (False Negative) represents false negatives, FP (False Positive) represents false positives, and TN (True Negative) represents true negatives. The True Positive Rate (TPR) = TP / (TP+FN), and the False Positive Rate (FPR) = FP / (FP+TN).
[0086] To verify the effectiveness of the method of this invention, this experiment selected two representative mainstream methods for comparative analysis. The first is the comparative patent method (CN115456248A), which represents the traditional technical route of convolutional neural network (CNN) + multi-source meteorological data, focusing on using convolutional layers to extract the spatial static features of radar echoes or meteorological fields; the second is the standard Transformer model, which represents the current mainstream sequence modeling technology route based on self-attention mechanism, focusing on capturing the global temporal dependencies in long sequence data. The lightning prediction model constructed in this invention was rigorously compared with the above two models under the same experimental conditions.
[0087] Table 1 shows a comparison of the core indicators of the method of the present invention and existing technologies in the task of lightning presence prediction.
[0088] Table 1: Comparison of core indicators of different methods in lightning presence prediction task
[0089] As shown in Table 1, the lightning prediction model of this invention demonstrates significantly superior overall performance compared to existing technologies in the lightning presence prediction task of this experimental scenario. Specifically, this invention achieves CSI (Success Index) and HSS (Skill Score) of 0.706 and 0.773, respectively, representing a 152% improvement in CSI compared to the patented method and an 85% improvement compared to the standard Transformer model. Furthermore, this invention successfully addresses the high false alarm rate of traditional methods, significantly reducing FAR (False Alarm Rate) to 0.081 and achieving POD (Point of Detection) of 0.753. The core reason for these performance differences lies in the fact that the patented method relies solely on fully connected layers (MLP) for classification, lacking the ability to perceive the spatiotemporal evolution context, while the lightning prediction model of this invention effectively captures global dependencies. Additionally, compared to the standard Transformer model, the dual-stream feature encoder structure employed in this invention's lightning prediction model achieves deep alignment and complementarity of visual and statistical heterogeneous features. This fully demonstrates that the lightning prediction model of the present invention, while ensuring high sensitivity, greatly improves the reliability and effectiveness of early warning.
[0090] Table 2 shows a comparison of the core indicators of the method of the present invention and the prior art in the tasks of lightning frequency level prediction and lightning intensity level prediction.
[0091] Table 2: Comparison of core indicators of different methods in lightning frequency level prediction and lightning intensity level prediction tasks
[0092] As shown in Table 2, the lightning prediction model of this invention demonstrates excellent comprehensive performance in the task of predicting the frequency and intensity of lightning. In particular, on the Macro-F1, a core metric for measuring class balance, this invention achieves scores of 0.787 and 0.783, respectively; compared to the comparative patent methods' scores of only 0.417 and 0.397, this invention achieves significant improvements of 88.7% and 97.2%, respectively; it also significantly outperforms the standard Transformer model. The lightning prediction model of this invention overcomes the deficiency of spatiotemporal context awareness in the MLP classification of the comparative patent methods by constructing a bidirectional global dependency; while the unique dual-stream feature encoder structure solves the problem that Transformer single-sequence modeling cannot deeply align visual and physical features.
[0093] The underlying reason for the aforementioned significant differences lies in the unique architectural advantages of this invention designed specifically for the physical characteristics of lightning. First, compared to patent CN115456248A, although it utilizes convolutional neural networks to extract spatial features, it primarily relies on fully connected layers (multilayer perceptron, MLP) for simple nonlinear mapping and classification. This architecture is essentially static and lacks the ability to explicitly model the temporal evolution and spatial neighborhood between spatial grids (i.e., "context-aware"), thus failing to perceive the dynamic migration path of thunderstorm clouds. In contrast, the bidirectional state space module introduced in this invention, through a forward and reverse sequence scanning mechanism, constructs a comprehensive global spatiotemporal context dependency while maintaining linear complexity, thereby accurately capturing the generation and dissipation trends of lightning.
[0094] Compared to the standard Transformer model, this invention constructs a dual-stream feature encoder structure. Unlike the Transformer, which simply treats data as a single sequence for implicit attention computation, this structure extracts macroscopic spatial morphology (heatmap) and microscopic physical extrema (58-dimensional statistical modal features) through parallel and independent "visual stream" and "statistical stream," respectively. It also achieves deep alignment and complementarity of heterogeneous modalities through an adaptive average pooling layer. This "visual + physical" dual-stream fusion mechanism provides the lightning prediction model with stronger physical constraints and feature orientation than the purely data-driven self-attention mechanism, thereby achieving a qualitative leap in performance in the highly unbalanced task of strong lightning prediction.
[0095] It should be noted that, to verify the potential of pure lightning data in nowcasting, all benchmark models in this experiment (including the CNN architecture compared to patent CN115456248A) were stripped of expensive multi-source meteorological environmental data and trained and tested on a unified single-modal (pure lightning) dataset. Furthermore, traditional CNNs and standard Transformer models, limited by their single input architecture, find it difficult to directly fuse the 58-dimensional statistical modal features constructed in this invention. Therefore, the experimental results not only reflect the inherent architectural advantages of the bidirectional state-space module in temporal capture but also highlight the irreplaceable nature of the dual-stream multimodal fusion architecture of this invention in effectively utilizing heterogeneous lightning features.
[0096] To thoroughly verify the effectiveness and indispensability of the proposed dual-stream feature encoder and bidirectional state-space core module in lightning presence prediction, this experiment designed a series of ablation experiments. Based on the constructed complete lightning prediction model, the experiments constructed various variant models by successively stripping the bidirectional state-space core module, the visual stream (visual feature encoder) processing the heatmap, and the statistical stream (statistical feature encoder) processing the 58-dimensional statistical modal features. These variant models were then compared and tested under unified experimental conditions.
[0097] Table 3 shows the impact of each variant model on the lightning presence prediction task.
[0098] Table 3: Ablation Experiments – Impact of Variant Models on Lightning Prediction Task
[0099] As shown in Table 3, the ablation experiment, which involved peeling away key components layer by layer, effectively verified the validity and indispensability of the dual-flow feature encoder and bidirectional state-space core module proposed in this invention. Specifically, when the bidirectional state-space core module was removed, the model performance plummeted, with the CSI dropping to only 0.384, demonstrating that the lack of a long-sequence spatiotemporal modeling mechanism would lead to the inability to capture the trend of thunderstorm formation and dissipation. Removing the visual flow, i.e., the visual feature encoder, which processes the heatmap, caused the false alarm rate (FAR) to surge to 0.229, indicating that visual modal features are crucial for suppressing false signals in non-core areas. Removing the statistical flow, i.e., the statistical feature encoder, which contains 58-dimensional statistical modal features, caused the CSI to drop to 0.528, indicating that the lack of physical extreme value constraints would significantly weaken the model's accuracy in discriminating complex convective weather. The complete lightning prediction model performed best across all metrics (CSI 0.706, FAR 0.081), demonstrating a deep synergistic effect among the three elements: "spatial perception of visual flow," "physical constraints of statistical flow," and "spatiotemporal evolution modeling of the bidirectional state space core module," which together ensured the high accuracy and robustness of the prediction.
[0100] Table 4 shows the impact of each variant model on the lightning frequency level prediction task and the lightning intensity level prediction task.
[0101] Table 4: Ablation Experiments – Impact of Variant Models on Lightning Frequency and Intensity Prediction Tasks
[0102] As shown in Table 4, the ablation experiments further revealed the unique contributions of each modality to the task of predicting the refined levels of lightning frequency and intensity. Experimental results show that the complete lightning prediction model significantly outperforms all variant models on the Macro-F1, a core indicator measuring the model's ability to distinguish between dense and destructive lightning. Although the variant model without a statistical encoder achieved an overall accuracy of 0.835 on the lightning intensity level prediction task, not significantly different from the complete lightning prediction model, its Macro-F1 index plummeted from 0.783 to 0.668. This phenomenon of inflated accuracy coupled with a sharp drop in Macro-F1 profoundly reveals the limitations of a single visual modality. While heatmaps can effectively identify "ordinary lightning," which constitutes the majority of samples, the model struggles to distinguish the few but crucial "destructive lightning" events in the absence of physical extreme value constraints. This fully demonstrates the decisive role of statistical modal features in solving the long-tail distribution problem. It compensates for the shortcomings of visual modal features in physical intensity perception, providing an irreplaceable quantitative basis for the model to distinguish between "ordinary risk" and "extreme risk." Furthermore, the removal of the bidirectional state space core module caused the Macro-F1 to drop to 0.525, further confirming the fundamental role of temporal evolution characteristics in judging the development potential of thunderstorms.
Claims
1. A short-time regional lightning prediction method based on multimodal fusion and bidirectional state space, characterized in that, This includes the model training phase and the real-time prediction phase; The model training phase includes: S11. Obtain historical lightning data for the target geographical area and perform data preprocessing to obtain a lightning event sequence arranged in chronological order; S12. Based on the temporal distribution characteristics of lightning events in the lightning event sequence, multiple simulated current moments for feature extraction are generated through a dynamic time window generation strategy. S13. For each sub-region of the target geographic region, select all lightning events that occurred in the sub-region from the lightning event sequence, and then determine a historical lightning activity area in the sub-region; perform high-resolution and low-resolution gridding processing on the historical lightning activity area corresponding to each sub-region to obtain the corresponding high-resolution grid array and low-resolution grid array, and map each lightning event that occurred in the sub-region to the corresponding spatial grid in the high-resolution grid array and low-resolution grid array respectively; S14. For each simulation time, for each sub-region corresponding to the high-resolution grid array, extract the corresponding visual modal features that reflect the spatial morphological distribution of lightning. The visual modal features are multi-channel spatiotemporal heat map tensors generated based on the high-resolution grid array. For each sub-region corresponding to the low-resolution grid array, extract the corresponding statistical modal features that reflect the microscopic physical properties of lightning. The statistical modal features are multi-dimensional statistical feature vectors. S15. Construct a lightning prediction model, which includes: The dual-stream feature encoder includes a parallel visual feature encoder and a statistical feature encoder; the visual feature encoder is used to encode spatial features of visual modal features to generate a visual feature sequence; the statistical feature encoder is used to perform high-dimensional semantic mapping of statistical modal features to generate a statistical feature sequence. The multimodal fusion module is used to fuse visual feature sequences with statistical feature sequences and embed spatial location information to generate multimodal fused feature sequences; The bidirectional state space core module consists of at least two stacked bidirectional state space modules with unified parameter configurations. The bidirectional state space module includes a forward scan branch and a backward scan branch, which are used to perform bidirectional spatiotemporal sequence modeling on multimodal fusion feature sequences and output high-level spatiotemporal feature sequences. A multi-task prediction head is used to output the probability of lightning presence, the probability distribution of lightning frequency level, and the probability distribution of lightning intensity level in parallel based on high-level spatiotemporal feature sequences. S16. Using the visual modal features and statistical modal features corresponding to each sub-region at each simulated current time as training samples, and using the multi-task labels of lightning events mapped to each spatial grid in the low-resolution grid array corresponding to the sub-region within a future time window after the simulated current time as multi-task supervision signals, supervise the training of the lightning prediction model to obtain the trained lightning prediction model; wherein, the multi-task labels include lightning presence labels, lightning frequency level labels, and lightning intensity level labels; The real-time prediction phase involves using a trained lightning prediction model to perform short-term regional lightning predictions for one or more target sub-regions within the target geographical region.
2. The short-time regional lightning prediction method based on multimodal fusion and bidirectional state space according to claim 1, characterized in that, The data preprocessing includes data cleaning and time-series arrangement; The data cleaning process is used to filter out lightning event data from the historical lightning data that are complete and valid in the five key fields of occurrence time, steepness, intensity, longitude and latitude, as valid data entries. The time sequence arrangement is used to arrange the lightning events corresponding to all the valid data entries in chronological order of their occurrence to form the lightning event sequence.
3. The short-time regional lightning prediction method based on multimodal fusion and bidirectional state space according to claim 1 or 2, characterized in that, Step S12 includes: The occurrence time distribution of lightning events in the lightning event sequence is detected to identify dense periods and sparse events on the time axis corresponding to the lightning event sequence. A dense period refers to a time interval containing at least two lightning events, and the time interval between any two consecutive lightning events is no greater than a first threshold. The dense period begins at the occurrence time of the first lightning event it contains and ends at the occurrence time of the last lightning event it contains. A sparse event refers to an isolated lightning event, where the time interval between this lightning event and both its preceding and following lightning events is greater than the first threshold. Based on the dense time periods and the sparse events, execute the dynamic time window generation strategy: For each dense period, starting from its start time, a series of time windows are generated by sliding at a fixed time step, with the start time of each time window serving as a simulation of the current moment; For each sparse event, its occurrence time is used as a simulation of the current moment, and a corresponding time window is generated.
4. The short-time regional lightning prediction method based on multimodal fusion and bidirectional state space according to claim 3, characterized in that, The process for determining the historical lightning activity area is as follows: based on the minimum and maximum longitude, minimum and maximum latitude of all lightning events occurring within the sub-region, a geographical boundary box is defined for the sub-region as the historical lightning activity area.
5. The short-time regional lightning prediction method based on multimodal fusion and bidirectional state space according to claim 2, characterized in that, The high-resolution meshing process is as follows: the historical lightning activity area is divided into M1×N1 spatial meshes to form the high-resolution mesh array; the low-resolution meshing process is as follows: the historical lightning activity area is divided into M2×N2 spatial meshes to form the low-resolution mesh array; wherein, M1 is a positive integer multiple of M2 and N1 is a positive integer multiple of N2; The method for mapping each lightning event to the corresponding spatial grid is as follows: the grid coordinates to which the lightning event belongs are determined based on the longitude and latitude of the lightning event.
6. The short-time regional lightning prediction method based on multimodal fusion and bidirectional state space according to claim 1, characterized in that, In step S14, a historical backtracking time window is defined based on the current simulated time. The historical backtracking time window is divided into a recent observation sub-window and a long-term reference sub-window. The multi-channel spatiotemporal heatmap tensor is constructed by splicing together spatiotemporal heatmaps from five channels along the channel dimension, specifically as follows: The recent frequency density map is generated by: counting the number of lightning events mapped to each spatial grid in the high-resolution grid array within the recent observation sub-window, generating a recent spatial count distribution matrix composed of these numbers, performing two-dimensional Gaussian smoothing on the recent spatial count distribution matrix, and then normalizing the value range to the [0, 1] interval to obtain the recent frequency density map. The maximum intensity distribution map is generated by extracting the absolute values of the intensity of all lightning events mapped to each spatial grid in the high-resolution grid array within the recent observation sub-window, taking the maximum value as the maximum intensity value of the spatial grid, generating a maximum intensity spatial distribution matrix composed of these maximum intensity values, and then normalizing the value range to the interval [0, 1] to obtain the maximum intensity distribution map. The maximum steepness distribution map is generated as follows: extract the absolute steepness values of all lightning events mapped to each spatial grid in the high-resolution grid array within the recent observation sub-window, take the maximum value as the maximum steepness value of the spatial grid, generate the maximum steepness spatial distribution matrix composed of these maximum steepness values, and then normalize the value range to the interval [0, 1] to obtain the maximum steepness distribution map. The time decay density map is generated as follows: the occurrence time of all lightning events mapped to each spatial grid in the high-resolution grid array within the recent observation sub-window is recorded; the time difference between the occurrence time of each lightning event and the current simulation time is calculated; and the time decay weighted value of each lightning event is obtained based on the exponential decay weighting mechanism of the time difference. The time decay weighted values of all lightning events in each spatial grid are accumulated to obtain the time decay value of that spatial grid, and a time decay spatial distribution matrix composed of these time decay values is generated. Then, the value range is normalized to the interval [0, 1] to obtain the time decay density map. The density difference trend map is generated as follows: the number of lightning events mapped to each spatial grid in the high-resolution grid array within the long-term reference sub-window is counted, a long-term spatial count distribution matrix composed of these numbers is generated, and the long-term spatial count distribution matrix is subjected to two-dimensional Gaussian smoothing to obtain a long-term background density map; the long-term background density map is subtracted from the recent frequency density map, and the value range is normalized to the interval [-1, 1] to obtain the density difference trend map.
7. The short-time regional lightning prediction method based on multimodal fusion and bidirectional state space according to claim 6, characterized in that, In step S14, any one of the spatial grids in the low-resolution grid array is defined as the target grid. The multi-dimensional statistical feature vector includes: Global time series features, used to reflect time series patterns, include hourly cycle codes and date cycle codes. The hourly cycle code refers to the feature pairs obtained by performing sine and cosine transformations on the hour value of the simulated current moment in a day. The date cycle code refers to the feature pairs obtained by performing sine and cosine transformations on the yearly day value of the simulated current moment in a year. Local counting features are used to reflect the frequency of lightning activity within the target grid. The local counting features include long-term activity counts and recent activity counts. The long-term activity count refers to the number of lightning events mapped to the target grid within the historical backtracking time window. The recent activity count refers to the number of lightning events mapped to the target grid within the recent observation sub-window. Local physical features, used to reflect the electromagnetic properties of lightning discharges within the target grid, include maximum intensity and average steepness. The maximum intensity is determined as follows: when the long-term activity count is 0, the maximum intensity is 0; when the long-term activity count is greater than 0, the maximum intensity refers to the maximum absolute value of the intensity of all lightning events mapped to the target grid within the historical backtracking time window. The average steepness is determined as follows: when the long-term activity count is 0, the average steepness is 0; when the long-term activity count is greater than 0, the average steepness refers to the average absolute value of the steepness of all lightning events mapped to the target grid within the historical backtracking time window. Local environmental features, used to reflect the spatial context of the target grid, include neighborhood counts and dynamic ratios. The neighborhood counts refer to the number of lightning events mapped to all spatial grids in the low-resolution grid array other than the target grid within the historical backtracking time window. The dynamic ratio is determined as follows: when the long-term activity count is 0, the dynamic ratio is determined to be 0; when the long-term activity count is greater than 0, the dynamic ratio is the ratio of the recent activity count to the long-term activity count.
8. The short-time regional lightning prediction method based on multimodal fusion and bidirectional state space according to claim 1, characterized in that, The visual feature encoder includes a first convolutional block, a first max pooling layer, a second convolutional block, a second max pooling layer, a third convolutional block, an adaptive average pooling layer, and a first flattening layer connected in sequence. The input of the first convolutional block is the visual modality feature, and the output of the first flattening layer is the visual feature sequence. The first convolutional block, the second convolutional block, and the third convolutional block have the same structure, each consisting of a convolutional layer, a batch normalization layer, and a ReLU activation function connected in sequence. The statistical feature encoder includes a first linear layer, a first normalization layer, a first ReLU activation function, a dropout layer, a second linear layer, a second normalization layer, a second ReLU activation function, and a second flattening layer connected in sequence. The input of the first linear layer is the statistical modal feature, and the output of the second flattening layer is the statistical feature sequence. The multimodal fusion module includes: The feature concatenation layer is used to concatenate the visual feature sequence and the statistical feature sequence in the channel dimension to form a multimodal hybrid feature sequence; The feature fusion layer is used to perform feature transformation and dimensionality compression on the multimodal hybrid feature sequence through a multilayer perceptron, so as to realize deep interaction of cross-modal information and form an interactive fusion feature sequence; The location encoding embedding layer is used to initialize a set of learnable spatial location encoding parameters, and to use a broadcast mechanism to superimpose the learnable spatial location encoding parameters onto the interactive fusion feature sequence to form the multimodal fusion feature sequence; The multi-task prediction head adopts a parallel decoupled design, including three independent prediction branches, specifically: The lightning presence prediction branch includes a first output linear layer and a Sigmoid activation function. The input of the first output linear layer is the high-level spatiotemporal feature sequence, and the output of the Sigmoid activation function is the lightning presence probability. The lightning frequency level prediction branch includes a second output linear layer and a first Softmax activation function. The input of the second output linear layer is the high-level spatiotemporal feature sequence, and the output of the first Softmax activation function is the lightning frequency level probability distribution. The lightning intensity level prediction branch includes a third output linear layer and a second Softmax activation function. The input of the third output linear layer is the high-level spatiotemporal feature sequence, and the output of the second Softmax activation function is the lightning intensity level probability distribution.
9. The short-time regional lightning prediction method based on multimodal fusion and bidirectional state space according to claim 8, characterized in that, The bidirectional state space module further includes a third normalization layer. The forward scanning branch includes a forward processing unit, and the reverse scanning branch includes a sequence flipping operation, a reverse processing unit, and a sequence flipping restoration operation. The forward processing unit and the reverse processing unit have the same structure but independent parameters. The input of the third normalization layer is the multimodal fusion feature sequence. One output of the third normalization layer passes through the forward processing unit to obtain the forward scanning feature sequence, and the other output of the third normalization layer passes through the sequence flipping operation, the reverse processing unit, and the sequence flipping restoration operation to obtain the reverse scanning feature sequence. The forward scanning feature sequence and the reverse scanning feature sequence are added element-wise, and the addition result is fused with the multimodal fusion feature sequence through residual connection to obtain the high-level spatiotemporal feature sequence.
10. The short-time regional lightning prediction method based on multimodal fusion and bidirectional state space according to claim 9, characterized in that, The implementation process of the forward processing unit and the reverse processing unit is as follows: the input sequence of the forward processing unit and the reverse processing unit passes through the first projection layer and the first SiLu activation function in one path, and the input sequence passes through the second projection layer, the one-dimensional convolutional layer, the second SiLu activation function, and the selective state space model in another path. The output of the first SiLu activation function is used as a gating signal and multiplied element-wise with the output of the selective state space model to realize the gating mechanism. The output of the gating mechanism passes through the third projection layer to obtain the output sequence of the forward processing unit and the reverse processing unit.
Citation Information
Patent Citations
Method for constructing thunderbolt prediction model based on convolutional neural network
CN115456248A