Hail short-impending prediction method and system based on multi-modal fusion and space-time attention
Patent Information
- Application Number
- CN202610270532.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-06
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-03-06
AI Technical Summary
[0008]本发明的目的在于提供一种基于多模态融合与时空注意力的冰雹短临预测方法及系统,以解决的技术问题在于克服现有冰雹短临预测方法中存在的时空分辨率不足、多源数据融合困难、样本极度不均衡以及预报时效性与精度难以兼顾等缺陷
1.本发明通过融合雷达局地细节、数值模式环境背景及历史事件惯性三类信息,并利用中心偏置注意力聚焦关键区域、双向交互实现模态深度互补,模型对冰雹事件的判别能力大幅增强。消融实验表明,中心偏置空间注意力模块使加权TS评分提升约40%,双向跨模态交互模块带来进一步增益。
Smart Images

Figure CN122218845B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of weather forecasting and artificial intelligence, specifically to a method and system for short-term hail prediction based on multimodal fusion and spatiotemporal attention. Background Technology
[0002] Hail is a typical example of severe convective weather, characterized by its sudden onset, small spatial scale, and severe destructiveness, posing a serious threat to agricultural production, transportation, urban operations, and public safety. Timely and accurate short-term forecasts (0-2 hours ahead) are crucial for effective disaster prevention and mitigation decision-making. However, accurate hail forecasting remains a major challenge in the meteorological field.
[0003] Current hail forecasting technologies mainly rely on the following three methods, but all of them have significant limitations: Physical methods based on numerical weather prediction models: These methods forecast by solving atmospheric dynamics and thermodynamic equations. While robust to large-scale weather pattern prediction, their inherent limitations in computational timeliness and uncertainties in parameterization schemes lead to insufficient spatiotemporal resolution and delayed response to sudden events when capturing small- to medium-scale, rapidly evolving convective systems such as hail. Forecast results typically cannot meet the refined operational requirements of 6-minute intervals and a 5-kilometer radius around the forecasting station.
[0004] Traditional short-term forecasting methods based on radar extrapolation and statistical experience: These methods primarily utilize linear or nonlinear extrapolation from radar echo images, combined with forecaster experience for judgment. Their advantage lies in their high real-time performance, but they are essentially continuous predictions of historical conditions, lacking in-depth modeling of the physical mechanisms of hail formation and development (especially environmental field conditions). Their forecasting ability for nonlinear processes such as hail initiation, intensification, and dissipation is limited, resulting in high false alarm and missed alarm rates.
[0005] Machine learning / deep learning methods based on a single data source: In recent years, some studies have attempted to apply deep learning models (such as convolutional neural networks) to learn forecast patterns directly from radar data. These methods have improved the ability to identify echo patterns to some extent. However, they often overlook the fact that hail events are the product of the combined effects of microscopic convective dynamics and macroscopic environmental thermal conditions. Relying solely on radar data, models struggle to distinguish between ordinary heavy precipitation and hail, and are unable to effectively suppress false alarms under unfavorable environmental conditions. Furthermore, the extremely sparse and unbalanced nature of hail samples in both space and time makes conventional models highly susceptible to being dominated by massive amounts of negative samples, making it difficult to effectively learn the key features of positive samples.
[0006] Furthermore, existing technologies face significant challenges in handling the fusion of multimodal and multiscale meteorological data: there is heterogeneity and scale mismatch between high spatiotemporal resolution radar data (e.g., 6 minutes / 0.01°) used to capture convective details and low spatiotemporal resolution numerical model data (e.g., 1 hour / 0.03°) used to describe environmental background. Simple data stitching or early fusion strategies often lead to unstable model training or fail to fully explore the deep physical relationships between modes.
[0007] Therefore, how to construct a short-term intelligent hail prediction model that can deeply integrate radar micro-dynamic information with the macro-environmental constraints of numerical models, effectively cope with extreme sample imbalance and spatiotemporal heterogeneity of data, and meet the real-time requirements of operations has become a key technical problem that urgently needs to be solved. Summary of the Invention
[0008] The purpose of this invention is to provide a short-term hail prediction method and system based on multimodal fusion and spatiotemporal attention, in order to overcome the shortcomings of existing short-term hail prediction methods, such as insufficient spatiotemporal resolution, difficulty in multi-source data fusion, extreme sample imbalance, and difficulty in balancing forecast timeliness and accuracy.
[0009] To achieve the above objectives, the present invention adopts the following technical solution: A short-term hail prediction method based on multimodal fusion and spatiotemporal attention includes the following steps: S1. Obtain radar data, numerical model forecast data and historical hail observation tag data corresponding to the target site, and perform spatial cropping processing on the radar data and numerical model forecast data centered on the target site, and perform temporal alignment and reconstruction on the cropped data to form a spatiotemporal sample. S2. Input the spatiotemporal samples into a multi-branch deep learning model: use the first feature extraction branch to extract spatiotemporal evolution features reflecting convection intensity and vertical structure from the radar data; use the second feature extraction branch to extract background field features reflecting environmental thermodynamic and microphysical conditions from the numerical model forecast data; use the third feature extraction branch to extract event persistence and temporal inertia features from the historical hail observation tag data; S3. The spatiotemporal evolution features and the background field features are sequentially subjected to center-biased spatial attention weighting and bidirectional cross-modal interaction to obtain enhanced features that fuse local convection details and environmental background constraints; and the enhanced features are spliced and gatedly fused with the temporal inertial features to generate a unified predictive representation. S4. Learn the corresponding time embedding vectors for a series of future forecast periods, and conditionally modulate the time embedding vectors with the unified prediction representation. Through the shared prediction network, output the probability of hail occurrence for each forecast period within the next 0 to 2 hours. Furthermore, the spatial cropping process specifically involves cropping a 61×61 grid area centered on the target station for radar data; and cropping a 33×33 grid area centered on the same station for numerical model data.
[0010] Furthermore, the time series reconstruction adopts the sliding window method: the first 20 consecutive time periods (6 minutes each) are used as the input time window, and the subsequent 20 time periods are used as the prediction time window.
[0011] Furthermore, the sample screening includes: marking samples with hail in the prediction window as positive samples; marking samples with hail in the input window but no hail in the prediction window as first-type negative samples (extinction type); and marking samples with no hail in both the front and back windows as second-type negative samples (calm type).
[0012] Furthermore, the modal independent normalization is as follows: for the eight physical channels of radar data and the four physical channels of numerical model data, minimum and maximum values are set according to their physical reasonable range or climate statistical range, and Min-Max normalization is performed.
[0013] Furthermore, the implementation of the center-biased spatial attention weighting includes: generating a Gaussian mask with the grid center as the peak value; learning a variable attention map from the input features; multiplying the two to obtain the center-weighted final attention map, and recalibrating the features.
[0014] Furthermore, the bidirectional cross-modal interaction is achieved through two parallel cross-attention mechanisms: the first mechanism uses radar features as queries and numerical mode features as keys / values to correct the radar features based on the environmental background; the second mechanism uses numerical mode features as queries and radar features as keys / values to enhance the environmental features based on radar information.
[0015] Furthermore, the training loss function of the multi-branch deep learning model is the FocalLoss function, which incorporates time-sensitive weights, and its expression is: ; in, For the sample size, For the total forecast lead time, For the first Preset weights for timeliness, To predict probabilities for the model, The true label; the preset weight The weighting is consistent with that used in the weighted threat score in business evaluation.
[0016] Furthermore, the present invention also provides a system for implementing the above method, comprising: a data preprocessing module, a multi-branch feature extraction module, a feature fusion and enhancement module integrating a center-biased spatial attention unit and a bidirectional cross-modal interaction unit, and a time-series-aware prediction module.
[0017] As can be seen from the above technical solutions, the present invention has the following technical advantages compared with the prior art: 1. This invention significantly enhances the model's ability to distinguish hail events by fusing three types of information: radar local details, numerical model environmental background, and historical event inertia. It also utilizes a central bias attention module to focus on key areas and bidirectional interaction to achieve modal depth complementarity. Ablation experiments show that the central bias spatial attention module improves the weighted TS score by approximately 40%, and the bidirectional cross-modal interaction module provides further gains.
[0018] 2. The targeted data processing flow and the loss function design of Focal Loss combined with time-effect weighting in this invention effectively alleviate the problem of extreme sample imbalance and directly align the model optimization objective with the core evaluation indicators of short-term forecasting operations.
[0019] 3. The model structure of this invention is lightweight and the inference speed is fast. A single sample prediction only takes about 3 milliseconds, and a round of forecasting for all stations in the country only takes 7-8 seconds, which fully meets the time constraints of business applications. Attached Figure Description
[0020] Figure 1 A flowchart illustrating the steps of a short-term hail forecasting method; Figure 2 A schematic diagram of spatial clipping during sample creation; Figure 3 A schematic diagram illustrating the construction of spatiotemporal samples using a time-sliding window during sample creation; Figure 4 This is a diagram illustrating the overall framework of a multi-branch deep learning model. Figure 5 Comparison of feature response heatmaps before and after Gaussian masking in a center-biased spatial attention module; Figure 6 This is a schematic diagram of the structure and data flow of the two-way cross-modal interaction module; Figure 7 This is a comparison chart of the overall performance with and without a central attention module in the model ablation experiment; Figure 8 This is a comparison chart of the overall performance of models with and without cross-modal interaction modules in the model ablation experiment. Detailed Implementation
[0021] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0022] like Figure 1 The hail short-term prediction method shown mainly includes four steps: multimodal data processing, multi-branch feature extraction, feature fusion and enhancement, and time-series-aware prediction. The core of this invention lies in constructing an end-to-end deep learning model that can receive and collaboratively process information from three heterogeneous data sources: radar, numerical models, and historical observations, and ultimately output a station hail occurrence probability sequence for the next 0-2 hours, incremented by 6 minutes.
[0023] Example 1: Multimodal Data Processing Specifically, the goal of the multimodal data processing step is to process the raw, heterogeneous meteorological data into a model-usable, spatiotemporally aligned, and sample-balanced normalized input. This step specifically includes: First, acquire multi-source data corresponding to the target meteorological station (forecast object), including radar data, numerical model data, and historical observation labels.
[0024] The radar data (RADAR) is acquired from radar data products of the area surrounding the target site, including combined reflectivity (CR), vertical cumulative liquid water content (VIL), and echo top height (ET) and echo bottom height (EB) at different reflectivity thresholds (20dBZ, 30dBZ, 45dBZ), totaling 8 physical channels; the spatiotemporal resolution is 6 minutes, and the horizontal grid spacing is 0.01°.
[0025] The numerical model data (NWP) was acquired from numerical weather prediction products for the area surrounding the target site, and includes four physical channels: 0°C altitude (HT0), -10°C altitude (HT10), -20°C altitude (HT20), and wet-bulb 0°C altitude (HTw0). The spatiotemporal resolution is 1 hour, and the horizontal grid spacing is 0.03°.
[0026] The historical observation tags include hail observation records (0 or 1) for 6 minutes at a time over a period of time at the target site.
[0027] Then, spatial cropping is performed on the radar data and numerical model data centered on the target site: For example... Figure 2As shown, for radar data, a 61×61 (0.01° resolution) square area is cropped centered on the grid point. This area covers the possible movement and impact area of a typical hail convective cell within a short transient period. For numerical model data, a 33×33 (0.03° resolution) square area is cropped centered on the same station; the overall geographical span of the grid is 0.96°×0.96° (approximately 106.6km×106.6km), that is, extending 0.48° (approximately 53.3km) east-west and north-south from the target station, meeting the coverage requirements of a large-scale environmental field.
[0028] like Figure 3 As shown, to address the problem of extremely sparse positive samples in the original data, which is based on "weather processes," this embodiment employs a sliding window method for station-level time series reconstruction. Using a "station-time window" as the basic sample unit, multi-source features within a continuous time window are extracted for each station, with the prediction target being whether hail events will occur within the next 0-2 hours. A time window is defined based on a continuous time series, taking the first 20 time periods (120 minutes in total, T-119 to T0) as the model input observation period, and the next 20 time periods (T+6 to T+120) as the model prediction label period. This 40-time-period window is slid in 6-minute increments, traversing all available data. For each sliding window position, its corresponding 8-channel radar data, 4-channel numerical model data (interpolated to 6-minute resolution in the time dimension), and station historical labels are packaged into a spatiotemporal sample. The sample structure is as follows: radar data shape (20, 61, 61, 8), numerical pattern data shape (5, 33, 33, 4), historical label shape (20, 1), target station latitude, target station longitude, and corresponding radar station number. In this way, the limited process-level data is effectively expanded into a larger set of station-level samples, providing sufficient sample support for deep learning model training.
[0029] Furthermore, to address extreme class imbalances and guide the model to learn key physical processes, this embodiment implements refined sample classification and screening: If hail is observed at least once within the predicted label period (future 20 hours), it is classified as a positive sample. This type of sample is used for the model to learn the characteristic patterns of hail formation. If hail is observed at least once within the input observation period (past 20 hours), but not within the predicted label period, it is classified as a first-class negative sample (Hard Negative / Disappearance type). This type of sample is used to force the model to learn the disappearance mechanism and removal characteristics of convective systems, avoiding the model simply predicting "there will be hail in the future" based on the inertia of "there was hail in the past", thereby effectively reducing the false alarm rate. If no hail is observed at the target station within a total of 40 hours before and after, it is classified as a second-class negative sample (Easy Negative / Calm type). This type of sample is used for the model to identify background characteristics of weather without convection or with ordinary precipitation.
[0030] Furthermore, data cleaning is performed based on physical rationality to remove hail samples with abnormal radar data CR channels: For positive samples, the maximum value of the 5x5 area above the station in the CR channel of all samples is calculated and denoted as CR_Area_max. If CR_Area_max is lower than 10dBZ, it is considered abnormal and removed; for the first type of negative samples, the threshold is set to 30dBZ.
[0031] In multi-source data fusion scenarios, different modalities exhibit significant differences in physical dimensions, numerical ranges, and statistical distributions. Directly inputting the raw data into a deep model can easily lead to unstable gradient updates, slow model convergence, and even a dominant effect of a particular modality on the loss function. Therefore, it is necessary to independently normalize data modalities with different physical dimensions and numerical ranges to ensure the stability of network training and preserve physical meaning. The normalization strategy described in this embodiment follows the following design principles: modal independence, with radar data and numerical model data using independent normalization parameters to avoid numerical interference between different physical variables; channel-level fine-tuning, setting minimum and maximum values for each radar or numerical model physical channel to preserve the relative physical meaning of each channel; physical prior constraints, with the normalization interval set based on historical statistics and a reasonable physical range, rather than relying entirely on sample statistics, to enhance the model's robustness to extreme samples; and consistency during the inference phase, using completely consistent normalization parameters during training and inference to ensure stable model input distribution.
[0032] Specifically, the raw numerical ranges of the eight physical channels of radar data (CR, EB20, EB30, EB45, ET20, ET30, ET45, and VIL) differ significantly. Regarding the... Each radar channel, its normalization process is defined as: ; In the formula: Indicates the first Each radar channel in time Spatial location The original observations; and These represent the preset minimum and maximum threshold values for this channel, respectively. This indicates a truncation operation, used to suppress outliers; , representing the numerical stability term; the upper and lower limits of the specific normalization interval for each channel of radar data are shown in Table 1 below: Table 1. Upper and lower limits of the normalization interval for each radar data channel.
[0033] The normalization process for numerical model data is consistent with that for radar data. However, since numerical model data primarily describes the large-scale environmental field surrounding the site, including four height- or thermodynamically related physical quantities (HT0, HT10, HT20, HTw0), their dimensions and range of variation are significantly larger than those of radar echo data. Therefore, a separate normalization interval is set for numerical model data. The upper and lower limits of the specific normalization intervals for each channel of the numerical model are shown in Table 2 below. Table 2. Upper and lower limits of the normalization interval for each channel in the numerical model.
[0034] By employing independent Min-Max normalization for each modality, the model retains physical meaning while effectively mitigating the problem of inconsistent numerical scales between different modalities, thus improving the training stability of multi-branch networks. Through truncation and physical prior constraints, the model's robustness to anomalous echoes and extreme environmental fields is enhanced. All input features are mapped to the [0,1] interval, facilitating subsequent numerical modeling of attention mechanisms and gating structures. The model is highly compatible with actual business inference processes and possesses good reproducibility and deployability.
[0035] Example 2: Construction of Deep Learning Models like Figure 4 As shown, this embodiment addresses the short-term prediction problem of hail event probability at a single site by constructing a multi-source spatiotemporal deep learning model that integrates radar data, numerical model data, and historical observation labels. The multi-source spatiotemporal deep learning model is designed with "site-centric decision-making" as its core principle. Through multi-branch feature extraction, spatiotemporal attention mechanisms, and cross-modal interaction structures, it achieves refined modeling of the evolution of local severe convection. The multi-branch feature extraction includes radar feature branches, numerical model feature branches, and historical observation branches.
[0036] The multi-source spatiotemporal deep learning model specifically includes: The input layer contains three types of core meteorological data: Radar data branch: accessing radar multi-channel volume scan sequences with dimensions (8, 20, 61, 61), directly carrying real-time intensity and vertical structure information of convective systems; Numerical model branch: accessing numerical model element sequences with dimensions (4, 5, 33, 33), providing thermal and microphysical background constraints for hail occurrence; Historical data branch: accessing historical tag sequences with dimensions (20, 1), recording the temporal correlation of events. The feature extraction layer employs independent feature extraction sub-modules tailored to the characteristics of different data. Radar branch: A 3D convolutional network extracts the spatiotemporal evolution features of radar echoes, then a central attention module filters edge noise, and finally, temporal attention converges the temporal information to obtain a refined representation of localized strong convection. Numerical model branch: A 3D convolutional network captures the spatial distribution and temporal trends of the environmental field, combined with a central attention module to highlight the environmental constraints of the target area, outputting structured features of the large-scale background. Historical data branch: A BiLSTM network encodes the temporal dependence of historical label sequences, extracting prior information such as "event persistence and recent occurrence trends," generating a 128-dimensional historical inertial feature vector. The interactive fusion layer, based on same-dimensional alignment, achieves deep complementarity of multi-source features through a multi-layer interactive mechanism: First, it establishes a bidirectional association between radar and numerical model features through cross-modal attention, realizing nonlinear fusion of the two types of information; then, it concatenates the fused features with the 128-dimensional vector of the historical real-time branch, and then uses a gated fusion module to adaptively suppress the noise dimension, obtaining a unified representation that takes into account "local details, environmental background, and historical inertia"; The output layer first learns a unique time embedding vector for each forecast lead time, encoding the forecast characteristics of different lead times; then, through a time-series gating module, it dynamically combines the fused representation with the time embedding to achieve conditional modulation for each lead time; finally, through a shared prediction head, it outputs the probability of hail occurrence for each lead time, ensuring the consistency of the discrimination logic and adapting to the prediction needs of different lead times.
[0037] The center-biased spatial attention module described in this embodiment utilizes a central region prior constructed from samples to improve the response to key regions and suppress edge noise, thereby enhancing the effective signal-to-noise ratio of radar / numerical model spatial features. The input supports two forms: one is intermediate features of the radar branch, a 5D tensor. Another type is the intermediate feature of the numerical pattern branch, a 4D tensor: The output has the same shape as the input, performing point-by-point weighting: And when the input is 5D, Transformed into Attention is calculated frame by frame, and then the shape is restored to ensure consistency with 4D implementation and controllable computational overhead.
[0038] The computation process begins with attention generation, which is performed on each spatial feature map. conduct The 2D convolution operation generates a single-channel spatial attention weight map and normalizes it using Sigmoid: Secondly, construct a Gaussian mask consistent with the spatial scale. The center has a large weight, focusing on core features; the edge has a small weight, reducing noise interference. ; in: It refers to the coordinate position within the grid; Indicates the coordinates of the grid center; The standard deviation of the Gaussian distribution is used to control the decay rate. It is a natural exponential function; Next, the learnable attention is multiplied by the fixed prior to obtain the biased attention: The output is the point-by-point recalibrated features: .
[0039] like Figure 5 The image shows the heatmap changes in feature responses before and after the introduction of a Gaussian mask in the center-biased spatial attention module. Left image (without Gaussian mask): The learnable spatial attention response is relatively dispersed, with many redundant signals in the edge regions; Middle image (Gaussian mask alone): The Gaussian distribution centered on the site directly strengthens the weights of the core region, forming a fixed prior of "high in the center, low at the edges"; Right image (final attention after fusion): Combining learnable attention with a Gaussian mask strengthens key information about the echo / environmental field around the site and provides interpretable spatial constraints for the model; it can also directly embed Conv3D / Conv2D structures without additional supervision signals, improving the signal-to-noise ratio of features without increasing training costs.
[0040] Furthermore, to break down the barriers between radar and numerical model information and achieve deep complementarity, this embodiment designs a bidirectional cross-modal interaction module (BCIM). Since traditional multimodal models often use simple feature splicing, it is difficult to characterize the asymmetric dependencies between modalities. By introducing bidirectional cross-modal interaction, the complementary relationship between "radar-numerical model" is explicitly modeled before fusion, enabling the two modalities to exchange information after dimensional alignment. The radar uses the numerical model for background correction, and the numerical model uses the radar for event response enhancement, thereby improving discrimination capability and robustness.
[0041] like Figure 6As shown, the Bidirectional Cross-Modal Interaction Module (BCIM) comprises two parallel cross-attention submodules. In the first submodule, using F_radar as the query and F_nwp as the key and value, the attention of radar features to environmental features is calculated to obtain the radar feature F_radar_corr after environmental background correction. This allows the model to determine the hail formation potential of the current radar echo based on large-scale environmental conditions (such as whether the 0°C layer height is suitable). In the second submodule, using F_nwp as the query and F_radar as the key and value, the attention of environmental features to radar features is calculated to obtain the environmental feature F_nwp_aug after radar information enhancement. This allows the model to select the most relevant environmental field configuration based on the current convection intensity.
[0042] Subsequently, F_radar_corr, F_nwp_aug, and F_hist are concatenated to obtain a joint feature vector of (B, 1152). This vector is passed through a fully connected layer with a gated mechanism to adaptively adjust the contribution weights of each modality feature, and finally outputs a unified and highly discriminative (B, 1024) dimensional fused feature vector F_fused.
[0043] Furthermore, in order to meet the multi-time-leadership forecasting requirements of "0-2 hours, 6 minutes" and to distinguish the focus of different lead time forecasts, this invention designs a time-aware output module (TAOM).
[0044] First, for each of the 20 forecast lead times (6, 12, ..., 120 minutes), an independent 1024-dimensional temporal embedding vector E_t is learned. Then, each E_t is modulated using a gating vector G generated by F_fused. ; Where: ⊙ represents element-wise addition or a specific fusion operation. This step enables conditional feature construction, allowing the same fusion state to produce differentiated representations under different forecast lead times.
[0045] Finally, all time-dependent Z_t values share the same lightweight multilayer perceptron (MLP) prediction head, and each value independently outputs the hail occurrence probability P_t corresponding to its time-dependent time through the sigmoid activation function, forming the final probability sequence (P_1, P_2, ..., P_20).
[0046] To address the sparse hail samples and operational evaluation requirements, the deep learning model described in this preferred embodiment employs the following combined loss function: ; FocalLoss is used to address class imbalance, with parameters set to α=0.7 and γ=2 to increase the model's attention to hard-to-classify positive samples. w_t is the time-sensitive weight vector, whose values are set according to the weighted threat score (Weighted TS) in the business (e.g., [0.10, 0.09, 0.09, 0.08, ..., 0.005]), so that the model optimization objective is directly aligned with the final business evaluation index.
[0047] Furthermore, the model training and optimization strategy includes: Dataset partitioning: Stratified block sampling was adopted (block size set to 120), and a positive to negative sample ratio of 1:3 was set; the total sample was divided into training set (approximately 70%), validation set (approximately 10%) and test set (approximately 20%) to ensure that data from the same weather event was not leaked across sets; Optimizer: The AdamW optimizer is used, and the learning rate is scheduled using a cosine annealing strategy, which enables fast search in the early stage of training and smooth convergence in the later stage, reducing oscillations and improving the final generalization performance. Regularization strategies include: adding Gaussian noise to the input data for data augmentation; using Dropout in the fully connected layers of the network; and implementing an early stopping strategy based on the weighted threat score (TS) metric on the validation set to prevent overfitting. Training configuration: The development machine is configured with a 16-core CPU, 120GB of memory, dual GPUs (4090-24GB), PyTorch 2.5.1, and CUDA 12.1; Hyperparameter settings: batch size set to 32; number of training epochs set to 50; early stopping steps set to 5; AdamW optimizer used. Default; initial learning rate lr is set to Weight decay is set to ; The learning rate scheduling CosineAnnealingLR uses T_max=20 and η_min=1e-6.
[0048] To verify the effectiveness of the embodiments of the present invention, a comprehensive experiment was conducted on a specified severe convective weather dataset. Table 3 below shows the ablation experiment results of the key modules: Table 3 Ablation Experiment Results of Key Modules
[0049] Experiments show that: Center-biased spatial attention (CB-SAM) improves weighted forecast accuracy (TS) by about 40%, and bidirectional cross-modal interaction (BCIM) further improves weighted TS by about 7%. The CB-SAM module, by focusing on the core area of the site, is the most critical factor in improving forecast accuracy (TS) and controlling false alarms and false misses. The BCIM module further improves the physical consistency and discriminative power of the model by fusing environmental field information.
[0050] Furthermore, the performance of this embodiment of the invention is verified through forecast accuracy and operational efficiency: Forecast accuracy: The model using the method of this invention achieved a weighted TS of 0.5637, a false alarm rate of 0.3139, and a missed alarm rate of 0.2551 on a large-scale self-built test set, demonstrating good forecast performance and robust generalization ability.
[0051] Operational efficiency: The model has approximately 7.64 million parameters. On a single NVIDIA RTX 4090 GPU, the inference time for 11,724 test samples is 36.3 seconds, with an average inference time of approximately 3 milliseconds per sample. Based on this, it is estimated that completing a hail probability forecast for the next two hours for all national-level meteorological stations (approximately 2,000) across the country would only take 7-8 seconds per round of calculation, fully meeting the timeliness requirements of operational real-time forecasting.
[0052] In summary, the technical solution provided by the embodiments of the present invention effectively solves the core problems of multimodal fusion, sample imbalance, and spatiotemporal modeling in short-term hail prediction through a complete technical chain from data processing and model architecture to training optimization. It achieves a balance between high accuracy and high real-time performance and has significant business application value.
[0053] The above-described embodiments are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A short-term hail prediction method based on multimodal fusion and spatiotemporal attention, characterized in that, Includes the following steps: S1. Obtain radar data, numerical model forecast data and historical hail observation tag data corresponding to the target site, and perform spatial cropping processing on the radar data and numerical model forecast data centered on the target site, and perform temporal alignment and reconstruction on the cropped data to form a spatiotemporal sample. S2. Input the spatiotemporal samples into a multi-branch deep learning model: use the first feature extraction branch to extract spatiotemporal evolution features reflecting convection intensity and vertical structure from the radar data; use the second feature extraction branch to extract background field features reflecting environmental thermodynamic and microphysical conditions from the numerical model forecast data; use the third feature extraction branch to extract event persistence and temporal inertia features from the historical hail observation tag data; S3. The spatiotemporal evolution features and the background field features are sequentially subjected to center-biased spatial attention weighting and bidirectional cross-modal interaction to obtain enhanced features that integrate local convection details and environmental background constraints. The enhanced features are then concatenated and gated with the temporal inertial features to generate a unified predictive representation. S4. Learn the corresponding time embedding vectors for each of the future forecast periods, and conditionally modulate the time embedding vectors with the unified prediction representation. Through the shared prediction network, output the probability of hail occurrence for each forecast period within the next 0 to 2 hours.
2. The hail short-term prediction method based on multimodal fusion and spatiotemporal attention according to claim 1, characterized in that, The spatial clipping process includes: For radar data, a square area with a side length of 61 grid points is cropped with the target site as the center. For numerical model forecast data, a square region with a side length of 33 grid points is cropped with the same target site as the center.
3. The hail short-term prediction method based on multimodal fusion and spatiotemporal attention according to claim 1, characterized in that, The temporal alignment and reconstruction of the cropped data includes: The input time window is constructed using multiple consecutive historical time periods, and the prediction time window is constructed using multiple consecutive future time periods. The length of the input time window is 20 time intervals, the length of the prediction time window is 20 time intervals, and the time interval between each time interval is 6 minutes.
4. The hail short-term prediction method based on multimodal fusion and spatiotemporal attention according to claim 1, characterized in that, Step S1 further includes filtering the spatiotemporal samples: The sample that shows hail occurring at least once within the future prediction time window is marked as a positive sample; Samples that experienced hail within the past input time window but did not experience hail within the future prediction time window are labeled as Class I negative samples; Samples in which no hail occurred within both the past input time window and the future prediction time window were labeled as negative samples of the second type.
5. The hail short-term prediction method based on multimodal fusion and spatiotemporal attention according to claim 1, characterized in that, Step S1 further includes performing modal-independent normalization on the radar data and numerical model forecast data respectively: For each physical channel of radar data, minimum and maximum values are set according to the reasonable range of their physical quantities, and Min-Max normalization is performed. For each physical channel of the numerical model forecast data, minimum and maximum values are set according to its climatological statistical range, and Min-Max normalization is performed.
6. The hail short-term prediction method based on multimodal fusion and spatiotemporal attention according to claim 1, characterized in that, Center-biased spatial attention weighting is achieved in the following way: Generate a Gaussian distribution mask with the same spatial size as the input feature map and a peak value at the center of the grid. A variable attention weight map is learned from the input feature map through convolutional layers; The variable attention weight map is multiplied element-wise with the Gaussian distribution mask to obtain the final attention map with center bias. The input features are weighted element-wise using the final attention map.
7. The hail short-term prediction method based on multimodal fusion and spatiotemporal attention according to claim 1, characterized in that, The bidirectional cross-modal interaction is achieved through two parallel cross-attention mechanisms: The first cross-attention mechanism: using the spatiotemporal evolution features as the query vector and the background field features as the key vector and value vector, the radar features after environmental background correction are calculated. The second cross-attention mechanism uses the background field features as the query vector and the spatiotemporal evolution features as the key vector and value vector to calculate the environmental features enhanced by radar information.
8. The short-term hail prediction method based on multimodal fusion and spatiotemporal attention according to claim 1, characterized in that, The conditional modulation is implemented through a gating mechanism: A set of gating coefficients is generated based on the unified prediction representation; The time embedding vector corresponding to each forecast lead time is multiplied element-wise with the gating coefficient to obtain the modulated lead time-specific vector; The modulated time-specific vector is combined with the unified prediction representation and input into a shared prediction network.
9. The hail short-term prediction method based on multimodal fusion and spatiotemporal attention according to claim 1, characterized in that, The multi-branch deep learning model is trained using the following loss function: The loss function is the Focal Loss function, which incorporates time-related weights, and its expression is as follows: ; in, For the sample size, For the total forecast lead time, For the first Preset weights for timeliness, To predict probabilities for the model, The true label; the preset weight The weighting is consistent with that used in the weighted threat score in business evaluation.
10. A short-term hail prediction system based on multimodal fusion and spatiotemporal attention, used to implement the method of any one of claims 1 to 9, characterized in that, The system includes: The data preprocessing module is used to perform the steps described in S1 of claim 1; A multi-branch feature extraction module is used to perform the step described in S2 of claim 1; The feature fusion and enhancement module integrates a center-biased spatial attention unit and a bidirectional cross-modal interaction unit, and is used to perform the steps described in S3 of claim 1. A time-aware prediction module is used to perform the steps described in S4 of claim 1.
Citation Information
Patent Citations
Deep learning thunder and lightning prediction method and system based on space-time attention mechanism, and storage medium
CN120030329A
Short-time rainfall prediction method based on radar image and reanalysis data fusion
CN120337179A