A storm wind speed field prediction method based on ST-LSTM space-time attention and physical perception driving
Patent Information
- Application Number
- CN202611090981.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-22
- Publication Date
- 2026-09-25
AI Technical Summary
然而,NWP模型对初始条件极为敏感,观测数据的微小误差会导致预报结果偏差累积
[0017]本发明的有益效果为:(1)通过跨模态多头注意力机制,实现了NWP数值预报数据与观测数据的深度语义对齐融合,充分利用了异构数据中的互补信息,提高了预测精度。
Smart Images

Figure CN122815577A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of interdisciplinary technology of weather forecasting and deep learning, and in particular to a method for predicting thunderstorm wind speed fields based on ST-LSTM spatiotemporal attention and physical perception. Background Technology
[0002] Thunderstorm wind gusts are strong gusts generated by thunderstorm weather systems, often accompanied by hail, heavy rain, and other severe convective weather, posing a serious threat to aviation, power supply, buildings, and the safety of people's lives and property. Accurately predicting the spatiotemporal distribution of thunderstorm wind speed fields is of great significance for disaster prevention and mitigation.
[0003] Currently, thunderstorm and strong wind forecasts mainly rely on the following methods: (1) Numerical Weather Prediction (NWP) method. This method is based on atmospheric dynamics and thermodynamics equations and uses supercomputers to simulate atmospheric evolution. However, NWP models are extremely sensitive to initial conditions, and small errors in observational data can lead to the accumulation of forecast biases. In addition, NWP models have limited spatial resolution, making it difficult to accurately capture the local fine structure of thunderstorms and strong winds, and their computational timeliness is insufficient, making it difficult to meet the needs of short-term and nowcasting.
[0004] (2) Traditional ConvLSTM sequence prediction method. Convolutional Long Short-Term Memory (ConvLSTM) networks, by replacing matrix multiplication in fully connected LSTM with convolution operations, can simultaneously model spatiotemporal dependencies and have been applied to meteorological sequence prediction. However, traditional ConvLSTM only maintains two state vectors, the hidden state h and the cell state c, and lacks the ability to explicitly model spatiotemporal association memory. When dealing with long sequence prediction tasks, spatiotemporal information decay is likely to occur.
[0005] (3) Simple splicing and fusion method. Existing methods usually fuse NWP forecast data and observation data by splicing channels, which fails to learn the deep semantic alignment relationship between different modal data, resulting in the complementary information in heterogeneous data not being fully utilized.
[0006] (4) Traditional loss function. Existing deep learning weather forecasting methods usually use mean squared error (MSE) or mean absolute error (MAE) as training targets. These loss functions cannot guarantee the spatial continuity and temporal consistency of the forecast results, which may lead to physically unreasonable abrupt changes or isolated noise in the predicted wind speed field.
[0007] Therefore, how to effectively integrate multi-source heterogeneous meteorological data, enhance the ability to model long-range spatiotemporal dependencies, and introduce physical consistency constraints to achieve high-precision prediction of thunderstorm wind speed fields is a problem that current technology urgently needs to solve. Summary of the Invention
[0008] The purpose of this invention is to provide a method for predicting the wind speed field of thunderstorms based on ST-LSTM spatiotemporal attention and physical perception, so as to achieve high-precision hourly grid-point prediction of the wind speed field of thunderstorms.
[0009] To achieve the above objectives, the present invention provides the following solution: A method for predicting thunderstorm wind speed fields based on ST-LSTM spatiotemporal attention and physical perception includes: Multi-source meteorological data are acquired and preprocessed to construct a spatiotemporal sequence dataset, wherein the spatiotemporal sequence dataset includes the numerical weather prediction (NWP) factor for past times, the observation labels for past times, the numerical weather prediction (NWP) factor for future times, and the terrain elevation for each time step. The spatiotemporal sequence dataset is input into a model driven by ST-LSTM spatiotemporal attention and physical perception for wind speed prediction. The model is trained by using a weighted sum of meteorological score loss and physical constraint loss as the training objective. The model includes a CBAM-enhanced ST-LSTM encoder and a CBAM-enhanced ST-LSTM decoder. After preprocessing the multi-source meteorological data to be predicted, it is input into the trained model, outputs the wind speed prediction results, and performs correction processing to obtain the final thunderstorm wind speed field prediction results.
[0010] Optionally, inputting the spatiotemporal sequence dataset into a model driven by ST-LSTM spatiotemporal attention and physical perception for wind speed prediction includes: The input data corresponding to each time step in the spatiotemporal sequence dataset is fused using a cross-modal multi-head attention mechanism to obtain fused input features; The fused input features are input into a CBAM-enhanced ST-LSTM encoder to generate the final hidden state and then CBAM attention weighting is applied. The weighted final hidden state is input into the CBAM-enhanced ST-LSTM decoder. The output of each decoding time step is input into the CBAM attention weighting and then sent to the output layer to obtain the wind speed prediction result.
[0011] Optionally, feature fusion can be performed to obtain fused input features, including: The NWP factors and observation labels of past times are input into the first multi-head cross-modal attention module. Attention is calculated using the observation labels of past times as the Query and the NWP factors of past times as the Key and Value, to obtain the fused past features. The NWP factor of the future time and the fused past feature are input into the second multi-head cross-modal attention module. Attention is calculated using the fused past feature as the Query and the NWP factor of the future time as the Key and Value to obtain the fused future feature. The fused past features, the fused future features, and the terrain elevation are stitched together and then subjected to convolutional projection to obtain the fused input features.
[0012] Optionally, the fused input features are input into a CBAM-enhanced ST-LSTM encoder to generate the final hidden state, including: Apply CBAM attention to the fused input features to obtain attention-weighted features; The attention-weighted features are input into two layers of ST-LSTM units for time-step encoding. The cell state is updated by standard LSTM gating, the spatiotemporal memory state is updated by spatiotemporal memory gating, and the cell state and spatiotemporal memory state are fused by 1×1 convolution to generate the final hidden state.
[0013] Optionally, updating the spatiotemporal memory state via spatiotemporal memory gating includes: ; ; in, Let t be the spatiotemporal memory state. For the portal to time and space oblivion This represents the spatiotemporal memory state at time t-1. As a spacetime input gate, As a candidate gate for spacetime, The hidden state at time t-1 Input at time t, For activation function, The input and hidden states are fed into a convolutional layer with three spatiotemporal memory gates, and ⊙ represents element-wise multiplication.
[0014] Optionally, the meteorological scoring loss includes Pearson correlation coefficient, multi-threshold weighted TS score, and exponential MAE penalty term; the physical constraint loss includes spatial smoothness constraint based on Laplacian operator, boundary constraint based on Sobel operator, and time consistency constraint.
[0015] Optionally, the correction processing for the wind speed prediction results includes: target-based denoising, quantile mapping frequency matching, conditional morphological dilation, and threshold adsorption processing.
[0016] Optionally, the CBAM attention weighting is implemented through a CBAM attention module including a channel attention submodule and a spatial attention submodule. The channel attention submodule obtains two channel description vectors through two parallel global pooling paths and feeds them into a shared two-layer fully connected network. The outputs of the two layers are added together and then a sigmoid activation function is used to generate channel attention weights. The spatial attention submodule calculates the average and maximum values along the channel dimension, concatenates the two to obtain a dual-channel feature map, and generates spatial attention weights through convolutional layers and a sigmoid activation function.
[0017] The beneficial effects of the present invention are as follows: (1) Through the cross-modal multi-head attention mechanism, the deep semantic alignment and fusion of NWP numerical prediction data and observation data are realized, making full use of the complementary information in the heterogeneous data and improving the prediction accuracy.
[0018] (2) By using the ST-LSTM three-state memory unit, an independent spatiotemporal memory state is added on the basis of the traditional ConvLSTM. By using the dual-gated structure to model short-term changes and long-term spatiotemporal correlations respectively, the modeling ability of thunderstorm wind evolution trend is significantly enhanced.
[0019] (3) By using the CBAM dual attention mechanism, the importance weights of meteorological element channels and spatial regions are adaptively learned, so that the model focuses on the most critical features and regions for thunderstorm and strong wind prediction.
[0020] (4) By using a physical constraint composite loss function, physical laws such as spatial continuity, boundary constraints and temporal consistency constraints are embedded into the training process to ensure that the prediction results meet the basic meteorological constraints, thereby improving the physical rationality and operational availability of the prediction.
[0021] (5) Through the four-stage statistical post-processing pipeline, the frequency deviation and spatial noise of the prediction results were effectively corrected, further improving the final forecast score. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart of a thunderstorm wind speed field prediction method based on ST-LSTM spatiotemporal attention and physical perception driven by an embodiment of the present invention. Figure 2 A schematic diagram of a multi-head cross-modal attention structure provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the ST-LSTM spatiotemporal memory unit structure according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the CBAM attention mechanism structure according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the physical constraint composite loss function structure according to an embodiment of the present invention; Figure 6 This is a flowchart of the deep learning-based strong wind prediction result post-processing pipeline according to an embodiment of the present invention. Figure 7 This is a schematic diagram of the improved Seq2Seq model architecture according to an embodiment of the present invention. Detailed Implementation
[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0026] like Figure 1 As shown, this embodiment proposes a method for predicting thunderstorm wind speed fields based on ST-LSTM spatiotemporal attention and physical perception, including: Multi-source meteorological data are acquired and preprocessed to construct a spatiotemporal sequence dataset, wherein the spatiotemporal sequence dataset includes the numerical weather prediction (NWP) factor for past times, the observation labels for past times, the numerical weather prediction (NWP) factor for future times, and the terrain elevation for each time step. The spatiotemporal sequence dataset is input into a model driven by ST-LSTM spatiotemporal attention and physical perception for wind speed prediction. The model is trained by using a weighted sum of meteorological score loss and physical constraint loss as the training objective. The model includes a CBAM-enhanced ST-LSTM encoder and a CBAM-enhanced ST-LSTM decoder. After preprocessing the multi-source meteorological data to be predicted, it is input into the trained model, outputs the wind speed prediction results, and performs correction processing to obtain the final thunderstorm wind speed field prediction results.
[0027] Specifically, the data sources in this implementation include three categories: (1) Numerical weather forecast (NWP) data, which includes 25 meteorological factors: CAPE (convective available potential energy), CIN (convective suppression energy), DCAPE (descending convective available potential energy), DT705 (700-500 hPa temperature difference), DT855 (850-500 hPa temperature difference), HTO (0°C layer height), LCL (lifted condensation height), LFC (free convection height), LI300 (300 hPa lifting index), LI500 (500 hPa lifting index), MLCIN (mixed layer convection). (1) Suppression energy), Q1000 (1000hPa specific humidity), Q850 (850hPa specific humidity), RH1000 / RH500 / RH700 / RH850 (relative humidity of each layer), SBCIN (convection suppression energy based on the surface), SWEAT (strong weather threat index), T700 (700hPa temperature), TdSfc850 (dew point difference from the surface to 850hPa), VP01 / VP03 / VP06 (1 / 3 / 6 hour cumulative precipitation), WS925 (925hPa wind speed); (2) Historical thunderstorm and strong wind observation tag data, the unit is 0.1 m / s, the value -9 indicates missing measurement, when processing, the missing measurement value is set to zero and the unit is converted to m / s; (3) SRTM topographic elevation data.
[0028] The preprocessing steps include: (1) Spatial resolution unification: unify all data to a 301×301 grid, use bilinear interpolation to scale NWP factors, and use the rasterio library to crop and bilinearly resample terrain data according to latitude and longitude range; (2) Outlier handling: set the -9999 marker value in NWP data to zero, set the -9 missing value in label data to zero, and uniformly perform NaN detection and replacement; (3) Normalization: based on the mean and standard deviation of each NWP factor, the mean and standard deviation of terrain data, and the mean and standard deviation of label data, perform Z-score standardization on the input data respectively.
[0029] The dataset was constructed as follows: spatiotemporal sequences were organized by case, with each sample containing data from 40 consecutive time steps. The first 20 steps were the input sequence (including past NWPs, past observation labels, future NWPs, and terrain), and the last 20 steps were the target sequence (observation labels). The dataset was randomly divided into a training set (64%), a validation set (16%), and a test set (20%) based on the cases.
[0030] Furthermore, inputting the spatiotemporal sequence dataset into a model driven by ST-LSTM spatiotemporal attention and physical perception for wind speed prediction includes: The input data corresponding to each time step in the spatiotemporal sequence dataset is fused using a cross-modal multi-head attention mechanism to obtain fused input features; The fused input features are input into a CBAM-enhanced ST-LSTM encoder to generate the final hidden state and then CBAM attention weighting is applied. The weighted final hidden state is input into the CBAM-enhanced ST-LSTM decoder. The output of each decoding time step is input into the CBAM attention weighting and then sent to the output layer to obtain the wind speed prediction result.
[0031] Furthermore, feature fusion is performed to obtain fused input features, including: The NWP factors and observation labels of past times are input into the first multi-head cross-modal attention module. Attention is calculated using the observation labels of past times as the Query and the NWP factors of past times as the Key and Value, to obtain the fused past features. The NWP factor of the future time and the fused past feature are input into the second multi-head cross-modal attention module. Attention is calculated using the fused past feature as the Query and the NWP factor of the future time as the Key and Value to obtain the fused future feature. The fused past features, the fused future features, and the terrain elevation are stitched together and then subjected to convolutional projection to obtain the fused input features.
[0032] Specifically, cross-modal multi-head attention feature fusion includes: For each input time step, the 52-channel input features are split into four modalities: past NWP (25 channels), past observation labels (1 channel), future NWP (25 channels), and terrain (1 channel).
[0033] The first cross-modal attention module processes past data: using past observation label features as the Query (projected from 1 channel to the hidden_channels dimension via a 1×1 convolution), and past NWP features as the Key and Value (projected from 25 channels to the hidden_channels dimension via a 1×1 convolution), cross-modal alignment weights are calculated using 4-head scaled dot product attention. Specifically, the Query, Key, and Value are reshaped into the form (b, num_heads, head_dim, H×W), and the attention matrix is calculated using einsum: attn = softmax( ),in =hidden_channels / num_heads. The attention output is projected back to the observation channel dimension via a 1×1 convolution, multiplied by the learnable parameter γ, and then concatenated with the residual of the original observation features to obtain the fused past features.
[0034] The second cross-modal attention module processes future data: using the fused past features as the query and the future NWP features as the key and value, it performs the same multi-head attention calculation to obtain the fused future features.
[0035] Finally, the fused past features (1-channel dimension), fused future features (1-channel dimension), and terrain features (1-channel) are concatenated into a 3-channel feature, which is then projected onto the hidden_channels dimension through a fusion layer of 3×3 convolution + BatchNorm + ReLU.
[0036] like Figure 2 As shown, the cross-modal multi-head attention module of this invention includes a Query projection layer, a Key projection layer, a Value projection layer, a multi-head attention computation unit, and a learnable residual gating. The core formula is: Attention(Q,K,V)= Output_NWP=GatedResidual(NWP, Attn(Obs,NWP)); Output_obs=GatedResidual(Obs, Attn(NWP,Obs)), where d is the feature dimension of each attention head.
[0037] The Query projection layer is a 1×1 convolutional layer (without bias) that projects the observed features (obs_channels) onto the hidden_channels dimension. The Key projection layer and Value projection layer are both 1×1 convolutional layers (without bias) that project the NWP features (nwp_channels) onto the hidden_channels dimension.
[0038] In multi-head mode, `hidden_channels` are evenly divided into 4 heads (`num_heads`), each with a dimension of `head_dim = hidden_channels / 4`. For each head, the Query, Key, and Value are reshaped to the form (`b`, `head_dim`, `H×W`), and the attention matrix is calculated using a scaled dot product. ; in, This is a scaling factor to prevent the dot product value from becoming too large, which could cause the gradient to vanish. The attention output is: ; The outputs of all heads are concatenated and then projected through a 1×1 convolutional layer to restore the obs_channels dimension.
[0039] The learnable parameter γ (initialized to 0) controls the degree of residual fusion: ; When γ=0, the output equals the original observed features (identity mapping); as training progresses, γ gradually increases, and the model learns an effective cross-modal fusion strategy. This design avoids the training instability problem caused by the scale mismatch of cross-modal features in the early stages of training.
[0040] Furthermore, the fused input features are input into the CBAM-enhanced ST-LSTM encoder to generate the final hidden state, including: Apply CBAM attention to the fused input features to obtain attention-weighted features; The attention-weighted features are input into two layers of ST-LSTM units for time-step encoding. The cell state is updated by standard LSTM gating, the spatiotemporal memory state is updated by spatiotemporal memory gating, and the cell state and spatiotemporal memory state are fused by 1×1 convolution to generate the final hidden state.
[0041] Specifically, the workflow of the CBAM-enhanced ST-LSTM encoder includes: First, apply CBAM attention to the fused input features: Channel attention part: Global average pooling (AdaptiveAvgPool2d(1)) and global max pooling (AdaptiveMaxPool2d(1)) are performed on the input features to obtain two C-dimensional vectors, which are fed into a shared two-layer MLP (the first layer has a channel compression ratio of 1 / 8 and the number of output channels is max(C / 8, 8), and the second layer is restored to C channels). The outputs of the two layers are added together and then activated by Sigmoid to obtain C-dimensional channel attention weights, which are multiplied with the input features channel by channel.
[0042] Spatial attention component: The mean and maximum values of the features are calculated along the channel dimension to obtain two H×W feature maps. These are then concatenated into a 2-channel feature map and H×W spatial attention weights are obtained through 7×7 convolution and Sigmoid activation. These weights are then multiplied by the features at each spatial position.
[0043] The CBAM-weighted features are fed into a two-layer ST-LSTM encoder. The number of hidden channels in the encoder is [64, 64].
[0044] The first layer ST-LSTM unit (input_dim=64, hidden_dim=64, kernel_size=3×3): For each time step, the input is concatenated with the hidden state of the previous layer, and then processed... The convolution yields a 4×64=256-channel output, which is then split into four 64-channel gating signals (input gate i, forget gate f, candidate gate g, and output gate). ), activated by Sigmoid and Tanh respectively. Simultaneously through... Convolution yields a 3×64=192-channel output, which is then split into three 64-channel spatiotemporal memory gating signals (spatiotemporal input gate i', spatiotemporal forget gate f', and spatiotemporal candidate gate g'). Cell state spatiotemporal memory state Output gate fusion and : Ultimately hidden state .in, Let t be the spatiotemporal memory state. For the portal to time and space oblivion This represents the spatiotemporal memory state at time t-1. As a spacetime input gate, As a candidate gate for spacetime, The hidden state at time t-1 Input at time t, For activation function, The input and hidden states are fed into a convolutional layer with three spatiotemporal memory gates, and ⊙ represents element-wise multiplication.
[0045] The second-layer ST-LSTM unit (input_dim=64, hidden_dim=64, kernel_size=3×3) takes the hidden state h1 of the first layer as input and performs the same computation. After the two layers of encoding are completed, CBAM attention is applied to the final hidden state h2 of the second layer.
[0046] The workflow of the CBAM-enhanced ST-LSTM decoder includes: The two layers of ST-LSTM units in the decoder are initialized with the final state of the encoder. The first layer of the decoder (input_dim=64, hidden_dim=64) takes the hidden state of the second layer of the encoder as input, and the second layer of the decoder (input_dim=64, hidden_dim=64) takes the hidden state of the first layer as input.
[0047] Autoregressive decoding has 20 time steps: In each time step, the two ST-LSTM layers update the state sequentially, and the hidden state of the second layer is fed into the output layer after being weighted by CBAM attention. The output layer consists of 3×3 convolutions (64→32 channels) + BatchNorm + ReLU + 1×1 convolutions (32→1 channel), and finally the output value is ensured to be non-negative by the Softplus activation function.
[0048] like Figure 3 As shown, the ST-LSTM unit adds an independent spatiotemporal memory state M and three spatiotemporal memory gates to the traditional ConvLSTM.
[0049] Traditional ConvLSTM The layer will input and hidden state After concatenation, the signals are mapped to 4×hidden_dim channels via convolution, generating four standard gated signals. The ST-LSTM of this invention further enhances this process. Layer, also After concatenation, the signals are mapped to 3×hidden_dim channels via convolution, generating three spatiotemporal memory-gated signals. Additionally, [further details are needed]. Layer will The convolutional mapping after concatenation is applied to the hidden_dim channel, which is used to modulate the output gate. Layers will The spliced 2×hidden_dim channel is projected as the hidden_dim channel and used to calculate the final hidden state.
[0050] The key innovation lies in the dual-path state update mechanism: the cell state C is updated through standard LSTM gating to capture short-term change features; the spatiotemporal memory state M is updated through independent gating to capture long-term spatiotemporal correlations. The two states are fused through 1×1 convolution, so that the final hidden state contains both short-term and long-term information.
[0051] During initialization, there are three states. , , All are initialized to zero tensors.
[0052] like Figure 4 As shown, the CBAM attention module contains sequentially connected channel attention submodules and spatial attention submodules.
[0053] The channel attention submodule comprises two parallel global pooling paths: an average pooling path (AdaptiveAvgPool2d(1), which compresses H×W features into a 1×1 global descriptor) and a max pooling path (AdaptiveMaxPool2d(1)). The two pooling results share a two-layer MLP: the first layer compresses the number of channels from C to max(C / 8, 8) and uses ReLU activation; the second layer restores the number of channels to C. The MLP outputs from both paths are element-wise summed and then activated by Sigmoid to generate a C-dimensional channel attention vector. Final feature = Input feature × Mc (broadcast multiplication).
[0054] The spatial attention submodule calculates the average and maximum values along the channel dimension, resulting in two H×W feature maps. These are concatenated into a 2-channel feature map and then activated by a 7×7 convolution (without bias) and a sigmoid function to generate the spatial attention weight map. Final feature = Channel attention-weighted feature × Ms.
[0055] The deployment locations of CBAM in the model include: (1) Input layer CBAM (cbam_input): applying attention to the fused input features, reduction=8; (2) Encoder CBAM (cbam_encoder): applying attention to the final output of the encoder, reduction=8; (3) Decoder CBAM (cbam_decoder): applying attention to the output of each decoding step, reduction=8.
[0056] Furthermore, the meteorological scoring loss includes Pearson correlation coefficient, multi-threshold weighted TS score, and exponential MAE penalty term; the physical constraint loss includes spatial smoothness constraint based on Laplacian operator, boundary constraint based on Sobel operator, and time consistency constraint.
[0057] Specifically, training uses the AdamW optimizer with a learning rate of 0.0005, weight decay of 1e-2, and gradient clipping threshold of 1.0. Due to the high model complexity, a strategy of batch size=2 combined with 16-step gradient accumulation (equivalent to batch size=32) is adopted. The learning rate scheduling uses the ReduceLROnPlateau strategy with a decay factor of 0.2 and a patience value of 5 epochs. The maximum number of training epochs is 200, and the early stopping patience value is 15 epochs.
[0058] The loss function is the combined loss: CombinedLoss = L_weather + 0.1 × L_physics.
[0059] The meteorological score loss L_weather is based on the severe convective weather forecast assessment formula: ; Time step weight A bell-shaped distribution is adopted: [0.0075, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 0.1, 0.09, 0.08, 0.07, 0.06, 0.05, 0.04, 0.03, 0.02, 0.0075, 0.005], with the middle time intervals (steps 9-10) having the largest weights.
[0060] Threshold weight = [0.1, 0.1, 0.1, 0.2, 0.2, 0.3], corresponding to 6 wind speed thresholds: 0.0, 5.5, 10.8, 13.9, 17.2, 24.5 m / s, with higher thresholds (thunderstorm winds and extreme wind levels) having greater weight.
[0061] Let be the Pearson correlation coefficient between the prediction and the target at step k. For the TS score of the i-th threshold at step k, L_weather is the mean absolute error of the active region (the region where the prediction or target is greater than 0). L_weather = -Scorep, meaning that maximizing the score is equivalent to minimizing the negative score.
[0062] The physical constraint loss L_physics consists of three components: (1) Spatial smoothness loss: The predicted sequence is reshaped to (b×T×C, 1, H, W), and the second spatial derivative is calculated by convolution with a 3×3 Laplacian kernel [[0,1,0],[1,-4,1],[0,1,0]]. The mean of the L2 norm is taken and the weight is 0.01.
[0063] (2) Boundary constraint loss: The horizontal and vertical boundary gradients of the prediction and the target are calculated using 3×3 Sobel_x kernel [[-1,0,1],[-2,0,2],[-1,0,1]] and Sobel_y kernel [[-1,-2,-1],[0,0,0],[1,2,1]], respectively. The mean of the L2 norm of the difference between the boundary gradients of the two is calculated with a weight of 0.005.
[0064] (3) Time consistency constraint loss: Calculate the difference between the prediction and the target at adjacent time steps (step t+1 minus step t), take the mean of the L2 norm of the difference, and weight it 0.01.
[0065] like Figure 5As shown, the composite loss function consists of two parts: meteorological score loss and physical constraint loss.
[0066] Weather score loss directly aligns with business assessment indicators The evaluation includes three levels: (1) The Pearson correlation coefficient rk measures the degree of linear correlation between the prediction and the target, through... The transformation maps the score to a positive range; (2) The multi-threshold weighted TS score is calculated as TS=TP / (TP+FP+FN) under the six key business wind speed thresholds, where TP, FP, and FN are the number of hit, false alarm, and missed alarm grid points, respectively, and the higher thresholds (17.2 m / s and 24.5 m / s) are given greater weight; (3) The exponential MAE penalty term exp(-MAEik) makes the prediction error and the TS score form a joint optimization objective.
[0067] Physical constraint loss constrains the physical rationality of the prediction results from three dimensions: (1) Spatial smoothness: the Laplacian operator penalizes the second spatial derivative of the prediction field to suppress isolated noise and unreasonable abrupt changes; (2) Boundary constraint: the Sobel operator constrains the prediction to be consistent with the boundary gradient direction of the target, maintaining the spatial structure characteristics of the wind speed field; (3) Temporal consistency constraint: the time constraint ensures that the prediction is consistent with the temporal change trend of the target, preventing physically unreasonable time jumps in the prediction sequence.
[0068] The total loss is: L_total = L_weather + 0.1 × L_physics, where 0.1 is the total weight of the physical constraints, ensuring that the physical constraints assist in the optimization of the weather score rather than dominate the training process.
[0069] Further correction processing of wind speed prediction results includes: target-based denoising, quantile mapping frequency matching, conditional morphological dilation, and threshold adsorption processing.
[0070] Specifically, such as Figure 6 As shown, the original model output (in int16 format with m / s precision, retaining 0.1 m / s precision) undergoes four stages of post-processing: Phase 1, Targeted Denoising: The OpenCV function `connectedComponentsWithStats` (8-connected) is used to perform connected component analysis on regions ≥17.2 m / s in the predicted wind speed field. All connected regions are traversed; regions with an area smaller than a preset threshold are identified as noise, and their values are reduced to 10.0 m / s (100 × 0.1) to prevent them from participating in subsequent high-wind-speed enhancement. Identifying and removing these isolated regions through connected component analysis ensures the spatial continuity of the prediction results.
[0071] Phase Two, Quantile Mapping Frequency Matching, addresses the systematic frequency bias in model output. Due to the regression characteristics of deep learning models, predicted value distributions are often more concentrated than the true value distribution (i.e., smaller variance), leading to frequency bias in high-frequency or low-frequency events. By mapping the 14-point distribution, the predicted cumulative distribution function is corrected to match the true value distribution of the training set, effectively restoring the predicted frequency of extreme events. Based on the training data, the true value distribution gt_dist and the predicted distribution pred_dist at the 14 quantile points are obtained: gt_dist = [0, 27, 47, 62, 86, 111, 127, 141, 160, 180, 249, 288, 330,410] (unit: 0.1 m / s); pred_dist = [0, 27, 43, 55, 76, 97, 112, 124, 144, 164, 208, 229, 253, 254] (unit: 0.1 m / s).
[0072] Using NumPy's linear interpolation function np.interp, the predicted values are mapped from pred_dist to gt_dist, correcting the cumulative distribution function of the predicted results and aligning the predicted distribution with the true distribution.
[0073] Phase Three, Conditional Morphological Dilation: For wind speed areas exceeding 16.5 m / s (165 × 0.1) after frequency matching, a morphological dilation is performed using a 3 × 3 rectangular structuring element. The newly added pixel value is the larger of the current region's maximum value and 166 (16.6 m / s), which expands the high-wind-speed area while avoiding creating excessively high values. By conditionally dilating high-wind-speed areas, the impact range of strong winds is expanded while maintaining reasonable wind speed values, thus improving the hit rate of strong wind events.
[0074] Phase Four, Threshold Absorption: Predicted values close to critical wind speed thresholds are absorbed to positions slightly above the thresholds. Near critical thresholds such as 17.2 m / s (blue gale warning standard) and 24.5 m / s (yellow gale warning standard), even small fluctuations in predicted values can lead to abrupt changes in warning levels. By absorbing values near the thresholds to positions slightly above them, the stability of warning decisions is enhanced. Specifically, values in the range of 168 to 172 (16.8-17.2 m / s) are absorbed to 172.5 (17.25 m / s), and values in the range of 240 to 245 (24.0-24.5 m / s) are absorbed to 245.5 (24.55 m / s). This operation ensures that predicted values trigger warnings near the critical thresholds of operational assessment.
[0075] The final result is cropped to the range [0, 700] (0-70 m / s) and rounded down to int16 for storage.
[0076] To verify the effectiveness of this invention, a systematic experiment was conducted in the "Intelligent Application Innovation Challenge of Severe Convective Weather Training Dataset" organized by the China Meteorological Administration.
[0077] Experimental data: The training set contains multiple thunderstorm and strong wind cases, each consisting of observations over 40 consecutive time steps on a 301×301 spatial grid with a time resolution of 1 hour. The 25 NWP factors are derived from ERA5 reanalysis data, and the terrain data is from the SRTM digital elevation model.
[0078] Baseline model: Standard ConvLSTM Seq2Seq model with [32, 32] hidden channels, trained using SevereWeatherLoss.
[0079] like Figure 7 As shown, the improved model (in this invention) is: ST-LSTM Seq2Seq + CBAM + cross-modal attention + physical constraint loss, with the number of hidden channels being [64, 64], 4-head cross-modal attention, and trained using CombinedLoss.
[0080] Experimental results show that the method of the present invention achieves significant improvements over the baseline model in the following aspects: (1) On the test sets of the competition A and B lists, the overall score Score_p of the method of this invention is significantly higher than that of the baseline ConvLSTM model, which verifies the synergistic effect of the four innovations.
[0081] (2) The comparison of TS scores under different wind speed thresholds shows that the method of the present invention has a particularly significant improvement at high thresholds (17.2 m / s and 24.5 m / s), indicating that physical constraint loss and statistical post-processing have obvious advantages in predicting extreme wind events.
[0082] (3) The comparison of TS scores at different forecast lead times shows that the decay rate of the method of the present invention is lower than that of the baseline model in long forecast lead times (15-20 steps), indicating that the ST-LSTM three-state memory unit effectively slows down the information decay in long sequence prediction.
[0083] (4) Ablation experiments show that the four innovations (CBAM attention, ST-LSTM, physical constraint loss, and cross-modal attention) each contribute positive benefits independently, and the combined use of these innovations yields the best results.
[0084] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A method for predicting thunderstorm wind speed fields based on ST-LSTM spatiotemporal attention and physical perception, characterized in that, include: Multi-source meteorological data are acquired and preprocessed to construct a spatiotemporal sequence dataset, wherein the spatiotemporal sequence dataset includes the numerical weather prediction (NWP) factor for past times, the observation labels for past times, the numerical weather prediction (NWP) factor for future times, and the terrain elevation for each time step. The spatiotemporal sequence dataset is input into a model driven by ST-LSTM spatiotemporal attention and physical perception for wind speed prediction. The model is trained by using a weighted sum of meteorological score loss and physical constraint loss as the training objective. The model includes a CBAM-enhanced ST-LSTM encoder and a CBAM-enhanced ST-LSTM decoder. After preprocessing the multi-source meteorological data to be predicted, it is input into the trained model, outputs the wind speed prediction results, and performs correction processing to obtain the final thunderstorm wind speed field prediction results.
2. The method for predicting thunderstorm wind speed fields based on ST-LSTM spatiotemporal attention and physical perception as described in claim 1, characterized in that, Inputting the aforementioned spatiotemporal sequence dataset into a model driven by ST-LSTM spatiotemporal attention and physical perception for wind speed prediction includes: The input data corresponding to each time step in the spatiotemporal sequence dataset is fused using a cross-modal multi-head attention mechanism to obtain fused input features; The fused input features are input into a CBAM-enhanced ST-LSTM encoder to generate the final hidden state and then CBAM attention weighting is applied. The weighted final hidden state is input into the CBAM-enhanced ST-LSTM decoder. The output of each decoding time step is input into the CBAM attention weighting and then sent to the output layer to obtain the wind speed prediction result.
3. The method for predicting thunderstorm wind speed fields based on ST-LSTM spatiotemporal attention and physical perception as described in claim 2, characterized in that, Feature fusion is performed to obtain the fused input features, including: The NWP factors and observation labels of past times are input into the first multi-head cross-modal attention module. Attention is calculated using the observation labels of past times as the Query and the NWP factors of past times as the Key and Value, to obtain the fused past features. The NWP factor of the future time and the fused past feature are input into the second multi-head cross-modal attention module. Attention is calculated using the fused past feature as the Query and the NWP factor of the future time as the Key and Value to obtain the fused future feature. The fused past features, the fused future features, and the terrain elevation are stitched together and then subjected to convolutional projection to obtain the fused input features.
4. The method for predicting thunderstorm wind speed fields based on ST-LSTM spatiotemporal attention and physical perception as described in claim 2, characterized in that, The fused input features are input into a CBAM-enhanced ST-LSTM encoder to generate the final hidden state, which includes: Apply CBAM attention to the fused input features to obtain attention-weighted features; The attention-weighted features are input into two layers of ST-LSTM units for time-step encoding. The cell state is updated by standard LSTM gating, the spatiotemporal memory state is updated by spatiotemporal memory gating, and the cell state and spatiotemporal memory state are fused by 1×1 convolution to generate the final hidden state.
5. The method for predicting thunderstorm wind speed fields based on ST-LSTM spatiotemporal attention and physical perception as described in claim 4, characterized in that, Updating the state of spacetime memory via spacetime memory gating includes: ; ; in, Let t be the spatiotemporal memory state. For the portal to time and space oblivion This represents the spatiotemporal memory state at time t-1. As a spacetime input gate, As a candidate gate for spacetime, The hidden state at time t-1 Input at time t, For activation function, The input and hidden states are fed into a convolutional layer with three spatiotemporal memory gates, and ⊙ represents element-wise multiplication.
6. The method for predicting thunderstorm wind speed fields based on ST-LSTM spatiotemporal attention and physical perception as described in claim 1, characterized in that, The meteorological scoring loss includes Pearson correlation coefficient, multi-threshold weighted TS score, and exponential MAE penalty term; the physical constraint loss includes spatial smoothness constraint based on Laplacian operator, boundary constraint based on Sobel operator, and time consistency constraint.
7. The method for predicting thunderstorm wind speed fields based on ST-LSTM spatiotemporal attention and physical perception as described in claim 1, characterized in that, Correction processing of wind speed prediction results includes: target-based denoising, quantile mapping frequency matching, conditional morphological dilation, and threshold adsorption processing.
8. The method for predicting thunderstorm wind speed fields based on ST-LSTM spatiotemporal attention and physical perception as described in claim 2, characterized in that, The CBAM attention weighting is implemented through a CBAM attention module that includes a channel attention submodule and a spatial attention submodule. The channel attention submodule obtains two channel description vectors through two parallel global pooling paths and feeds them into a shared two-layer fully connected network. The outputs of the two layers are added together and then a sigmoid activation function is used to generate channel attention weights. The spatial attention submodule calculates the average and maximum values along the channel dimension, concatenates the two to obtain a dual-channel feature map, and generates spatial attention weights through convolutional layers and a sigmoid activation function.