Short temporary rainfall prediction method based on multi-source multi-temporal-spatial-feature fusion

Through multi-source data fusion and multi-level feature extraction methods, the problem of insufficient single-source data in existing methods is solved, efficient feature expression and accurate prediction of radar echoes and meteorological satellite data are achieved, and the accuracy of short-term precipitation prediction is improved.

CN120635694APending Publication Date: 2025-09-12SOUTHEAST UNIV
View PDF 0 Cites 11 Cited by

Patent Information

Application Number
CN202510646540.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing precipitation prediction methods mostly rely on single-source radar echo data, ignore important information from meteorological satellite data, and fail to fully explore the spatiotemporal correlations between multi-source data, resulting in low accuracy in identifying complex precipitation scenarios.

Method used

A multi-source data framework that fuses radar echoes and meteorological satellite data is adopted. Through multi-spatial and multi-temporal aggregators, multi-scale patch embedding, cross-scale perception optimization and multi-temporal self-attention modules, combined with Transformer and CNN blocks, multi-granularity feature extraction and hierarchical information transmission are achieved. A bidirectional bridging fusion module is designed to alleviate the differences in multi-source data and enhance feature expression.

Benefits of technology

The accuracy of short-term precipitation forecasts has been significantly improved, especially the key success index and Heidecker skill score in heavy precipitation forecasts have been significantly improved, and the model's ability to extract and fuse features in the spatial and temporal dimensions of radar echo sequences has been enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635694A_ABST
    Figure CN120635694A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source multi-temporal-spatial feature fusion short and temporary rainfall prediction method (M4Cester), which realizes layered temporal-spatial feature extraction through a multi-space multi-time aggregator (MMA), and comprises a multi-scale patch embedding (MSPE) module, a cross-scale perception optimization (CPR) module and a multi-time self-attention (MTS) module. The first and second spatial features are respectively used for capturing local-global spatial features, dynamically balancing cross-scale information and mining a time sequence dependency relationship; meanwhile, a bidirectional bridging fusion module (MSFM) is designed, bidirectional alignment and enhancement of radar and satellite features are realized by using a cross attention mechanism, and modal difference is relieved through residual connection. Experiments on a weather data set in the Yangtze River Delta region show that the key success index (CSI) reaches 0.267 and the HSS reaches 0.372 in heavy rainfall (greater than or equal to 50dBZ) prediction by the method, which are obviously improved compared with the existing advanced model, particularly, the limitation of single-source data is effectively overcome in the prediction of a convection initiation (CI) event, and a high-precision solution is provided for short and temporary rainfall prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of spatiotemporal sequence prediction in computer vision, and in particular relates to a short-term precipitation prediction method based on the fusion of multi-source and multi-spatiotemporal features. Background Art

[0002] Short-term precipitation forecasting has long been a hot topic in meteorological research. It is primarily used to accurately predict the intensity and scale of precipitation within a short timeframe. It significantly impacts the daily lives of many people and is crucial for weather guidance in agriculture, commerce, transportation, and other sectors. Due to the inherent complexity of the atmosphere and related dynamical processes, real-time, large-scale precipitation forecasting is a very challenging problem. Traditional numerical weather prediction algorithms are widely used in precipitation forecasting. However, the prediction process involves simulating complex atmospheric physics equations, and therefore typically requires extensive computing resources and hours of computation time.

[0003] Another class of precipitation nowcasting methods is based on radar echo extrapolation. These methods ignore the complex atmospheric dynamics and simply predict future echo maps from past radar echo sequences. Optical flow-based methods have proven effective for extrapolating echo maps. However, they have limitations due to the assumptions of Lagrangian persistence and smooth motion fields. Furthermore, they are unsupervised and fail to leverage abundant echo data. Recently, supervised deep learning techniques have been used to capture the complex nonlinear spatiotemporal patterns of precipitation nowcasting. First, several RNN / LSTM (Long-Short-Term Memory)-based methods were proposed for weather forecasting. Subsequently, Shi et al. combined convolution with recurrent architectures to propose the convolutional LSTM (ConvLSTM) model, using convolutional operations instead of fully connected layers for LSTM state-to-state transitions. They further proposed the trajectory GRU (TrajGRU) model, which utilizes a subnetwork to actively learn the positional variation structure of recurrent connections, achieving good prediction accuracy on the real-world precipitation benchmark HKO-7. In addition, the Predictive Recurrent Neural Network (PredRNN) and its improved version (PredRNN++) have also been validated in the radar echo prediction challenge. Gradient highway units and a new recurrent structure called causal LSTM are used to alleviate the difficulties of gradient propagation and achieve better spatiotemporal prediction learning on synthetic and real datasets. However, these RNN-based methods often suffer from the vanishing gradient problem and require memory-bandwidth-limited computation.

[0004] In addition to including recurrence, the time series of frames can also be processed as part of the convolutional architecture and encoded along the channel dimension. A CNN method is also proposed to transform the convective storm nowcasting problem into a classification problem. They further proposed a UNet-based FCN (Fully Convolution Networks) model for precipitation nowcasting. Experimental results show that a simple custom FCN can achieve almost the same performance as TrajGRU. Alternatively, an attention module and depth-wise separable convolution are proposed in SmaAt-UNet. Compared with the original UNet, this method achieved better prediction performance on a real precipitation map dataset in the Netherlands, while using only a quarter of the trainable parameters. However, convolution is generally an effective method for extracting local features. Some current Transformer models require global spatiotemporal features, which is very beneficial to further improve the ability of heavy rain forecasting.

[0005] Unlike convolutional operations, which focus on local features, the self-attention mechanism in the Transformer effectively captures a global view and exploits global features. Recent research has demonstrated the advantages of combining convolution and the Transformer for precipitation nowcasting. For example, the proposed AA-TransUnet network, which combines a U-Net and a Transformer, was used for precipitation forecasting in the Netherlands, and the Rainformer model, which combines a Swin-Transformer, a U-Net, and a grid fusion unit, has achieved good results. However, existing methods still have significant shortcomings. First, most studies rely solely on single-source radar echo data, ignoring important information such as cloud top temperature and water vapor distribution contained in meteorological satellite data. This information is crucial for the formation, development, and evolution of precipitation systems. The fusion of multi-source data can more comprehensively characterize precipitation processes. Second, existing models fail to fully represent the spatial and temporal characteristics of radar echo sequences and fail to effectively exploit the spatiotemporal correlations between multi-source data. This results in low accuracy in recognizing complex precipitation scenarios, particularly processes such as the onset of convection and the evolution of heavy precipitation. Therefore, it is urgent to propose a new method to achieve deep fusion of multi-source data and efficient extraction of spatiotemporal features to meet the practical application needs of high-precision short-term precipitation prediction. Summary of the Invention

[0006] Objective of the invention: The objective of the present invention is to provide a short-term precipitation prediction method (M4Caster) based on the fusion of multi-source and multi-spatiotemporal features. Firstly, in view of the shortcomings of existing single-source data radar echo extrapolation and pure convolution / recurrent networks in spatiotemporal feature modeling, a multi-source data framework of radar echo and meteorological satellite data is adopted; then, by introducing a multi-spatial multi-temporal aggregator (MMA), which includes a multi-scale patch embedding (MSPE) module, the original pure convolution module is replaced by a parallel structure of Transformer and CNN blocks to realize multi-granularity feature extraction from local texture to global context, obtain more precipitation feature information, and enhance the multi-scale prediction capability of the model; then, a multi-level fusion structure composed of cross-scale perception optimization (CPR) and multi-temporal self-attention (MTS) modules is used to realize the inter-level information transmission of local, medium and global features, fully combining CNN's attention to local information and Transformer's attention to local information. The model takes advantage of global dependency to strengthen the model's feature expression of the spatial dimension of radar echo sequences. At the same time, a cross-temporal Transformer algorithm is adopted to mine the temporal dependency of radar echo sequences through the multi-temporal self-attention (MTS) module, learn the temporal variation trend of short-term precipitation, and enhance the model's feature expression in the temporal dimension. In addition, a bidirectional bridging fusion module (MSFM) is designed to use the cross-attention mechanism to achieve bidirectional alignment and enhancement of radar and satellite features, and alleviate modal differences through residual connections. Ultimately, the improved model significantly enhances the feature extraction and fusion capabilities of the spatial and temporal dimensions of radar echo sequences and multi-source data, greatly improving the accuracy of short-term precipitation prediction, especially in the prediction of heavy precipitation (≥50dBZ). The critical success index (CSI) reaches 0.267 and the Hydeke skill score (HSS) reaches 0.372, which are significantly improved compared with existing advanced models.

[0007] Technical solution: To achieve this purpose, the present invention adopts the following technical solution:

[0008] The present invention provides a method for predicting short-term precipitation by integrating multiple sources and multiple spatiotemporal features, comprising the following steps:

[0009] S1: Preprocess radar and satellite precipitation datasets, perform denoising and normalization, and divide the datasets to obtain cleaned training samples;

[0010] S2: Using the radar satellite encoder sub-block containing the multi-spatial multi-temporal aggregator (MMA) and the bidirectional bridge fusion module (MSFM), we obtain feature extraction networks at different stages to extract local, medium, and global fusion features at different levels;

[0011] S3: Multi-scale patch embedding (MSPE), cross-scale perception optimization (CPR), and multi-temporal self-attention (MTS) modules are used to obtain multi-level satellite and radar features at each stage, realize information transfer between different levels, capture local-global spatial features, dynamically balance cross-scale information, and explore temporal dependencies, thereby strengthening the model's feature representation of the spatial dimension of radar echo sequences and satellite sequences.

[0012] S4: Design a bidirectional bridging fusion module (MSFM) to obtain the fusion module of radar and satellite data. It uses the cross-attention mechanism to achieve bidirectional alignment and enhancement of radar and satellite features, and alleviates the differences in multi-source data through residual connections.

[0013] S5: Add the residual attention module (ResATT) decoder to obtain a short-term precipitation prediction model based on the fusion of radar echo and satellite data. M4Caster is trained and tested to obtain radar echo extrapolation prediction results with a 1-hour forecast time.

[0014] As a further technical solution of the present invention, in step S1, the data preprocessing method includes image denoising, satellite data channel selection, and normalization. The data preprocessing method of the present invention is:

[0015] S11: Obtain the raw data of radar and satellite precipitation maps. We divide the dataset into several precipitation events, of which 28,953 precipitation events are selected as the training set and the remaining 7,992 precipitation events are used as the validation set;

[0016] S12: De-noising the radar data: radar echo intensity values ​​below 15 dBZ and above 70 dBZ are usually noise or non-precipitation areas such as hail. The radar echo intensity of radar image points below 15 dBZ and above 70 dBZ is set to 0.

[0017] S13: Channel selection and unified spatial resolution of satellite data: First, the correlation between each satellite channel and the radar echo is evaluated using the Pearson correlation analysis method, and channels 8 (WV), 11 (MI), 13 (IR), and 15 (I2) with high correlation are selected for subsequent modeling;

[0018] S14: Normalization processing: For satellite observation data, the maximum and minimum values ​​of channels 8, 11, 13, and 15 are counted. For radar data, the range is 0-70, and then the normalization formula is applied to map it to the interval [0,1].

[0019] As a further technical solution of the present invention, in step S2, the radar satellite encoder sub-block including the multi-spatial multi-temporal aggregator (MMA) and the bidirectional bridge fusion module (MSFM) is used to extract local, medium, and global fusion features at different levels.

[0020] S21: Multiple downsampling (Conv) operations and multi-spatial multi-temporal aggregator (MMA) modules are stacked in the encoder part. Each downsampling operation halves the input image size and doubles the number of feature maps. The encoding process is described as: STF i+1 = Radar satellite encoder Conv(STF i ), where STF i (i=0, 1, 2, 3) represents the spatiotemporal features at different scales, STF0 is the input radar echo or satellite sequence, and i is the downsampling stage number;

[0021] S22: Enhance spatiotemporal representation capabilities by balancing the local processing focused on by convolutional operations and the global context captured by Transformer operations through the Multi-Spatial Multi-Temporal Aggregator (MMA) module;

[0022] S23: A two-way bridge mechanism is used between radar and satellite encoders to exchange information between multi-source radar and satellite data to capture richer cloud phase information.

[0023] S24: Adopts an encoder-predictor architecture, with the encoder followed by the same number of decoders, and finally uses bilinear upsampling in the prediction network to halve the number of feature maps and double the size of the feature maps.

[0024] As a further technical solution of the present invention, in step S3, multi-scale patch embedding (MSPE), cross-scale perception optimization (CPR) and multi-temporal self-attention (MTS) modules are added to obtain multi-level satellite and radar features at each stage:

[0025] S31: Separate the first pure convolutional network structure of the encoder of the original UNet network and replace it with a multi-scale patch embedding (MSPE) with three branches, two of which are Transformer blocks and one is a CNN block. Connect the output of the previous layer to the input of the multi-scale feature embedding module. This module parallelizes multiple different transformations on the same map, that is, parallel convolutional blocks and Transformer blocks. Use this module to perform multi-scale learning on fixed-size images and extract local, medium, and global features.

[0026] S32: The outputs of the three branches of the multi-scale feature module in S32 are spliced ​​to obtain the features of the same image at different scales, and the features are output to the next layer of the network, the cross-scale perception optimization (CPR) module, for processing;

[0027] S32: The output features from the convolution branch of the new multi-scale feature embedding module in S32 are recorded as local features X l , the output feature of the branch after two 3×3 convolutions and the Transformer block is recorded as the medium feature X m , the output features of the branch after three 3×3 convolutions and the Transformer block are recorded as global features X g ;

[0028] S33: local feature X l Splice to medium feature X m , the medium feature X m Spliced ​​to the global feature X g middle;

[0029] S34: Connect the output ends of the three branches of S33 to the input ends of the multi-temporal self-attention (MTS) module respectively;

[0030] S35: Connect the three branches of S34 to the output features of the multi-temporal self-attention (MTS) module, respectively, and record them as X' l , X' m , X' g , the algorithm formula of the multi-temporal self-attention block is:

[0031]

[0032] in It is an adaptive masking strategy;

[0033] S36: The outputs of the three branches in S35 are concatenated and sent to the model structure of the next layer through skip connection and convolution operation.

[0034] As a further technical solution of the present invention, in step S4, a method for adding a bidirectional bridging fusion module (MSFM) to achieve bidirectional alignment and enhancement of radar and satellite features is as follows:

[0035] S41: Multi-source fusion module (MSFM) consists of multiple stacked RS blocks, each of which contains four sub-blocks: radar encoder, satellite encoder, radar-to-satellite representation fusion, and satellite-to-radar representation fusion;

[0036] S42: The radar-to-satellite representation fusion sub-block uses radar features as queries and satellite features as keys and values, and fuses radar features into satellite features through a cross-attention mechanism:

[0037]

[0038] Among them, the radar characteristic X r and satellite feature X s It is divided into h heads for multi-head attention operation, and the segmentation of the i-th head W i Q is the query projection matrix of the i-th head, and is the projection matrix of keys and values;

[0039] S43: The satellite-to-radar representation fusion sub-block is the opposite. It uses satellite features as queries and radar features as keys and values, and fuses satellite features into radar features through a cross-attention mechanism:

[0040]

[0041] S44: The fused features are combined with the original features through residual connections to enhance the learning ability of the network. This bidirectional design ensures the exchange of roles between query and information variables, making the fusion process alternately centered on radar and satellite features, effectively addressing the heterogeneity gap and enhancing the prediction ability of convective inception and the migration of precipitation systems outside the radar detection range.

[0042] As a further technical solution of the present invention, in step S5, a residual attention module (ResATT) decoder is added to train the M4Caster based on the encoder-decoder structure, and the short-term precipitation prediction method based on multi-source multi-temporal and multi-spatial feature fusion is obtained as follows:

[0043] S51: Download the pre-trained default parameters from the UNet official website and perform weight fine-tuning operations to load the obtained parameters into our improved UNet network. The optimized objective function is:

[0044] The decoder architecture integrates a residual attention module (ResATT), which includes a temporal attention module (TAM) and a spatial attention module (SAM), to achieve sequential adaptive enhancement of spatiotemporal features.

[0045] S52: The temporal attention module extracts channel features through average pooling and maximum pooling, generates attention weights on the channel dimension through a feedforward neural network and a sigmoid function, and calibrates the channel importance of the input features;

[0046] S53: The spatial attention module performs a pooling operation on features along the channel dimension and generates a spatial attention map through 7×7 convolution and sigmoid activation function to achieve adaptive enhancement of spatial position;

[0047] S54: Use bilinear upsampling in the prediction network to gradually double the size of the feature map, and combine global skip connections to fuse the upsampled features with the multi-scale features of each stage of the encoder to achieve step-by-step prediction from high to low levels;

[0048] S55: Finally, a 1×1×1 3D convolutional layer is used to generate radar echo predictions for the next six frames, covering a forecast time of 0-60 minutes. The model is trained using a weighted loss function to obtain high-precision short-term precipitation prediction results.

[0049] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, a short-term precipitation prediction method based on the fusion of multi-source and multi-temporal and multi-spatial features is implemented.

[0050] A computer-readable storage medium stores computer instructions, which, when executed by a processor, implement a short-term precipitation prediction method that integrates multiple sources and multiple spatiotemporal features.

[0051] Beneficial effect: The purpose of the present invention is to provide a short-term precipitation prediction method (M4Caster) based on the fusion of multi-source and multi-spatiotemporal features. First, in view of the shortcomings of existing single-source data radar echo extrapolation and pure convolution / recurrent networks in spatiotemporal feature modeling, a multi-source data framework of radar echo and meteorological satellite data is adopted; then, by introducing a multi-spatial multi-temporal aggregator (MMA), which includes a multi-scale patch embedding (MSPE) module, the original pure convolution module is replaced by a parallel structure of Transformer and CNN blocks to realize multi-granularity feature extraction from local texture to global context, obtain more precipitation feature information, and enhance the multi-scale prediction ability of the model; then, a multi-level fusion structure composed of cross-scale perception optimization (CPR) and multi-temporal self-attention (MTS) modules is used to realize the inter-level information transmission of local, medium and global features, fully combining CNN's attention to local information and Transformer's attention to local information. The model takes advantage of global dependency to strengthen the model's feature expression of the spatial dimension of radar echo sequences. At the same time, a cross-temporal Transformer algorithm is adopted to mine the temporal dependency of radar echo sequences through the multi-temporal self-attention (MTS) module, learn the temporal variation trend of short-term precipitation, and enhance the model's feature expression in the temporal dimension. In addition, a bidirectional bridging fusion module (MSFM) is designed to use the cross-attention mechanism to achieve bidirectional alignment and enhancement of radar and satellite features, and alleviate modal differences through residual connections. Ultimately, the improved model significantly enhances the feature extraction and fusion capabilities of the spatial and temporal dimensions of radar echo sequences and multi-source data, greatly improving the accuracy of short-term precipitation prediction, especially in the prediction of heavy precipitation (≥50dBZ). The critical success index (CSI) reaches 0.267 and the Hydeke skill score (HSS) reaches 0.372, which are significantly improved compared with existing advanced models. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 This is a flow chart of the short-term precipitation prediction method based on multi-source data fusion of the present invention.

[0053] Figure 2 This is a schematic diagram of data preprocessing in the present invention.

[0054] Figure 3 This is a schematic diagram of the network structure of the short-term precipitation prediction method based on multi-source data fusion of the present invention.

[0055] Figure 4 This is a detailed schematic diagram of the architecture of the multi-space multi-time aggregator (MMA) of the present invention.

[0056] Figure 5 This is a schematic diagram of the cross-attention mechanism of the bidirectional bridging fusion module (MSFM) of the present invention.

[0057] Figure 6This is a schematic diagram of the structure of the multi-scale feature embedding module (MSPE) of the present invention.

[0058] Figure 7 This is a schematic diagram of the information flow of the multi-level feature fusion module (CPR) of the present invention.

[0059] Figure 8 This is a schematic diagram of the multi-temporal self-attention (MTS) module of the present invention.

[0060] Figure 9 Schematic diagram of the spatiotemporal weighting mechanism of the residual attention module (ResATT) of the present invention,

[0061] Figure 10 This is a schematic diagram of the results of the present invention and the baseline model under different thresholds of 1h prediction time.

[0062] Figure 11 This is a schematic diagram of a typical ordinary precipitation event prediction scenario in the Yangtze River Delta region of the present invention.

[0063] Figure 12 This is a schematic diagram of a typical heavy rainfall event prediction scenario in the Yangtze River Delta region in the present invention.

[0064] Figure 13 Schematic diagram for comparing predictions of convection initiation (CI) events according to the present invention. DETAILED DESCRIPTION

[0065] The technical solution of the present invention is further described below in conjunction with specific implementation methods and drawings.

[0066] Example: This specific embodiment discloses a short-term precipitation prediction method based on the fusion of multi-source and multi-temporal and multi-spatial features. Figures 1 to 12 As shown, the following steps are included:

[0067] S1: Before the samples are input into the network for training, they need to go through a series of preprocessing operations. The radar and satellite precipitation datasets are preprocessed, denoising and normalization are performed, and the datasets are divided to obtain cleaned training samples (such as Figure 2 shown);

[0068] S2: Use a multi-space multi-time aggregator (MMA) (such as Figure 3 as shown) and the bidirectional bridging fusion module (MSFM) (as shown Figure 4 The radar satellite encoder sub-block shown in Figure 2 is used to obtain feature extraction networks at different stages to extract local, medium and global fusion features at different levels (such as Figure 5 shown);

[0069] S3: Adopting Multi-Scale Patch Embedding (MSPE) module (e.g. Figure 6 As shown in ), Cross-Scale Perception Optimization (CPR) module (as shown in Figure 7) and the Multi-temporal Self-Attention (MTS) module (as Figure 8 As shown in the figure, the multi-level satellite and radar features of each stage are obtained respectively, and the information transmission between different levels is realized. This captures the local-global spatial features, dynamically balances the cross-scale information, and mines the temporal dependency relationship to strengthen the model's feature expression of the spatial dimension of the radar echo sequence and the satellite sequence.

[0070] S4: Design a bidirectional bridging fusion module (MSFM) to obtain the fusion module of radar and satellite data. It uses the cross-attention mechanism to achieve bidirectional alignment and enhancement of radar and satellite features, and alleviates the differences in multi-source data through residual connections.

[0071] S5: Add residual attention module (ResATT) decoder (such as Figure 9 As shown in the figure), a short-term precipitation prediction model based on the fusion of radar echo and satellite data was obtained. The radar echo extrapolation prediction results with a 1-hour forecast time were obtained by training and testing M4Caster (as shown in the figure). Figures 10 to 13 shown).

[0072] In step S1, the data preprocessing method includes image denoising and normalization. The data preprocessing method of the present invention is:

[0073] S11: Obtain the raw data of radar and satellite precipitation maps. We divide the dataset into several precipitation events, of which 28,953 precipitation events are selected as the training set and the remaining 7,992 precipitation events are used as the validation set;

[0074] S12: De-noising the radar data: radar echo intensity values ​​below 15 dBZ and above 70 dBZ are usually noise or non-precipitation areas such as hail. The radar echo intensity of radar image points below 15 dBZ and above 70 dBZ is set to 0.

[0075] S13: Channel selection and unified spatial resolution of satellite data: First, the correlation between each satellite channel and the radar echo is evaluated using the Pearson correlation analysis method, and channels 8 (WV), 11 (MI), 13 (IR), and 15 (I2) with high correlation are selected for subsequent modeling;

[0076] S14: Normalization processing: For satellite observation data, the maximum and minimum values ​​of channels 8, 11, 13, and 15 are counted. For radar data, the range is 0-70. Then, the normalization formula is applied to map it to the interval [0,1] to obtain training samples and test samples with a resolution of 300*300. These samples are then applied to the training of the M4aster network. The 6-frame historical radar sequence and the 4*6-frame satellite sequence are used as the input of the model. In addition, the 6-frame future radar sequence is used as the ground truth (GT), R t-j+1 ,...,R t is a radar echo sequence diagram at j consecutive time stamps input into the network, S t-j+1 ,...,S t is a 4-channel satellite sequence graph at j consecutive time stamps input to the network.

[0077] In step S2, the radar satellite encoder sub-block including the multi-spatial multi-temporal aggregator (MMA) and the bidirectional bridge fusion module (MSFM) is used to extract the local, medium and global fusion features at different levels as follows:

[0078] S21: Multiple downsampling (Conv) operations and multi-spatial multi-temporal aggregator (MMA) modules are stacked in the encoder part. Each downsampling operation halves the input image size and doubles the number of feature maps. The encoding process is described as: STF i+1 = Radar satellite encoder Conv(STF i ), where STF i (i=0, 1, 2, 3) represents the spatiotemporal features at different scales, STF0 is the input radar echo or satellite sequence, and i is the downsampling stage number;

[0079] S22: Enhance spatiotemporal representation capabilities by balancing the local processing focused on by convolutional operations and the global context captured by Transformer operations through the Multi-Spatial Multi-Temporal Aggregator (MMA) module;

[0080] S23: A two-way bridge mechanism is used between radar and satellite encoders to exchange information between multi-source radar and satellite data to capture richer cloud phase information.

[0081] S24: Adopts an encoder-predictor architecture, with the encoder followed by the same number of decoders, and finally uses bilinear upsampling in the prediction network to halve the number of feature maps and double the size of the feature maps.

[0082] In step S3, the multi-scale patch embedding (MSPE), cross-scale perception optimization (CPR) and multi-temporal self-attention (MTS) modules are added to obtain the multi-level satellite and radar features at each stage as follows:

[0083] S31: Separate the first pure convolutional network structure of the encoder of the original UNet network and replace it with a multi-scale patch embedding (MSPE) with three branches, two of which are Transformer blocks and one is a CNN block. Connect the output of the previous layer to the input of the multi-scale feature embedding module. This module parallelizes multiple different transformations on the same map, that is, parallel convolutional blocks and Transformer blocks. Use this module to perform multi-scale learning on fixed-size images and extract local, medium, and global features.

[0084] S32: The outputs of the three branches of the multi-scale feature module in S32 are spliced ​​to obtain the features of the same image at different scales, and the features are output to the next layer of the network, the cross-scale perception optimization (CPR) module, for processing;

[0085] S32: The output features from the convolution branch of the new multi-scale feature embedding module in S32 are recorded as local features X l , the output feature of the branch after two 3×3 convolutions and the Transformer block is recorded as the medium feature X m , the output features of the branch after three 3×3 convolutions and the Transformer block are recorded as global features X g ;

[0086] S33: local feature X l Splice to medium feature X m , the medium feature X m Spliced ​​to the global feature X g middle;

[0087] S34: Connect the output ends of the three branches of S33 to the input ends of the multi-temporal self-attention (MTS) module respectively;

[0088] S35: Connect the three branches of S34 to the output features of the multi-temporal self-attention (MTS) module, respectively, and record them as X' l , X' m , X' g , the algorithm formula of the multi-temporal self-attention block is:

[0089]

[0090] in It is an adaptive masking strategy;

[0091] S36: The outputs of the three branches in S35 are concatenated and sent to the model structure of the next layer through skip connection and convolution operation.

[0092] In step S4, the method of adding a bidirectional bridging fusion module (MSFM) to achieve bidirectional alignment and enhancement of radar and satellite features is as follows:

[0093] S41: Multi-source fusion module (MSFM) consists of multiple stacked RS blocks, each of which contains four sub-blocks: radar encoder, satellite encoder, radar-to-satellite representation fusion, and satellite-to-radar representation fusion;

[0094] S42: The radar-to-satellite representation fusion sub-block uses radar features as queries and satellite features as keys and values, and fuses radar features into satellite features through a cross-attention mechanism:

[0095]

[0096] Among them, the radar characteristic X r and satellite feature X s It is divided into h heads for multi-head attention operation, and the segmentation of the i-th head W i Q is the query projection matrix of the i-th head, and is the projection matrix of keys and values;

[0097] S43: The satellite-to-radar representation fusion sub-block is the opposite. It uses satellite features as queries and radar features as keys and values, and fuses satellite features into radar features through a cross-attention mechanism:

[0098]

[0099] S44: The fused features are combined with the original features through residual connections to enhance the learning ability of the network. This bidirectional design ensures the exchange of roles between query and information variables, making the fusion process alternately centered on radar and satellite features, effectively addressing the heterogeneity gap and enhancing the prediction ability of convective inception and the migration of precipitation systems outside the radar detection range.

[0100] In step S5, a residual attention module (ResATT) decoder is added to train the M4Caster based on the encoder-decoder structure, and the short-term precipitation prediction method based on the fusion of multi-source and multi-temporal features is obtained as follows:

[0101] S51: Download the pre-trained default parameters from the UNet official website and perform weight fine-tuning operations to load the obtained parameters into our improved UNet network. The optimized objective function is:

[0102] The decoder architecture integrates a residual attention module (ResATT), which includes a temporal attention module (TAM) and a spatial attention module (SAM), to achieve sequential adaptive enhancement of spatiotemporal features.

[0103] S52: The temporal attention module extracts channel features through average pooling and maximum pooling, generates attention weights on the channel dimension through a feedforward neural network and a sigmoid function, and calibrates the channel importance of the input features;

[0104] S53: The spatial attention module performs a pooling operation on features along the channel dimension and generates a spatial attention map through 7×7 convolution and sigmoid activation function to achieve adaptive enhancement of spatial position;

[0105] S54: Use bilinear upsampling in the prediction network to gradually double the size of the feature map, and combine global skip connections to fuse the upsampled features with the multi-scale features of each stage of the encoder to achieve step-by-step prediction from high to low levels;

[0106] S55: Finally, a 1×1×1 3D convolutional layer is used to generate radar echo predictions for the next six frames, covering a forecast time of 0-60 minutes. The model is trained using a weighted loss function (see the weights in the following formula) to obtain a high-precision short-term precipitation forecast result.

[0107]

[0108] Our model is compared with the baseline model on the Yangtze River Delta dataset, as shown in Tables 1, 2, and 3. The best performance is highlighted in bold. Compared with other SOTA models, our model achieves the best performance level on four evaluation metrics (CSI, HSS, FAR, MSE, SSIM).

[0109] Table 1 Comparison results of CSI and HSS between the present invention (M4Caster) and other algorithms

[0110]

[0111] Table 2 Comparison results of FAR, MSE, and SSIM between the present invention (M4Caster) and other algorithms

[0112]

[0113] Table 3 Ablation results of each component of the present invention (M4Caster)

[0114]

[0115] Table 4 Comparison of parameters and inference time between the present invention (M4Caster) and the baseline model

[0116]

[0117] The model has 9.69M parameters and a single inference time of 6.179 milliseconds (predicting 6 frames, covering a 0-60 minute forecast timeframe), making it suitable for deployment in real-time business systems. Implemented using PyTorch on a workstation equipped with an NVIDIA RTX 3090 GPU, it uses a small batch training strategy of 18 sequences and selects the model with the lowest validation loss as the final prediction model.

[0118] In summary, the short-term precipitation prediction method (M4Caster) based on the fusion of multi-source and multi-spatiotemporal features proposed in this paper breaks through the limitations of traditional single-source data and single network architecture in spatiotemporal feature modeling by constructing a deep fusion framework of radar echoes and meteorological satellite data, combined with innovative structures such as the multi-space multi-time aggregator (MMA) and the bidirectional bridging fusion module (MSFM). Experiments on meteorological datasets in the Yangtze River Delta region show that compared with advanced models such as PredRNN++ and Rainformer, M4Caster achieves optimal performance in core evaluation indicators such as the critical success index (CSI), Hydeke skill score (HSS), false alarm rate (FAR), and mean square error (MSE). Especially in the heavy precipitation (≥50dBZ) prediction scenario, the CSI is improved to 0.267, an increase of 33.5% compared to PredRNN++'s 0.2, and the HSS reaches 0.372, an increase of 24% compared to PredRNN++'s 0.3; compared with Rainformer, M4Caster achieves an increase of approximately 4.3% and approximately 3.05% in CSI and HSS indicators, respectively. At the same time, ablation experiments verify the key role of components such as MMA and MSFM in improving model performance. This method not only significantly improves the accuracy of short-term precipitation prediction, but also enhances the robustness of the model under complex weather conditions through the complementarity of multi-source data and adaptive learning of dynamic features, providing a more reliable and practical solution for 0-1 hour short-term precipitation prediction.

[0119] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person familiar with the technology can understand and think of any changes or replacements within the technical scope disclosed by the present invention, which should be included in the scope of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A short-term precipitation prediction method based on the fusion of multi-source and multi-temporal and multi-spatial features, characterized by: The following steps are involved: S1: Preprocess radar and satellite precipitation datasets, perform denoising and normalization, and divide the datasets to obtain cleaned training samples; S2: Using the radar satellite encoder sub-block containing the multi-spatial multi-temporal aggregator (MMA) and the bidirectional bridge fusion module (MSFM), we obtain feature extraction networks at different stages to extract local, medium, and global fusion features at different levels; S3: Multi-scale patch embedding (MSPE), cross-scale perception optimization (CPR), and multi-temporal self-attention (MTS) modules are used to obtain multi-level satellite and radar features at each stage, realize information transfer between different levels, capture local-global spatial features, dynamically balance cross-scale information, and explore temporal dependencies, thereby strengthening the model's feature representation of the spatial dimension of radar echo sequences and satellite sequences. S4: Design a bidirectional bridging fusion module (MSFM) to obtain the fusion module of radar and satellite data. It uses the cross-attention mechanism to achieve bidirectional alignment and enhancement of radar and satellite features, and alleviates the differences in multi-source data through residual connections. S5: Add the residual attention module (ResATT) decoder to obtain a short-term precipitation prediction model based on the fusion of radar echo and satellite data. M4Caster is trained and tested to obtain radar echo extrapolation prediction results with a 1-hour forecast time.

2. The method for predicting short-term precipitation based on radar echo pattern extrapolation according to claim 1, characterized in that: In step S1, the data preprocessing method is: S11: Obtain the raw data of radar and satellite precipitation images, divide the dataset into several precipitation events, select 28,953 precipitation events as the training set, and the remaining 7,992 precipitation events as the validation set; S12: Denoising the radar data: setting the radar echo intensity of radar image points below 15 dBZ and above 70 dBZ to 0; S13: Channel selection and unified spatial resolution of satellite data: First, the correlation between each satellite channel and the radar echo is evaluated using the Pearson correlation analysis method, and channels 8 (WV), 11 (MI), 13 (IR), and 15 (I2) with high correlation are selected for subsequent modeling; S14: Normalization processing: For satellite observation data, the maximum and minimum values ​​of channels 8, 11, 13, and 15 are counted. For radar data, the range is 0-70, and then the normalization formula is applied to map it to the interval [0,1].

3. The short-term precipitation prediction method based on multi-source and multi-temporal and multi-spatial feature fusion according to claim 1 is characterized in that: In step S2, the radar satellite encoder sub-block including the multi-spatial multi-temporal aggregator (MMA) and the bidirectional bridge fusion module (MSFM) is used to extract the local, medium and global fusion features at different levels as follows: S21: Multiple downsampling (Conv) operations and multi-spatial multi-temporal aggregator (MMA) modules are stacked in the encoder part. Each downsampling operation halves the input image size and doubles the number of feature maps. The encoding process is described as: STF i+1 = Radar satellite encoder Conv(STF i ), where STF i (i=0, 1, 2, 3) represents the spatiotemporal features at different scales, STF0 is the input radar echo or satellite sequence, and i is the downsampling stage number; S22: Enhance spatiotemporal representation capabilities by balancing the local processing focused on by convolutional operations and the global context captured by Transformer operations through the Multi-Spatial Multi-Temporal Aggregator (MMA) module; S23: A two-way bridge mechanism is used between radar and satellite encoders to exchange information between multi-source radar and satellite data to capture richer cloud phase information. S24: Adopts an encoder-predictor architecture, with the encoder followed by the same number of decoders, and finally uses bilinear upsampling in the prediction network to halve the number of feature maps and double the size of the feature maps.

4. The method for short-term precipitation prediction based on multi-source and multi-temporal and multi-spatial feature fusion according to claim 1 is characterized in that: In step S3, multi-scale patch embedding (MSPE), cross-scale perception optimization (CPR) and multi-temporal self-attention (MTS) modules are added to obtain multi-level satellite and radar features at each stage: S31: Separate the first pure convolutional network structure of the encoder of the original UNet network and replace it with a multi-scale patch embedding (MSPE) with three branches, two of which are Transformer blocks and one is a CNN block. Connect the output of the previous layer to the input of the multi-scale feature embedding module. This module parallelizes multiple different transformations on the same map, that is, parallel convolutional blocks and Transformer blocks. Use this module to perform multi-scale learning on fixed-size images and extract local, medium, and global features. S32: splice the outputs of the three branches of the multi-scale feature module in S31 to obtain features of the same image at different scales, and output the features to the next layer of the network, the cross-scale perception optimization (CPR) module, for processing; S33: The output features are derived from the convolution branch of the new multi-scale feature embedding module in S32, denoted as local features X l , the output feature of the branch after two 3×3 convolutions and the Transformer block is recorded as the medium feature X m , the output features of the branch after three 3×3 convolutions and the Transformer block are recorded as global features X g ; S34: local feature X l Splice to medium feature X m , the medium feature X m Spliced ​​to the global feature X g middle; S35: Connect the three branches of S34 to the output features of the multi-temporal self-attention (MTS) module, respectively, and record them as X' l , X' m , X' g , the algorithm formula of the multi-temporal self-attention block is: in It is an adaptive masking strategy; S36: The outputs of the three branches in S35 are concatenated and sent to the model structure of the next layer through skip connection and convolution operation.

5. The method for short-term precipitation prediction based on multi-source and multi-temporal and multi-spatial feature fusion according to claim 1 is characterized in that: In step S4, the method of adding a bidirectional bridging fusion module (MSFM) to achieve bidirectional alignment and enhancement of radar and satellite features is as follows: S41: Multi-source fusion module (MSFM) consists of multiple stacked RS blocks, each of which contains four sub-blocks: radar encoder, satellite encoder, radar-to-satellite representation fusion, and satellite-to-radar representation fusion; S42: The radar-to-satellite representation fusion sub-block uses radar features as queries and satellite features as keys and values, and fuses radar features into satellite features through a cross-attention mechanism: Among them, the radar characteristic X r and satellite feature X s It is divided into h heads for multi-head attention operation, and the segmentation of the i-th head W i Q is the query projection matrix of the i-th head, and is the projection matrix of keys and values; S43: The satellite-to-radar representation fusion sub-block is the opposite. It uses satellite features as queries and radar features as keys and values, and fuses satellite features into radar features through a cross-attention mechanism: S44: The fused features are combined with the original features through residual connections to enhance the learning ability of the network. This bidirectional design ensures the exchange of roles between query and information variables, making the fusion process alternately centered on radar and satellite features, effectively addressing the heterogeneity gap and enhancing the prediction ability of convective inception and the migration of precipitation systems outside the radar detection range.

6. The method for short-term precipitation prediction based on multi-source and multi-temporal and multi-spatial feature fusion according to claim 1 is characterized in that: In step S5, a residual attention module (ResATT) decoder is added to train the M4Caster based on the encoder-decoder structure, and the short-term precipitation prediction method based on the fusion of multi-source and multi-temporal features is obtained as follows: S51: Download the pre-trained default parameters from the UNet official website and perform weight fine-tuning operations to load the obtained parameters into the improved UNet network. The optimized objective function is: The decoder architecture integrates a residual attention module (ResATT), which includes a temporal attention module (TAM) and a spatial attention module (SAM), to achieve sequential adaptive enhancement of spatiotemporal features. S52: The temporal attention module extracts channel features through average pooling and maximum pooling, generates attention weights on the channel dimension through a feedforward neural network and a sigmoid function, and calibrates the channel importance of the input features; S53: The spatial attention module performs a pooling operation on features along the channel dimension and generates a spatial attention map through 7×7 convolution and sigmoid activation function to achieve adaptive enhancement of spatial position; S54: Use bilinear upsampling in the prediction network to gradually double the size of the feature map, and combine global skip connections to fuse the upsampled features with the multi-scale features of each stage of the encoder to achieve step-by-step prediction from high to low levels; S55: Finally, a 1×1×1 3D convolutional layer is used to generate radar echo predictions for the next six frames, covering a forecast time of 0-60 minutes. The model is trained using a weighted loss function to obtain high-precision short-term precipitation prediction results.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method for predicting short-term precipitation by fusing multi-source and multi-temporal and multi-spatial features as described in any one of claims 1 to 6 above is implemented.

8. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed by a processor, a short-term precipitation prediction method based on multi-source and multi-temporal and multi-spatial feature fusion is implemented as described in any one of claims 1-6.

Citation Information

Cited By

  • Short temporary rainfall forecasting method, device and equipment based on multi-source data and memory

    CN120821005A

  • Multi-time scale fusion wind speed prediction method based on dual-encoder UNet model

    CN120822434A

  • A multi-time scale fusion wind speed prediction method based on a double-encoder UNet model

    CN120822434B

  • Short temporary rainfall intensity prediction and identification method based on space-time composite attention

    CN121033697A

  • Low-vision dangerous road section intelligent early warning method and device based on multi-radar perception

    CN121438605A