Precipitation prediction method based on multi-scale physical Transform

By combining the multi-scale physical Transformer network and the spatiotemporal physical consistency loss function, the problems of error accumulation and low computational efficiency of traditional models in extreme precipitation forecasting are solved, and high-precision and real-time extreme precipitation forecasting is achieved.

CN120653919APending Publication Date: 2025-09-16HANGZHOU DIANZI UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510715914.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Traditional numerical weather forecast models have problems of error accumulation, physical inconsistency and low computational efficiency in the short-term and nowcasting of extreme precipitation events. It is difficult to take into account both the large-scale background and the details of local heavy precipitation, and high-resolution forecasts are difficult to achieve real-time performance.

Method used

A precipitation prediction method based on multi-scale physical Transformer is adopted. By constructing a multi-scale physical information Transformer network, combining a lightweight residual adapter, STRMT encoder and generative decoder, and using a multi-scale spatiotemporal physical consistency loss function for end-to-end training, the global and local spatiotemporal features can be captured and the forecast accuracy can be improved.

Benefits of technology

It improves the forecast accuracy and real-time performance of extreme precipitation events, ensures the physical consistency and spatiotemporal continuity of forecast results, and reduces computing resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653919A_ABST
    Figure CN120653919A_ABST
Patent Text Reader

Abstract

The invention discloses a rainfall prediction method based on multi-scale physical Transform, and the method comprises the steps: firstly collecting continuous historical radar observation data from a meteorological radar system, and carrying out the preprocessing of the data; secondly, according to the processed observation data, on the basis of a Transform network, capturing global semantics to obtain aggregation features; and then motion field and intensity residual field prediction is carried out based on the aggregation features to obtain a deterministic prediction sequence. And finally, constructing a generative network based on a multi-scale adapter, obtaining a rainfall probabilistic prediction map according to a deterministic prediction sequence, constructing a multi-scale space-time physical consistency loss function, and carrying out end-to-end training. According to the method, the problems of error accumulation, physical inconsistency, low calculation efficiency and the like in a traditional forecasting method are effectively solved, so that a brand new and accurate technical path is provided for real-time forecasting of the extreme rainfall event.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of meteorological forecasting, data mining and deep learning technology, and specifically relates to an extreme precipitation nowcasting method based on multi-scale physical information Transformer and spatiotemporal feature adaptation. Background Art

[0002] With the accelerating pace of global climate change and urbanization, the frequency and intensity of extreme precipitation events are increasing, posing severe challenges to socioeconomic security and disaster prevention and mitigation. Traditional numerical weather prediction (NWP) models, relying on physical formulas and numerical calculations, offer high accuracy for large-scale weather forecasts. However, in short-term nowcasting, particularly for capturing nonlinear processes such as localized convection and heavy precipitation, they often suffer from high rates of false alarms and omissions. Furthermore, they are limited by their inability to simulate small and medium-scale systems. Furthermore, the high computational resources required for high-resolution forecasts by traditional models make them difficult to meet the requirements of real-time forecasting.

[0003] In recent years, deep learning-based weather forecasting methods have gradually become a research hotspot. Although early convolutional recurrent networks (such as ConvLSTM) have made some progress in capturing spatiotemporal dynamics, they often suffer from problems such as ambiguous forecasts and missing local details when dealing with extreme precipitation events. Some methods have begun to incorporate physical prior knowledge into deep learning models in an attempt to improve physical consistency and predictive stability, but significant challenges remain in multi-scale feature extraction, fusion of global and local information, and computational efficiency. Existing technologies face the following major problems in accurately simulating the evolution of precipitation systems: First, they fail to capture precipitation characteristics at different spatial scales, making it difficult to simultaneously account for both large-scale background and localized details of heavy rainfall; second, the integration of physical constraints and data-driven methods is not close enough, resulting in deviations in forecast results in terms of physical plausibility and spatiotemporal continuity; and third, high-resolution forecasts require high computing resources, making real-time forecasts difficult to achieve.

[0004] Based on the above background, the current extreme precipitation nowcasting technology urgently needs to improve the modeling capabilities of multi-scale spatiotemporal characteristics while maintaining physical interpretability, and take into account computational efficiency while ensuring forecast accuracy. Summary of the Invention

[0005] To address the aforementioned issues in existing technologies, the present invention provides a multi-scale physics-based Transformer precipitation prediction method. This method effectively captures the global and local spatiotemporal characteristics of precipitation while ensuring physical consistency, thereby improving forecast accuracy and real-time performance. The specific technical solution includes the following steps:

[0006] (1) Collection and preprocessing of historical radar data. Collect continuous historical radar observation data X from the meteorological radar system. in , serving as the basis for modeling the evolution of precipitation fields. Data preprocessing includes normalization and noise filtering of raw radar data, providing a stable data foundation for subsequent model input.

[0007] (2) Construction of an evolutionary network based on multi-scale physical information transformer. Based on the processed observation data, the global semantics are captured based on the transformer network to obtain the aggregated feature G. Before entering the feature extraction process, the historical radar reflectivity sequence X obtained in (1) is first in The input is fine-tuned into a lightweight residual adapter to ensure initial consistency of global features. The adapter first compresses the number of channels through a 1×1 convolutional layer to produce an intermediate feature A1. It then performs block-wise compression, SE channel attention, on A1: It first performs global average pooling to obtain a description vector, then sequentially passes through two layers of full connections, plus ReLU and Softmax to generate channel weights, and uses these weights to recalibrate A1 to obtain A2. Finally, A2 is element-wise added to the original input to output the fine-tuned feature F0:

[0008] F0=A2+X in

[0009] After obtaining F0, it is sent to the downsampling part of the multi-scale spatiotemporal feature STRMT encoder, which is divided into three stages. In the first stage, F0 is obtained by passing it through two layers of 3×3 convolution + BN + ReLU. The outputs of the DWConv and SATRM modules are then fed in parallel. SATRM consists of two submodules: the Multi-Head Hybrid Convolution (MHMC) and the Scale-Aware Aggregation (SAA). MHMC extracts multi-scale features using multiple kernels, while SAA assigns fusion weights based on global context. The outputs of the DWConv and SATRM modules are concatenated channel-wise and then subjected to 1×1 convolution for dimensionality reduction, yielding the first-stage output F1. The second and third stages are similar to the first, except that the features are first divided into blocks using PatchEmbedding and linearly mapped to a higher dimension. These blocks are then passed through DWConv and SATRM in parallel, resulting in concatenation and dimensionality reduction to yield F2 and F3, respectively.

[0010] After downsampling, F3 is input into the upsampling UPerHead module, which uses multi-scale pyramid pooling (PPM) to capture global semantics and fuse feature maps of different resolutions. Finally, it upsamples layer by layer to restore the original resolution and outputs the aggregated feature G.

[0011] (3) Prediction of motion field and intensity residual field. The aggregated feature G is processed in two parallel paths: one is sent to the motion decoder and the other is sent to the intensity decoder. The motion decoder consists of three layers of 3×3 convolution + BN + ReLU, and the last layer outputs a 2-channel motion field. Channels 0 and 1 correspond to the horizontal and vertical displacement components respectively; the intensity decoder has the same structure, but the last layer outputs a 1-channel intensity residual field. Get v t+1 and s t+1 Afterwards, based on the two-dimensional physical continuity equation;

[0012] Generate the reflectivity at the next moment by iterative advection and residual superposition:

[0013]

[0014] Among them, x t is the input reflectivity sequence at time t, p is the coordinate, and Δt is the time step;

[0015] Repeat this process to generate a length of T ′ The deterministic prediction sequence of:

[0016] X pred ={x t+1 (p),x t+2 (p),…,x t+T′ (p)}.

[0017] (4) Construction of a generative network based on a multi-scale adapter. in With X pred Concatenate in the channel dimension to get the joint input tensor X joint , first through double 3×3 convolution to fuse history and prediction information, output intermediate feature F0. Subsequently, F0 is downsampled three times, and after each downsampling, a multi-scale adapter module is embedded - including 1×1Conv compression, DWConv extraction of spatial features, SE channel attention recalibration and residual addition to generate F1, F2, F3. The parallel noise projector first convolves the random Gaussian noise Z~N(0,1) and then passes it through four linear projection modules to obtain the noise feature, which is then spliced ​​with F3 to form The decoder is fed into a generative decoder, which consists of an upsampling and a generation block connected in series. During the upsampling process, the generation block is executed twice in the first and third stages, and once in the second stage. Each generation block incorporates motion field and intensity residual field condition information for feature control. The decoder finally outputs a 1-channel probabilistic precipitation forecast map through LeakyReLU and convolution.

[0018] The generation module contains three parallel branches, namely, a convolutional network, a linear mapping of the deterministic prediction sequence at position p, and a local attention mechanism. Finally, the outputs of the three are added and fused. (5) Construction and optimization of multi-scale spatiotemporal physical consistency loss function. In order to ensure the rationality and spatiotemporal continuity of the forecast results in a physical sense, the present invention designs a multi-scale spatiotemporal physical consistency loss function. The loss function includes:

[0019] Balanced pixel-level loss: Balanced mean square error (BMSE) and balanced mean absolute error (BMAE) are used to assign different weights to areas with different precipitation intensities.

[0020] Multi-scale spatial loss: The forecast image is decomposed into low-frequency and high-frequency sub-bands based on discrete wavelet transform, and the structural similarity SSIM indicator is used to calculate the spatial error of each sub-band.

[0021] Multi-scale temporal loss: Constrains the temporal continuity of the forecast sequence by performing a pooling operation on the differences between consecutive time frames.

[0022] The total loss function is defined as a dynamically weighted sum of balanced pixel-level loss, multi-scale spatial loss, and multi-scale spatial loss.

[0023] (6) Parameter optimization and forecast result generation. By maximizing the above loss function and performing end-to-end training on the network parameters, we achieve the optimal matching of the parameters of the evolution network and the generation network while taking into account the fusion of physical constraints and multi-scale features. Ultimately, we obtain extreme precipitation field prediction results with high resolution (e.g., 2 km grid) and a 3-hour forecast timeliness.

[0024] This invention organically combines multi-scale physical priors with Transformer structure, lightweight spatiotemporal adapter and multi-scale physical consistency loss for the first time. By jointly optimizing global dynamics and local details, it effectively solves the problems of error accumulation, physical inconsistency and low computational efficiency existing in traditional forecasting methods, thus providing a new and accurate technical path for real-time forecasting of extreme precipitation events. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 This is a schematic diagram of the overall framework of the extreme precipitation nowcasting method of the present invention;

[0026] Figure 2 This is a diagram of the evolutionary network architecture based on the scale-aware Transformer and a diagram of the generative network architecture with a multi-scale adapter in the present invention;

[0027] Figure 3 This is a diagram of the architecture of the STRMT encoder of the present invention. DETAILED DESCRIPTION

[0028] In order to describe the present invention more specifically, the technical solution of the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0029] The multi-scale physical Transformer precipitation prediction method (hereinafter referred to as "this method") of this embodiment mainly includes the following steps, and is combined with Figure 1 (Schematic diagram of MPFormer overall architecture), Figure 2 (evolutionary network architecture based on scale-aware Transformer and generative network architecture with multi-scale adapter) and Figure 3 (Architecture diagram of STRMT encoder) for explanation.

[0030] Step 1: Collection and preprocessing of historical radar data

[0031] When implementing this method, firstly, historical observation data is continuously collected from the weather radar system. The data sequence is recorded as:

[0032]

[0033] Among them, X in Represents the precipitation field data from time -T0 to time 0. The collected data is preprocessed by normalization, noise filtering and other operations to ensure data stability and validity, providing a high-quality data foundation for subsequent model input.

[0034] Step 2: Evolutionary network construction based on scale-aware Transformer

[0035] Before entering the encoding process, the preprocessed historical radar data sequence X in Perform an overall residual adaptation process to fine-tune the global feature distribution in the initial stage of the network, so as to better adapt to the subsequent multi-scale extraction process. Specifically, the input sequence X in First, the number of channels is compressed through a 1×1 convolution layer. This step can reduce the feature dimension and aggregate preliminary channel information. Then, the compressed feature tensor A1 is sent to the channel attention (Squeeze-and-Excitation) module. The module first performs global average pooling on A1 to obtain the channel description vector, and then generates channel recalibration weights through two layers of full connection, ReLU and Softmax. Finally, these weights are applied to A1 to obtain the recalibrated feature A2. Finally, A2 is compared with the original input Add element by element to form the residual connection output F0:

[0036] F0=A2+X in

[0037] In this way, the residual adapter not only preserves the original data information but also injects channel-level adaptive feature enhancement.

[0038] After fine-tuning the features of the residual adapter, feature F0 is used as the input to the STRMT encoder. The STRMT encoder's downsampling process is divided into three stages, each consisting of preprocessing, parallel convolution and modulation, and feature fusion. The stages are sequentially connected to ensure that multi-scale information is gradually refined.

[0039] In the first stage, F0 is sent to the "pre-convolution group". The purpose of this step is to do preliminary local smoothing and channel mixing. Then, the output The information is fed into two branches simultaneously: one is the depthwise separable convolution (DWConv) branch, which is used to efficiently extract spatial details; the other is the sequential temporal residual modulation module (SATRM) branch, which is used for cross-scale context fusion. The DWConv branch performs spatial convolution with small kernels on each channel, while the SATRM branch contains two submodules, MHMC and SAA. MHMC is responsible for extracting multi-scale features in parallel using multiple convolution kernels, and SAA is responsible for dynamically assigning weights to each scale based on global significance. The outputs of the two branches are then concatenated in the channel dimension and then subjected to dimensionality reduction using a 1×1 convolution to obtain the final output F1 of the first stage, which is then passed to the second stage.

[0040] The second and third stages receive F1 as input and first send it to the block coding module. Block coding divides the feature map into several fixed-size image blocks and performs linear mapping on each block, thereby converting the spatial representation into a sequence of block-level feature vectors and outputting high-dimensional features. Then, The outputs are then fed into the DWConv and SATRM branches in parallel: the DWConv branch continues to extract local spatial features, while the SATRM branch continues to fuse global-local information across scales. The outputs of the two branches are concatenated again in the channel direction and subjected to a 1×1 convolution to reduce the number of channels back to the original level, resulting in the second-stage output F2. This stage introduces block-level structural information through block encoding, and combined with parallel DWConv and SATRM, further enriches the multi-scale representation of features. The continuous concatenation of these three stages ensures the gradual refinement of local fine-grained information to global large-scale semantics, resulting in features that are both rich in detail and broad in context.

[0041] The specific implementation steps of the MHMC module and SAA module are as follows:

[0042] (2.1) Use the multi-head hybrid convolution (MHMC) module to perform multi-scale processing on the input features. The calculation formula of the MHMC module is:

[0043]

[0044] in, Indicates that the input x is processed according to different convolution kernel sizes k i Perform depth-wise separable convolution operation, Concat means concatenating the features of each scale along the channel direction. M , proceed to Section 2.2 for scale-aware aggregation.

[0045] (2.2) Inherited from F generated in 2.1 M , the scale-aware aggregation (SAA) module is used to convolute the features F M The fusion is performed and the expression ability of local features is further improved by introducing the residual enhancement path. The enhanced features can be expressed as follows:

[0046] Z enhanced =M⊙V+Residual-EnhancedPath(F M )

[0047] Among them, M is the scale-aware weight matrix generated by the SAA module, V is the feature matrix representing global information obtained through linear transformation, the symbol ⊙ represents element-level multiplication, and Residual-EnhancedPath(X) is the residual path that enhances local information.

[0048] After downsampling, the feature F2 is fed into the UPerHead module for upsampling and cross-scale feature fusion. UPerHead first captures global contextual information through multi-scale pyramid pooling (PPM) and fuses it with feature maps of different resolutions. It then sequentially upsamples to the original resolution and outputs the final aggregated feature G. This step restores spatial scale and integrates multi-scale context, making it ideal as decoder input.

[0049] After the above processing is completed, step 2 outputs the aggregated feature G and enters the "Step 3: Prediction of Motion Field and Intensity Residual" link.

[0050] Figure 2 The middle left part shows the evolutionary network architecture based on scale-aware Transformer. Its core is to use the above modules to capture the cross-scale evolution characteristics of the precipitation system.

[0051] Step 3: Prediction of motion field and intensity residuals

[0052] After completing the feature aggregation in step 2 and obtaining the final context feature tensor G, this step immediately feeds this tensor into the "Joint Motion-Intensity Decoder," which predicts the velocity field (motion field) and reflectivity intensity residual field for the future moment, and iteratively generates the reflectivity prediction for the next moment based on the physical continuity equation. The entire process can be divided into the following steps: parallel branch preparation, motion decoder details, intensity decoder details, physical synthesis and iteration, and result splicing and consolidation.

[0053] First, two copies of the same feature tensor G are made: one is fed into the "motion decoder" and the other into the "intensity decoder." This parallel branching setup ensures that motion and intensity information are learned independently in the same context, avoiding error accumulation due to sequence dependence and enabling optimal coupling between the two in subsequent physical synthesis.

[0054] The motion decoder branch consists of three convolutional encoding layers and one output layer. All layers use a 3×3 convolution kernel with BatchNorm and ReLU activation. Then, a 3×3 convolution (number of channels = 2) is input, followed by BatchNorm, and finally the predicted velocity field is directly output without activation:

[0055]

[0056] Here, channel 0 and channel 1 correspond to the horizontal and vertical velocity components, respectively, and H and W represent height and width, respectively. BatchNorm helps stabilize the value, and the linear output ensures that the velocity field can be positive or negative.

[0057] In the intensity decoder branch, the convolutional structure is also completely symmetrical with the motion decoder, except that the last layer is changed to a linear output to allow the positive and negative residuals:

[0058]

[0059] At this time, the residual at each position can be positive to indicate enhancement or negative to indicate attenuation, which perfectly corresponds to the design of NowcastNet.

[0060] Get sports field v t+1 With the residual field s t+1 Finally, the “iterative advection and residual superposition” module proposed by the meteorological forecast network NowcastNet is used to implement the physical continuity equation numerically. First, the reflectivity field x at the current moment is calculated. t Perform the advection operation:

[0061]

[0062] Where p = (i, j) is the pixel coordinate; Δt is the time step; v t+1(p) is the velocity vector at position p; the advection operation is accurately performed through bilinear interpolation, thus avoiding discretization bias. Then, the advection result is added point by point to the intensity residual field:

[0063]

[0064] In this way, the reflectivity update at each step takes into account both physical transport (advection) and local increases and decreases (residuals).

[0065] In practical applications, the above advection and residual superposition operations will be iteratively performed multiple times as needed to generate a continuous prediction sequence of length T′:

[0066] X pred ={x t+1 (p),x t+2 (p),…,x t+T′ (p)}.

[0067] Each iteration starts with the latest x t+k As the input of the next moment, it ensures temporal consistency and physical coherence.

[0068] After all iterations are completed, the complete deterministic prediction sequence X is obtained. pred At the end of step 3, change X pred With the original input sequence X in The features are spliced ​​in the channel dimension to form joint features for use in "Step 4: Generate Network and Uncertainty Compensation", which starts the next stage of supplementary adjustment of small-scale details and prediction uncertainty.

[0069] Step 4: Generative network construction with multi-scale adapter

[0070] After completing the motion-intensity joint decoding of the previous stage, and obtaining the deterministic prediction sequence X at the future moment pred After this, we enter this step, the "Generative Network" phase. This network further refines the local structure of the predicted image and introduces reasonable uncertainty to compensate for microscale variations in precipitation that cannot be modeled deterministically. The entire generative network consists of three parts: a generative encoder, a noise projector, and a generative decoder. Architecturally, it follows a typical U-Net encoder-decoder framework, with a multi-scale adapter module embedded in the encoder and multiple generative modules connected in series in the decoder to reconstruct the output image layer by layer.

[0071] First, the original historical radar data sequence X in and the deterministic prediction sequence X generated in the previous step pred Perform channel-level splicing and synthesize it into a joint tensor X jointThis tensor retains the spatiotemporal information of past real observations and future predictions, providing input conditions for the generation network in this stage. The specific splicing method is:

[0072]

[0073] Where T in and T out are the input and output sequence lengths, respectively, and H and W are the spatial dimensions. This tensor is then fed into the generative encoder.

[0074] In the encoder generation part, we first perform a set of double convolution operations on X joint Perform two consecutive 3×3 convolutions and output the intermediate feature tensor F0. Double convolution not only extracts primary spatial structural features but also fuses historical and predicted information into a shared representation space.

[0075] Next, F0 undergoes three layer-by-layer downsampling steps. Each downsampling step consists of a spatial downsampling step followed by a multi-scale adapter module. The adapter module is responsible for enhancing local detail and adjusting regional sensitivity in feature maps at different scales. The output of each downsampling + adapter step is denoted as F1, F2, and F3, representing the hierarchical features from shallow to deep layers.

[0076] While the encoder is executing, the other input is a random Gaussian noise tensor The shape of this tensor matches the encoder's deep layer output. The noise is first transformed through a 3×3 convolution, then passed through four "projection modules" in sequence. Each projection module consists of a linear transformation layer followed by an activation function (ReLU or GELU), which maps the random noise into a representation aligned with the deep semantic space. The four layers stacked together gradually increase the noise dimension and enhance its expressive power.

[0077] The projection result Proj(Z) is concatenated with the generated encoder output F3 in the channel dimension to synthesize a joint representation:

[0078]

[0079] This representation combines deterministic representation with random perturbations, injecting necessary uncertainty and diversity into the subsequent generation process.

[0080] Next, It is fed into the Generator Decoder, the core of which consists of three upsampling steps connected in series with the Generation Block.

[0081] Each generation module contains the following key components:

[0082] The backbone is a standard 3×3 convolution + BN + ReLU;

[0083] The external condition control signal is output by the motion-intensity decoder module and affects the generation result through the conditional fusion mechanism;

[0084] Local attention mechanisms (such as group normalization + spatial attention) are used to emphasize precipitation boundaries or strong convection areas.

[0085] In specific deployment:

[0086] The first upsampling stage: perform upsampling, then enter the generation module, generate the module again (twice in total), and output feature G1;

[0087] The second upsampling stage: performs upsampling and then enters the generation module (once), outputting G2;

[0088] The third upsampling stage: performs upsampling and then enters the generation module twice, outputting G3.

[0089] During the entire generation process, each generation module will introduce the motion-intensity information obtained from the evolutionary network as an external condition. Specifically, the condition is embedded into the module through gating to regulate the spatial propagation direction and detail generation intensity, thereby achieving high-fidelity synthesis consistent with the conditionality of the previous step.

[0090] The final output G3 of the decoder is generated and passed through a leaky correction activation unit (LeakyReLU) and a 3×3 convolution layer to compress the channel into 1 channel to obtain the final probabilistic reflectivity prediction map:

[0091]

[0092] This output, as the final result of the generative network, provides a detailed, uncertainty-aware distributional forecast map for precipitation forecasting. The final result is fed into the evaluation module and combined with the loss function design for training optimization.

[0093] Step 5: Design and optimization of multi-scale spatiotemporal physical consistency loss function

[0094] To ensure that the forecast results have good physical significance and spatiotemporal continuity, this method designs a multi-scale spatiotemporal physical consistency loss function. The loss function mainly consists of the following parts:

[0095] (5.1) Balancing pixel-level losses

[0096] Different weights are assigned to different precipitation intensity areas and described by the balanced mean square error (BMSE) and balanced mean absolute error (BMAE). The formulas are:

[0097]

[0098] Among them, w i is the weight corresponding to pixel i, N is the total number of pixels, y i represents the true value, y i ′ represents the predicted value.

[0099] (5.2) Multi-scale spatial loss

[0100] Storm images have obvious multi-scale characteristics, so we introduce a multi-scale spatial loss function based on discrete wavelet transform (DWT) to better capture the multi-scale features in the image.

[0101] First, we use DWT to decompose the input image into subbands of different scales, including one low-frequency subband and three high-frequency subbands. Then, we calculate the structural similarity index (SSIM) between the predicted image and the target image on each subband and use it as part of the spatial loss.

[0102] Specifically, we define the spatial loss as:

[0103]

[0104] in, and are the low-frequency and high-frequency subbands of the predicted image and target image after the k-th layer wavelet decomposition, respectively, k is the weight coefficient of the kth layer, L ssim is the SSIM loss, defined as:

[0105] L ssim =1-SSIM(x,y)

[0106] The calculation formula of SSIM is as follows:

[0107]

[0108] Among them, μ x ,μ y are the means of x and y, σ x ,σ y are the standard deviations of x and y, σ xy is the covariance of x and y, and C1 and C2 are constants.

[0109] (5.3) Multi-scale time loss

[0110] Temporal consistency is crucial for short-term precipitation forecasting. Inspired by recent work on video denoising, we design a loss function based on temporal frame variation constraints to directly constrain the precipitation intensity variation between adjacent frames.

[0111] Specifically, we calculate the difference between the predicted image and the target image at adjacent moments and use it as part of the temporal loss. To avoid the "double penalty" problem [5], we use pooling operations at different scales to extract multi-scale features of the image and calculate the temporal loss at each scale.

[0112] Specifically, we define the time loss as:

[0113]

[0114] Among them, P k is the k-th layer pooling operation, λ k is the weight coefficient of the kth layer, x t and y t are the predicted image and target image at time t respectively.

[0115] (5.4) Total loss function

[0116] The weighted combination of the above losses gives the total loss function:

[0117] L total =λ1L bmse +λ2L bmae +λ3L spatial +λ4L temporal

[0118] Step 6: Parameter optimization and forecast result generation

[0119] By maximizing the aforementioned objective function, all network parameters are jointly optimized end-to-end, achieving effective synergy between the evolving and generating networks. Once trained, the system is able to generate high-resolution (e.g., 2km grid) extreme precipitation forecasts with a three-hour forecast timeframe based on given historical radar data.

[0120] Figure 1 The overall architecture of this implementation is shown in the figure. The left side of the figure shows the data acquisition and preprocessing module, which then enters the scale-aware Transformer-based evolutionary network, and then generates preliminary predictions through motion field and intensity decoders. The right side shows the generation network module, which refines the preliminary predictions through a multi-scale adapter and ultimately outputs the forecast results. Figure 2 The internal structure of the evolutionary and generative networks is detailed. The left half shows the evolutionary network of the scale-aware Transformer, whose core consists of the MHMC and SAA modules and the residual enhancement path. The right half shows the generative network with a multi-scale adapter, clearly showing how the adapter module adjusts and fuses features after each downsampling.

[0121] Experimental verification

[0122] Experimental Setup: This experiment uses 2021 radar reflectivity data from the US Multi-Radar / Multi-Sensor (MRMS) system, covering North America from 20°–55°N and 60°–130°W, with a spatial resolution of 1 km × 1 km and a temporal resolution of 2 minutes. The data is divided chronologically into a training set (January 1–20), a validation set (January 21–23), and a test set (January 24, all day). The input sequence length is 9 frames, and the prediction length is 20 frames.

[0123] Comparison method: Four representative baselines are selected for comparison: U-Net: a classic multi-scale convolutional network, improved with the weighted loss of Ravuri et al.; pySTEPS: a radar echo extrapolation method based on optical flow, using the default configuration; NowcastNet: a conditional generation framework integrating physical conservation laws.

[0124] Evaluation indicators: Five indicators are used, including the multi-threshold domain expansion critical success index (CSIN), Heidke skill score (HSS), power spectral density error (PSD), mean absolute error (MAE) and root mean square error (RMSE), with thresholds of 0.5 mm / h, 7 mm / h and 20 mm / h respectively.

[0125] Experimental results: Table 1 lists the performance comparison of each method on the test set.

[0126] Table 1 Comparison of key indicators of various methods

[0127] method CSIN HSS PSD MAE RMSE U-Net 0.2322 0.0530 2.0703 3.2626 11.0816 pySTEPS 0.0495 0.0030 4.3572 0.5112 3.1621 NowcastNet 0.2583 0.0671 32.3486 0.5886 1.5955 The present invention 0.3316 0.0741 49.9206 0.5808 1.6347

[0128] Result analysis: As shown in Table 1, the proposed method reaches 0.3316 and 0.0741 on CSIN and HSS, respectively, which are 7.33% and 10.5% higher than NowcastNet respectively; it is significantly higher than other methods on PSD, indicating better restoration of multi-scale structures; MAE and RMSE are further reduced, verifying the advantage of the model in numerical prediction accuracy.

[0129] The detailed description of the above embodiment is only one specific implementation of the present invention. Those skilled in the art may make various modifications and substitutions to the embodiment without departing from the spirit and scope of the present invention. All such modifications and substitutions shall fall within the scope of protection of the present invention.

Claims

1. The precipitation prediction method based on multi-scale physical Transformer is characterized by: The steps include: Step 1: Collect continuous historical radar observation data from the weather radar system and preprocess it; Step 2: Based on the processed observation data, the global semantics are captured and aggregated features are obtained based on the Transformer network; Step 3: Predict the motion field and intensity residual field based on the aggregated features to obtain a deterministic prediction sequence; Step 4: Construct a generative network based on a multi-scale adapter to obtain a probabilistic precipitation forecast map from the deterministic forecast sequence; Step 5: Construct and optimize the multi-scale spatiotemporal physical consistency loss function and perform end-to-end training by maximizing the loss function.

2. The precipitation prediction method based on multi-scale physical Transformer according to claim 1 is characterized in that: The specific implementation process of step 2 is as follows: First, the reflectivity sequence of historical radar observation data is input into the lightweight residual adapter for fine-tuning. The lightweight residual adapter first compresses the number of channels through a convolutional layer to generate intermediate features. Then, block compression SE channel attention is performed on the intermediate features. Finally, the features obtained by SE channel attention are element-wise added to the original input reflectivity sequence to output the fine-tuned features. After obtaining the fine-tuned features, they are fed into the downsampling part of the multi-scale spatiotemporal feature STRMT encoder. This part is divided into three stages. In the first stage, the fine-tuned features are convolved and then input into the depth-wise separable convolution DWConv and temporal residual modulation SATRM modules in parallel. SATRM contains two sub-modules: multi-head hybrid convolution MHMC and scale-aware aggregation SAA. MHMC extracts multi-scale features based on multiple cores, and SAA assigns fusion weights based on the global context. The outputs of the depthwise separable convolution (DWConv) and temporal residual modulation (SATRM) modules are concatenated in the channel direction and then subjected to convolutional dimensionality reduction to obtain the output of the first stage. The second and third stages are the same as the first stage, except that the input features are first divided into blocks through block coding and linearly mapped, and then passed through DWConv and SATRM in parallel. After concatenation and dimensionality reduction, the output features of the second and third stages are obtained respectively. The output features of the third stage are input into the upsampling UPerHead module, which uses multi-scale pyramid pooling (PPM) to capture global semantics and fuse feature maps of different resolutions. Finally, it upsamples layer by layer to restore the original resolution and outputs the aggregated features G.

3. The precipitation prediction method based on multi-scale physical Transformer according to claim 2 is characterized in that: The multi-head hybrid convolution MHMC includes several depth-separable convolutions of different scales, and splices the output features of all depth-separable convolutions to obtain spliced ​​features; The scale-aware aggregation SAA is implemented as follows: a scale-aware weight matrix is ​​generated, the splicing features are linearly transformed to obtain a feature matrix representing global information, the scale-aware weight matrix is ​​element-wise multiplied with the feature matrix representing global information, and the matrix is ​​added to the splicing features through residual connection to obtain enhanced features.

4. The precipitation prediction method based on multi-scale physical Transformer according to claim 3 is characterized in that: The prediction of the motion field and the intensity residual field is specifically as follows: The aggregated features G are processed in two parallel paths: one path is fed into the motion decoder, and the other path is fed into the intensity decoder. The motion decoder consists of a convolutional network, and the last layer outputs a 2-channel motion field, where channels 0 and 1 correspond to the horizontal and vertical displacement components respectively. The intensity decoder has the same structure as the motion decoder, but the last layer outputs a 1-channel intensity residual field: After obtaining the 2-channel motion field and the 1-channel intensity residual field, based on the two-dimensional physical continuity equation, the reflectivity of the next moment is generated by iterative advection and residual superposition. This process is repeated to form a system with a length of T from all the next moment reflectivities. ′ Deterministic prediction sequence.

5. The precipitation prediction method based on multi-scale physical Transformer according to claim 4 is characterized in that: The motion decoder branch includes a convolutional coding layer and an output layer, and performs motion field prediction based on the aggregated feature G; The intensity residual branch uses the same convolutional structure as the motion decoder, and changes the output to a linear one in the last layer to allow the residual to be positive or negative.

6. The precipitation prediction method based on multi-scale physical Transformer according to claim 5, characterized in that: The specific implementation process of step 4 is as follows: The reflectivity sequence of the input historical radar observation data and the deterministic prediction sequence are spliced ​​in the channel dimension to obtain a joint input tensor. The historical and predicted information are first fused through double convolution to output the intermediate fusion feature. Subsequently, the intermediate fusion features are downsampled three times, and a multi-scale adapter module is embedded after each downsampling - including 1×1 convolution compression, DWConv extraction of spatial features, SE channel attention recalibration and residual addition to generate three downsampled features; the parallel noise projector convolves the random Gaussian noise and then passes it through four linear projection modules to obtain the noise features and the last downsampled features, which are then spliced ​​into the generative decoder. The generative decoder is composed of upsampling and generation modules in series. During the upsampling process, the generation module is executed twice in the first and third stages, and once in the second stage. Each generation module incorporates motion field and intensity residual field condition information for feature regulation; the decoder finally outputs a 1-channel probabilistic prediction map through convolution, that is, the precipitation prediction map.

7. The precipitation prediction method based on multi-scale physical Transformer according to claim 6 is characterized in that: The generation module contains three parallel branches, namely a convolutional network, a linear mapping of the deterministic prediction sequence at position p, and a local attention mechanism. Finally, the outputs of the three are added and fused.

8. The precipitation prediction method based on multi-scale physical Transformer according to claim 7 is characterized in that: The construction of the multi-scale spatiotemporal physical consistency loss function is specifically as follows: The multi-scale spatiotemporal physical consistency loss function is a dynamic weighted sum of balanced pixel-level loss, multi-scale spatial loss, and multi-scale spatial loss; The balanced pixel-level loss: using balanced mean square error (BMSE) and balanced mean absolute error (BMAE) to assign different weights to areas with different precipitation intensities; The multi-scale spatial loss: the forecast image is decomposed into low-frequency and high-frequency sub-bands based on discrete wavelet transform, and the structural similarity SSIM index is used to calculate the spatial error of each sub-band; The multi-scale time loss constrains the temporal continuity of the forecast sequence by performing a pooling operation on the differences between consecutive time frames.

Citation Information

Cited By

  • Atmospheric pollutant concentration prediction method based on ICLU-CWGAN model

    CN121210987A

  • Numerical weather forecast result calibration and correction method, system and equipment, medium and program product

    CN121679766A