A short-term precipitation forecast method and system based on diffusion focus fusion

By combining input diffusion denoising and backbone focusing with output conditional diffusion enhancement, the problems of noise influence and key area focusing in short-term precipitation forecasts of deep learning models are solved, improving the noise robustness and prediction accuracy of the forecasts, and generating high-resolution precipitation forecast images with clear details.

CN122063558BActive Publication Date: 2026-07-28NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING UNIV OF INFORMATION SCI & TECH
Filing Date
2026-04-21
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

Existing deep learning models are prone to oversmoothing predictions in short-term precipitation forecasts, resulting in blurred output echo boundaries, missing details, difficulty in focusing on key strong echo areas, and input noise and errors affecting robustness, as well as insufficient transparency and interpretability.

Method used

The input diffusion denoising module suppresses noise, the backbone prediction network focuses on key regions by weighting gradient information, and the output conditional diffusion enhancement module restores high-resolution details. By combining a 3D U-shaped network and a 3D convolutional neural network structure, the input quality and prediction accuracy are improved.

Benefits of technology

It improves noise robustness, enhances forecast accuracy in areas of strong convection, achieves high-resolution detail reconstruction, and generates high-quality forecast images with rich textures and clear structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122063558B_ABST
    Figure CN122063558B_ABST
Patent Text Reader

Abstract

The application discloses a short-term precipitation prediction method and system based on diffusion focus fusion, which comprises obtaining a historical radar echo image sequence as input; firstly, time-space joint denoising is carried out through an input diffusion denoising module to suppress observation noise, and a high-quality feature sequence after cleaning is output; then, the high-quality feature sequence is input into a time-space prediction backbone network to extract spatial features and generate dynamic weights based on gradient information, so that the network focuses on key areas such as strong echoes, and a preliminary prediction sequence is output through time series modeling; finally, the preliminary sequence is input into an output conditional diffusion enhancement module as a conditional constraint to guide the reverse denoising process, and a high-resolution prediction sequence with detail enhancement and clear boundaries is generated. The application significantly improves the precision and detail performance of short-term prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of radar echo extrapolation technology, specifically to a short-term precipitation forecasting method and system based on diffusion-focusing fusion. Background Technology

[0002] In recent years, deep learning-based short-term precipitation forecasting methods have been developed. These methods automatically learn precipitation patterns and implicit motion laws by training historical radar echo sequences end-to-end, which to some extent makes up for the reliance of traditional extrapolation on manual features and simplification assumptions. However, existing deep learning models still face several shortcomings in meteorological applications, both in terms of engineering and performance: First, training targets are often driven by pixel-level errors (such as MSE / MAE), which can easily lead to "oversmoothed" predictions, resulting in blurred output echo boundaries, loss of detail, and difficulty in preserving the structural texture of the core area of ​​heavy precipitation, especially under longer forecast lead times. Second, radar observation data may have quality issues such as noise, missing data, and clutter during acquisition and mosaicking, and input errors will be amplified during time propagation, leading to a decrease in prediction robustness. Third, precipitation echo fields have significant spatial inhomogeneity, with key strong echo areas often accounting for a small proportion and changing rapidly. Conventional networks are easily dominated by large areas of weak echoes / no echoes when learning globally on an equal footing, making it difficult to effectively focus on key areas, thus affecting the performance of operational indicators under heavy precipitation thresholds. Fourth, the prediction process of deep models lacks transparency and interpretability, which is not conducive to the understanding and verification of the contribution sources of key features by the operational side. Summary of the Invention

[0003] Purpose of the invention: The purpose of this invention is to provide a short-term precipitation forecasting method and system based on diffusion-focusing fusion. At the input end, diffusion denoising is used to improve the input quality. During the main prediction process, gradient information is used to activate and weight the feature map to focus on key spatiotemporal regions. At the output end, conditional diffusion is used to achieve super-resolution reconstruction and detail enhancement of the prediction results. This solves the problems of unstable input radar echo sequence quality, easy blurring of output prediction results, and difficulty in focusing and modeling key strong echo regions.

[0004] Technical solution: The present invention provides a short-term precipitation forecasting method based on diffusion-focusing fusion, comprising the following steps:

[0005] (1) Obtain the radar echo image sequence at historical moments as the input sequence;

[0006] (2) Input the input sequence into the pre-trained input diffusion denoising module. Through the forward noise addition and reverse denoising process, the noise in the input sequence is suppressed to obtain a cleaned high-quality spatiotemporal feature sequence.

[0007] (3) The cleaned high-quality spatiotemporal feature sequence is input into the pre-trained focused spatiotemporal prediction backbone network. The backbone network first extracts the spatial feature map through convolution operation, and then generates dynamic weights based on gradient information to weight the feature map, so that the network focuses on key evolution regions such as strong echo regions. Then, the recurrent neural network models the temporal dependency relationship and outputs the preliminary predicted radar echo sequence.

[0008] (4) The preliminary predicted radar echo sequence is used as conditional information and input into the pre-trained output conditional diffusion enhancement module. During the reverse denoising generation process, the preliminary predicted sequence is used as a guide to generate a high-resolution radar echo prediction sequence with enhanced details and clear boundaries.

[0009] Furthermore, in step (2), the input diffusion denoising module adopts a three-dimensional U-shaped network structure, which gradually adds noise to the input sequence through the forward diffusion process, and then removes the noise step by step through the reverse denoising process, thereby realizing data reconstruction and noise suppression in the spatiotemporal dimension.

[0010] Furthermore, in step (3), the dynamic weights generated based on gradient information are specifically: spatial attention weights are generated by using the gradient information of the feature map through the loss function, and the weights are applied to the original feature map to enhance the feature expression of the region sensitive to prediction error and improve the model's ability to model nonlinear evolution processes such as precipitation generation and dissipation and intensity abrupt changes.

[0011] Furthermore, in step (4), the output conditional diffusion enhancement module adopts a three-dimensional convolutional neural network structure. In the reverse denoising process, the preliminary prediction sequence is used as a conditional constraint to guide the generation process to maintain the overall structural consistency while restoring high-frequency detail information and alleviating the problems of blurry and missing boundaries in the forecast image.

[0012] The present invention discloses a short-term precipitation forecasting system based on diffusion-focusing fusion, comprising: Data acquisition module: used to acquire radar echo image sequences from historical moments as input sequences; Input diffusion denoising module: It is used to input the input sequence into the pre-trained input diffusion denoising module. Through forward noise addition and backward noise reduction processes, the noise in the input sequence is suppressed to obtain a cleaned high-quality spatiotemporal feature sequence. Focusing Spatiotemporal Prediction Backbone Network: This network is used to input the cleaned high-quality spatiotemporal feature sequences into the pre-trained focusing spatiotemporal prediction backbone network. The backbone network first extracts spatial feature maps through convolution operations, and then generates dynamic weights based on gradient information to weight the feature maps, so that the network focuses on key evolution regions such as strong echo regions. Then, it models the temporal dependencies through recurrent neural networks and outputs the preliminary predicted radar echo sequences. Output conditional diffusion enhancement module: This module takes the initially predicted radar echo sequence as conditional information and inputs it into the pre-trained output conditional diffusion enhancement module. During the reverse denoising generation process, it uses the initially predicted sequence as a guide to generate a high-resolution radar echo prediction sequence with enhanced details and clear boundaries.

[0013] Furthermore, in the input diffusion denoising module, a three-dimensional U-shaped network structure is adopted. Noise is gradually added to the input sequence through a forward diffusion process, and then the noise is removed step by step through a reverse denoising process, so as to realize data reconstruction and noise suppression simultaneously in the spatiotemporal dimensions.

[0014] Furthermore, in the spatiotemporal prediction backbone network, the dynamic weights generated based on gradient information are specifically: spatial attention weights are generated by using the gradient information of the feature map through the loss function, and these weights are applied to the original feature map to enhance the feature representation of regions sensitive to prediction errors and improve the model's ability to model nonlinear evolution processes such as precipitation generation and dissipation and abrupt changes in intensity.

[0015] Furthermore, in the output conditional diffusion enhancement module, a three-dimensional convolutional neural network structure is adopted. During the reverse denoising process, the preliminary prediction sequence is used as a conditional constraint to guide the generation process to maintain the overall structural consistency while restoring high-frequency detail information and alleviating the problems of blurry and missing boundaries in the forecast image.

[0016] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages: First, it improves noise robustness by effectively suppressing radar observation clutter and errors through input-end diffusion denoising, ensuring high-quality input features; Second, it enhances the forecast accuracy of strong convection areas by using a gradient dynamic weighting mechanism to enable the model to accurately focus on the core area of ​​heavy precipitation, overcoming the interference of large-area weak echoes and improving the modeling ability for nonlinear sudden precipitation; Third, it achieves high-resolution detail reconstruction by effectively solving the defects of "oversmoothing" and blurred boundaries in conventional deep learning model predictions through a conditional diffusion mechanism, generating high-quality forecast images with rich textures and clear structures. Attached Figure Description

[0017] Figure 1 This is a flowchart of the present invention. Detailed Implementation

[0018] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0019] like Figure 1 As shown, this embodiment of the invention provides a short-term precipitation forecasting method based on diffusion-focusing fusion, comprising the following steps:

[0020] Step 1: Obtain a sequence of multiple consecutive original radar echo images as an input sequence (Input), where each frame is two-dimensional grid data. The input sequence is used to predict the future sequence of multiple radar echo images (Output).

[0021] Step 2: Input the input sequence into the input data cleaning module for denoising. The input data cleaning module uses a 3D U-Net backbone network to extract and reconstruct features from the spatiotemporal sequence using 3D convolutions; the 3D U-Net output layer can output tensor results with shapes of (8, 1, 10, 480, 480), and the intermediate decoding process can restore the dimensions to (8, 64, 10, 480, 480).

[0022] The formula for the forward noise addition process in the input data cleaning module is as follows:

[0023]

[0024] in This represents the true conditional probability distribution of the positive noise-adding process. Represents a normal distribution (Gaussian distribution). During the noise addition process, the first Radar echo data at time step, This represents the noise level at each step, typically increasing with time t. During the noise addition process, the first Radar echo data at time step, This represents the identity matrix. Through multiple recursive steps, the data is finally transformed into standard Gaussian noise in step T:

[0025]

[0026] in This represents the total number of steps in the noise-adding process. This represents the original historical radar echo image sequence with observation noise (i.e., the input sequence).

[0027] For radar echo data, noise addition and propagation occur simultaneously in both spatial and temporal dimensions. Therefore, noise... It is a 3D Gaussian noise with a shape consistent with the input data.

[0028] The reverse diffusion process is a step-by-step denoising process that recovers the original data from pure noise. Each step utilizes 3D U-Net to predict the true distribution of the data and progressively removes noise, thereby generating true samples. The following formula represents this:

[0029]

[0030] in, This represents the conditional probability distribution of the inverse denoising process learned and fitted by the 3D U-Net model. Represented by 3D U-Net in the 1st The mean of the data predicted at each time step. Representative by the first The variance of the time step data This represents all learnable network parameters in the 3D U-Net model. At each step, it represents the mean of the 3D U-Net predictions at the current time step. and variance Then, this information is used to recover the data from the noise.

[0031] Step 3: Input the denoised image sequence into the backbone prediction network. The backbone prediction network first extracts spatial features through convolutional layers to generate feature maps. Next, the gradient activation mapping module calculates the gradient of the target output with respect to the convolutional layer feature map through backpropagation, generating channel-level attention weights. The weights are then multiplied by the original feature map to obtain the weighted feature map. This enhances the representation of key precipitation regions. The weighted feature map, after downsampling by a pooling layer, is passed to a GRU layer. The GRU layer models temporal dependencies through a gating mechanism, capturing long-range dependencies in the time series. The mathematical formula for the GRU layer is as follows.

[0032] The input layer receives the input sequence. Input the updated gate vector and resetting the gating vector In this context, the formula for updating the gating vector is:

[0033]

[0034] The formula for resetting the gating vector is:

[0035]

[0036] in, and It is the weight matrix in the calculation process. and It is the bias vector.

[0037] Candidate hidden state The calculation formula is:

[0038]

[0039] Current hidden state The calculation formula is:

[0040]

[0041] Finally, the output of the GRU layer is mapped to the preliminary precipitation prediction results through a fully connected layer.

[0042] Step 4: Input the preliminary prediction sequence output from Step 3 as conditional information c into the output data augmentation module. As a guiding variable, it is used in the inverse denoising stage with noisy image sequences. The data is concatenated through channels and input into the output data enhancement module, where it is enhanced through a conditional diffusion process. The output data enhancement module progressively adds noise during forward diffusion and removes noise progressively through a 3D U-Net during reverse denoising. Conditional information 'c' is introduced at each step to constrain the generation process. While the results predicted by the backbone prediction network (step 3) may have issues such as ambiguity and missing details, they provide the core structure and macroscopic movement trends of future precipitation evolution. The output data enhancement module uses this preliminary prediction result as a conditional... The purpose of constraining the generation process is to ensure that the final generated high-resolution image sequence is highly consistent with the initial prediction in terms of macroscopic structure. At the same time, the generation capability of the diffusion model is used to supplement the missing high-frequency textures and boundary details, so that the output results are consistent with the initial prediction and have higher resolution and clearer details.

[0043] The formula for the forward noise addition process of the output data augmentation module is expressed as follows:

[0044]

[0045] Here, This represents the noise level at each step, typically increasing with time t. Through recursive application across multiple steps, the data eventually transforms into standard Gaussian noise at step T.

[0046]

[0047] The reverse diffusion process is from noise data The process of progressively denoising and generating high-quality data. The module uses conditional information 'c' to guide the generation process at each step.

[0048]

[0049] At each step, the mean of the 3D U-Net prediction data at the current time step. and variance Then, this information is used to recover the data from the noise.

[0050] [5]. Output the enhanced short-term precipitation forecast results and perform operational evaluation based on the dBZ threshold: when dBZ is between 20 and 35, it corresponds to light to moderate precipitation; when dBZ is between 35 and 45, it corresponds to moderate to heavy precipitation; when dBZ ≥ 45, it corresponds to heavy precipitation scenario.

[0051] experiment:

[0052] Dataset: The HKO-7 real radar echo dataset was used as the experimental dataset. The dataset information is shown in Table 1.

[0053] Table 1 HKO-7 Dataset ;

[0054] Data Preprocessing: To effectively improve the accuracy and computational efficiency of precipitation nowcasting, enabling the model to better cope with complex and ever-changing meteorological forecasting needs, some data preprocessing is required on the original HKO-7 dataset. This mainly includes: image resolution adjustment, as the high resolution of the original radar echo images leads to excessive computational overhead, so the resolution of all images needs to be uniformly adjusted (downsampling) to enable spatiotemporal feature learning under reasonable computational resources; and data standardization, which normalizes pixel values ​​to a standard range and filters out abnormal noise through linear mapping, thereby accelerating the convergence of the neural network and making the prediction results more robust.

[0055] Experimental platform:

[0056] CPU: Intel Xeon Gold 5220R

[0057] GPU: Nvidia A100 40G *2

[0058] Memory: 256G

[0059] System: Ubuntu 20.04.5

[0060] Experimental results:

[0061] The model evaluation metrics include two aspects: image quality metrics and service evaluation metrics. Experiments compared the Diffusion-Focusing Convolutional Gated Network (DF-CGNet) from this invention with other comparative models such as Convolutional Long Short-Term Memory Network (ConvLSTM), Trajectory Gated Recurrent Unit (TrajGRU), PredRNN++ (PredRNN++), Nowcastnet, and Trajectory PredRNN+ (trajPredRNN+) on the HKO-7 dataset through unified training and testing. As shown in Table 2, in terms of image quality metrics, DF-CGNet from this invention achieved the best overall results: mean absolute error of 286.4, mean squared error of 243.6, structural similarity index of 0.858, and peak signal-to-noise ratio of 30.25. Compared to Nowcastnet, DF-CGNet reduced the mean absolute error and mean squared error by approximately 6.62% and 7.87%, respectively, and improved the structural similarity index and peak signal-to-noise ratio by approximately 1.06% and 2.13%, respectively. The results show that the present invention, through the synergistic effect of input-end diffusion denoising, backbone focusing modeling, and output-end conditional diffusion enhancement, can reduce errors while maintaining higher structural consistency and better image quality.

[0062] Furthermore, to verify the operational applicability of this invention under different precipitation intensity scenarios, the prediction results were binarized and evaluated according to dBZ thresholds at dBZ≥20, dBZ≥35, and dBZ≥45, and operational indicators such as CSI, HSS, POD, and FAR were calculated. As shown in Tables 3 to 5, the DF-CGNet of this invention exhibits superior overall operational performance under all three thresholds. Using NowcastNet as a reference: when dBZ ≥ 20, DF-CGNet's CSI improved from 0.803 to 0.812, HSS from 0.741 to 0.796, and FAR decreased from 0.208 to 0.201; when dBZ ≥ 35, CSI improved from 0.578 to 0.596, POD from 0.762 to 0.784, and FAR decreased from 0.501 to 0.488; when the heavy precipitation threshold dBZ ≥ 45, CSI improved from 0.362 to 0.397, demonstrating the invention's stronger ability to characterize key areas of heavy precipitation, while FAR decreased from 0.672 to 0.662, which helps reduce false alarms. These results indicate that the invention has better operational evaluation performance in both weak to moderate precipitation and heavy precipitation scenarios, with a more pronounced advantage at the heavy precipitation threshold.

[0063] Table 2 Comparison of Model Accuracy and Performance ;

[0064] Table 3. Comparison of operational indicators when the precipitation intensity threshold is dBZ≥20 ;

[0065] Table 4 Comparison of operational indicators when the precipitation intensity threshold is dBZ≥35 ;

[0066] Table 5 Comparison of operational indicators when the precipitation intensity threshold is dBZ≥45 .

Claims

1. A short-term precipitation forecasting method based on diffusion-focusing fusion, characterized in that, Includes the following steps: (1) Obtain the radar echo image sequence at historical moments as the input sequence; (2) Input the input sequence into the pre-trained input diffusion denoising module. Through the forward noise addition and reverse denoising process, the noise in the input sequence is suppressed to obtain a cleaned high-quality spatiotemporal feature sequence. (3) The cleaned high-quality spatiotemporal feature sequence is input into the pre-trained focused spatiotemporal prediction backbone network. The backbone network first extracts spatial feature maps through convolutional operations, and then generates dynamic weights based on gradient information to weight the feature maps, so that the network focuses on the strong echo region. Then, the recurrent neural network models the temporal dependency relationship and outputs the preliminary predicted radar echo sequence. The backbone prediction network first extracts spatial features through convolutional layers to generate feature maps. Next, the gradient activation mapping module calculates the gradient of the target output with respect to the convolutional layer feature map through backpropagation, generating channel-level attention weights. The weights are then multiplied by the original feature map to obtain the weighted feature map. To enhance the characteristic representation of key precipitation areas; The weighted feature map is downsampled by the pooling layer and then passed to the GRU layer. The GRU layer models temporal dependencies through a gating mechanism to capture long-range dependencies in the time series. The input layer receives the input sequence. Input the updated gate vector and resetting the gating vector In this context, the formula for updating the gating vector is: ; The formula for resetting the gating vector is: ; in, and It is the weight matrix in the calculation process. and It is the bias vector; Candidate hidden state The calculation formula is: ; Current hidden state The calculation formula is: ; Finally, the output of the GRU layer is mapped to the preliminary precipitation prediction results through a fully connected layer; (4) The preliminary predicted radar echo sequence is used as conditional information and input into the pre-trained output conditional diffusion enhancement module. During the reverse denoising generation process, the preliminary predicted sequence is used as a guide to generate a high-resolution radar echo prediction sequence with enhanced details and clear boundaries.

2. The short-term precipitation forecasting method based on diffusion-focusing fusion according to claim 1, characterized in that, In step (2), the input diffusion denoising module adopts a three-dimensional U-shaped network structure. It gradually adds noise to the input sequence through the forward diffusion process, and then removes the noise step by step through the reverse denoising process, so as to realize data reconstruction and noise suppression in the spatiotemporal dimension at the same time.

3. The short-term precipitation forecasting method based on diffusion-focusing fusion according to claim 1, characterized in that, In step (4), the output conditional diffusion enhancement module adopts a three-dimensional convolutional neural network structure. In the reverse denoising process, the preliminary prediction sequence is used as a conditional constraint to guide the generation process to maintain the overall structural consistency while restoring high-frequency detail information.

4. A short-term precipitation forecasting system based on diffusion-focusing fusion, implemented using the method described in any one of claims 1-3, characterized in that, include: Data acquisition module: used to acquire radar echo image sequences from historical moments as input sequences; Input diffusion denoising module: It is used to input the input sequence into the pre-trained input diffusion denoising module. Through forward noise addition and backward noise reduction processes, the noise in the input sequence is suppressed to obtain a cleaned high-quality spatiotemporal feature sequence. Focusing Spatiotemporal Prediction Backbone Network: This network is used to input the cleaned high-quality spatiotemporal feature sequence into the pre-trained focusing spatiotemporal prediction backbone network. The backbone network first extracts spatial feature maps through convolution operations, then generates dynamic weights based on gradient information to weight the feature maps, so that the network focuses on strong echo regions. Finally, it models the temporal dependency relationship through a recurrent neural network and outputs the preliminary predicted radar echo sequence. Output conditional diffusion enhancement module: This module takes the initially predicted radar echo sequence as conditional information and inputs it into the pre-trained output conditional diffusion enhancement module. During the reverse denoising generation process, it uses the initially predicted sequence as a guide to generate a high-resolution radar echo prediction sequence with enhanced details and clear boundaries.

5. A short-term precipitation forecasting system based on diffusion-focusing fusion according to claim 4, characterized in that, In the input diffusion denoising module, a three-dimensional U-shaped network structure is adopted. Noise is gradually added to the input sequence through a forward diffusion process, and then the noise is removed step by step through a reverse denoising process, so as to realize data reconstruction and noise suppression in the spatiotemporal dimensions at the same time.

6. A short-term precipitation forecasting system based on diffusion-focusing fusion according to claim 4, characterized in that, In the output conditional diffusion enhancement module, a three-dimensional convolutional neural network structure is adopted. During the reverse denoising process, the preliminary prediction sequence is used as a conditional constraint to guide the generation process to maintain the overall structural consistency while restoring high-frequency detail information.