Rainfall prediction method and system, electronic equipment and storage medium
Through the PFTF precipitation forecast model, the encoder and decoder combined with the attention mechanism of the TripFormer module is solved by solving the problem of difficult to achieve short-term precipitation forecast in the existing technology, and the accurate prediction of complex meteorological phenomena and effective capture of precipitation characteristics are achieved.
Patent Information
- Application Number
- CN202510070118.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-27
AI Technical Summary
It is difficult for the prior art to achieve short-term precipitation forecasts, and there are huge uncertainties in numerical weather forecasts based on physical models and problems of increased computational demand.
Using a PFTF precipitation forecast model, which includes an encoder, a decoder and multiple TripFormer modules, captures spatiotemporal features through an optimized attention mechanism and performs deep convolution operations on multiple scales to achieve precipitation forecasts.
The understanding and forecasting accuracy of complex meteorological phenomena is improved, the accurate prediction of short-term precipitation is achieved, and the characteristics of rainfall events and intermittent spatial distribution patterns are more effectively captured through the weighted structural loss function.
Smart Images

Figure CN120045871A_ABST
Abstract
Description
Background Art
[0002] Accurate weather forecasting is crucial for many aspects of people's lives. Traditional short-term precipitation forecasting methods mainly rely on numerical weather forecasting calculated based on weather dynamics modeling. However, numerical weather forecasting based on high-resolution physical models requires hours of calculations on supercomputers. Deep learning networks have strong feature extraction capabilities and can directly map information such as images and meteorological data into a high-dimensional feature space, extracting abstract features through neurons, and thus making predictions through continuous training and learning. However, in the past, deep learning was usually used for short-term precipitation forecasting by the extrapolation method of radar echo maps, and short-term precipitation forecasting could not be achieved.
[0003] In the past few decades, weather forecasting has mainly relied on numerical weather models, which use the framework of geophysical fluid dynamics to simulate the evolution of the atmospheric state by combining various physical parameterizations. However, due to many reasons, there are still huge uncertainties. For example, precipitation prediction involves complex, multi-scale interactions between aerosols, clouds, radiation, and large-scale meteorological conditions. Due to the lack of understanding of many physical processes (such as ice-phase and mixed-phase cloud physics), it is still difficult to use physical models for precipitation simulation. In addition, with finer grid resolutions and more realistic physical parameterizations, the demand for computing power and data storage is also increasing. Compared with the physical models of numerical weather forecasting (NWP), meteorological researchers train deep neural networks to find the relationship between weather data at a specific time point and weather data at the target time point. The extrapolation method using radar echo data does not fully utilize the rich sources in weather data sets, such as numerical grids and meteorological station observation data. The main reason is that weather is a chaotic system affected by many factors such as vegetation, geographical contours, and human factors, resulting in the fact that limited information cannot well simulate the chaotic behavior of weather. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a precipitation prediction method, system, electronic device, and storage medium in view of the deficiencies of the prior art, specifically as follows:
[0005] 1) In the first aspect, the present invention provides a precipitation prediction method, and the specific technical solution is as follows:
[0006] Construct a PFTF precipitation forecasting model. The PFTF precipitation forecasting model includes an encoder, a decoder, and multiple TripFormer modules. The multiple TripFormer modules are sequentially arranged between the encoder and the decoder. The encoder is used to: extract multi-scale features from precipitation sample data and perform enhancement processing. Each TripFormer module is used to: capture the dependencies in the channel dimension and spatial dimension of the features received by each through optimizing the attention mechanism to obtain spatio-temporal features. The decoder is used to: perform depth convolution operations on the received features at multiple scales to obtain the precipitation forecasting result corresponding to the precipitation sample data. Among them, the features received by the decoder are: the features obtained after processing the spatio-temporal features output by the last TripFormer module;
[0007] Train the PFTF precipitation forecasting model to obtain a trained PFTF precipitation forecasting model;
[0008] Input precipitation target data into the trained PFTF precipitation forecasting model to obtain the precipitation forecasting result corresponding to the precipitation target data.
[0009] The beneficial effects of a precipitation prediction method provided by the present invention are as follows:
[0010] The encoder can extract and enhance multi-scale features to improve the understanding and forecasting accuracy of complex meteorological phenomena. The TripFormer module improves the ability to capture spatio-temporal features by optimizing the attention mechanism, while reducing the computational burden. The decoder performs depth convolution operations on the received features at multiple scales, enhancing the feature generation ability in the cascade expansion path. Therefore, the present invention can achieve accurate prediction of short-term precipitation.
[0011] Based on the above solution, a precipitation prediction method of the present invention can also be improved as follows.
[0012] Further, it further includes:
[0013] Combine the continuous weighted mean square error and the difference divergence regularization term to construct a weighted structure loss function;
[0014] Use the weighted structure loss function as the loss function used when training the PFTF precipitation forecasting model.
[0015] The beneficial effect of adopting the above further solution is that it can more effectively capture the characteristics of rainfall events and the intermittent spatial distribution pattern of precipitation, and perform weighting for different regions and precipitation intensities, which can enhance the perception ability of complex precipitation phenomena while optimizing the performance of the PFTF precipitation forecasting model.
[0016] Further, the encoder includes a first ConvSC layer and a first multi-scale convolution module, which are arranged in sequence. The decoder includes a second ConvSC layer and a second multi-scale convolution module, which are arranged in sequence.
[0017] Further, the TripFormer module includes a first normalization layer, an attention module, a second normalization layer, and an MLP with non-linear activation, which are arranged in sequence. The attention module includes a channel attention calculation branch and two spatial dimension interaction branches.
[0018] 2) In a second aspect, the present invention also provides a precipitation prediction system, and the specific technical solution is as follows:
[0019] It includes a model construction module, a model training module, and a precipitation prediction module;
[0020] The model construction module is used to: construct a PFTF precipitation prediction model. The PFTF precipitation prediction model includes an encoder, a decoder, and multiple TripFormer modules. The multiple TripFormer modules are arranged in sequence between the encoder and the decoder. The encoder is used to: extract multi-scale features from precipitation sample data and perform enhancement processing. Each TripFormer module is used to: capture the dependencies in the channel dimension and spatial dimension of the features received by each through optimizing the attention mechanism to obtain spatio-temporal features. The decoder is used to: perform depth convolution operations on the received features at multiple scales to obtain the precipitation prediction result corresponding to the precipitation sample data. Among them, the features received by the decoder are: the features obtained after processing the spatio-temporal features output by the last TripFormer module;
[0021] The model training module is used to: train the PFTF precipitation prediction model to obtain a trained PFTF precipitation prediction model;
[0022] The precipitation prediction module is used to: input precipitation target data into the trained PFTF precipitation prediction model to obtain the precipitation prediction result corresponding to the precipitation target data.
[0023] Based on the above solution, a precipitation prediction system of the present invention can also be improved as follows.
[0024] Further, it further includes a weighted structure loss function construction module, which is used to: combine the continuous weighted mean square error and the difference divergence regularization term to construct a weighted structure loss function;
[0025] The model training module is used to: use the weighted structure loss function as the loss function when training the PFTF precipitation prediction model.
[0026] Furthermore, the encoder includes a first ConvSC layer and a first multi-scale convolution module, which are arranged in sequence, and the decoder includes a second ConvSC layer and a second multi-scale convolution module, which are arranged in sequence.
[0027] Furthermore, the TripFormer module includes a first normalization layer, an attention module, a second normalization layer, and an MLP with non-linear activation, which are arranged in sequence. The attention module includes a channel attention calculation branch and two spatial dimension interaction branches.
[0028] 3) In a third aspect, the present invention also provides an electronic device, which includes a processor. The processor is coupled to a memory, and at least one computer program is stored in the memory. The at least one computer program is loaded and executed by the processor to enable the electronic device to implement any one of the above precipitation prediction methods.
[0029] 4) In a fourth aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements any one of the above precipitation prediction methods.
[0030] It should be noted that for the beneficial effects obtained by the technical solutions of the second to fourth aspects of the present invention and their corresponding possible implementation manners, reference may be made to the technical effects of the first aspect and its corresponding possible implementation manners described above, and details are not elaborated herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention:
[0032] Figure 1 It is a schematic flowchart of a precipitation prediction method according to an embodiment of the present invention;
[0033] Figure 2 It is a schematic diagram of the network structure of the PFTF precipitation forecast model;
[0034] Figure 3 It is a schematic diagram of the network structure of the MSC;
[0035] Figure 4 It is a schematic diagram of the network structure of the TripFormer module;
[0036] Figure 5 It is a schematic diagram of the network structure of the MSCB;
[0037] Figure 6 It is a schematic diagram of the structure of a precipitation prediction system according to an embodiment of the present invention;
[0038] Figure 7 The structural schematic diagram of an electronic device according to an embodiment of the present invention. Specific embodiments
[0039] The principles and features of the present invention are described below. The examples given are only used to explain the present invention and are not intended to limit the scope of the present invention.
[0040] The technical solution of the present invention and how the technical solution of the present invention solves the above technical problems will be described in detail below with specific embodiments. These several specific embodiments can be combined with each other. For the same or similar concepts or processes, they may not be repeated in some embodiments. The embodiments of the present invention will be described below with reference to the accompanying drawings.
[0041] As Figure 1 shown, a precipitation prediction method according to an embodiment of the present invention includes the following steps:
[0042] S1. Construct a PFTF precipitation forecast model.
[0043] As Figure 2 shown, the PFTF precipitation forecast model includes an encoder, a decoder, and a plurality of TripFormer modules. The plurality of TripFormer modules are sequentially arranged between the encoder and the decoder. The encoder is used to: extract multi-scale features from precipitation sample data and perform enhancement processing. Each TripFormer module is used to: capture the dependencies of the channel dimension and the spatial dimension in the features received by each through optimizing the attention mechanism to obtain spatio-temporal features. The decoder is used to: perform depth convolution operations on the received features at multiple scales to obtain a precipitation forecast result corresponding to the precipitation sample data, where the features received by the decoder are: the features obtained after processing the spatio-temporal features output by the last TripFormer module.
[0044] Figure 2 In, Conv1 represents the first convolutional layer, Conv2 represents the second convolutional layer, ConvSC1 represents the first ConvSC layer, and ConvSC2 represents the second ConvSC layer.
[0045] Among them, after the output of the encoder is processed by the first convolutional layer, the output of the first convolutional layer is used as the input of the first TripFormer module. The output of the last TripFormer module is input into the second convolutional layer, and the output of the second convolutional layer is used as the input of the decoder.
[0046] To systematically analyze the structure of the model in spatio-temporal prediction learning, the present invention proposes a PFTF precipitation forecasting model. The encoder (also referred to as the Multi-Scale Encoder) in the PFTF precipitation forecasting model includes a first ConvSC layer and a first multi-scale convolution module, and the first ConvSC layer and the first multi-scale convolution module are arranged in sequence. Specifically:
[0047] The first ConvSC layer extracts local features in the input data (the feature map corresponding to the precipitation sample data) through convolution operations. That is to say, the precipitation sample data is input into the first ConvSC layer in the form of a feature map, and the ConvSC layer is used to extract the spatial features of the input data. The ConvSC layer captures the local features in the input data through convolution operations and enhances the non-linear expression ability of the model through normalization and activation functions. The encoder has the function of downsampling, which can reduce the size of the feature map corresponding to the precipitation sample data to gradually reduce the spatial resolution and obtain higher feature abstraction.
[0048] Figure 3 In this, Conv3 represents the third convolutional layer, Conv4 represents the fourth convolutional layer, Conv5 represents the fifth convolutional layer, Conv6 represents the sixth convolutional layer, Conv7 represents the seventh convolutional layer, Conv8 represents the eighth convolutional layer, Conv9 represents the ninth convolutional layer, Conv10 represents the tenth convolutional layer, Conv11 represents the eleventh convolutional layer, Conv12 represents the twelfth convolutional layer, MSCA represents the MSCA module, and StarReLU represents the StarReLU layer.
[0049] Among them, the first multi-scale convolution module can be MSC (Multi-Scale Convolution), which enhances the feature representation by aggregating context information of different scales, such as Figure 3As shown in the figure, the MSC includes a third convolutional layer, a StarReLU layer, an MSCA module, and a fourth convolutional layer. The output of the third convolutional layer is used as the input of the first ConvSC layer. The StarReLU layer enhances the flexibility and expressive power of feature representation by introducing a non-linear transformation. The output of the first ConvSC layer is used as the input of the MSCA module. The MSCA module includes a fifth convolutional layer, a sixth convolutional layer, a seventh convolutional layer, an eighth convolutional layer, a ninth convolutional layer, a tenth convolutional layer, an eleventh convolutional layer, and a twelfth convolutional layer. The MSCA module includes three branches. The sixth convolutional layer and the seventh convolutional layer form one branch, the eighth convolutional layer and the ninth convolutional layer form one branch, and the tenth convolutional layer and the eleventh convolutional layer form one branch. The output of the first ConvSC layer is specifically used as the input of the fifth convolutional layer. The output of the fifth convolutional layer is processed through each branch respectively, and the outputs of each branch are added together. Then the added result is used as the input of the twelfth convolutional layer. The convolution kernel size of the twelfth convolutional layer is 5×5, which can aggregate the local spatial information in the added result, further enhancing the model's ability to capture local features. Moreover, through the three branches, the capture of multi-scale context is realized, so as to extract richer semantic information at different scales. The output of the twelfth convolutional layer is multiplied by the output of the first ConvSC layer, and the multiplied result is input into the fourth convolutional layer. The convolution kernel size of the fourth convolutional layer is 1×1, which models the relationship between different channels. Its output is directly used as the attention weight to re-weight the input data. The expression is as follows:
[0050]
[0051] Among them, F represents the input feature (the feature map corresponding to the precipitation sample data), Att and Out are the attention weight and the output respectively, represents the element-wise matrix multiplication operation, DWConv represents the depth convolution, Scale i (i ∈ {1, 2, 3}) represents the i-th branch, Scale i is the residual connection.
[0052] Among them, the convolution kernel size of the fifth convolutional layer is 5×5, the convolution kernel size of the sixth convolutional layer is 1×7, the convolution kernel size of the seventh convolutional layer is 7×1, the convolution kernel size of the eighth convolutional layer is 1×11, the convolution kernel size of the ninth convolutional layer is 11×1, the convolution kernel size of the tenth convolutional layer is 21×1, and the convolution kernel size of the eleventh convolutional layer is 1×21.
[0053] Figure 4 In, Norm1 represents: the first normalization layer, Norm2 represents: the second normalization layer. Triplet Attention represents: the attention module.
[0054] As shown Figure 4 in the figure, the TripFormer module includes a first normalization layer, an attention module, a second normalization layer, and an MLP with non-linear activation, which are arranged in sequence. The attention module includes a channel attention calculation branch and two spatial dimension interaction branches. Specifically:
[0055] After the first normalization layer normalizes the received data X, it is input into the attention module, and the data X is added to the output of the attention module, that is, the first normalization layer is residually connected to the attention module for mixing channels. This process can be expressed by the following expression:
[0056] Y = Token Mixer(Norm(X)) + X
[0057] where Norm(X) represents that the first normalization layer normalizes the received data X, Token Mixer(Norm(X)) represents that the attention module processes the normalized data, and Y represents the addition result obtained by adding the data X to the output of the attention module.
[0058] The addition result Y is input into the second normalization layer. After the second normalization layer normalizes the addition result Y, it is input into an MLP (Multi-Layer Perceptron) with non-linear activation. The output of the MLP is added to the addition result Y, and this addition result is used as the output of the TripFormer module. This process can be expressed by the following expression:
[0059] Z = σ(Norm(Y)W 1 )W 2 + Y
[0060] where Z represents the addition result obtained by adding the output of the MLP to the addition result Y, Norm(Y) represents that the second normalization layer normalizes the addition result Y, σ represents the non-linear activation function, and W 1 and W 2 are weight matrices.
[0061] Among them, the attention module of the TripFormer module includes a channel attention calculation branch and two spatial dimension interaction branches. The two spatial dimension interaction branches are respectively: a channel-spatial width interaction branch and a channel-spatial height interaction branch. Specifically:
[0062] 1) The channel attention calculation branch includes a first pooling layer, a thirteenth convolutional layer, and a first BN layer (Batch Normalization Layer) arranged in sequence. The channel attention calculation branch is responsible for capturing the dependencies of the input features (the output of the first normalization layer) in the channel dimension (C). The channel attention calculation branch also applies element-wise multiplication (i.e., multiplying element by element) to the output of the first BN layer to obtain the output of the channel attention calculation branch. Figure 4 In it, C-Pool represents the first pooling layer, Conv13 represents the thirteenth convolutional layer, and BN1 represents the first BN layer.
[0063] 2) The channel-spatial width interaction branch includes a second pooling layer, a fourteenth convolutional layer, and a second BN layer arranged in sequence. The spatial dimension interaction branch calculates the attention weights of the input features (the output of the first normalization layer) in the width (W) and height (H) dimensions respectively, so as to establish the dependency relationship in the spatial dimension. The spatial dimension interaction branch also applies element-wise multiplication (i.e., multiplying element by element) to the output of the second BN layer to obtain the output of the spatial dimension interaction branch. Figure 4 In it, Z-Pool1 represents the second pooling layer, Conv14 represents the fourteenth convolutional layer, and BN2 represents the second BN layer.
[0064] The channel-spatial width interaction branch is responsible for capturing the interaction between the channel C and the spatial W dimensions. The input features are first transformed into H×C×W dimensional features, and then the Z-Pool operation is performed in the H dimension. After that, through a series of operations (Conv and BN), it finally becomes C×H×W dimensional features for element-wise multiplication (referring to operating on the corresponding elements in vectors, matrices, or tensors).
[0065] 3) The channel-spatial height interaction branch includes a third pooling layer, a fifteenth convolutional layer, and a third BN layer arranged in sequence, which is responsible for traditional spatial attention weight calculation to enhance the spatial information encoding ability. Specifically, the input features first go through the Z-Pool operation, which aggregates global information through the pooling operation. Then, the input passes through the fifteenth convolutional layer to extract local spatial features. Finally, the attention weights are generated through the Sigmoid activation function, and these weights are used to re-weight the input features, thereby enhancing the information interaction between channels and spatial dimensions. The spatial attention branch also applies element-wise multiplication (i.e., multiplying element by element) to the output of the third BN layer to obtain the output of the spatial attention branch. Figure 4 In it, Z-Pool2 represents the third pooling layer, Conv15 represents the fifteenth convolutional layer, and BN3 represents the third BN layer.
[0066] The channel-spatial height interaction branch is responsible for capturing the interaction between the channel C and spatial H dimensions. The input features are first transformed into features of dimension W×H×C, and then a Z-Pool operation is performed on the W dimension. After that, through a series of operations (Conv and BN), it finally becomes features of dimension C×H×W for element-wise multiplication.
[0067] The outputs of the channel attention calculation branch, the channel-spatial width interaction branch, and the channel-spatial height interaction branch are added and averaged to fuse the interaction information between different dimensions, obtaining the output of the attention module.
[0068] Figure 4 In it, the convolutional kernels of the thirteenth convolutional layer, the fourteenth convolutional layer, and the fifteenth convolutional layer are all 7×7.
[0069] Multiple TripFormer modules are stacked between the encoder and the decoder, aiming to utilize the attention mechanism of the three branches (the channel attention calculation branch, the channel-spatial width interaction branch, and the channel-spatial height interaction branch) to effectively capture the dependencies of the channel dimension and the spatial dimension in the input features by processing information of different dimensions in parallel. Through such a design architecture, it can effectively and clearly perceive the spatial correlation, temporal correlation, and spatio-temporal correlation, thus laying a solid foundation for the efficient operation and accurate output of the model.
[0070] Among them, the decoder (which can also be called the EfficientDecoder decoder) includes a second ConvSC layer and a second multi-scale convolutional module, and the second ConvSC layer and the second multi-scale convolutional module are arranged in sequence.
[0071] Among them, the second ConvSC layer is responsible for performing depth convolution operations to improve the quality of feature extraction. The second ConvSC layer implements the convolution operation through BasicConv2d, and whether to enable upsampling is determined by the upsampling parameter. In the case of enabling upsampling, the second ConvSC layer quadruples the number of output channels, and then uses PixelShuffle to double the spatial dimension of the feature map to complete the upsampling effect. If upsampling is not enabled, only standard convolution operations are used. To improve the stability and performance of the network, GroupNorm and SiLU activation functions are included in BasicConv2d for normalization and non-linear transformation, and whether to apply normalization and activation simultaneously can be controlled by parameters. In addition, the weight initialization in BasicConv2d uses truncated normal distribution, and the bias is initialized to zero.
[0072] To effectively preserve spatial-related features, a residual connection is introduced between the last convolutional layer and the first convolutional layer of the second ConvSC layer to enhance information transmission and ensure the preservation of spatial features during the decoding process.
[0073] Among them, the second multi-scale convolutional module is MSCB (Multi-scale Depth-wise Convolution, which is a modified convolutional neural network structure). As Figure 5 shown, MSCB further enhances the feature extraction ability by performing depthwise convolutional operations at multiple scales. In addition, the MSCB module uses the channel shuffle mechanism to strengthen the interaction between channels, effectively making up for the deficiency of depthwise convolution in modeling channel dependencies. Specifically, the EfficientDecoder first provides a basis for subsequent feature fusion and decoding by the ConvSC layer responsible for performing convolutional operations and upsampling operations, gradually increasing the spatial resolution of the feature map.
[0074] MSCB enhances the feature extraction and representation ability through multiple steps. First, MSCB expands the channels of the input features through 1×1 convolution (pointwise convolution), and the expansion factor is set to 2 to increase the channel capacity. Immediately following is the batch normalization layer to ensure the stability of the network and accelerate convergence. Then, the ReLU6 activation function is used for non-linear transformation to effectively expand the feature space and lay the foundation for multi-scale information capture.
[0075] In addition, MSCB captures context information at different scales by parallelly executing multi-scale depthwise convolutional modules (MSDC). Since depthwise convolution only operates in the spatial dimension and cannot directly capture the dependencies between channels, a channel shuffle operation is introduced. By rearranging the channels, the interaction between channels is enhanced, thereby further improving the feature expression ability. The features processed by multi-scale convolution are restored to the original dimension through another pointwise convolutional layer and BN layer, explicitly encoding the dependencies between channels to optimize the feature fusion effect.
[0076] Overall, the design of MSCB not only enhances the feature extraction and representation ability but also shows significant advantages in capturing multi-scale context information, thereby improving the network's adaptability to complex tasks and overall performance. The data processing process of MSCB can be expressed as:
[0077] MSCB(x) = BN(PWC2(CS(MSDC(R6(BN(PWC1(x)))))))
[0078] Among them, the data processing process of multiple MSDCs with different kernel sizes in parallel can be expressed as:
[0079]
[0080] Among them, DWCB ks (X) = R6(BN(DWCks(X))). Here, DWCB ks (.) is a depth convolution with a kernel size of ks, and BN and R6 are batch normalization and ReLU6 activation respectively. In addition, in the present invention, MSDC uses the recursively updated input X, where the input X is residually connected to the previous DWCB ks (.) to obtain better regularization.
[0081] S2. Train the PFTF precipitation forecasting model to obtain a trained PFTF precipitation forecasting model;
[0082] Among them, based on the ERA5 dataset, the PFTF precipitation forecasting model is trained. The ERA5 dataset is the fifth-generation atmospheric reanalysis dataset of the global climate from January 1950 to the present by the ECMWF (European Centre for Medium-Range Weather Forecasts). The ERA5 dataset is produced by the Copernicus Climate Change Service (C3S) of the ECMWF. The ERA5 dataset provides hourly estimates of a large number of atmospheric, land, and ocean climate variables. These data cover the Earth on a 30-kilometer grid and use 137 altitude levels from the surface to 80 kilometers to resolve the atmosphere, including information on the uncertainty of all variables when reducing the spatial and temporal resolution. The ERA5 dataset combines model data with observational data from all over the world to form a globally complete and consistent dataset, replacing its predecessor ERA-Interim reanalysis.
[0083] WeatherBench is an open-source project developed by the Pangeo Data team, aiming to provide standardized datasets and benchmark testing tools for meteorological and climate science research. This project is based on the reanalysis data of the ERA5 dataset and re-grids it into different resolutions (5.625°, 2.8125°, and 1.40625°). The present invention uses the 5.625° resolution, with a time resolution of 1 hour, including 32 latitudes and 64 longitudes.
[0084] S3. Input the precipitation target data into the trained PFTF precipitation forecasting model to obtain the precipitation forecasting result corresponding to the precipitation target data.
[0085] Among them, the precipitation target data is the data specified by the user, and this precipitation target data is also input into the trained PFTF precipitation forecasting model in the form of a feature map.
[0086] Optionally, in the above technical solution, it further includes:
[0087] Construct a weighted structural loss function by combining the continuous weighted mean square error and the difference divergence regularization term;
[0088] Use the weighted structural loss function as the loss function when training the PFTF precipitation forecasting model.
[0089] Precipitation intensity can change rapidly. Therefore, constructing the loss function by calculating the differences between adjacent time steps can effectively capture these subtle and transient changes, enabling the model to better learn the dynamic characteristics of precipitation changes. In addition, the construction method of the KL divergence also supports the model to generate smooth forecasting results, because maintaining a high similarity between predicted values at similar times can significantly improve the short-term prediction accuracy of the model. On the other hand, by regularizing the differences between predictions, the model can better avoid overfitting during training, which is particularly important for improving the generalization ability to future data. Compared with traditional loss functions (such as the mean square error), conventional losses may not fully reflect the dynamic characteristics of time series data, while the loss function constructed by the following method is more suitable for prediction tasks with time-varying characteristics such as short-term precipitation.
[0090] Based on the L2 loss function, dynamically weight the difference divergence regularization term to constrain the differences between adjacent time steps. The specific calculation is as follows:
[0091] First, use the following formula to calculate the differences between the predicted values and the true values at adjacent time steps:
[0092] gap p = pred[:, 1:] - pred[:, :-1]
[0093] gap t = true[:, 1:] - true[:, :-1]
[0094] where, gap p represents the difference between the predicted values of two adjacent time steps, and gap t represents the difference between the true values of two adjacent time steps. pred[:, 1:] represents the predicted value of the previous time step among two adjacent time steps, pred[:, :-1] represents the predicted value of the next time step among two adjacent time steps, true[:, 1:] represents the true value of the previous time step among two adjacent time steps, and true[:, :-1] represents the true value of the next time step among two adjacent time steps.
[0095] Next, perform temperature scaling and Softmax calculation on the calculated gap p and gap t , and then take the logarithm of them. The expression is as follows:
[0096]
[0097] Among them, τ is the temperature parameter, which is usually used for Softmax output.
[0098] Then, the KL divergence loss is calculated using the following formula to measure the distribution difference between the predicted difference and the true difference. Finally, the mean value of the KL divergence losses of all samples is obtained to get the final loss loss of the difference dispersion regularization term gap :
[0099]
[0100] Among them, M is the number of samples after difference calculation.
[0101] Combining the main loss function L2 and the difference dispersion regularization term, the weighted structure loss function is obtained:
[0102] Loss = L2 + λ·loss gap
[0103] Among them, λ = 0.1 is the regularization coefficient, which is used to control the weights of the main loss term and the regularization term.
[0104] The present invention uses the hourly global precipitation data from January 1, 2014 to December 31, 2018 in WeatherBench. To ensure the effectiveness of the dataset and the stability of model training, the present invention divides the dataset according to years. The data from 2014 to 2016 is used for model training, the data in 2017 is used as the validation set to evaluate the model performance, and the data in 2018 is used as the test set for the final evaluation of the model. In the precipitation forecasting task, the present invention uses the precipitation data of the first 12 time steps as the input to predict the precipitation situation in the next 12 time steps, so as to realize the forecasting of the precipitation in the next 12 hours. Specifically:
[0105] Download the dataset of 5.625dge from WeatherBench, and perform preprocessing to make its input shape become (16, 12, 1, 32, 64), where 16 represents the batch size, 12 represents 12 time steps, 1 represents the channel size, and 32 and 64 represent the longitude and latitude. Since for the tp used as the label, the original data is similar to a sparse matrix, so the first step is to change the unit from m to mm, and the second step is to regularize the data.
[0106] Next, the dataset is divided. In the present invention, the precipitation data of the first 12 time steps is used as input to predict the precipitation in the next 12 time steps, so as to realize the prediction of the precipitation in the next 12 hours. And the dataset is divided according to the year. The data from 2014 to 2016 is used for model training, the data in 2017 is used as the validation set to evaluate the model performance, and the data in 2018 is used as the test set for the final evaluation of the model.
[0107] Suppose a batch of input tensors B ∈ R B×T×C×H×W , and the number of sequences B = |B|. In the encoder and decoder, the present invention reshapes the sequential input data B×T×C×H×W into (B×T)×C×H×W to only consider spatial correlation. In the TripFormer module, the present invention reshapes the feature B×T×C×H×W into B×(T×C)×H×W so that the frames are arranged in the channel dimension. In terms of the loss function, the loss function is constructed by calculating the difference between adjacent time steps and the KL divergence. On the other hand, by regularizing the difference between predictions, the model can better avoid overfitting during training, which is particularly important for improving the generalization ability to future data.
[0108] The PFTF model of the present invention is compared with deep learning-related models in recent years. Specifically, the present invention selects advanced model architectures such as ConvLSTM, PredRNN, SimVP, UniFormer, HorNet, and MogaNet for experimental comparison. The comparison results are shown in Table 1. Through this comparison, the present invention aims to evaluate the performance and efficiency of the proposed method in processing time series data. And it can be seen from the table that the PFTF model of the present invention has a significant improvement in performance compared with other models.
[0109] Table 1:
[0110]
[0111] The meteorological evaluation indicators used in the present invention include RMSE, MSE, and MAE, and the specific calculation formulas are as follows:
[0112]
[0113] Among them, f is the model prediction, and t is the true value in the ERA5 dataset. L(j) is the latitude weight factor exponent at the jth latitude, used to adjust the area difference in different latitude regions.
[0114] Since the final precipitation prediction result is a multi-dimensional array, in terms of evaluating the model performance, the results can be seen from RMSE, MSE, and MAE, as Figure 5 shown.Figure 5 Shows the performance of five deep learning models in the precipitation forecasting task, including ConvLSTM, PredRNN, SimVP, MogaNet, and the newly proposed PFTF precipitation forecasting model of the present invention. The performance of each model is shown through three subgraphs: True precipitation, Predicted precipitation by the model, and Prediction error (|True - Pred|). The vertical axis represents precipitation in millimeters, and the values on the color bar represent the intensity of precipitation, where the darker the color, the greater the precipitation. Traditional recurrent models ConvLSTM and PredRNN have larger prediction errors in areas with higher rainfall. In contrast, non-recurrent models SimVP and MogaNet, as well as the PFTF precipitation forecasting model, show smaller prediction errors in the precipitation forecasting task. The PFTF precipitation forecasting model is closest to the true value in precipitation prediction, especially in areas with higher precipitation, and its prediction accuracy is significantly better than other models. This indicates that the PFTF precipitation forecasting model can effectively integrate various key information when dealing with precipitation prediction problems, thus providing more accurate forecasting results.
[0115] In the above embodiments, although the steps are numbered S1, S2, etc., these are only specific embodiments given by the present invention. Those skilled in the art can adjust the execution order of S1, S2, etc. according to the actual situation, and this is also within the protection scope of the present invention. It can be understood that in some embodiments, it may include some or all of the above embodiments.
[0116] As Figure 6 shown, a precipitation prediction system 200 according to an embodiment of the present invention includes a model construction module 201, a model training module 202, and a precipitation prediction module 203;
[0117] The model construction module 201 is used to: construct a PFTF precipitation forecasting model, where the PFTF precipitation forecasting model includes an encoder, a decoder, and multiple TripFormer modules. The multiple TripFormer modules are sequentially arranged between the encoder and the decoder. The encoder is used to: extract multi-scale features from precipitation sample data and perform enhancement processing. Each TripFormer module is used to: capture the dependencies in the channel dimension and spatial dimension of the features received by each of them through an optimized attention mechanism to obtain spatio-temporal features. The decoder is used to: perform depth convolution operations on the received features at multiple scales to obtain a precipitation forecasting result corresponding to the precipitation sample data, where the features received by the decoder are: the features obtained after processing the spatio-temporal features output by the last TripFormer module;
[0118] The model training module 202 is used to: train the PFTF precipitation forecasting model to obtain a trained PFTF precipitation forecasting model;
[0119] The precipitation prediction module 203 is used to: input precipitation target data into the trained PFTF precipitation forecasting model to obtain a precipitation forecasting result corresponding to the precipitation target data.
[0120] Optionally, in the above technical solution, a weighted structural loss function construction module is further included. The weighted structural loss function construction module is used to: combine the continuous weighted mean square error and the difference divergence regularization term to construct a weighted structural loss function;
[0121] The model training module 202 is used to: use the weighted structural loss function as the loss function when training the PFTF precipitation forecasting model.
[0122] Optionally, in the above technical solution, the encoder includes a first ConvSC layer and a first multi-scale convolution module, which are arranged in sequence. The decoder includes a second ConvSC layer and a second multi-scale convolution module, which are arranged in sequence.
[0123] Optionally, in the above technical solution, the TripFormer module includes: an attention module and a two-layer MLP with a non-linear activation. The attention module includes a channel attention calculation branch and two spatial dimension interaction branches.
[0124] It should be noted that the beneficial effects of the precipitation prediction system 200 provided in the above embodiment are the same as those of the above precipitation prediction method, and will not be elaborated here. In addition, when the system provided in the above embodiment implements its functions, only the above division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the system can be divided into different functional modules according to the actual situation to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiment belong to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be elaborated here.
[0125] Among them, the precipitation prediction system of the present invention can be a computer program (including program code) running in a computer device. For example, the precipitation prediction system of the present invention is an application software, which can be used to execute the corresponding steps in the precipitation prediction method of the present invention.
[0126] In some embodiments, the precipitation prediction system of the present invention can be implemented in a combination of software and hardware. As an example, the precipitation prediction system of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the precipitation prediction method of the present invention. For example, a processor in the form of a hardware decoding processor can employ one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0127] Among them, the modules involved in the embodiments of the present invention can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the module itself in some cases.
[0128] An electronic device according to an embodiment of the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned precipitation prediction method is implemented. That is to say, an electronic device according to an embodiment of the present invention may include, but is not limited to: a processor and a memory; the memory is used to store the computer program; the processor is used to execute the precipitation prediction method shown in any embodiment of the present invention by calling the computer program.
[0129] In an alternative embodiment, an electronic device is provided, as Figure 7 shown, Figure 7 the electronic device 4000 shown includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as connected through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 can be used for data interaction between the electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in actual applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present invention.
[0130] The processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in connection with the disclosure of the present invention. The processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0131] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 7 only a thick line is used to represent the bus 4002 in the figure, but it does not mean that there is only one bus or one type of bus.
[0132] The memory 4003 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0133] The memory 4003 is used to store the application program code (computer program) for implementing the solution of the present invention, and is controlled by the processor 4001 for execution. The processor 4001 is used to execute the application program code stored in the memory 4003 to implement the content shown in the foregoing method embodiments.
[0134] Among them, the electronic device may also be a terminal device, and the terminal device may be any device on which an application can be installed, including at least one of a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a smart TV, and a smart vehicle device.
[0135] It should be noted that Figure 7 the electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.
[0136] A computer-readable storage medium according to an embodiment of the present invention has a computer program stored thereon, and when the computer program is executed by a processor, the above-mentioned precipitation prediction method is implemented.
[0137] Optionally, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0138] In an exemplary embodiment, a computer program product or a computer program is further provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the above-mentioned precipitation prediction method.
[0139] Computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0140] It should be understood that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of methods and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0141] The computer-readable storage medium provided by the embodiments of the present invention may be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0142] The above computer-readable storage medium stores one or more programs, which, when executed by the electronic device, cause the electronic device to perform the method shown in the above embodiments.
[0143] The above description is only a preferred embodiment of the present invention and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present invention is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the present invention.
[0144] It should be noted that the terms "first", "second", etc. in the specification and claims of this application are used to distinguish similar objects, and do not represent a limitation on a specific order or sequence. In appropriate cases, the order of use of similar objects can be interchanged so that the embodiments of this application described here can be implemented in an order other than the illustrated or described order.
[0145] Those skilled in the art know that the present invention can be implemented as a system, method, or computer program product. Therefore, the present invention can be specifically implemented in the following forms: it can be completely hardware, can also be completely software (including firmware, resident software, microcode, etc.), and can also be in the form of a combination of hardware and software, generally referred to as "circuit", "module", or "system" in this article. In addition, in some embodiments, the present invention can also be implemented in the form of a computer program product in one or more computer-readable media, which contains computer-readable program code.
[0146] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A precipitation prediction method, characterized in that: include: A PFTF precipitation forecast model is constructed. The PFTF precipitation forecast model includes an encoder, a decoder and a plurality of TripFormer modules. The plurality of TripFormer modules are sequentially arranged between the encoder and the decoder. The encoder is used to extract multi-scale features from precipitation sample data and perform enhancement processing. Each TripFormer module is used to capture the dependency of channel dimensions and spatial dimensions in the respective received features by optimizing the attention mechanism to obtain spatiotemporal features. The decoder is used to perform deep convolution operations on the received features at multiple scales to obtain precipitation forecast results corresponding to the precipitation sample data. The features received by the decoder are features obtained after processing the spatiotemporal features output by the last TripFormer module. Training the PFTF precipitation forecast model to obtain a trained PFTF precipitation forecast model; The precipitation target data is input into the trained PFTF precipitation forecast model to obtain the precipitation forecast result corresponding to the precipitation target data.
2. A precipitation prediction method according to claim 1, characterized in that: Also includes: Combining the continuous weighted mean square error and the difference dispersion regularization term, a weighted structural loss function is constructed; The weighted structural loss function is used as the loss function used when training the PFTF precipitation forecast model.
3. A precipitation prediction method according to claim 1 or 2, characterized in that: The encoder includes a first ConvSC layer and a first multi-scale convolution module, which are arranged in sequence. The decoder includes a second ConvSC layer and a second multi-scale convolution module, which are arranged in sequence.
4. A precipitation prediction method according to claim 1 or 2, characterized in that: The TripFormer module includes a first normalization layer, an attention module, a second normalization layer and an MLP with nonlinear activation, which are arranged in sequence. The attention module includes a channel attention calculation branch and two spatial dimension interaction branches.
5. A precipitation prediction system, characterized in that: It includes model building module, model training module and precipitation prediction module; The model building module is used to: build a PFTF precipitation forecast model, the PFTF precipitation forecast model includes an encoder, a decoder and multiple TripFormer modules, the multiple TripFormer modules are sequentially arranged between the encoder and the decoder, the encoder is used to: extract multi-scale features from precipitation sample data and perform enhancement processing, each TripFormer module is used to: optimize the attention mechanism to capture the dependency of channel dimensions and spatial dimensions in the respective received features to obtain spatiotemporal features, the decoder is used to: perform deep convolution operations on the received features at multiple scales to obtain precipitation forecast results corresponding to the precipitation sample data, wherein the features received by the decoder are: features obtained after processing the spatiotemporal features output by the last TripFormer module; The model training module is used to: train the PFTF precipitation forecast model to obtain a trained PFTF precipitation forecast model; The precipitation prediction module is used to: input precipitation target data into the trained PFTF precipitation forecast model to obtain precipitation forecast results corresponding to the precipitation target data.
6. A precipitation prediction system according to claim 5, characterized in that: Also included is a weighted structural loss function construction module, wherein the weighted structural loss function construction module is used to: construct a weighted structural loss function by combining a continuous weighted mean square error and a difference dispersion regularization term; The model training module is used to: use the weighted structural loss function as the loss function used when training the PFTF precipitation forecast model.
7. A precipitation prediction system according to claim 5 or 6, characterized in that: The encoder includes a first ConvSC layer and a first multi-scale convolution module, which are arranged in sequence. The decoder includes a second ConvSC layer and a second multi-scale convolution module, which are arranged in sequence.
8. A precipitation prediction system according to claim 5 or 6, characterized in that: The TripFormer module includes a first normalization layer, an attention module, a second normalization layer and an MLP with nonlinear activation, which are arranged in sequence. The attention module includes a channel attention calculation branch and two spatial dimension interaction branches.
9. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements a precipitation prediction method according to any one of claims 1 to 4 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the precipitation prediction method according to any one of claims 1 to 4 is implemented.