Rainfall prediction method based on improved large-kernel self-convolution neural network

Through the large-core autoconvolution neural network combined with multi-core fusion and channel attention mechanism, the problems of global information loss and insufficient temporal feature modeling in the existing methods are solved, and more accurate precipitation prediction is achieved, especially in complex space-time tasks.

CN120294877APending Publication Date: 2025-07-11CHENGDU UNIV OF INFORMATION TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510431922.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing deep learning methods for precipitation prediction have problems such as global information loss and insufficient temporal feature modeling capabilities in spatiotemporal modeling, especially the methods based on convolutional neural networks are difficult to capture global spatial information and key temporal features simultaneously.

Method used

The large-core autoconvolution neural network is adopted, combined with the large-core autoconvolution module, multi-core fusion module and channel attention mechanism, and the spatial feature capture capability is enhanced through large-core convolution. The multi-core fusion module combines different convolution kernel feature maps, and the channel attention mechanism CAM further fuses the characteristics to ensure the integrity of global information.

Benefits of technology

Effectively capture global information, improves the accuracy and reliability of precipitation prediction, especially in complex climate conditions, showing significant performance advantages, which is significantly better than the current mainstream spatio-temporal modeling methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120294877A_ABST
    Figure CN120294877A_ABST
Patent Text Reader

Abstract

The invention discloses a rainfall prediction method based on an improved large-kernel self-convolution neural network, and the method comprises the steps: S1, carrying out the preprocessing of a rainfall data set, and generating a training set and a test set; s2, inputting the training set into a large kernel self-convolution neural network to obtain a preliminarily trained large kernel self-convolution neural network; s3, judging whether the current training round number is greater than a training round number threshold value or not, if not, adding 1 to the current training round number and entering S4, and if yes, generating a trained large-kernel self-convolution neural network and entering S5; s4, inputting the test set into the preliminarily trained large-kernel self-convolution neural network, judging whether the mean square error index of the obtained rainfall prediction information and the real rainfall information is higher than the best mean square error index or not, if yes, storing the current model parameters, updating the best mean square error index, and returning to S2; if not, returning to S2; and S5, inputting the test set into the trained large-kernel self-convolution neural network to obtain a final rainfall prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of precipitation prediction, and particularly relates to a precipitation prediction method based on an improved large kernel self-convolution neural network. Background Art

[0002] Precipitation prediction is a typical spatio-temporal modeling task and has a wide range of applications in the meteorological service community. In reality, precipitation prediction plays an important role in fields such as agriculture, resource management, and disaster warning. Accurate precipitation prediction can effectively avoid disaster risks and reduce the consumption of human and material resources. At present, precipitation prediction mainly includes physics-based numerical prediction methods and data-driven deep learning methods. However, since numerical prediction methods are difficult to effectively reveal the spatio-temporal correlation mechanism of precipitation, data-driven methods provide a new modeling paradigm for precipitation forecasting. Especially in recent years, deep learning methods have shown increasingly powerful data fitting capabilities in visual and natural language processing tasks, and deep learning methods for precipitation forecasting have received more and more attention.

[0003] Existing deep learning methods for precipitation prediction can be roughly divided into convolution-based, recurrent-based, and Transformer-based methods. Among them, convolution-based methods focus on spatial feature learning. The classic neural network architecture is UNet, which consists of encoder and decoder components for downsampling and upsampling respectively. In practice, the encoder-decoder paradigm of UNet-based models is prone to information loss problems, especially downsampling in the encoder. In addition, for precipitation prediction, convolution-based methods mainly focus on spatial feature learning and may have difficulty effectively considering the learning of temporal representation. Recurrent-based methods connect historical information through recurrent units, providing a classic neural network for temporal representation learning. A general setting for precipitation prediction can combine recurrent-based methods with convolution-based methods to pursue better spatio-temporal learning. For example, ConvLSTM extends LSTM by adding convolution in the input-to-state and state-to-state transitions. PredRNNv2 extends the inner transition function of the LSTM memory state to introduce spatio-temporal memory flow. For this paradigm, the combination of recurrent units and convolution can reasonably be considered spatio-temporal correlation modeling. However, like convolution-based methods, this combination cannot overcome their limitations because using small kernels in stacked layers limits the receptive field to local details, thus restricting their ability to capture global information. Recently, Transformer-based methods have provided a new spatio-temporal modeling paradigm, where self-attention is used for temporal learning and patch embedding is used for spatial learning. For example, FourCastNet, Informer, and Rainformer adopt Transformer-based architectures for precipitation prediction. Theoretically, patch embedding is similar to convolution downsampling, where partitioning the input disrupts the original spatial structure, leading to potential information loss.

[0004] Deficiencies of the prior art:

[0005] 1. Deep learning methods based on CNN:

[0006] Deep learning methods based on convolutional neural networks (CNNs) have demonstrated powerful spatial feature extraction capabilities in spatio-temporal modeling tasks, such as precipitation prediction. However, these methods have some inherent problems. First, the local receptive field of CNNs limits their ability to capture global spatial information. Although this problem can be partially solved by expanding the convolutional kernel (such as RepLKNet and UniRepLKNet), the computational burden increases, and the optimization difficulty also rises. In addition, CNNs are better at capturing spatial features, while their ability to model temporal features is relatively weak, which is particularly significant in spatio-temporal modeling tasks because temporal dynamics are crucial for accurate prediction. Traditional encoder-decoder structures such as UNet also often result in information loss due to downsampling operations, further weakening the retention of global information and the modeling of fine-grained spatio-temporal relationships. Therefore, how to improve the modeling and optimization efficiency of temporal features while maintaining the spatial feature extraction ability has become a key challenge for CNN-based methods in spatio-temporal prediction tasks.

[0007] 2. Deep learning methods based on RNNs:

[0008] Recurrent-based methods learn temporal representations through recurrent units and can effectively connect historical information. Classic architectures such as RNNs, LSTMs, and GRUs have been widely applied to time series modeling tasks. These methods focus on the learning of temporal features and are particularly suitable for processing time series data with long-term dependencies. However, although recurrent-based methods have advantages in temporal modeling, they usually have limitations in capturing spatial features. To address this, many studies have attempted to combine recurrent and convolutional-based methods for better spatio-temporal modeling. For example, ConvLSTM achieves joint modeling of space and time by introducing convolutional operations in the state transition. Further developments such as convolutional tensor train LSTM with low-rank factorization and PredRNNv2 have improved the learning ability of spatio-temporal representations. However, these methods often rely on stacking multiple small convolutional kernel layers, resulting in their receptive fields being limited to capturing fine-grained information and unable to fully model global spatial information. This limitation may affect the overall performance of the model in complex spatio-temporal prediction tasks (such as precipitation forecasting).

[0009] 3. Deep learning methods based on Transformers:

[0010] Transformer is a classic neural network architecture that was initially widely used in the field of natural language processing and has demonstrated excellent capabilities in time series data modeling with its self-attention mechanism. By performing grid processing on spatial data, Transformer has gradually been applied to spatio-temporal modeling tasks. For example, Vision Transformer (ViT) divides a two-dimensional image into a series of flattened patches and then embeds them into a one-dimensional latent space for processing, demonstrating its potential in spatio-temporal representation learning. Swin Transformer introduces a shifted window mechanism to improve efficiency and performs well in spatial representation. FourCastNet improves spatio-temporal modeling by combining a shared MLP and a frequency soft threshold strategy. However, the patch embedding operation used by these methods in the spatial dimension is similar to convolution. Although it effectively simplifies the processing, it somewhat destroys the original spatial structure and may lead to the problem of global information loss. This loss of global information may affect the performance of the model for complex spatio-temporal tasks such as precipitation forecasting. Summary of the Invention

[0011] Aiming at the above deficiencies in the prior art, a precipitation prediction method based on an improved large-kernel self-convolution neural network provided by the present invention constructs a network structure specifically for precipitation prediction tasks. This precipitation prediction network uses a large-kernel self-convolution module with the original feature path as the convolution kernel, which can effectively capture global information and solve the problem of global information loss in existing spatio-temporal modeling methods.

[0012] To achieve the above invention purpose, the technical solution adopted by the present invention is: a precipitation prediction method based on an improved large-kernel self-convolution neural network, including the following steps:

[0013] S1. Obtain a precipitation data set, preprocess the precipitation data set to generate a training set and a test set, and initialize the number of training epochs;

[0014] S2. Input the training set into the large-kernel self-convolution neural network, train the large-kernel self-convolution neural network according to the prediction time step to obtain a preliminarily trained large-kernel self-convolution neural network;

[0015] S3. Determine whether the current number of training epochs is greater than the training epoch threshold. If not, add 1 to the current number of training epochs and enter S4. If so, generate a trained large-kernel self-convolution neural network according to the saved model parameters and enter S5;

[0016] S4. Input the test set into the initially trained large-kernel self-convolution neural network, and determine whether the mean square error index between the obtained precipitation prediction information and the true precipitation information is higher than the best mean square error index. If so, save the current model parameters, update the best mean square error index, replace the current training set, and return to S2; if not, replace the current training set and return to S2;

[0017] S5. Input the test set into the trained large-kernel self-convolution neural network to obtain the final precipitation prediction result.

[0018] Furthermore: In S1, the method for preprocessing the precipitation data set includes the following sub-steps:

[0019] S11. Crop the data in the precipitation data set to or size;

[0020] S12. Normalize the cropped data, and divide the normalized data according to a preset convention ratio to generate a training set and a test set.

[0021] Furthermore: In S2, the large-kernel self-convolution neural network includes a large-kernel self-convolution module, a multi-kernel fusion module, a channel attention mechanism CAM, an encoder structure, and a decoder structure connected in sequence;

[0022] The large-kernel self-convolution module includes a normalization layer, the first convolution of and the second convolution of the first convolution of and the second convolution of

[0023] The multi-kernel fusion module includes the third convolution of

[0024] and the dilated convolution with a kernel size half of that of the large-kernel self-convolution kernel; The channel attention mechanism CAM includes an average pooling layer, a maximum pooling layer, the third convolution of the third convolution of

[0025] the Relu activation function layer, and the Sigmoid function layer. Among them, both the average pooling layer and the maximum pooling layer are connected to

[0026] Among them, the large kernel self-convolution module is used to effectively enhance the network's ability to capture spatial features, the multi-kernel fusion module is used to combine feature maps from different convolutional kernels, and the channel attention mechanism CAM is used to further fuse features.

[0027] The beneficial effects of the above further scheme are as follows: The large kernel self-convolution module is used to effectively enhance the network's ability to capture spatial features, thereby ensuring the accurate extraction of target features. The multi-kernel fusion module is used to combine feature maps from different convolutional kernels, improving the network's ability to model complex spatio-temporal relationships. The channel attention mechanism CAM is used to further fuse features, ensuring the integrity of global information and reducing information loss. The overall architecture aims to improve the accuracy and reliability of precipitation prediction, especially in applications under complex climate conditions. Through these innovations, the proposed method demonstrates significant performance advantages in various precipitation prediction tasks.

[0028] Furthermore: In S2, the method for training the large kernel self-convolution neural network includes the following sub-steps:

[0029] S21. Input the training set into the large kernel self-convolution module to obtain the large kernel self-convolution kernel parameters;

[0030] S22. Input the large kernel self-convolution kernel parameters into the multi-kernel fusion module to obtain the first feature map after multi-kernel fusion operation;

[0031] S23. Input the first feature map into the channel attention mechanism CAM to obtain the channel attention weights;

[0032] S24. Perform downsampling operation on the channel attention weights, input the result after downsampling into the encoder to obtain the encoded features, and input the encoded features into the decoder to obtain the predicted precipitation map, completing the training of the large kernel self-convolution neural network.

[0033] Furthermore: S21 includes the following sub-steps:

[0034] S211. Use the data in the training set as the input data, perform mean pooling operation on the input data according to one-fourth of the input size through the normalization layer to obtain the global feature information;

[0035] S212. Split the global feature information by channel, and respectively pass the split results through the first convolution of and the second convolution of

[0036] to perform convolution operations to obtain the first feature map and the second feature map;

[0037] The beneficial effects of the above further solution are as follows: Mean pooling helps extract the overall statistical features of the image and ensure the integrity of global information. To alleviate the discontinuity in direct convolution operations, the present invention performs a large-kernel self-convolution operation on the global feature information to obtain large-kernel self-convolution parameters.

[0038] Further: The S22 includes the following sub-steps:

[0039] S221. Input the large-kernel self-convolution parameters into the third convolution and the dilated convolution with a kernel size half of the large-kernel self-convolution kernel size to obtain the convolution kernel parameters of the third convolution and the dilated convolution;

[0040] S222. Fuse the large-kernel self-convolution parameters with the convolution kernel parameters of the third convolution and the dilated convolution to obtain the first feature map after the multi-kernel fusion operation.

[0041] The beneficial effects of the above further solution are as follows: Design a multi-kernel fusion strategy to integrate convolution kernels with different attribute settings, and make full use of the feature information extracted by various convolution operations. Specifically, select different dilation rates and kernel sizes to meet the extraction requirements of multi-scale features, thereby improving the model's recognition ability for precipitation phenomena of different sizes.

[0042] Further: The S23 includes the following sub-steps:

[0043] S231. Input the first feature map into the average pooling layer and the maximum pooling layer respectively to obtain the second feature map and the third feature map;

[0044] S232. Pass the second feature map and the third feature map sequentially through the third convolution and the Relu activation function layer to obtain the first feature weight map and the second feature weight map;

[0045] S233. Perform weighted summation on the first feature weight map and the second feature weight map, and input the summation result into the Sigmoid function layer to obtain the channel attention weight.

[0046] The beneficial effects of the above further solution are as follows: To enhance the representation ability of temporal features, the present invention introduces a channel attention mechanism to extract key temporal features, thereby further improving the accuracy of spatio-temporal representation.

[0047] The beneficial effects of the present invention are as follows:

[0048] (1) The present invention provides a precipitation prediction method based on an improved large-kernel self-convolution neural network. By adopting the precipitation prediction network architecture of the improved large-kernel self-convolution neural network, it can effectively capture global information and achieve accurate modeling of complex spatio-temporal relationships. This network is particularly suitable for tasks that require maintaining the integrity of global information in precipitation forecasting scenarios and can achieve more accurate predictions on time and space scales.

[0049] (2) Aiming at the problem that traditional convolution methods are difficult to simultaneously process global information and key time features, the present invention designs a module combining self-convolution and channel attention mechanism. The large-kernel self-convolution module enhances the network's ability to capture spatial features through large-kernel convolution, while the channel attention mechanism CAM effectively extracts key time information, improving the network's spatio-temporal modeling ability.

[0050] (3) The present invention proves through a large number of experiments that the test results on multiple precipitation data sets show that the precipitation prediction network based on the large-kernel self-convolution strategy performs excellently in maintaining global information and the accuracy of spatio-temporal representation, significantly outperforming current mainstream spatio-temporal modeling methods, especially having better prediction performance in complex precipitation environments. Brief Description of the Drawings

[0051] Figure 1 is a flowchart of a precipitation prediction method based on an improved large-kernel self-convolution neural network of the present invention;

[0052] Figure 2 is a schematic diagram of the structure of the large-kernel self-convolution neural network of the present invention;

[0053] Figure 3 is a schematic diagram of the generation process of large-kernel self-convolution parameters;

[0054] Figure 4 is a comparison chart of the results of large-kernel self-convolution and ordinary self-convolution;

[0055] Figure 5 is a schematic diagram of the structure of the multi-core fusion strategy;

[0056] Figure 6 is the visualization result chart of the present invention using the past t time steps to predict time steps and time steps on the three data sets of CHIRPS, ERA5, and WeatherBench;

[0057] Figure 7 is a schematic diagram of the results of the time series experiment of the present invention. Detailed Embodiments

[0058] The specific embodiments of the present invention will be described below to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

[0059] As Figure 1 shown, in an embodiment of the present invention, a precipitation prediction method based on an improved large-kernel self-convolutional neural network includes the following steps:

[0060] S1. Obtain a precipitation data set, preprocess the precipitation data set to generate a training set and a test set, and initialize the number of training rounds;

[0061] S2. Input the training set into the large-kernel self-convolutional neural network, train the large-kernel self-convolutional neural network according to the prediction time step to obtain a preliminarily trained large-kernel self-convolutional neural network;

[0062] S3. Determine whether the current number of training rounds is greater than the training round threshold. If not, add 1 to the current number of training rounds and enter S4. If so, generate a trained large-kernel self-convolutional neural network according to the saved model parameters and enter S5;

[0063] S4. Input the test set into the preliminarily trained large-kernel self-convolutional neural network, and determine whether the mean square error index between the obtained precipitation prediction information and the true precipitation information is higher than the best mean square error index. If so, save the current model parameters, update the best mean square error index, replace the current training set and return to S2; if not, replace the current training set and return to S2;

[0064] S5. Input the test set into the trained large-kernel self-convolutional neural network to obtain the final precipitation prediction result.

[0065] In this embodiment, the precipitation data set can be selected from CHIRPS, ERA5, and WeatherBench meteorological data sets, and these data sets provide precipitation data at different time and space scales. Among them, CHIRPS and ERA5 are precipitation data, and WeatherBench covers 40 years of meteorological data from 1979 to 2018. It includes variables such as surface air pressure, 2-meter temperature, 10-meter wind speed, and total precipitation.

[0066] In the above S1, the method for preprocessing the precipitation data set includes the following sub-steps:

[0067] S11. Crop the data in the precipitation data set to or size;

[0068] In this embodiment, the specific situation of meteorological data is viewed through the PanoplyWin software, and the required precipitation variables are selected. The data is trimmed to or size, and the total data resolution is , where 20 is the number of input step numbers;

[0069] S12. Normalize the trimmed data, and divide the normalized data according to a preset convention ratio to generate a training set and a test set.

[0070] In this embodiment, the normalization process can detect the validity of the data, determine whether there are extreme values, and remove them if any.

[0071] As Figure 2 shown, in this embodiment, the input data of the large kernel self-convolution neural network proposed by the present invention has t time steps, and outputs predictions for the future time steps. Among them, the resolution of the original image is , 20 represents the number of channels, and 64 represents the height and width of the image respectively.

[0072] In S2, the large kernel self-convolution neural network includes a large kernel self-convolution module, a multi-kernel fusion module, a channel attention mechanism CAM, an encoder structure, and a decoder structure connected in sequence;

[0073] The large kernel self-convolution module includes a normalization layer, the first convolution of , and the second convolution of . The normalization layer is respectively connected to the first convolution of and the second convolution of

[0074] The multi-kernel fusion module includes the third convolution of

[0075] and a dilated convolution with a kernel size half of that of the large kernel self-convolution; The channel attention mechanism CAM includes an average pooling layer, a maximum pooling layer, the third convolution of , a Relu activation function layer, and a Sigmoid function layer. Among them, both the average pooling layer and the maximum pooling layer are connected to the third convolution of

[0076] The third convolution of

[0077] The encoder structure adopts large-kernel self-convolution operation and incorporates multi-core fusion and channel attention mechanism. ConvBlock1 and ConvBlock2 are replaceable basic convolution modules. ConvBlock1 uses standard convolution operation and convolution kernels, which are suitable for extracting local features and effectively capturing image details. ConvBlock2 combines Dropout, LayerNorm, and LeakyRelu, and uses convolution kernels, enhancing the learning ability for complex features. This design solves the problem of neuron "death" through LeakyRelu, improving the performance of the model in complex image structures. The overall architecture is flexible, and the decoder part includes operations such as concat, dilated convolution, and transposed convolution for feature recovery and fusion.

[0078] The core improvements and innovations of the large-kernel self-convolution neural network designed in the present invention lie in the large-kernel self-convolution module, the multi-core fusion module based on the multi-core fusion strategy, and the channel attention mechanism CAM. Among them, the large-kernel self-convolution module is used to effectively enhance the network's ability to capture spatial features, thereby ensuring accurate extraction of target features. The multi-core fusion module is used to combine feature maps from different convolution kernels, improving the network's ability to model complex spatio-temporal relationships. The channel attention mechanism CAM is used to further fuse features, ensuring the integrity of global information and reducing information loss. The overall architecture aims to improve the accuracy and reliability of precipitation prediction, especially in applications under complex climate conditions. Through these innovations, the proposed method has shown significant performance advantages in various precipitation prediction tasks.

[0079] In step S2, the method for training the large-kernel self-convolution neural network includes the following sub-steps:

[0080] S21: Input the training set into the large-kernel self-convolution module to obtain large-kernel self-convolution kernel parameters;

[0081] S22: Input the large-kernel self-convolution kernel parameters into the multi-core fusion module to obtain the first feature map after multi-core fusion operation;

[0082] S23: Input the first feature map into the channel attention mechanism CAM to obtain channel attention weights;

[0083] S24: Perform downsampling operation on the channel attention weights, input the result after downsampling operation into the encoder to obtain encoded features, and input the encoded features into the decoder to obtain the predicted precipitation map, completing the training of the large-kernel self-convolution neural network.

[0084] The step S21 includes the following sub-steps:

[0085] S211. Use the data in the training set as input data, and perform average pooling operation on the input data according to one-fourth of the input size through the normalization layer to obtain global feature information;

[0086] S212. Split the global feature information by channel, and respectively pass the split results through the first convolution of and the second convolution of

[0087] to perform convolution operations, obtaining a first feature map and a second feature map;

[0088] In this embodiment, average pooling helps to extract the overall statistical features of the image and ensure the integrity of global information. The size of pooling is passed in the form of hyperparameters. Since the encoding and decoding operations will cause the image size to change continuously, the kernel size corresponding to the large kernel self-convolution will change continuously, mainly taking one-fourth of the input size at this time. In addition, in order to alleviate the discontinuity in the direct convolution operation, the present invention performs a large kernel self-convolution operation on the global feature information to obtain large kernel self-convolution parameters.

[0089] In this embodiment, the generation process of the large kernel self-convolution parameters is as Figure 3 shown, performing operations such as channel splitting, convolution, and residual connection on the global feature information. Specifically, the global feature information is divided into two parts in the channel dimension, and convolution is used to further abstract the self-convolution kernel representation to enhance its ability.

[0090] Figure 4 A comparison example between large kernel self-convolution and ordinary self-convolution is given. In this embodiment, the inputs of two time steps are selected for comparison. The bar chart in the figure shows the mean square error comparison between the outputs of the two methods at each time step and the input. The difference between this method and the standard self-convolution lies in performing channel splitting and convolution operations. Since the pooling operation is performed in the local area of the data, directly performing self-convolution on its result will cause discontinuity in the convolution result. The error results in the figure show that the large kernel self-convolution effectively alleviates this discontinuity through channel splitting and increasing convolution operations. Compared with the standard self-convolution, the method adopted by the present invention pays more attention to the retention of global information. And the present invention reduces the error to one-third of the original.

[0091] The S22 includes the following sub-steps:

[0092] S221. Input the large kernel self-convolution parameters into the third convolution of

[0093] S222. Fuse the large kernel self-convolution parameters with the convolution kernel parameters of the third convolution and the dilated convolution to obtain the first feature map after the multi-kernel fusion operation.

[0094] In this embodiment, the multi-kernel fusion module performs a multi-kernel fusion operation on the large kernel self-convolution parameters. During this process, the convolution kernel parameters of different convolutions are retained and parameter sharing is performed, that is, at an appropriate ratio, the saved kernel parameters are fused into the large kernel self-convolution kernel at intervals. The fusion strategy is mainly centered to ensure the consistency of global features and at the same time enhance the ability to capture information at different scales.

[0095] Design a multi-kernel fusion strategy to integrate convolution kernels with different attribute settings and make full use of the feature information extracted by various convolution operations. Specifically, select different dilation rates and kernel sizes to meet the extraction requirements of multi-scale features, thereby improving the model's recognition ability for precipitation phenomena of different sizes.

[0096] Perform a multi-kernel fusion operation on the input data. Specifically, the input passes through the third convolution and the dilated convolution with a kernel size half of that of the large kernel self-convolution kernel, and the kernel parameters of each convolution operation are retained. Perform a parameter sharing operation on the feature maps obtained by different convolution operations to generate a first feature map, which combines information at multiple scales and can better capture the spatial variation characteristics of precipitation.

[0097] During the generation of the fused feature map, reduce the model complexity through an appropriate parameter sharing strategy and improve the calculation efficiency. At the same time, ensure that the convolution kernel parameters are continuously optimized through backpropagation during the training process to improve the generalization ability and prediction accuracy of the model.

[0098] During the generation of the fused feature map, reduce the model complexity through an appropriate parameter sharing strategy and improve the calculation efficiency. At the same time, ensure that the convolution kernel parameters are continuously optimized through backpropagation during the training process to improve the generalization ability and prediction accuracy of the model.

[0099] As Figure 5 shown, in this embodiment, the present invention gives an example of multi-kernel fusion of a small dilated convolution kernel (such as ), a large kernel (such as ), and a normal convolution kernel (such as ). The motivation for multi-kernel fusion is similar to the structural reparameterization technique, which is a method of equivalently transforming the model structure by transforming parameters. Here, the present invention defines the fusion method as parameter sharing, that is, parameter substitution is performed between kernels. The figure shows the combination of and The convolution kernel parameters are successively substituted. The small kernel parameters are used to replace the parameters distributed in the center of the large kernel. Secondly, to achieve the extraction form of multiple dilated convolutions, the present invention replaces them in a form with one interval at a time. Finally, the two fused parameters are fused into the self-convolution of the large kernel to achieve the final multi-kernel fusion. This method allows the receptive field to be expanded without increasing the number of parameters, thereby enabling the capture of a larger range of context information and enhancing the ability of the large kernel self-convolution to capture global information details. In addition, by enriching the large-scale convolution kernel, information loss can be further reduced.

[0100] The S23 includes the following sub-steps:

[0101] S231: Input the first feature map into the average pooling layer and the maximum pooling layer respectively to obtain the second feature map and the third feature map;

[0102] S232: Pass the second feature map and the third feature map successively through the third convolution and the Relu activation function layer to obtain the first feature weight map and the second feature weight map;

[0103] S233: Perform weighted summation on the first feature weight map and the second feature weight map, and input the summation result into the Sigmoid function layer to obtain the channel attention weight.

[0104] In this embodiment, the channel attention mechanism CAM effectively divides the importance ratio of different channels through this adaptive weight allocation, ensuring the model's attention to key temporal information; in this way, the network can more effectively capture the information crucial for precipitation prediction in the time dimension, thereby enhancing the time dependence and overall performance of the model.

[0105] In the S24, downsampling (pooling) operation is performed on the channel attention weight, and this whole is used as the encoding operation. The entire network has four encoders and four decoders, and the precipitation is predicted through an effective encoder and decoder structure. In the encoder, four feature extraction operations are performed; in the decoder, upsampling operation is performed through transposed convolution to restore the information. To enhance the information flow, residual connections are adopted between the encoder and the decoder to promote the effective interaction of information. This design allows the model to retain important features during the information transmission process and alleviates the potential problem of gradient disappearance. In addition, through this structure, the network can more flexibly handle the precipitation prediction task and adapt to diverse input patterns. Overall, the collaborative work of the encoder and the decoder not only improves the efficiency of feature extraction but also enhances the model's adaptability to complex precipitation patterns, thereby achieving more accurate precipitation prediction.

[0106] In this embodiment, the methods of S21 - S23 focus on reducing information loss during the downsampling process. By using various information retention strategies such as average pooling and large - kernel self - convolution, information loss is effectively alleviated while global information is also taken into account. This design ensures that the encoder and decoder can continuously extract effective information from the image, thus achieving more accurate precipitation forecasts. Through residual connections, the network can learn deeper - level features and accurately restore important information during the de - convolution process, further improving the accuracy and reliability of the prediction. The design of the overall structure enables the model to maintain high robustness and generalization ability when facing complex precipitation patterns.

[0107] At the output stage of the large - kernel self - convolution neural network, after being processed by the network, precipitation prediction information at different time steps is obtained. To generate the predicted precipitation map, the network uses convolution operations to map the feature map to the required output dimension. This design not only simplifies the calculation process but also ensures the output consistency of the model at each time step. Through this stage, the model can accurately predict the precipitation amount at a specific time point, providing reliable data support for practical applications.

[0108] In this embodiment, the training round threshold of the present invention is set to 200.

[0109] In S4, the test set is input into the trained large - kernel self - convolution network for testing. The mean square error (MSE) index between the obtained precipitation prediction information and the real precipitation information is judged. This index is specifically the mean square error loss, and this loss function provides a feedback signal for the model, guiding it to optimize parameters during the training process. Through the backpropagation mechanism, the model can continuously adjust its internal weights to reduce the prediction error. In addition, since the large - kernel self - convolution kernel is directly obtained based on the input data rather than through training optimization, this feature enables the model to quickly adapt to different precipitation patterns, improving the prediction speed and effectively enhancing the ability to capture complex precipitation phenomena. This efficient design makes the model more adaptable and robust in practical applications;

[0110] In S5, the trained large - kernel self - convolution neural network is the network with the best performance during the test process. The test set is input into this network, and the root mean square error (RMSE), mean absolute deviation (MAD), mean absolute error (MAE), and relative absolute error (RAE) are calculated, and the final detection effect diagram is saved.

[0111] RMSE emphasizes larger errors and is suitable for scenarios that are sensitive to extreme prediction failures. On the other hand, MAD and MAE provide an intuitive understanding of the errors and reflect the stability of the overall deviation of the model. As a standardized metric, RAE facilitates comparison with other models and helps evaluate the relative accuracy of predictions. By combining these metrics, a deeper understanding of the model's performance can be gained from different perspectives, which is helpful for identifying potential improvement directions.

[0112] To further illustrate the effectiveness of the method of the present invention, the method of the present invention is compared with other existing methods. For a fair comparison, the official released codes of other methods are used and their experimental settings are followed, where all methods are implemented in the same computing environment and quantitative and qualitative analyses are carried out simultaneously. The 10 methods are specifically as follows: Method 1 is the UNet method, which was originally developed for biomedical image segmentation but has also shown great potential in the field of precipitation forecasting in recent years. It models spatio-temporal correction through the classical encoder-decoder neural network paradigm; Method 2 is the FourCastNet method, which is a Transformer-based time series prediction method that improves spatio-temporal representation by designing an adaptive Fourier neural operator with a shared MLP and a frequency soft threshold strategy; Method 3 is TAU, which decomposes spatio-temporal attention into two parts: intra-frame static attention and inter-frame dynamic attention. This method improves the existing framework, learns temporal and spatial dependencies separately, and proposes a parallelizable time series attention unit in the time dimension to achieve efficient video prediction; Method 4 is SimVP, which learns spatial features through an encoder-decoder architecture and further considers introducing a translator to introduce temporal evolution; Method 5 is SmaAt-UNet, which adds a convolutional block attention module to the original UNet, converts conventional convolutional operations into depthwise separable convolutions, reduces the number of parameters, and achieves a certain lightweight effect. Method 6 is WAST, which combines wavelet transform and self-attention mechanism for spatio-temporal prediction. Method 7 is ConvLSTM, which extends LSTM by adding convolutions in the input-to-state and state-to-state transitions. Method 8 is GRU, which uses a control mechanism to manage information flow and long-term dependencies. Method 9 is PredRNNv2, which introduces spatio-temporal memory flow in ConvLSTM by extending the inner transition function of the memory state.

[0113] Since the present invention is designed for precipitation forecasting in the meteorological field, tests are carried out on three datasets, namely CHIRPS, ERA5, and WeatherBench. Tables 1, 2, and 3 respectively give the quantitative comparison results of the root mean square error (RMSE), mean absolute deviation (MAD), mean absolute error (MAE), and relative absolute error (RAE) metrics for the future two prediction steps of 10 different network structures on these three datasets.

[0114] Table 1 Performance Metrics on the CHIRPS Dataset, Bold Indicates the Best

[0115] Adopt the method RMSE MAE MAD RAE Method 1 8.619 5.024 2.512 0.842 Method 2 8.709 5.498 2.748 0.915 Method 3 8.600 5.202 2.601 0.900 Method 4 8.684 5.210 2.605 0.938 Method 5 8.824 5.171 2.585 0.922 Method 6 8.628 5.310 2.655 0.883 Method 7 10.500 5.208 2.603 7.372 Method 8 8.758 5.509 2.754 0.936 Method 9 8.634 5.206 2.603 0.899 The method of the present invention 8.540 5.132 2.566 0.841

[0116] Table 2 Performance Metrics on the ERA5 Dataset, Bold Indicates the Best

[0117] Adopt the method RMSE MAE MAD RAE Method 1 0.574 0.489 0.244 0.330 Method 2 0.658 0.564 0.281 0.410 Method 3 0.582 0.495 0.247 0.334 Method 4 0.577 0.493 0.246 0.335 Method 5 0.602 0.522 0.261 0.355 Method 6 0.584 0.507 0.253 0.388 Method 7 0.997 0.805 0.402 0.779 Method 8 0.618 0.531 0.265 0.365 Method 9 0.582 0.501 0.250 0.364 The method of the present invention 0.574 0.492 0.245 0.327

[0118] Table 3 Performance Metrics on the WeatherBench Dataset, Bold Indicates the Best

[0119] Adopt the method RMSE MAE MAD RAE Method 1 0.889 0.432 0.217 0.408 Method 2 0.949 0.501 0.249 0.485 Method 3 0.910 0.445 0.222 0.426 Method 4 0.916 0.454 0.213 0.454 Method 5 0.940 0.465 0.225 0.446 Method 6 0.903 0.461 0.230 0.555 Method 7 1.423 0.609 0.305 1.095 Method 8 0.913 0.445 0.22 0.431 Method 9 0.903 0.442 0.223 0.446 The method of the present invention 0.880 0.420 0.209 0.402

[0120] Comparison with CNN-based Methods: As shown in Table 2, among the convolutional baselines, UNet has the best performance on the ERA5 dataset, with RMSE, MAE, MAD, and RAE values of 0.574, 0.489, 0.244, and 0.330 respectively. UNet adopts a classic encoder-decoder architecture with four upsampling and four downsampling operations. However, during the downsampling process, a certain degree of information loss is inevitable. The present invention proposes a new self-convolution method to mitigate the information loss usually associated with downsampling by retaining global information. The RMSE, MAE, MAD, and RAE values of this method are 0.574, 0.492, 0.245, and 0.327 respectively. Overall, the indicators are comparable, and only the RAE has a significant improvement. However, the present invention adopts a novel idea to reduce the degree of information loss and reduce the relative absolute error between data.

[0121] Comparison with RNN-based Methods: As shown in Table 3, for the CNN-based baseline, ConvLSTM uses historical data to predict future time steps. However, as the sequence length increases, its ability to utilize early historical information weakens, resulting in poor prediction performance. Figure 6 The visualized predictions are shown, showing a significant error. The model only captures the general outline of the precipitation area but cannot accurately predict the specific rainfall within the area. These methods effectively capture temporal dependencies through continuous recurrent memory and make full use of historical information. However, over time, the inevitable loss of historical details limits their prediction ability. In addition, these methods mainly focus on the availability of historical information and do not fully address the integrity of global information, which is crucial for reducing the risk of information loss at each time step. To overcome this limitation, the present invention modifies the convolutional operation so that for the Transformer-based baseline, since its unique patch operation destroys the spatial structure of the original precipitation data, the present invention selects FourCastNet as the representative model of this category for comparison.

[0122] Comparison of Transformer-based methods: For similar Transformer-based methods, since their unique patch operation destroys the spatial structure of the original precipitation data, the present invention selects FourCastNet as the representative model of this category for comparison. The results show that this method may lead to the loss of original information and, similar to previous methods, cause serious global information loss. In contrast, the method of the present invention effectively alleviates this problem by retaining global information, thus greatly improving the prediction accuracy.

[0123] To evaluate the ability of the present invention to retain information over time, a time series prediction experiment was conducted. Using the WeatherBench dataset as a case study, the present invention further verified the performance of the proposed method. Specifically, the present invention gradually increased the prediction time to 20 time points, and the input was still the historical data of 20 time points. To evaluate the fitting of the model to the data, the present invention used the R2 score as an evaluation metric. Figure 7 For specific results, generally, as the prediction range expands, the prediction performance of the model decreases, and the R2 score decreases accordingly. As Figure 7 shown, over time, the prediction accuracy of ConvLSTM drops rapidly, and its R2 score approaches zero. Compared with other models, the method of the present invention has obvious advantages in long-term prediction, and can maintain excellent performance and effectively fit the data even over a long period of time, which reflects its robustness and stability.

[0124] The effectiveness of each proposed module was verified through ablation experiments. The evaluation datasets for the ablation experiments were still the CHIRPS, ERA5, and WeatherBench datasets.

[0125] Table 4 Results of ablation experiments, bold indicates the best

[0126] Adopt the method CHIRPS (RMSE) ERA5 (RMSE) Weather-Bench (RMSE) Standard Convolution 8.695 0.792 1.133 Large Kernel Convolution 8.615 0.742 1.056 Large Kernel Self-Convolution 8.609 0.666 0.955 Large Kernel Self-Convolution&Multi-Kernel Fusion 8.554 0.651 0.905 Large Kernel Self-Convolution&Channel Attention 8.565 0.665 0.911 Ours 8.538 0.573 0.879

[0127] The quantitative results are shown in Table 4. In this part, the present invention follows the same experimental settings as before. Specifically, the present invention compares the performance of different convolutional operations (such as standard convolution, large kernel convolution, and large kernel self-convolution) in a standard encoder-decoder model. The kernel sizes of the large kernel convolution and the large kernel self-convolution are kept the same. The results show that the large kernel convolution effectively compensates for the limited receptive field of the standard convolution by expanding the kernel size. In addition, the large kernel self-convolution transfers the focus from local features to global features, thus significantly improving the performance. Moreover, the present invention finds that the self-convolution significantly improves the performance by solving the limitation of the standard convolution in capturing global information. Integrating channel attention and multi-kernel fusion can further improve the performance. The complementarity of these methods (proven by their combined effects) highlights the effectiveness of the method proposed by the present invention.

[0128] The beneficial effects of the present invention are as follows: The present invention provides a precipitation prediction method based on an improved large kernel self-convolution neural network, adopting the precipitation prediction network architecture of the improved large kernel self-convolution neural network, which can effectively capture global information and achieve accurate modeling of complex spatio-temporal relationships. This network is particularly suitable for tasks that require maintaining the integrity of global information in precipitation forecast scenarios and can achieve more accurate predictions on time and space scales.

[0129] Aiming at the problem that traditional convolution methods are difficult to process global information and key time features simultaneously, the present invention designs a module combining self-convolution and channel attention mechanism. The large kernel self-convolution module enhances the network's ability to capture spatial features through large kernel convolution, while the channel attention mechanism CAM effectively extracts key time information and improves the network's spatio-temporal modeling ability.

[0130] The present invention proves through a large number of experiments that the test results on multiple precipitation datasets show that the precipitation prediction network based on the large kernel self-convolution strategy performs excellently in maintaining global information and the accuracy of spatio-temporal representation, significantly outperforming the current mainstream spatio-temporal modeling methods, especially having better prediction performance in complex precipitation environments.

[0131] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as "center", "thickness", "upper", "lower", "horizontal", "top", "bottom", "inner", "outer", "radial", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of technical features. Therefore, the features defined by "first", "second", "third" may explicitly or implicitly include one or more of such features.

Claims

1. A precipitation prediction method based on an improved large-kernel self-convolution neural network, characterized in that It includes the following steps: S1. Obtain a precipitation dataset, preprocess the precipitation dataset to generate a training set and a test set, and initialize the number of training rounds; S2. Input the training set into the large-kernel self-convolution neural network, train the large-kernel self-convolution neural network according to the prediction time step to obtain a preliminarily trained large-kernel self-convolution neural network; S3. Determine whether the current number of training rounds is greater than the training round threshold. If not, increment the current number of training rounds by 1 and proceed to S4. If so, generate a trained large-kernel self-convolution neural network based on the saved model parameters and proceed to S5; S4. Input the test set into the preliminarily trained large-kernel self-convolution neural network, and determine whether the mean square error index between the obtained precipitation prediction information and the true precipitation information is higher than the best mean square error index. If so, save the current model parameters, update the best mean square error index, replace the current training set, and return to S2; if not, replace the current training set and return to S2; S5. Input the test set into the trained large-kernel self-convolution neural network to obtain the final precipitation prediction result.

2. The precipitation prediction method based on the improved large kernel self-convolution neural network according to claim 1, wherein In S1, the method for preprocessing the precipitation dataset includes the following sub-steps: S11. Crop the data in the precipitation dataset to the size of or size; S12. Normalize the cropped data, and divide the normalized data according to a preset agreed ratio to generate a training set and a test set.

3. The precipitation prediction method based on the improved large kernel self-convolution neural network according to claim 1, characterized in that In S2, the large-kernel self-convolution neural network includes a large-kernel self-convolution module, a multi-kernel fusion module, a channel attention mechanism CAM, an encoder structure, and a decoder structure connected in sequence; The large core self-convolution module includes a normalization layer, the first convolution of and the second convolution of The normalization layer is respectively connected to the first convolution of and the second convolution of The multi-core fusion module includes a third convolution with the same size as and a dilated convolution with a kernel size that is half of the large kernel self-convolution; The channel attention mechanism CAM includes an average pooling layer, a maximum pooling layer, the third convolution of , the Relu activation function layer and the Sigmoid function layer. Among them, both the average pooling layer and the maximum pooling layer are connected to the third convolution of The encoder structure includes 4 encoders connected in sequence, the decoder structure includes 4 decoders connected in sequence, and a residual connection is used between the encoder and the decoder; Among them, the large-kernel self-convolution module is used to effectively enhance the network's ability to capture spatial features, the multi-kernel fusion module is used to combine feature maps from different convolution kernels, and the channel attention mechanism CAM is used to further fuse features.

4. The precipitation prediction method based on the improved large kernel self-convolution neural network according to claim 3, characterized in that In S2, the method for training the large-kernel self-convolution neural network includes the following sub-steps: S21. Input the training set into the large-kernel self-convolution module to obtain large-kernel self-convolution kernel parameters; S22. Input the large-kernel self-convolution kernel parameters into the multi-kernel fusion module to obtain the first feature map after multi-kernel fusion operation; S23. Input the first feature map into the channel attention mechanism CAM to obtain channel attention weights; S24. Perform a downsampling operation on the channel attention weights, input the result after the downsampling operation into the encoder to obtain encoded features, and input the encoded features into the decoder to obtain a predicted precipitation map, completing the training of the large-kernel self-convolution neural network.

5. The precipitation prediction method based on the improved large kernel self-convolution neural network according to claim 4, wherein, S21 includes the following sub-steps: S211. Use the data in the training set as input data, perform average pooling operation on the input data according to one-fourth of the input size through a normalization layer to obtain global feature information; S212. Split the global feature information by channel, and respectively perform convolution operations on the split results through the first convolution of and the second convolution of to obtain a first feature map and a second feature map; and the first convolution of the second convolution to perform a convolution operation to obtain a first feature map and a second feature map; S213. Connect the first feature map and the second feature map to obtain large-kernel self-convolution parameters.

6. The precipitation prediction method based on the improved large kernel self-convolution neural network according to claim 4, characterized in that S22 includes the following sub-steps: S221. Input the large-core self-convolution parameters into the third convolution of and the dilated convolution with a kernel size half of the large-core self-convolution kernel size respectively to obtain the convolution kernel parameters of the third convolution and the dilated convolution; ​ S222. Fuse the large-kernel self-convolution parameters with the convolution kernel parameters of the third convolution and dilated convolution to obtain the first feature map after multi-kernel fusion operation.

7. The precipitation prediction method based on the improved large-kernel self-convolution neural network according to claim 4, characterized in that S23 includes the following sub-steps: S231. Input the first feature map into the average pooling layer and the maximum pooling layer respectively to obtain a second feature map and a third feature map; S232. Sequentially pass the second feature map and the third feature map through the third convolution and ReLU activation function layer to obtain a first feature weight map and a second feature weight map; S233. Perform weighted summation on the first feature weight map and the second feature weight map, and input the result of the summation into the Sigmoid function layer to obtain the channel attention weight.

Citation Information

Cited By

  • Parameter sharing lightweight model picture classification method based on convolutional neural network

    CN120932005A