Lightweight short-time rainfall prediction method based on context attention fusion
By introducing modules such as deformable convolution, context-dilation awareness, and neighboring layer feature fusion, a lightweight short-term precipitation prediction model is constructed, which solves the problems of structural complexity and deployment adaptability of existing models and achieves efficient and accurate precipitation prediction.
Patent Information
- Application Number
- CN202511358042.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-09-23
AI Technical Summary
Existing short-term precipitation prediction models are complex in structure, lack spatiotemporal modeling capabilities, have limited contextual information fusion, and poor deployment adaptability, making them difficult to effectively deploy on resource-constrained small and medium-sized meteorological stations and mobile terminals.
A lightweight short-term precipitation prediction method based on contextual attention fusion is adopted. Deformable convolution module DC, context expansion perception module C, neighboring layer feature fusion module F and SMamba module are introduced to enhance feature extraction and temporal modeling capabilities, reduce model parameters, and build an efficient image temporal prediction model.
It improves prediction accuracy and boundary clarity, reduces computational costs, and has good real-time prediction performance and engineering availability, making it suitable for deployment in resource-constrained scenarios.
Smart Images

Figure CN120871304A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a lightweight short-term precipitation prediction method based on contextual attention fusion, belonging to the field of short-term precipitation prediction technology. Background Technology
[0002] Short-term precipitation forecasting typically refers to predicting the evolution of precipitation within the next 0-2 hours. This task requires high temporal and spatial resolution for early warning response and is widely used in areas such as traffic management, agricultural production, urban flood control, and meteorological disaster early warning. Therefore, developing accurate and efficient short-term precipitation forecasting models is of great significance for improving public safety and ensuring the smooth operation of society.
[0003] Traditional forecasting methods mainly include numerical weather prediction (NWP) and image extrapolation. NWP uses atmospheric physical models combined with initial field data for numerical integration calculations, making it suitable for medium- and long-term weather forecasts. However, this method is highly dependent on computational resources, has limited forecast update frequency, requires a supercomputing platform, and faces accuracy bottlenecks at short timescales of 0-2 hours due to rapid error growth. Image extrapolation methods (such as optical flow) infer future trends based on the spatial displacement of cloud clusters in radar image sequences, offering advantages in real-time performance and speed. However, it has weak modeling capabilities for changes in precipitation intensity and struggles to accurately address the highly nonlinear and locally sudden evolution characteristics.
[0004] With the accumulation of meteorological data and the development of artificial intelligence technology, data-driven deep learning forecasting models are becoming increasingly mature. Among them, Convolutional Neural Networks (CNNs), with their advantages in image recognition and spatial structure modeling, are widely used in precipitation prediction tasks for radar image sequences. In particular, as one of the most commonly used CNN architectures, UNet, due to its symmetrical structure, controllable parameters, and strong adaptability, is widely used in image segmentation and reconstruction tasks, and also provides an effective basic structure for meteorological image sequence prediction.
[0005] However, traditional UNet models also have certain limitations in precipitation forecasting tasks. On the one hand, their spatial feature modeling capabilities are limited; traditional convolutional kernels struggle to adaptively handle the nonlinear deformation of complex boundaries, affecting the edge sharpness of predicted images, and they lack explicit temporal modeling mechanisms, failing to fully utilize time-dependent information. On the other hand, the UNet structure is not sufficiently robust in its mechanisms for complex feature extraction and cross-scale, cross-level feature interaction, limiting its ability to characterize complex image structures and details. Furthermore, most current deep learning models, while pursuing prediction accuracy, often involve a large number of parameters and high computational resource consumption, making them difficult to deploy on small and medium-sized weather stations, edge devices, or mobile terminals with limited computing power, thus restricting their rapid integration and promotion in resource-constrained scenarios. Therefore, designing a lightweight short-term precipitation forecasting model that combines high-precision prediction performance with good deployment capabilities has become a critical issue that urgently needs to be addressed. Summary of the Invention
[0006] The technical problem to be solved by this invention is to provide a lightweight short-term precipitation prediction method based on contextual attention fusion, which solves the problems of complex model structure, insufficient spatiotemporal modeling capability, limited contextual information fusion and poor deployment adaptability in the existing technology.
[0007] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: A lightweight short-term precipitation prediction method based on contextual attention fusion includes the following steps: Step 1: Acquire a continuous sequence of weather radar images. Using a sliding window strategy, slice the acquired continuous weather radar image sequence according to a preset step size. The window contains slices from the previous slice to the next slice. The input image sequence and the subsequent frames The output image sequence consists of frames, and the window length is [missing information]. and The sum of these samples, controlled by a preset step size, generates several samples from a continuous sequence of weather radar images. All samples together constitute a weather radar image dataset. Step 2: Preprocess the meteorological radar image dataset by dividing the preprocessed dataset into a training set, a validation set, and a test set. Step 3: Construct a lightweight short-term precipitation prediction model based on contextual attention fusion, which is used to predict the precipitation in the future preset second time period using the historical preset first time period of the current moment. The historical preset first time period and the future preset second time period are continuous. The model takes the input image sequence as input and the generated prediction image sequence as output. The lightweight short-term precipitation prediction model based on contextual attention fusion includes an encoder, a decoder, and a bottleneck layer connecting the encoder and decoder. Step 4: Use the training set and validation set to train and validate the lightweight short-term precipitation prediction model based on contextual attention fusion constructed in Step 3, and obtain the trained lightweight short-term precipitation prediction model. Step 5: Input the input image sequence from the test set into the trained lightweight short-term precipitation prediction model to obtain the short-term precipitation prediction results.
[0008] Compared with the prior art, the present invention, employing the above technical solution, has the following technical effects: 1. This invention introduces a deformable convolution module (DC) to enhance the ability to perceive and extract features from nonlinear and complex precipitation images.
[0009] 2. The present invention designs a context expansion perception module C, which uses parallel dilated convolution to capture rich contextual information and improve the model’s sensitivity to complex spatial structures and irregular boundary features.
[0010] 3. The present invention designs a neighboring layer feature fusion module F, which combines channel attention and spatial attention mechanisms to fuse features from different levels in skip connections, thereby enhancing the model’s attention to important image feature regions and suppressing the transmission of redundant features.
[0011] 4. This invention introduces the SMamba module for modeling long-range time dependencies, enhancing the modeling of long-term time-series dependency features.
[0012] 5. The lightweight short-term precipitation prediction model designed in this invention significantly reduces model parameters (approximately 1 / 3 of the original UNet), improving prediction accuracy while reducing computational costs, and possessing excellent real-time prediction performance and engineering usability. Attached Figure Description
[0013] Figure 1 This is a flowchart of the lightweight short-term precipitation prediction method based on contextual attention fusion according to the present invention; Figure 2 This is an architecture diagram of the lightweight short-term precipitation prediction model based on contextual attention fusion constructed in this invention; Figure 3 This is a schematic diagram of the deformable convolution module DC in this invention; Figure 4 This is a schematic diagram of the structure of the context extension sensing module C in this invention; Figure 5 This is a schematic diagram of the structure of the adjacent layer feature fusion module F in this invention; Figure 6 This is a schematic diagram of the SMamba module introduced in the bottleneck layer in this invention. Detailed Implementation
[0014] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0015] like Figure 1 As shown, this invention proposes a lightweight short-term precipitation prediction method based on contextual attention fusion, applicable to image sequence prediction and spatiotemporal dynamic modeling tasks. This method integrates deformable convolution, context-dilated perception mechanism, cross-layer feature fusion mechanism, and SMamba module to achieve multi-scale, strongly nonlinear, and long-term dependent information perception capabilities, constructing an efficient and deployable image temporal prediction model. The proposed improved network not only effectively improves the prediction accuracy and boundary clarity of radar image sequences but also reduces the number of model parameters, exhibiting good real-time prediction performance and engineering usability. The specific steps are as follows: S1: Data input, acquiring a continuous sequence of weather radar images as raw input data, image size is... A sliding window strategy is adopted, based on the step size. A meteorological radar image dataset is constructed from the original sequence slices. Each sample consists of the input sequence. and predicted output sequence Composition. By setting a time dimension index, any... and The combined input-output sequence partitioning method adapts to different prediction duration requirements. Furthermore, to adapt to the model input format, a unified tensor layout transformation strategy is adopted to transform the original image dimensions from NHWT (N is the number of samples, H is the image height, W is the image width, and T is the time step) to NTHW format, i.e. This allows it to be directly used in time series convolution processing within the model.
[0016] In this embodiment, the acquired continuous weather radar image sequence is the VIL sequence from the SEVIR dataset, containing more than 10,000 samples, and the original image size is [size missing]. The time interval is 10 minutes.
[0017] S2: Data preprocessing involves normalizing and cropping the radar images in the dataset to improve the efficiency and accuracy of model training and inference. The specific processing is as follows: The input radar image sequence is normalized, and the pixel values of each radar image are scaled to [value missing]. The interval, the formula is as follows: , in, These are the original pixel values. and These represent the minimum and maximum values of pixels in the image, respectively.
[0018] To adapt to the model input, the original size of the input is set to... Radar images cropped to Pixel size.
[0019] S3: Model Building. Construct a lightweight short-term precipitation prediction model based on contextual attention fusion, containing a four-layer encoder-decoder structure, such as... Figure 2 As shown, different layers are entered through downsampling or upsampling. The preprocessed precipitation sequence... The input model passes through a deformable convolutional module (DC), a context-dilated perceptual module (C), a neighboring layer feature fusion module (F), and a bottleneck layer SMamba module. It combines depthwise separable convolution (DSC) with a skip connection mechanism to extract, fuse, and reconstruct image features, and finally outputs the prediction results for future frames.
[0020] First, a deformable convolutional module (DC) is used at the initial input to extract nonlinear spatial features of the input image sequence, by learning spatial offsets. To achieve spatial adaptive sampling, the module's process is as follows: Figure 3 As shown, its expression is: , in, The coordinates of the current convolution kernel center. For feature map at location The response value at that position, i.e., the output at that position after deformable convolution; For the routine K × K In convolution, the first The relative coordinates of each core point with respect to the core. For the corresponding convolution weights, For the number of core points, This indicates bilinear interpolation sampling.
[0021] Afterwards, the feature map undergoes a 2x downsampling and depthwise separable convolution (DSC) before entering the second layer of the encoder. At this point, the feature map size is halved, and the number of feature channels doubles. The resulting feature map is then input into the context-expanded perceptual module C, whose structure is as follows: Figure 4As shown, by using a dual-branch dilated convolutional structure to retain multiple receptive fields and dynamically expanding the receptive fields using dilated convolutions with different dilation rates, multi-scale contextual information of image features can be captured, thereby integrating local and global features while maintaining spatial resolution. Subsequently, through feature fusion and residual connections, the image feature representation capability and model training stability are effectively improved. The specific process is as follows: For input features Using expansion rate respectively of Dilated convolution yields two branch features: Branch 1: , Branch 2: , Dual-branch feature splicing and fusion: , , The output after residual connection is enhanced feature: , , in, and These represent expansion rates of 4 and 6, respectively. Hollow convolution.
[0022] Similarly, after another 2x downsampling and depthwise separable convolution (DSC), the image enters the next layer of the encoder, where the context-dilation perceptual module C captures multi-scale contextual information of the image features. After the fourth layer, following the 2x downsampling and depthwise separable convolution (DSC), the image enters the bottleneck layer. In this layer, the SMamba module is used to process deeper, more abstract features. The workflow of this module is as follows: Figure 6 As shown. A lightweight sequence modeling structure, SMamba, is introduced into the bottleneck layer between the encoder and decoder to model long-range temporal dependencies in image sequences, enhancing the model's global perception and nonlinear modeling capabilities in the deep feature extraction stage. The processing flow is as follows: Input batch normalization: , Will Flattened into a sequence: , Sequence modeling: , Restored to 2D feature map: , Generate the final output: , in, This indicates the batch size, which is the number of samples input into the model at one time. Indicates the number of image channels. These represent the image height and width, respectively. This indicates the input to the SMamba module. Perform batch normalization. Indicates shape reshaping. Indicates dimension permutation. This indicates the input 2D feature map. Flattened according to spatial location, the length is... sequence, For the real number field, Indicates adaptation to 2D feature maps Sequence modeler, Indicates the sequence The 2D feature map obtained after dimensional displacement and shape reshaping. This indicates the GELU activation operation. This indicates the output of the SMamba module.
[0023] Subsequently, the features obtained from adjacent layers of the encoder are input into the adjacent layer feature fusion module F, the structure of which is as follows: Figure 5 As shown in the diagram, this module integrates channel attention and spatial attention mechanisms. Through progressive processing including scale alignment, feature stitching, channel attention, spatial attention, and feature compression, it effectively fuses image features from adjacent layers at different scales, thereby enhancing the perception of local image details and global image information across different levels and scales. The specific process is as follows: Scale alignment (based on whether the number of channels is the same): ; Two feature maps concatenated: .
[0024] First, perform channel attention-weighted fusion: For Perform global average pooling separately Similar to global max pooling (MaxPool). The two vectors obtained as channel indices, describing the global distribution of the channels, are then input into a two-layer MLP network with shared parameters for processing. , This is the channel vector obtained after pooling. These are the weight matrices for the first and second fully connected layers, respectively. The outputs of the two layers are summed element-wise, and then activated by a sigmoid function. Generate channel attention weights Afterwards, Channel-by-channel weighted output .
[0025] Then perform spatial attention weighted fusion: on the channel dimension, the output of the channel attention module is... respectively through average pooling and max pooling , for The number of channels is used to obtain two spatial attention maps, which are then stacked. The splicing characteristics, after Spatial weight map is obtained after convolution, batch normalization (BN), and sigmoid activation. Then, the output is weighted element by element. Finally, after Convolution, BN, and ReLU compress the channels to restore them to their original channel dimensions, and the output is... .
[0026] in, For the current layer features, For upper-level features, the + sign indicates element-wise addition by channel. This indicates channel-dimensional concat. This indicates element-wise multiplication. This represents the sigmoid activation function. This represents a multilayer perceptron consisting of two fully connected layers and ReLU. For channel attention fusion output, The spatial attention fusion output, the final module fusion result is .
[0027] Next, in the decoder stage, the output of the bottleneck layer SMamba module is upsampled (Up) and concatenated with the output of the neighboring layer feature fusion module F. This concatenation is then processed by depthwise separable convolution (DSC) to obtain the fourth layer decoder features. Subsequently, the fourth layer decoder features are upsampled (Up) again and concatenated with the output of the neighboring layer feature fusion module F. This process is repeated until the first layer is decoded, at which point the image size is restored to the input size. Finally, the image is processed... The convolution outputs the prediction result, and the model construction is complete.
[0028] S4: Model output. After model training is complete, the preprocessed radar image sequence to be predicted will be output. Input the constructed network model, output the radar predicted image sequence for several future time steps. .
[0029] Based on the error between the real image sequence and the predicted result, the model performance is evaluated using multiple metrics, including mean squared error (MSE), precision, accuracy, F1 score, critical success index (CSI), Heidegger skill score (HSS), and false alarm rate (FAR). The calculation methods are as follows: , , , , , , , in, Indicates the number of samples. Represents the actual ground value. Recall represents the model's predicted value. By binarizing the predicted and real images, the number of true positives (TP, true value = 1, predicted value = 1), false positives (FP, true value = 0, predicted value = 1), true negatives (TN, true value = 0, predicted value = 0), and false negatives (FN, true value = 1, predicted value = 0) in the confusion matrix is calculated.
[0030] To further verify the effectiveness and advancement of the method described in this invention, the following experiments were conducted: The experiment used the publicly available meteorological radar image dataset SEVIR, and adopted the VIL (Vertical Cumulative Liquid Water Content) image sequence as input data. Using the radar image sequence from the previous 60 minutes (6 frames in total, with a 10-minute time interval) as input, a lightweight short-term precipitation forecasting model based on image sequence modeling constructed in this invention was used to predict the image sequence for the next 60 minutes (6 frames).
[0031] To evaluate the computational efficiency of the model, this embodiment was compared with several existing advanced models in terms of the number of parameters, and the statistical results are shown in Table 1.
[0032] Table 1 Comparison of Model Sizes
[0033] As shown in Table 1, the lightweight structure proposed in this invention exhibits significant advantages in both model size and computational overhead. Regarding the number of parameters, the model contains only approximately 5.86M learnable parameters, representing reductions of 66.1%, 57.5%, and 57.0% compared to UNet, SAR-UNet, and SimVP, respectively, significantly lowering storage and deployment costs. In terms of computational complexity, the model's floating-point operations (FLOPs) are 29.48 GFLOPs, a reduction of 67.1% and 51.3% compared to SimVP and UNet, respectively, demonstrating good inference efficiency and resource adaptability. Although slightly higher than SAR-UNet in FLOPs, subsequent experiments verified that this invention performs better in prediction stability and spatiotemporal awareness, balancing performance and efficiency, and demonstrating its potential in practical deployment scenarios.
[0034] To further evaluate the predictive performance of the model, MSE, Precision, Accuracy, F1, CSI, HSS, and FAR were assessed with a threshold of 0. The results were compared with those of several state-of-the-art models, and the statistical results are shown in Table 2.
[0035] Table 2 Performance comparison of each model on the SEVIR dataset (threshold 0)
[0036] in, This indicates that a higher index means better performance. The smaller the indicator, the better the performance; the best result is shown in bold. As shown in Table 2, under the condition of a threshold of 0, the model described in this invention achieves at least a 4% performance improvement on most evaluation indicators except MSE, demonstrating better predictive stability and judgment accuracy. In particular, the improvement is especially significant on practically meaningful indicators such as HSS, Accuracy, and CSI, effectively verifying the effectiveness and practical value of the method described in this invention. Combining the model size comparison data in Table 1, although the FLOPs of this invention are not the lowest among the compared models, it achieves a good trade-off between performance and computational efficiency with a relatively low number of parameters (approximately 5.86M). Especially noteworthy is its effective control of model complexity while maintaining prediction accuracy, balancing computational efficiency and expressive power, giving it good resource adaptability and inference efficiency, demonstrating its potential value in practical application deployment.
[0037] Based on the same inventive concept, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the aforementioned lightweight short-term precipitation prediction method based on contextual attention fusion.
[0038] Based on the same inventive concept, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the aforementioned lightweight short-term precipitation prediction method based on contextual attention fusion.
[0039] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0040] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0041] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0042] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0043] The above embodiments are merely illustrative of the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications made to the technical solutions based on the technical concept proposed in this invention shall fall within the scope of protection of this invention.
Claims
1. A lightweight short-term precipitation forecasting method based on contextual attention fusion, characterized in that, Includes the following steps: Step 1: Acquire a continuous sequence of weather radar images. Using a sliding window strategy, slice the acquired continuous weather radar image sequence according to a preset step size. The window contains slices from the previous slice to the next slice. The input image sequence and the subsequent frames The output image sequence consists of frames, and the window length is [missing information]. and The sum of these samples, controlled by a preset step size, generates several samples from a continuous sequence of weather radar images. All samples together constitute a weather radar image dataset. Step 2: Preprocess the meteorological radar image dataset by dividing the preprocessed dataset into a training set, a validation set, and a test set. Step 3: Construct a lightweight short-term precipitation prediction model based on contextual attention fusion, which is used to predict the precipitation in the future preset second time period using the historical preset first time period of the current moment. The historical preset first time period and the future preset second time period are continuous. The model takes the input image sequence as input and the generated prediction image sequence as output. The lightweight short-term precipitation prediction model based on contextual attention fusion includes an encoder, a decoder, and a bottleneck layer connecting the encoder and decoder. Step 4: Use the training set and validation set to train and validate the lightweight short-term precipitation prediction model based on contextual attention fusion constructed in Step 3, and obtain the trained lightweight short-term precipitation prediction model. Step 5: Input the input image sequence from the test set into the trained lightweight short-term precipitation prediction model to obtain the short-term precipitation prediction results.
2. The lightweight short-term precipitation prediction method based on contextual attention fusion according to claim 1, characterized in that, The specific process of step 2 is as follows: Normalization is performed on each weather radar image in each sample, that is, the pixel values of each weather radar image are scaled down to... interval; The normalized weather radar images are cropped to a preset size to obtain a preprocessed weather radar image dataset. The preprocessed weather radar image dataset is then divided into a training set, a validation set, and a test set.
3. The lightweight short-term precipitation prediction method based on contextual attention fusion according to claim 1, characterized in that, In step 3, in the encoder, the model input is processed by a deformable convolutional module to obtain the output of the deformable convolutional module; the output of the deformable convolutional module is downsampled by 2 times, and then sequentially processed by a first depthwise separable convolutional module and a first context-division dilation perceptron module to obtain the output of the first context-division dilation perceptron module; the output of the first context-division dilation perceptron module and the output of the deformable convolutional module are fused by a first neighboring layer feature fusion module to obtain the first fused feature. The output of the first depthwise separable convolutional module is downsampled by a factor of 2, and then sequentially passed through the second depthwise separable convolutional module and the second context-division dilation perceptron to obtain the output of the second context-division dilation perceptron. The output of the second context-division dilation perceptron and the output of the first context-division dilation perceptron are then fused by the second neighboring layer feature fusion module to obtain the second fused feature. The output of the second depthwise separable convolutional module is downsampled by a factor of 2, and then passed through the third depthwise separable convolutional module to obtain its output. The output of the third depthwise separable convolutional module is then fused with the output of the second context-division extended sensing module through the third neighboring layer feature fusion module to obtain the third fused feature. ; The bottleneck layer includes the SMamba module. The output of the third depthwise separable convolutional module is downsampled by 2 times, and then passed sequentially through the fourth depthwise separable convolutional module and the SMamba module to obtain the output of the SMamba module. In the decoder, the output of the SMamba module is upsampled by a factor of 2 before being fused with the third feature. The skip connection is used to obtain the third skip connection result. The third skip connection result is then passed through the fifth depthwise separable convolutional module to obtain the output of the fifth depthwise separable convolutional module. The output of the fifth depthwise separable convolutional module is upsampled by a factor of 2 before being combined with the second fused feature. The skip connection is used to obtain the second skip connection result. The second skip connection result is then passed through the sixth depthwise separable convolutional module to obtain the output of the sixth depthwise separable convolutional module. The output of the sixth depthwise separable convolutional module is upsampled by a factor of 2 before being fused with the first feature. Skip connections are used to obtain the first skip connection result. The first skip connection result is then passed through the seventh depthwise separable convolutional module to obtain the output of the seventh depthwise separable convolutional module. The output of the seventh depthwise separable convolutional module is upsampled by a factor of 2, then concatenated with the output of the deformable convolutional module. The concatenated result is then sequentially processed by the eighth depthwise separable convolutional module and... Convolutional layers produce the model's output.
4. The lightweight short-term precipitation prediction method based on contextual attention fusion according to claim 3, characterized in that, The first and second context-expanded sensing modules have the same structure, both including two branches. The first branch includes a dilation rate of 4. Dilated convolution, batch normalization, and ReLU activation function; the second branch includes a dilation rate of 6. Dilated convolution, batch normalization, and ReLU activation function; the expressions for each context-dilation-aware module are as follows: , , , , , , in, This represents the input to each context-extended sensing module. and These represent expansion rates of 4 and 6, respectively. Hollow convolution, Indicates batch normalization. Represents the ReLU activation function. Indicates the first branch feature Second branch features splicing, express convolution, Indicates the splicing result conduct The results obtained from convolution, batch normalization, and ReLU activation operations Indicates to conduct The results obtained from convolution, batch normalization, and ReLU activation operations Indicates to and Enhanced features obtained by performing residual connections.
5. The lightweight short-term precipitation prediction method based on contextual attention fusion according to claim 3, characterized in that, The first, second, and third neighboring layer feature fusion modules have the same structure, all including a channel attention module and a spatial attention module, fusing image features of different scales from adjacent layers. The specific process of each neighboring layer feature fusion module is as follows: , in, Indicates the lower-level features of two adjacent layers. Indicates the upper-layer features of two adjacent layers. Features obtained by subsampling This indicates a max pooling operation. This represents the convolution operation. Indicates batch normalization. Represents the ReLU activation function. This represents the features obtained by splicing the channel dimensions. , They are respectively , The number of channels, Indicates channel-dimensional splicing; right Perform channel attention-weighted fusion, specifically: Perform adaptive max pooling separately and adaptive average pooling Using multilayer perceptron The results of max pooling and average pooling are processed separately. The processed results are then summed element-wise and activated by a sigmoid function to obtain the channel weights. , ,use right Perform channel-by-channel weighting to obtain the output of the channel attention module. , The plus sign indicates that the elements are added one by one along the channel. This indicates the sigmoid activation operation. This indicates element-wise multiplication; right Perform spatial attention-weighted fusion, specifically: Average pooling and max pooling are performed separately to obtain two spatial attention maps. These two spatial attention maps are then concatenated, and then... The spatial weight map is obtained after convolution, batch normalization, and sigmoid activation. , ,use right Features are obtained by element-wise weighting. , ,right conduct Convolution, batch normalization, and ReLU activation operations yield the outputs of the feature fusion modules for each neighboring layer. , ,in, This represents the result of concatenating two spatial attention maps. express convolution, express convolution.
6. The lightweight short-term precipitation prediction method based on contextual attention fusion according to claim 3, characterized in that, The expression for the SMamba module is as follows: Input batch normalization: , Will Flattened into a sequence: , Sequence modeling: , Restored to 2D feature map: , Generate the final output: , in, Indicates batch size. Indicates the number of image channels. These represent the image height and width, respectively. This indicates the input to the SMamba module. Perform batch normalization. Indicates shape reshaping. Indicates dimension permutation. This indicates the input 2D feature map. Flattened according to spatial location, the length is... sequence, For the real number field, Indicates adaptation to 2D feature maps Sequence modeler, Indicates the sequence The 2D feature map obtained after dimensional displacement and shape reshaping. This indicates the GELU activation operation. This indicates the output of the SMamba module.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the lightweight short-term precipitation forecasting method based on contextual attention fusion as described in any one of claims 1 to 6.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the lightweight short-term precipitation forecasting method based on contextual attention fusion as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Short temporary rainfall prediction method and device
CN114139690A
Crop disease segmentation method, system and equipment based on multi-scale fusion and CBAM-ResNet50 and medium
CN116543282A
Short-time heavy rainfall forecasting method fusing self-attention module and Unet model
CN117008217A
Short-time heavy rainfall minute-level forecasting method based on SFGAN-ARPreRNN model and multi-layer radar data
CN117741821A
Short temporary rainfall prediction method based on complex frequency domain depth attention voxel flow
CN118795575A
Cited By
Short-critical extreme rainfall prediction method based on abnormal driving residual dynamic diffusion
CN122131269A
A short-term extreme precipitation prediction method based on abnormal driving residual dynamic diffusion
CN122131269B