Rainfall image estimation method based on multi-scale convolution attention decoder
Through the precipitation image estimation method based on multi-scale convolutional attention decoder, the limitations of existing satellite precipitation inversion methods in spatiotemporal resolution and accuracy are solved, and high-precision and high-resolution precipitation estimation are achieved, especially in complex cloud systems and ground-free observations.
Patent Information
- Application Number
- CN202510661698.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing satellite precipitation inversion methods have limitations in the spatial and temporal resolution and accuracy of data. Especially in the case of non-convective precipitation or high clouds, it is difficult to achieve accurate precipitation estimation for all periods and all scenarios.
A precipitation image estimation method based on multi-scale convolutional attention decoder (MCAD) is proposed. By constructing an MCAD model, combining multi-scale information extraction and attention mechanism, the precipitation information characteristics in satellite remote sensing data are fully explored to achieve accurate portrayal of precipitation mode.
The end-to-end reconstruction of multi-source data is realized, which significantly improves the precipitation recognition accuracy of complex cloud systems, and can automatically output high-resolution precipitation data in areas covered by ground-free rainfall meter, realizing millimeter-level precipitation intensity estimation.
Smart Images

Figure CN120182855A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of satellite remote sensing data analysis and processing, and particularly to a precipitation image estimation method based on a multi-scale convolutional attention decoder. Background Art
[0002] Precipitation is one of the key meteorological elements affecting the global water cycle and energy exchange. In the monitoring of disastrous weather processes, it is of great significance to timely understand the spatial distribution of precipitation. It is not only crucial for agriculture, ecological environment and water resource management, but also directly related to the formation and development of extreme weather events such as typhoons, heavy rains and droughts. The spatio-temporal distribution of precipitation is extremely complex and is jointly affected by multiple factors such as the climate system, atmospheric dynamic processes and terrain. Compared with temperature data, precipitation has the characteristics of instantaneousness, strong instability, uneven time and space distribution, etc. Therefore, the overall data changes are very prominent and the differences are large. At present, satellite remote sensing technology can achieve large-scale precipitation monitoring and provide more continuous precipitation estimation data, becoming an important means for precipitation research in recent years. However, there are still certain limitations in the spatio-temporal resolution and accuracy of traditional satellite precipitation retrieval methods. Therefore, exploring high-precision and high-resolution precipitation estimation methods and improving the accuracy and reliability of precipitation reconstruction have important scientific significance and application value.
[0003] In recent years, satellite remote sensing technology and business products have become increasingly mature, and using remote sensing satellites for precipitation estimation tasks has become a trend. According to the differences in bands, the relevant methods for estimating and reconstructing precipitation by means of remote sensing satellites can generally be divided into the following three categories. (1) Methods based on visible light and infrared. By means of the radiation information and reflection information of the cloud top, the possibility of precipitation occurrence is judged. At the same time, by analyzing relevant information such as cloud thickness and cloud top temperature, the probability of precipitation occurrence and the duration of precipitation are determined, and finally the precipitation amount is estimated. (2) Methods based on microwave. Microwave precipitation estimation algorithms can generally be mainly classified into three categories, namely radiation algorithms, scattering algorithms and multi-band inversion algorithms. (3) Methods based on multi-sensor combination. Researchers have gradually begun to use the sensors carried by various satellites to obtain relevant information for estimating global precipitation.
[0004] These three traditional precipitation estimation methods based on remote sensing satellite data have shown good performance in precipitation reconstruction tasks. However, due to the high spatio-temporal heterogeneity of precipitation elements themselves, the above methods still have certain limitations. The infrared-based indirect inference method is prone to inaccurate precipitation estimation, especially in the case of non-convective precipitation or high clouds. The infrared method is interfered under cloudy or thick cloud cover, which limits its applicability to precipitation in all time periods and all scenarios. At present, the microwave band can generally only be used to detect precipitation data information over the ocean, and it is difficult to effectively carry out the detection of precipitation data over the land, which greatly limits its application in precipitation monitoring in a wider area. The multi-sensor combination method requires integrating data from different sources. The time resolution, spatial resolution, and observation principles of these sensors are different, resulting in high complexity in data matching and fusion and may introduce additional uncertainties. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the present invention proposes a precipitation image estimation method based on a multi-scale convolutional attention decoder, which is characterized in that a precipitation estimation model based on a multi-scale convolutional attention decoder MCAD is constructed. The precipitation features are extracted by the decoder MCAD. MCAD decodes the features extracted by the encoder and outputs the estimated precipitation image. MCAD adopts the architecture of a multi-scale convolutional attention module MCAM. MCAD combines multi-scale information extraction and an attention mechanism to fully mine the precipitation information features in satellite remote sensing data, specifically including: Step 1: Collect and organize the data set and perform preprocessing to construct a training set and a test set; Step 11: Collect and organize FY-4B satellite data and GPM IMRGE precipitation products as the data set, where the GPM IMRGE precipitation product is used as the label data; Step 12: Data preprocessing. Perform brightness temperature calibration and geometric correction on the remote sensing data collected by the FY-4B satellite. Crop the geometrically corrected remote sensing data according to the target area range to obtain 256×256 data blocks; Step 13: Divide the training set and the test set according to a preset ratio; Step 2: Construct a precipitation image estimation model based on a multi-scale convolutional attention decoder MCAD. The precipitation estimation model uses the U-Net model as the basic framework to perform precipitation reconstruction tasks based on an encoding-decoding architecture. The encoder extracts features through four convolutional layers and downsampling operations. The decoder replaces the traditional convolutional layer with a multi-scale convolutional attention module MCAM. The multi-scale convolutional attention module MCAM includes a channel attention module CAB, a spatial attention module SAB, and a multi-scale convolutional module MCB connected in sequence. The decoder outputs the predicted precipitation image through four multi-scale convolutional attention modules MCAM; Step 3: Input the training set into the constructed precipitation estimation model for training. The training process specifically includes: Step 31: Input the remote sensing data of channels 9 to 15 of FY-4B satellite in the training set into the encoder for feature extraction. The remote sensing data is successively subjected to feature extraction through four encoding layers to obtain the first encoding feature, the second encoding feature, the third encoding feature, and the fourth encoding feature respectively; Step 32: Input the four encoding features into the corresponding decoding layers for decoding. Specifically, the first encoding feature is input into the first decoding layer, the second encoding feature is input into the second decoding layer, the third encoding feature is input into the third decoding layer, and the fourth encoding feature is input into the fourth decoding layer. Each decoding layer includes a multi-scale convolutional attention module MCAM and an upsampling layer; Step 33: The fourth encoding feature passes through the fourth decoding layer to obtain the fourth decoding feature. After the fourth decoding feature is fused with the third encoding feature, it is input into the third decoding layer to output the third decoding feature. After the third decoding feature is fused with the second encoding feature, it is input into the second decoding layer to output the second decoding feature. After the second decoding feature is fused with the first encoding feature, it is input into the first decoding layer to output the first decoding feature, thus obtaining the predicted precipitation image; Step 34: Calculate the loss function between the predicted precipitation image and the true label image of the GPM IMERG precipitation product in the training set. Adjust the parameters of the precipitation estimation model according to the loss value until the loss value is less than the preset threshold or reaches the preset number of iterations, then stop the training and save the trained model; Step 4: Input the test set into the trained precipitation estimation model for testing.
[0006] According to a preferred implementation manner, the multi-scale convolutional attention module MCAM in Step 32 first performs feature screening through CAB and SAB, and then extracts the final precipitation information through the multi-scale convolutional module MCB. The processing of MCAM specifically includes: Step 321: The decoded feature first passes through the channel attention module CAB. CAB includes two branches, an adaptive max-pooling branch and an adaptive average-pooling branch. Among them, adaptive max-pooling is used to extract the most significant responses in each channel, and adaptive average-pooling captures the global statistical features of each channel; the features output by both are fused after subsequent convolution and ReLU activation, and a channel weight is generated through the Sigmoid function. Finally, it is fused with the original feature in an element-wise weighted form, and the fused feature is output to the spatial attention module SAB; Step 322: The Spatial Attention Module (SAB) extracts key spatial features through operations of obtaining the channel maximum value and the channel average value respectively. First, SAB performs max-pooling on the input fused features in the channel dimension to obtain a single-channel maximum value feature map. Meanwhile, SAB performs average-pooling on the input fused features in the channel dimension to calculate the global mean of each channel and obtains a single-channel average value feature map. The maximum value feature map and the average value feature map are fused, followed by a large convolution operation and a Sigmoid activation function to output spatial attention features; Step 323: The spatial attention features are input into the Multi-scale Convolution Module (MCB). MCB uses Multi-scale Depthwise Separable Convolution (MDC) to extract features with different receptive fields. MDC adopts three convolution kernels with different scales. The small convolution kernel is used to capture local texture information, the medium convolution kernel is used to identify medium-scale cloud systems, and the large convolution kernel is used to model the large-scale cloud distribution. Then, the feature maps under different-scale convolution kernels are respectively passed through batch normalization and ReLU6 activation functions. Next, the features extracted by the depthwise separable convolutions with different scales are fused by element-wise addition. The fused features go through a channel shuffle operation.
[0007] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. The present invention realizes end-to-end reconstruction of multi-source data collaboration. The proposed method breaks through the empirical threshold limit of traditional physical inversion algorithms, realizes end-to-end precipitation field reconstruction through a deep learning framework, utilizes the collaborative observation advantages of multiple spectral channels of FY-4B satellite, effectively fuses multi-dimensional features such as cloud top brightness temperature and water vapor content, significantly improves the precipitation recognition accuracy of complex cloud systems, and in areas without ground rain gauges, completes model training in a self-supervised manner, automatically outputs high-resolution precipitation data, realizes millimeter-level precipitation intensity estimation, and effectively solves the problem of missing precipitation data values caused by the lack of station sites.
[0008] 2. The present invention realizes the design of a multi-scale feature fusion architecture. A Multi-scale Convolution Attention Decoder (MCAD) architecture is constructed. A multi-scale feature pyramid is built through parallel depthwise separable convolutions to synchronously capture the spatial distribution features of large-scale stratiform cloud precipitation and small-scale convective precipitation, and solves the technical bottleneck that it is difficult for traditional single-scale convolution kernels to take into account different-scale precipitation patterns.
[0009] 3. The present invention proposes a channel-spatial dual attention optimization mechanism. A Channel Attention Module (CAB) is designed to dynamically calculate the contribution weights of each infrared channel of FY-4B, strengthening the feature expression of water vapor sensitive channels (channels 9-11); combined with the Spatial Attention Module (SAB) to focus on the spatial distribution of the core precipitation area, and enhance the accuracy of precipitation intensity estimation through feature map recalibration.
[0010] 4. The present invention proposes a cross-scale feature interaction mechanism. In the decoding stage, a multi-scale convolution attention module (MCAM) is introduced, in which the multi-scale convolution module (MCB) contains a parallel structure of deep separable convolution (MDC). Through the parallel operation of convolution kernels of different scales (1×1, 3×3, 5×5), local textures, convective clouds and large-scale precipitation structures are captured respectively, and precipitation feature extraction and fusion at resolutions ranging from 4km to 16km are realized. This mechanism effectively enhances the model's response to complex structures such as the typhoon eyewall and heavy precipitation boundaries, while suppressing the interference signals of non-precipitation clouds, and improving the robustness and accuracy of heavy rain level precipitation discrimination.
[0011] 5. The present invention proposes a supervision and constraint mechanism for multi-source remote sensing collaboration. The multi-spectral remote sensing capability of the 9th to 15th channels of the Fengyun-4B satellite (FY-4B) is fully utilized to integrate physical features such as water vapor layer, cloud top brightness temperature, and cloud thickness into the data, and the GPMIMERG precipitation product is used as a supervisory signal (label data) to achieve end-to-end supervised training of the model. This design combines the physical complementary characteristics of infrared and microwaves. On the one hand, it utilizes the high spatiotemporal continuity of infrared remote sensing, and on the other hand, it uses the high precision of microwave inversion of precipitation to make up for the shortcomings of a single payload in inverting precipitation phase (solid / liquid) and structural details, thereby effectively improving the physical rationality and quantitative accuracy of the overall precipitation estimation. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is a processing flow chart of the method of the present invention; Figure 2 Flow chart of AGRI data preprocessing of the radiation imager of the present invention; Figure 3 This is a schematic diagram of GLT geometric correction; Figure 4 It is a schematic diagram of the MCAD network model structure of the present invention; Figure 5 It is a schematic diagram of the structure of the MCAM of the present invention; Figure 6 It is a structural diagram of multi-scale depth-separable convolution MDC. DETAILED DESCRIPTION
[0013] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings. It should be understood that these descriptions are only exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, the description of well-known structures and technologies is omitted to avoid unnecessary confusion of the concept of the present invention.
[0014] The following is a detailed description with reference to the accompanying drawings.
[0015] Aiming at the deficiencies of the existing technology, based on the deep learning method, the present invention proposes a precipitation image estimation model based on a Multiscale Convolutional Attention Decoder (MCAD), aiming to utilize the remote sensing data of channels 9 to 15 of the FY-4B satellite to construct a high-precision precipitation image reconstruction method. This method adopts the architecture of a Multiscale Convolutional Attention Module (MCAM), combines a Multiscale Convolutional Module (MCB), a Channel Attention Module (CAB), and a Spatial Attention Module (SAB), fully excavates the precipitation information characteristics in the satellite observation data, and finally realizes the accurate characterization of the precipitation pattern, further improving the accuracy of precipitation reconstruction. The method of the present invention will be elaborated in detail below with reference to the accompanying drawings. Figure 1 It is the processing flow chart of the method of the present invention.
[0016] Step 1: Collect and organize the data set and perform preprocessing to construct a training set and a test set.
[0017] Step 11: Collect and organize the FY-4B satellite data and the GPM IMRGE precipitation product as the data set. Collect and organize the GPM IMERG Final Run data in the FY-4B satellite data and the GPM IMRGE precipitation product data from June 2022 to July 2024. The FY-4B data used is the data of channels 9 to 15 collected by the Geostationary Radiation Imager (AGRI) carried on it. Among them, channels 9-11 belong to the vertical water vapor detection layer. The brightness temperature data of this layer respectively reflect the high, medium, and low water vapor changes. By identifying the brightness temperature changes, the characteristics of upper-air moisture convergence can be identified. When sudden aggregation of upper-air moisture is found, it indicates that strong thunderstorms may be forming; channels 12-13 belong to the cloud top thermal characteristics layer. The data of this layer reflect the cloud top and surface temperature changes, directly measure the cloud top temperature, and quickly locate the area with a sudden drop in temperature; the data of channel 14 reflects the cloud top height, and the data of channel 15 reflects the cloud top height and water vapor. These two channels belong to the cloud physical property layer. In land areas where radar monitoring is not available (such as mountainous areas), by analyzing the cloud layer thickness and humidity, the blind spots of traditional monitoring means can be compensated. After processing the cloud map data of the 7 channels, its spatial resolution is 4 km, the time resolution is 15 minutes, the GPM IMERG spatial resolution is 0.1 degree, and the time resolution is 30 minutes.
[0018] Step 12: Data preprocessing. Perform brightness temperature calibration operation and geometric correction on the remote sensing data collected by the FY-4B satellite, and crop the geometrically corrected image according to the target area range, and finally generate 256×256 data blocks.
[0019] Figure 2It is the flowchart of the data preprocessing of the AGRI radiometric imager. Data preprocessing is a key step in using remote sensing satellite data for model training to ensure the accuracy and reliability of subsequent digital image processing. The method of the present invention mainly uses the remote sensing image data with a resolution of 4 km for the full disk of the AGRI L1 level on the FY-4B satellite. The flowchart of the AGRI data preprocessing is as Figure 2 shown. First, after downloading the satellite data of the AGRI L1 level with a resolution of 4 km for the full disk of FY-4B at the National Satellite Meteorological Center, calibration, projection, and longitude and latitude conversion operations are carried out to ensure that its data quality is consistent with the labeled data. Then, the GPM IMERG precipitation data downloaded from the NASA GPM official website is cropped to the same longitude and latitude area size as FY-4B. Finally, the two are aligned in time resolution for experimental use.
[0020] Generally speaking, remote sensing images may have geometric distortions due to various factors such as the imaging characteristics of the instrument itself, the angular differences during observation, or the deviation of the satellite orbit. Such distortions will cause the objects in the image to deviate from the actual situation in terms of size, position, and shape. Therefore, after performing the brightness temperature calibration operation on the multi-channel remote sensing data collected by Fengyun satellites, geometric correction work must also be carried out to achieve the precise matching of the remote sensing image with the actual geographical location.
[0021] Figure 3 is the schematic diagram of GLT geometric correction. To successfully complete the geometric correction task of remote sensing data, first, download the longitude and latitude lookup table (GLT file) that matches the original FY-4B L1 level disk data (resolution of 4 km) from the Fengyun satellite remote sensing data service network. Among them, the GLT file contains the row number (m) and column number (n) information corresponding to each pixel in the target projection coordinate system. Then, construct an initial image P with the same projection and size as GLT, Figure 3 where the upper left image P in it is a blank image with the same size as GLT, used to store the correction result. Then, parse the GLT file, read the row number m and column number n corresponding to each pixel respectively, and obtain two row-column mapping diagrams ( Figure 3 the upper image m and image n in it). For each pixel position in the image P, find the corresponding gray value in the original image Q by looking up the row-column coordinates (m, n) given by the GLT file. Then, assign the pixel value at the specified (m, n) position in the original image Q back to the current position of the initial image P to achieve the re-mapping of the pixel gray value. Finally, the corrected image P is obtained, as Figure 3As shown in the lower left, imaging distortion is eliminated to ensure accurate geolocation of each pixel. The corrected image P is then spatially cropped according to the target area range (18.05°N - 43.55°N, 97.05°E - 122.55°E). The cropped result forms a rectangular area of 256×256 pixels, corresponding to a resolution of 0.1°×0.1°. The final brightness temperature data is saved as a NetCDF (.nc) format file for use in subsequent precipitation estimation models.
[0022] In the invention, the GPM IMERG precipitation product is used as the true label value for network model training. Its temporal resolution is 30 minutes and its spatial resolution is 0.1°×0.1°. Therefore, the FY-4B single-frame cloud image data closest to the GPM half-hourly data is selected, and through bilinear interpolation, an input dataset with a temporal resolution of 30 minutes and a spatial resolution of 0.1°×0.1° is finally obtained.
[0023] Step 13: Divide the training set and the test set according to the set ratio.
[0024] The ratio of the training set to the test set is roughly set to 5:1. A total of 29,197 FY4B and GPM IMERG from June 2022 to July 2024 are selected as the training set. To verify the robustness of the network structure, an equal amount of data is randomly selected from each quarter during the same period, a total of 5,840 FY4B and GPM IMERG are selected as the test set.
[0025] Step 2: Construct a precipitation estimation model based on the multi-scale convolutional attention decoder MCAD. The precipitation estimation model uses the U-Net model as the basic framework to perform precipitation reconstruction tasks in an encoder-decoder architecture. The encoder extracts features through four convolutional layers and downsampling operations. The decoder replaces the traditional convolutional layer with a multi-scale convolutional attention module (MCAM). The multi-scale convolutional attention module MCAM includes a channel attention module CAB, a spatial attention module SAB, and a multi-scale convolutional module MCB connected in sequence. The decoder outputs the predicted precipitation image through four multi-scale convolutional attention modules MCAM. Replacing the traditional convolutional layer with the multi-scale convolutional attention module (MCAM) can achieve multi-level optimization of features, thereby improving the precipitation estimation accuracy and finally outputting the precipitation reconstruction image.
[0026] The overall structure of the MCAD network model is as Figure 4As shown in the figure. The innovation of the MCAD module lies in the multi-scale convolution module (MCB). Through parallel multi-scale depthwise separable convolutions (MDC), it captures precipitation information at different scales and enhances the model's perception of local and global precipitation patterns. The channel attention module (CAB) enhances the feature representation of key precipitation regions through a channel adaptive weight mechanism. The spatial attention module (SAB) extracts spatial information through average pooling and max pooling and optimizes the precipitation distribution pattern using large kernel convolutions, improving the model's ability to identify precipitation structures. The precipitation estimation model MCAD of the present invention fully combines multi-scale information extraction and attention mechanisms, enabling the model to more accurately depict precipitation patterns and improve the reliability and stability of precipitation estimation.
[0027] Step 3: Input the training set into the constructed precipitation estimation model for training. The training process specifically includes: Step 31: Input the remote sensing data of channels 9 to 15 of Fengyun satellite 4B in the training set into the encoder for feature extraction processing. The input data size is 256*256*7, where 7 represents 7 channels. The remote sensing data sequentially passes through four encoding layers for feature extraction, obtaining the first encoded feature, the second encoded feature, the third encoded feature, and the fourth encoded feature respectively.
[0028] Step 32: Input the four encoded features into the corresponding decoding layers for decoding. Specifically, the first encoded feature is input into the first decoding layer, the second encoded feature is input into the second decoding layer, the third encoded feature is input into the third decoding layer, and the fourth encoded feature is input into the fourth decoding layer. Each decoding layer includes a multi-scale convolutional attention module MCAM and an upsampling layer.
[0029] Among them, the multi-scale convolutional attention module MCAM first performs feature screening through the channel attention module CAB and the spatial attention module SAB, and then extracts the final precipitation information through the multi-scale convolution module MCB. The processing flow of MCAM specifically includes.
[0030] Step 321: The decoded feature first passes through the channel attention module CAB. CAB includes two branches, the adaptive max pooling branch and the adaptive average pooling branch. Among them, adaptive max pooling is used to extract the most significant responses in each channel; adaptive average pooling captures the global statistical features of each channel; the features output by both are fused after subsequent convolution and ReLU activation, and a channel weight is generated through the Sigmoid function. Finally, it is fused with the original feature in an element-wise weighted form, and the fused feature is output to the spatial attention module SAB.
[0031] Step 322: The Spatial Attention Module (SAB) extracts key spatial features through operations of obtaining the channel maximum value and the channel average value respectively. First, SAB performs max pooling on the input fused features in the channel dimension to obtain a single-channel maximum value feature map. Meanwhile, SAB performs average pooling on the input fused features in the channel dimension to calculate the global mean of each channel and obtains a single-channel average value feature map. The maximum value feature map and the average value feature map are fused, followed by a large convolution operation and a Sigmoid activation function to output the spatial attention features.
[0032] Step 323: The spatial attention features are input into the Multi-scale Convolution Module (MCB). MCB uses Multi-scale Depthwise Separable Convolution (MDC) to extract features with different receptive fields. MDC adopts three convolution kernels with different scales. The small convolution kernel is used to capture local texture information, the medium convolution kernel is used to identify medium-scale cloud systems, and the large convolution kernel is used to model the large-scale cloud layer distribution. Then, the feature maps under different-scale convolution kernels are respectively passed through batch normalization and ReLU6 activation functions. Next, the features extracted by different-scale depthwise separable convolutions are fused through element-wise addition. The fused features are subjected to a channel shuffle operation.
[0033] Step 33: The fourth encoded feature is passed through the fourth decoding layer to obtain the fourth decoded feature. After fusing the third encoded feature, the fourth decoded feature is input into the third decoding layer to output the third decoded feature. After fusing the second encoded feature, the third decoded feature is input into the second decoding layer to output the second decoded feature. After fusing the first encoded feature, the second decoded feature is input into the first decoding layer to output the first decoded feature, thus obtaining the predicted precipitation image. Step 34: Calculate the loss function between the predicted precipitation image and the true label image of the GPM IMERG precipitation product in the training set. Adjust the parameters of the precipitation estimation model through the feedback of the loss value until the loss value is less than the preset threshold or reaches the preset number of iterations, then the training stops and the trained model is saved.
[0034] Step 4: Input the test set into the trained precipitation estimation model for testing.
[0035] Figure 5 is the specific structural schematic diagram of MCAM. The following combines Figure 5 to describe in detail the specific processing flow of the multi-scale convolutional attention module MCAM in step 32.
[0036] The multi-scale convolutional attention module MCAM includes three sub-modules: the multi-scale convolution module (MCB), the channel attention module (CAB), and the spatial attention module (SAB). Through multi-scale feature fusion and attention mechanism, it enhances the model's focusing ability on key regions and is suitable for associative modeling of complex cloud system structures and precipitation intensities in precipitation estimation. The structure of MCAM is as Figure 5As shown on the upper left, the inside of the module adopts a cascaded method. First, feature screening is carried out through CAB and SAB, and then the final precipitation information is extracted through the multi-scale convolution module MCB. Next, the three internal modules will be introduced one by one.
[0037] The Channel Attention Block (CAB) realizes the saliency modeling in the channel dimension by constructing a "dual-path pooling structure (max pooling and average pooling)", guiding the model to dynamically judge the importance of different channels, enhancing the feature expression of precipitation-related channels, and thus improving the accuracy of precipitation estimation. Its structure is as Figure 5 shown on the lower left.
[0038] For the 7 infrared channels of FY-4B, their contributions to the precipitation reconstruction task are different: channels 9 to 11 are water vapor sensitive channels, reflecting the humidity changes in the high, middle, and low altitudes, and playing a key indicating role in the formation process of heavy precipitation; channels 12 to 13 mainly characterize the cloud top brightness temperature and thermal characteristics; channels 14 to 15 reveal the cloud layer structure and thickness information.
[0039] CAB performs global feature extraction in the channel dimension through the dual-pooling strategy of adaptive average-max pooling, effectively suppressing the noise features from interference channels such as surface thermal radiation and improving the channel discriminability. Among them, adaptive max pooling is used to extract the most significant responses in each channel, highlighting the local extrema in the precipitation features and enhancing the model's response ability to heavy precipitation areas; adaptive average pooling captures the global statistical features of each channel, helping the model understand the large-scale precipitation pattern. The features output by both are fused after subsequent convolution and ReLU activation, and the channel weights are generated through the Sigmoid function. Finally, they are multiplied with the original features in an element-wise weighted form to achieve the enhancement of key channels and the suppression of non-key channels. CAB helps the model reduce redundant features, improve the accuracy and computational efficiency of feature representation, and enhance the response speed in the inference stage.
[0040] CAB extracts the global saliency response and statistical mean of each channel by performing max pooling and average pooling operations on the input feature map in the spatial dimension respectively, and then concatenates the two and inputs them into a shared fully connected network for non-linear transformation to generate the weight vector corresponding to each channel. This process is the "dynamic calculation of channel weights", which reflects that the model adaptively calculates the contribution degree of each channel according to the input image, rather than a fixed setting.
[0041] In addition, during network training, the water vapor sensitive channels (9-11) usually exhibit high response values in precipitation regions. Therefore, during the max / average pooling stage, they naturally occupy higher weights in statistical features and obtain larger attention scores after Sigmoid mapping. During backpropagation, this module automatically strengthens the channel features that contribute to precipitation prediction, thereby achieving focused enhancement and selective highlighting of channels 9-11. Finally, all channel weights are applied to the input feature map through element-wise multiplication to achieve channel selection and redundancy suppression, improving the generalization ability and inference efficiency of the model.
[0042] The Spatial Attention Block (SAB) draws on the multi-axis spatial attention mechanism, aiming to dynamically identify and strengthen the key precipitation regions in the input feature map, effectively improving the recognition ability of the precipitation spatial distribution. Its structure is as Figure 5 shown on the upper right side above.
[0043] SAB extracts key spatial feature information by performing max pooling and average pooling operations on the input feature map in the channel dimension. Specifically, SAB first performs max pooling and average pooling on the activation values of all channels at each spatial position respectively to generate two single-channel spatial feature maps, highlighting the significant response regions and the overall spatial response trend respectively. Max pooling emphasizes the strong activation regions, highlighting the significant features of strong precipitation regions, increasing the attention to high precipitation intensity regions, retaining the most critical parts of the cloud structure, helping to distinguish the morphology of precipitation cloud clusters, and highlighting the local high-brightness regions in the cloud image, such as the center of strong convective cloud clusters, corresponding to short-term heavy precipitation; average pooling retains the overall precipitation distribution trend, enabling the model to focus on the overall precipitation distribution trend, being able to identify both local precipitation intensity and large-scale precipitation systems, and at the same time reflecting the overall cloud distribution trend, such as the wide coverage of stratiform clouds, associated with continuous precipitation.
[0044] Subsequently, the two are concatenated and input into a large kernel convolution to capture local and global spatial context, and a spatial attention map is generated through Sigmoid activation. This spatial attention map is then multiplied element-wise with the input feature map to complete the "re-calibration" operation on the original feature map: that is, assigning higher response values to the core precipitation regions while suppressing the background interference regions, enabling the model to focus on the spatially more physically meaningful structural regions, thereby improving the model's perception ability of complex precipitation structures, enhancing the attention to high-intensity regions, reducing misjudgments in non-precipitation regions, and improving the robustness of spatial feature expression.
[0045] Multiscale Convolutional Block (MCB). The traditional convolutional layer synchronously extracts spatial and channel features through a multi-channel convolutional kernel. Although it can capture rich feature information, its large number of parameters leads to a significant consumption of computing resources. To solve this problem, depthwise separable convolution decomposes the operation process into two steps. First, it independently performs spatial convolution on each input channel (depthwise convolution), and then integrates cross-channel information through 1×1 convolution (pointwise convolution). This design compresses the parameter scale to 1 / 8 of the traditional method, significantly reducing the computational amount while retaining the independence of multi-channel information.
[0046] MCB is based on the inverted residual block (IRB) of MobileNetV2 and uses multiscale depthwise separable convolution (MDC) instead of traditional convolution to be responsible for extracting precipitation information at different scales from the input features, so as to reduce the computational amount and enhance the expressive ability of the model. Its structure is as Figure 5 shown on the lower right side below.
[0047] Among them, the parallel multiscale depthwise separable convolution MDC can extract features under different receptive fields. Its structure is as Figure 6 shown. The small kernel (DWC1×1) is used to capture local textures, such as the edges of cloud clusters. The medium kernel (DWC3×3) is used to identify medium-scale cloud systems, such as cumulonimbus clusters. The large kernel (DWC5×5) can be used to model large-scale cloud distributions, such as frontal systems. Therefore, the model can simultaneously capture the local bright cloud tops corresponding to short-term heavy precipitation in small-scale features and the wide-area cloud cover of continuous precipitation in large-scale features. Then, the feature maps at different scales are respectively passed through normalization and the improved activation function ReLU6 to reduce the computational complexity and make the model operation performance more efficient. Then, the features extracted by each scale of the depthwise separable convolution MDC branches are fused through element-wise addition to ensure the effective integration of information at different scales and improve the model's perception ability of multi-scale precipitation features. However, the parallel structure of the multi-scale branches may lead to an isolation effect of channel information before fusion. To enhance the feature interaction between scales and enable more sufficient fusion of information at different scales, the present invention introduces a channel shuffle operation: first, the fused feature map is divided into multiple subgroups along the channel dimension, then the internal channels of each group are transposed and rearranged, and then recombined into a complete feature map, so as to realize the information mixing and intercommunication between channels at different scales. This operation has a simple and efficient calculation process and does not require new parameters, which helps to improve the model's perception ability and generalization ability of the spatial structure of precipitation cloud clusters, helps the model to more comprehensively understand the spatial structure of precipitation cloud clusters, and finally enhances the accuracy of precipitation estimation.
[0048] Experimental environment setup. The initial weights of the multi-scale convolutional attention decoder (MCAD) proposed in the present invention follow a normal distribution, and the AdamW optimizer is used to optimize its parameters to minimize the loss function of the model. The loss function adopts the mean squared error (MSE) loss. The initial learning rate is set to 0.0002, and an adaptive algorithm is used to dynamically adjust the learning rate. The Patience is set to 5 and the Factor is set to 0.5, that is, after 5 rounds of training, if the loss does not change significantly, the learning rate is halved. The training batch size is 8 and the number of training epochs is 100.
[0049] The evaluation metrics used in the present invention to test the experimental results are the root mean square error (RMSE) and the Pearson correlation coefficient (CC).
[0050] Root mean square error (RMSE). RMSE is used to calculate the error magnitude between the real data and the estimated data. In the present invention, it is used to evaluate the error between the reconstructed precipitation of the model and the label value GPM IMERG data, with the unit of millimeters (mm). The value range of RMSE is [0, +∞). The smaller the RMSE value, the smaller the error and the better the model performance; conversely, the worse the model performance.
[0051] Pearson correlation coefficient (CC). CC is used to reflect the correlation degree between the reconstructed precipitation value and the label value GPM IMERG data. The value range of CC is from -1 to 1, where a negative value represents negative correlation and a positive value represents positive correlation. The larger the absolute value of the numerical value, the stronger the correlation.
[0052] The method proposed in the present invention is compared with the prior art methods, and the experimental results are shown in Table 1. Taking the internationally most advanced precipitation product GPM IMERG as the ground truth, the algorithm proposed in the project is quantitatively compared with the prior deep learning methods and reanalysis products under the same conditions of time, space range, resolution, etc. Among them, the reanalysis data CFSv2 (Climate Forecast System Version 2) is a climate and weather reanalysis and prediction system launched by the National Centers for Environmental Prediction (NCEP) of the United States, and is widely used in the fields of global meteorology and climate research. CFSv2 realizes seamless extension in the time range, can provide reanalysis data and real-time climate prediction data simultaneously, its data resolution is further optimized, and the prediction accuracy of key meteorological elements such as temperature, precipitation and wind speed is improved. The FY-4A QPE (Quantitative Precipitation Estimation) product of Fengyun-4A satellite is developed based on the data collected by the Fengyun-4A geostationary meteorological satellite in China. This product is mainly dedicated to monitoring and analyzing the precipitation distribution and precipitation intensity in China and its surrounding areas. The FY-4A QPE product integrates the infrared and visible light data obtained by the multi-channel scanning imaging radiometer (AGRI), combines the ground precipitation observation data and meteorological models, and uses advanced algorithms to achieve high-precision estimation of precipitation.
[0053] As can be seen from Table 1, the precipitation estimation capabilities of all deep learning models are superior to those of CFSv2 and FY-4A QPE. Among them, the MCAD model proposed in this invention performs the best. In terms of RMSE, the error of the MCAD model is the lowest, only 0.4662 mm, which is 35.14% lower than that of CFSv2, indicating that its precipitation estimation error is the smallest and it can more accurately restore the spatial distribution of precipitation. The U-Net method and the Trans-UNet method also show strong precipitation reconstruction capabilities, with RMSEs dropping to 0.5142 mm and 0.5029 mm respectively, and the improvement amplitudes compared to CFSv2 being 28.46% and 30.04% respectively. The RMSE values of the ARM-cGAN method and FY-4A QPE are 0.5833 mm and 0.6170 mm respectively. Although both are superior to CFSv2, the errors are still relatively large, indicating that there are certain limitations in their precipitation estimation. In terms of the CC index, all deep learning models have significantly improved the correlation between precipitation estimation and the reference data. Among them, the CC value of the MCAD model reaches 0.8324, with an improvement amplitude of 59.74%, indicating that it has the strongest fitting ability for the spatial distribution of precipitation. The CC values of Trans-UNet and U-Net reach 0.7995 and 0.7850 respectively, showing relatively close performance, and both are significantly superior to ARM-cGAN (0.7164) and FY-4A QPE (0.6747). In contrast, the CC value of CFSv2 is the lowest, only 0.5211, indicating that the correlation between its precipitation estimation result and the actual precipitation is weak and its performance is poor. The above experimental data prove the effectiveness and superiority of the proposed method. Among them, RMSE-Gain and CC-Gain refer to the improvement degrees of the corresponding objective evaluation indicators.
[0054] Table 1 Quantitative Comparison of Experimental Results
[0055] It should be noted that the above specific embodiments are exemplary. Those skilled in the art can come up with various solutions inspired by the disclosed content of this invention, and these solutions also belong to the disclosure scope of this invention and fall within the protection scope of this invention. Those skilled in the art should understand that the specification and drawings of this invention are illustrative and do not constitute a limitation on the claims. The protection scope of this invention is defined by the claims and their equivalents.
Claims
1. A precipitation image estimation method based on a multi-scale convolutional attention decoder, characterized in that, Construct a precipitation estimation model based on the multi-scale convolutional attention decoder MCAD. The decoder MCAD extracts precipitation features, decodes the features extracted by the encoder, and outputs the estimated precipitation image. MCAD adopts the architecture of the multi-scale convolutional attention module MCAM. MCAD combines multi-scale information extraction and attention mechanism to fully exploit the precipitation information features in satellite remote sensing data, specifically including: Step 1: Collect and organize the dataset and perform preprocessing to construct the training set and test set; Step 11: Collect and organize FY-4B satellite data and GPM IMRGE precipitation products as the dataset, where the GPM IMRGE precipitation product is used as the label data; Step 12: Data preprocessing. Perform brightness temperature calibration and geometric correction on the remote sensing data collected by FY-4B satellite. Crop the geometrically corrected remote sensing data according to the target area range to obtain 256×256 data blocks; Step 13: Divide the training set and test set according to a preset ratio; Step 2: Construct a precipitation image estimation model based on the multi-scale convolutional attention decoder MCAD. The precipitation estimation model uses the U-Net model as the basic framework and performs precipitation reconstruction tasks based on the encoding-decoding architecture. The encoder extracts features through four convolutional layers and downsampling operations. The decoder replaces the traditional convolutional layer with the multi-scale convolutional attention module MCAM. The multi-scale convolutional attention module MCAM includes a channel attention module CAB, a spatial attention module SAB, and a multi-scale convolutional module MCB connected in sequence. The decoder outputs the predicted precipitation image through four multi-scale convolutional attention modules MCAM; Step 3: Input the training set into the constructed precipitation estimation model for training. The training process specifically includes: Step 31: Input the remote sensing data of channels 9 to 15 of FY-4B in the training set into the encoder for feature extraction processing. The remote sensing data sequentially passes through four encoding layers for feature extraction to obtain the first encoding feature, the second encoding feature, the third encoding feature, and the fourth encoding feature respectively; Step 32: Input the four encoding features into the corresponding decoding layers for decoding. Specifically, the first encoding feature is input into the first decoding layer, the second encoding feature is input into the second decoding layer, the third encoding feature is input into the third decoding layer, and the fourth encoding feature is input into the fourth decoding layer. Each decoding layer includes a multi-scale convolutional attention module MCAM and an upsampling layer; Step 33: The fourth encoding feature passes through the fourth decoding layer to obtain the fourth decoding feature. The fourth decoding feature is fused with the third encoding feature and then input into the third decoding layer to output the third decoding feature. The third decoding feature is fused with the second encoding feature and then input into the second decoding layer to output the second decoding feature. The second decoding feature is fused with the first encoding feature and then input into the first decoding layer to output the first decoding feature, obtaining the predicted precipitation image; Step 34: Calculate the loss function between the predicted precipitation image and the true label image of the GPM IMERG precipitation product in the training set, adjust the parameters of the precipitation estimation model according to the loss value until the loss value is less than the preset threshold or reaches the preset number of iterations, stop the training, and save the trained model; Step 4: Input the test set into the trained precipitation estimation model for testing.
2. The precipitation image estimation method according to claim 1, characterized in that, In the multi-scale convolutional attention module MCAM in Step 32, feature screening is first performed through CAB and SAB, and then the final precipitation information is extracted through the multi-scale convolutional module MCB. The processing of MCAM specifically includes: Step 321: The decoded features first pass through the channel attention module CAB. CAB includes two branches, the adaptive max pooling branch and the adaptive average pooling branch. Among them, adaptive max pooling is used to extract the most significant responses in each channel, and adaptive average pooling captures the global statistical features of each channel; the features output by both are fused after subsequent convolution and ReLU activation, and channel weights are generated through the Sigmoid function. Finally, they are fused with the original features in an element-wise weighted form, and the fused features are output to the spatial attention module SAB; Step 322: The spatial attention module SAB extracts key spatial features by obtaining the channel maximum value and channel average value operations respectively. SAB first performs max pooling on the input fused features in the channel dimension to obtain a single-channel maximum value feature map. At the same time, SAB performs average pooling on the input fused features in the channel dimension to calculate the global mean of each channel and obtain a single-channel average value feature map. The maximum value feature map and the average value feature map are fused, and then a large convolution operation and the Sigmoid activation function are performed to output the spatial attention features; Step 323: The spatial attention features are input into the multi-scale convolutional module MCB. MCB uses multi-scale depthwise separable convolution MDC to extract features with different receptive fields. MDC uses three convolutional kernels with different scales. The small convolutional kernel is used to capture local texture information, the medium convolutional kernel is used to identify medium-scale cloud systems, and the large convolutional kernel is used to model the large-scale cloud distribution. Then, the feature maps under different-scale convolutional kernels pass through batch normalization and the ReLU6 activation function respectively. Next, the features extracted by the depthwise separable convolutions with different scales are fused by element-wise addition. The fused features undergo a channel shuffle operation.
Citation Information
Patent Citations
Rainfall prediction method based on convolution
CN116307267A
Mode rainfall forecast correction method based on deep learning and wind cloud satellite
CN116840941A
Deep learning-based short-time quantitative rainfall numerical prediction correction method
CN118153628A
High-resolution remote sensing image change detection method based on multi-scale convolution decoding, electronic equipment and storage medium
CN119418216A
Cited By
Scanning microwave radiometer precipitation inversion method based on weight attention mechanism
CN121305393A