Cloud type inversion method and system based on cavity window attention
Through the hole window attention model, the shortcomings of traditional methods in processing multi-scale cloud structures are solved, high-precision cloud type inversion is achieved, and feature extraction and classification capabilities are enhanced, which is suitable for the identification of complex cloud scenes.
Patent Information
- Application Number
- CN202510970699.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-09-23
AI Technical Summary
Traditional cloud type inversion methods have difficulty handling multi-scale cloud structures, leading to misjudgments in complex scenes. In addition, deep learning models have shortcomings in feature fusion, which affects the accuracy and reliability of cloud type inversion.
A cloud type inversion method based on hole window attention is adopted. Through the hole window partition module and the weighted attention recovery module, multi-dimensional spectral image features and geographic features are combined to extract multi-scale features and perform feature fusion. The encoder-decoder structure is used for pixel-level classification.
It improves the multi-scale feature extraction capability, enhances global perception and local detail retention, integrates geospatial information, improves the accuracy and robustness of cloud type classification, and supports efficient computing and end-to-end pixel-level classification.
Smart Images

Figure CN120689769A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cloud type inversion of remote sensing images, and in particular relates to a cloud type inversion method and system based on hole window attention. Background Art
[0002] In the field of remote sensing image recognition, cloud type inversion methods based on satellite cloud imagery have long been a hot topic of interest and research for remote sensing researchers. Traditional cloud type inversion methods primarily rely on spectral and textural features, analyzing spectral features such as reflectance and brightness temperature in different bands, combined with texture analysis (such as the gray-level co-occurrence matrix), to distinguish different cloud types. In recent years, machine learning has been widely applied to cloud type inversion. Through feature extraction and pattern recognition, it can better handle complex cloud scenes. Deep learning has further improved the accuracy of cloud type inversion.
[0003] However, traditional methods typically only extract features at a single scale and struggle to process multi-scale cloud structures. They are prone to misjudgment when handling complex scenes (such as thin clouds, translucent clouds, and cloud and snow distinctions). Some deep learning models lack the ability to fuse features, resulting in loss of detail when restoring image resolution and a failure to fully utilize multi-band information for feature extraction, which compromises the reliability and accuracy of cloud type inversion results. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a cloud type inversion model and system based on hole window attention, which can improve the ability of cloud type inversion.
[0005] To achieve the above object, the present invention provides a cloud type inversion method based on hole window attention, comprising the following steps:
[0006] (1) Obtain satellite remote sensing image data, such as Landsat 8 and FY-4A, perform preprocessing operations such as normalization, data annotation, and data enhancement, and divide the data into training and test sets;
[0007] (2) For the preprocessed data, the multidimensional spectral image features are extracted, and the image latitude and longitude information is converted into spatial features through the position encoding module. The two are combined to obtain fusion features, which are input into the convolution layer to extract basic features such as image edges and textures;
[0008] (3) Stacking multiple hole window attention (WDWA) modules, each WDWA module contains the hole window partitioning module (ODWP) and weighted attention recovery module (WDAR) operations, introducing the information of neighboring local windows, balancing the attention distribution, and obtaining multi-scale feature representation;
[0009] (4) The decoder part gradually upsamples the low-resolution features output by the encoder to the same spatial resolution as the input image through feature upsampling, and finally maps the features to the category space through the convolution layer, outputting the classification probability of each pixel;
[0010] (5) Use the training set to train the network model, select a suitable optimizer, adopt a learning rate decay strategy, calculate the loss function through forward propagation, and then update the network parameters through back propagation, and use the test data set to learn and train the model.
[0011] Furthermore, step (1) includes the following steps:
[0012] (1.1) Based on the study area and time range, select appropriate satellite remote sensing images, such as Landsat 8 and FY-4A, and download remote sensing image data from the official website;
[0013] (1.2) Perform pixel normalization on remote sensing image data, scaling the original image pixel intensity values to between 0 and 1. After normalization, accurately identify and label the cloud areas in the image, clearly distinguishing between cloud and non-cloud areas. For the identified clouds, assign different category labels based on their morphology, texture, and other characteristics, thereby forming high-quality training labels.
[0014] (1.3) Perform data augmentation operations on the labeled training data, including random rotation, flipping, cropping, and scaling, to artificially increase data diversity and simulate various possible image changes. Random rotation simulates the effects of images taken at different angles; flipping enhances the model's ability to adapt to left-right or top-bottom symmetry; cropping focuses on local areas in the image to better capture detailed features; and scaling helps the model adapt to image data at different resolutions.
[0015] (1.4) Based on the size of the dataset, the dataset is divided into a 70% training set and a 30% test set. The division ratio mainly considers the adequacy of model training, the reliability of model evaluation, the representativeness of the dataset, and the balance between computing resources and time costs.
[0016] Furthermore, step (2) includes the following steps:
[0017] (2.1) Perform normalization, denoising and other preprocessing operations on the multidimensional spectral image, and extract image features from the image. The normalization formula is as follows:
[0018]
[0019] Where x is the original data, is the normalized data, and are the minimum and maximum values of the data respectively;
[0020] (2.2) The latitude and longitude information of each remote sensing image is obtained through the geographic coordinate system, and the sine and cosine functions are used for position encoding to provide a unique vector representation for each position in the sequence. The specific sine and cosine encoding formulas are as follows:
[0021] For each dimension 2i, the positional encoding PE is defined as:
[0022]
[0023] For each dimension 2i+1, the positional encoding PE is defined as:
[0024]
[0025] Among them, t represents the position index, i represents the dimension index, and d model Represents the dimension of the model embedding vector, the sine function is applied to all even dimensions, indexed as 2i; and the cosine function is applied to all odd dimensions, indexed as 2i+1;
[0026] (2.3) Use principal component analysis to reduce the dimensionality of spectral features, or use linear transformation to increase the dimensionality of geographic features, ensuring that the dimensions of spectral features and geographic features remain consistent before concatenating them;
[0027] (2.4) Use a fully connected layer or linear transformation to map the spectral feature vector and geographic feature vector to this unified dimension to obtain the adjusted feature vector;
[0028] (2.5) The adjusted spectral feature vector and geographic feature vector are concatenated along the channel dimension to obtain the final fusion feature;
[0029] (2.6) Build a 3-layer convolutional network and input the fused features into the convolution layer to achieve downsampling and extract image features such as edges and textures.
[0030] Furthermore, step (3) includes the following steps:
[0031] (3.1) Multiple hole window attention modules are stacked. In each hole window attention module, the input features are first processed by the hole window partitioning module to split the input features into local windows with a fixed window size and expansion rate r. The number of local windows is:
[0032]
[0033] Where H and W are the width and height of the feature map, h and w are the height and width of the local window, h ≥ 1 and w ≥ 1, ensuring a minimum 1 × 1 pixel window. and It must be an integer, otherwise the window cannot be evenly divided, resulting in incomplete coverage of the feature map. The dilation rate r should be a positive integer. When r=0, it is equivalent to no dilation; when r=1, the dilated window degenerates into a regular sliding window (step size = 1). The dilated window slides in four directions: upper left, upper right, lower left, and lower right to cover all areas of the input feature; if the dilated window exceeds the boundary of the input feature, reflection padding is used to ensure that all local windows have the same size;
[0034] (3.2) Each local window Embedded into query (Q), key (K) and value (V) through a linear layer, the expression is as follows:
[0035]
[0036]
[0037]
[0038] in, 、 and is the weight matrix, 、 and is the bias term;
[0039] Q and K are used to calculate the attention weights in multi-head self-attention:
[0040]
[0041] in, is the dimension of K, The function normalizes the attention weights into a probability distribution; the attention weights are used to perform a weighted summation on V:
[0042]
[0043] in, is the output feature of the i-th local window; the output features of all local windows are concatenated into the final output feature map:
[0044]
[0045] The output features provide richer feature representations for subsequent analysis tasks;
[0046] (3.3) The output of the hole window partitioning module is passed to the weighted attention restoration module, which performs an inverse operation on the output misaligned local attention window to restore it to the same shape as the input feature, aligns the misaligned window group, and fuses the attention using a weighted summation method, where the weight is calculated based on the number of overlaps, balancing the attention of overlapping and non-overlapping areas, building long-range relationships, and enhancing the model's global perception ability;
[0047] (3.4) The output of each hole window attention module is used as the input of the next hole window attention module to obtain multi-scale feature representations and enhance the understanding and expression capabilities of the input features;
[0048] (3.5) After being processed by multiple hole window attention modules, the final output features contain rich multi-scale information and long-distance dependencies, which serve as subsequent input.
[0049] Furthermore, step (4) includes the following steps:
[0050] (4.1) The decoder uses bilinear interpolation for upsampling, generating new pixel values by calculating the linear combination of adjacent pixels, smoothly enlarging the image, and gradually upsampling the low-resolution features output by the encoder to the same spatial resolution as the input image;
[0051] (4.2) After each upsampling, the resolution of the feature image will increase, but the semantic information of the feature may be lost. It is necessary to ensure that the spatial resolution of the upsampled features is consistent with that of the encoder part, and to perform feature splicing in the channel dimension;
[0052] (4.3) Use a 3×3 convolution kernel to perform a convolution operation on the fused features to extract a higher-level feature representation. The output features of the convolution layer will serve as the input for the next upsampling.
[0053] (4.4) A 1×1 convolution kernel is used to reduce the number of channels of the feature map to the number of categories, and the feature map is mapped to the category space. The softmax activation function is used to map the output features to the classification probability distribution. The output features of each pixel are normalized to a probability value, indicating the probability that the pixel belongs to each category. The category label of each pixel can be determined by taking the category with the highest probability value.
[0054] Furthermore, step (5) includes the following steps:
[0055] (5.1) Model Training: The network model was trained using the preprocessed training data. The initial learning rate was set to 0.002, the batch size was 32, the number of training epochs was set to 100, and the Adam optimizer was used, which generally performs well due to its adaptive learning rate and momentum properties. The cross entropy loss function was used, and the network parameters were optimized using the backpropagation algorithm.
[0056] (5.2) Model Validation: The model is trained on a test dataset. The test dataset is input into the model and the model’s prediction results are obtained through forward propagation. Evaluation metrics such as accuracy, recall, precision, and F1 score are used to measure the model’s performance in inverting cloud types.
[0057] The present invention also discloses a cloud type inversion system based on hole window attention, comprising:
[0058] Data preprocessing module: responsible for satellite data download, normalization, cloud category labeling, data enhancement and data set division;
[0059] Feature encoding module: includes spectral feature extraction (normalization + PCA dimensionality reduction), longitude and latitude sine and cosine position encoding, feature dimension unification, and a three-layer convolutional network to achieve spectral and spatial feature fusion and basic feature extraction;
[0060] Hole Window Attention Module: It consists of ODWP and WDAR. ODWP implements multi-directional hole window partitioning and multi-head self-attention calculation, while WDAR balances regional attention through inverse operations and weight calculation.
[0061] Decoder module: achieves resolution restoration and pixel classification through bilinear interpolation upsampling, feature splicing, 3×3 convolution and 1×1 convolution-softmax;
[0062] Model training and verification module: Integrates the Adam optimizer, cross entropy loss function, and multi-indicator evaluation system to complete model training and performance verification.
[0063] Furthermore, in the above-mentioned hole window attention module, WDAR calculates the weight by the number of overlaps, with the formula: reverse attention value = attention value / number of overlaps, to balance the attention of different regions.
[0064] Furthermore, in the above feature encoding module, the position encoding dimension d model It is kept consistent with the spectral feature dimension through linear transformation to ensure the effectiveness of feature splicing.
[0065] Compared with the prior art, the present invention has the following beneficial effects:
[0066] 1. Improve multi-scale feature extraction capabilities
[0067] Through the hole window partitioning module (ODWP) combined with the dynamic adjustment of the dilation rate (r), cloud structures of different scales (such as thin clouds, thick clouds, and broken clouds) in remote sensing images can be effectively captured, overcoming the limitations of single-scale features of traditional methods.
[0068] 2. Enhance global perception and local detail preservation
[0069] The weighted attention recovery module (WDAR) balances the attention distribution of the local window and the global context through inverse operations and overlapping area weight calculations, avoiding information loss during feature fusion and significantly improving the classification accuracy of complex cloud scenes (such as cloud-snow confusion areas).
[0070] 3. Integrate geospatial information
[0071] Sinusoidal and cosine position encoding is used to convert longitude and latitude information into spatial features and then spliced with spectral features, which solves the problem of traditional methods ignoring geographic spatial correlation. It is particularly suitable for cross-regional cloud classification of large-scale remote sensing images.
[0072] 4. Efficient computing and low resource consumption
[0073] The sliding mechanism and reflection filling strategy of the hole window reduce redundant calculations, reduce the number of parameters and improve the inference speed compared with the ordinary sliding window or full attention mechanism.
[0074] 5. Strong robustness
[0075] Through data enhancement (rotation, flipping, reflection filling) and dynamic weight fusion, it has stronger adaptability to image noise, partial occlusion and boundary areas, and improves the F1 score.
[0076] 6. End-to-end pixel-level classification
[0077] The encoder-decoder structure combined with multi-scale feature upsampling achieves high-precision pixel-level cloud classification and supports the output of probability distribution maps (such as cloud / non-cloud, cloud type), which directly serves meteorological analysis or surface monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] Figure 1 This is a network structure diagram of a cloud type inversion model based on hole window attention provided in Example 1 of the present invention;
[0079] Figure 2 This is a structural diagram of a hole window attention module in a cloud type inversion model based on hole window attention provided in Example 1 of the present invention. DETAILED DESCRIPTION
[0080] The present invention will be further described below in conjunction with the accompanying drawings. The following examples are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention.
[0081] Example 1:
[0082] This embodiment discloses a cloud type inversion model based on hole window attention, please refer to Figure 1 and Figure 2 , the specific training process includes the following steps:
[0083] Step (1) Obtain satellite remote sensing image data, such as Landsat 8, FY-4A and other satellite remote sensing image data, perform preprocessing operations such as normalization, data annotation, and data enhancement, and divide the data into training and test sets.
[0084] Step (1) includes the following steps:
[0085] (1.1) Select appropriate satellite remote sensing images based on the study area and time range, such as Landsat 8, FY-4A, etc., and download remote sensing image data from the official website;
[0086] (1.2) Perform pixel normalization on remote sensing image data, scaling the original image pixel intensity values to between 0 and 1. After normalization, accurately identify and label cloud areas in the image, clearly distinguishing between cloud and non-cloud areas. For identified clouds, assign different category labels (such as stratus, cumulus, cirrus, etc.) based on their morphology, texture, and other characteristics, thereby forming high-quality training labels.
[0087] (1.3) Perform data augmentation operations on the labeled training data, including random rotation, flipping, cropping, and scaling, to artificially increase the diversity of the data and simulate various possible image changes. Random rotation simulates the effects of images taken from different angles; flipping enhances the model's ability to adapt to left-right or top-bottom symmetry; cropping focuses on local areas in the image to better capture detailed features; and scaling helps the model adapt to image data at different resolutions.
[0088] (1.4) Based on the size of the dataset, the dataset is divided into a 70% training set and a 30% test set. The division ratio mainly considers the adequacy of model training, the reliability of model evaluation, the representativeness of the dataset, and the balance between computing resources and time costs.
[0089] Step (2) extracts multidimensional spectral image features from the preprocessed data, and converts the image latitude and longitude information into spatial features through the position encoding module. The two are combined to obtain fusion features, which are then input into the convolution layer to extract basic features such as image edges and textures.
[0090] Step (2) includes the following steps:
[0091] (2.1) Perform normalization, denoising and other preprocessing operations on the multidimensional spectral image, and extract image features from the image. The normalization formula is as follows:
[0092]
[0093] Where x is the original data, is the normalized data, and are the minimum and maximum values of the data respectively;
[0094] (2.2) The latitude and longitude information of each remote sensing image is obtained through the geographic coordinate system, and the sine and cosine functions are used for position encoding to provide a unique vector representation for each position in the sequence. The specific sine and cosine encoding formulas are as follows:
[0095] For each dimension 2i, the positional encoding PE is defined as:
[0096]
[0097] For each dimension 2i+1, the positional encoding PE is defined as:
[0098]
[0099] Among them, t represents the position index, i represents the dimension index, and d model Represents the dimension of the model embedding vector, the sine function is applied to all even dimensions (index 2i), and the cosine function is applied to all odd dimensions (index 2i+1);
[0100] (2.3) Use methods such as principal component analysis (PCA) to reduce the dimensionality of spectral features, or use methods such as linear transformation to increase the dimensionality of geographic features, ensuring that the dimensions of spectral features and geographic features are consistent before splicing them;
[0101] (2.4) Use a fully connected layer (or linear transformation) to map the spectral feature vector and the geographic feature vector to this unified dimension to obtain the adjusted feature vector;
[0102] (2.5) The adjusted spectral feature vector and geographic feature vector are concatenated along the channel dimension to obtain the final fusion feature;
[0103] (2.6) Build a 3-layer convolutional network and input the fused features into the convolution layer to achieve downsampling and extract image features such as edges and textures.
[0104] Step (3) stack multiple hole window attention (WDWA) modules in the encoder. Each WDWA module contains the hole window partitioning module (ODWP) and the weighted attention recovery module (WDAR) operations, which introduces the information of neighboring local windows, balances the attention distribution, and obtains multi-scale feature representation.
[0105] Step (3) includes the following steps:
[0106] (3.1) Stack multiple hole window attention (WDWA) modules. In each WDWA module, the input features are first processed by the hole window partitioning module (ODWP) to split the input features into local windows with a fixed window size and dilation rate r. The number of local windows is:
[0107]
[0108] Where H and W are the width and height of the feature map, h and w are the height and width of the local window, h ≥ 1 and w ≥ 1, ensuring a minimum 1 × 1 pixel window. and It must be an integer, otherwise the window cannot be evenly divided, resulting in incomplete coverage of the feature map. The dilation rate r should be a positive integer. When r=0, it is equivalent to no dilation; when r=1, the dilated window degenerates into a regular sliding window (step size = 1). The dilated window slides in four directions (upper left, upper right, lower left, and lower right) to cover all areas of the input feature. If the dilated window exceeds the boundary of the input feature, reflection padding is used to ensure that all local windows have the same size. Reflection padding is a common boundary processing technique. When the convolution kernel or window slides to the boundary of the feature map, the blank area is filled by mirroring the boundary pixel values to keep the output feature map at a fixed size.
[0109] (3.2) Each local window Embedded into query (Q), key (K) and value (V) through a linear layer, the expression is as follows:
[0110]
[0111]
[0112]
[0113] in, 、 and is the weight matrix, 、 and is the bias term.
[0114] Q and K are used to calculate the attention weights in multi-head self-attention:
[0115]
[0116] in, is the dimension of K, The function normalizes the attention weights into a probability distribution. Use the attention weights to perform a weighted summation on V:
[0117]
[0118] in, is the output feature of the i-th local window. The output features of all local windows are concatenated into the final output feature map:
[0119]
[0120] The output features provide richer feature representations for subsequent analysis tasks;
[0121] (3.3) The output of the ODWP operation is passed to the weighted attention recovery module (WDAR) operation, which performs an inverse operation on the output misaligned local attention window, restores it to the same shape as the input feature, aligns the misaligned window group, and fuses the attention using a weighted summation method, where the weight is calculated based on the number of overlaps, balances the attention of overlapping and non-overlapping areas, builds long-range relationships, and enhances the model's global perception ability;
[0122] (3.4) The output of each WDWA module is used as the input of the next WDWA module to obtain multi-scale feature representations and enhance the understanding and expression capabilities of the input features;
[0123] (3.5) After being processed by multiple WDWA modules, the final output features contain rich multi-scale information and long-distance dependencies, which serve as subsequent input.
[0124] Step (4) Through feature upsampling, the low-resolution features output by the encoder are gradually upsampled to the same spatial resolution as the input image. Finally, the features are mapped to the category space through the convolution layer, and the classification probability of each pixel is output;
[0125] Step (4) includes the following steps:
[0126] (4.1) The decoder uses bilinear interpolation for upsampling, generating new pixel values by calculating the linear combination of adjacent pixels, smoothly enlarging the image, and gradually upsampling the low-resolution features output by the encoder to the same spatial resolution as the input image;
[0127] (4.2) After each upsampling, the resolution of the feature image will increase, but the semantic information of the feature may be lost. It is necessary to ensure that the spatial resolution of the upsampled features is consistent with that of the encoder part, and to perform feature splicing in the channel dimension;
[0128] (4.3) Use a 3×3 convolution kernel to perform a convolution operation on the fused features to extract a higher-level feature representation. The output features of the convolution layer will serve as the input for the next upsampling.
[0129] (4.4) A 1×1 convolution kernel is used to reduce the number of channels of the feature map to the number of categories, and the feature map is mapped to the category space. The softmax activation function is used to map the output features to the classification probability distribution. The output features of each pixel are normalized to a probability value, indicating the probability that the pixel belongs to each category. The category label of each pixel can be determined by taking the category with the highest probability value.
[0130] Step (5) Use the training set to train the network model, select a suitable optimizer, adopt a learning rate decay strategy, calculate the loss function through forward propagation, and then update the network parameters through back propagation, and use the test data set to learn and train the model.
[0131] Step (5) includes the following steps:
[0132] (5.1) Model Training: The network model was trained using the preprocessed training data. The initial learning rate was set to 0.002, the batch size was 32, and the number of training epochs was set to 100. The Adam optimizer was used, which generally performs well due to its adaptive learning rate and momentum properties. Cross entropy was used as the loss function, and the network parameters were optimized using the backpropagation algorithm.
[0133] (5.2) Model Validation: The model is trained on a test dataset. The test dataset is input into the model and the model’s prediction results are obtained through forward propagation. Evaluation metrics such as accuracy, recall, precision, and F1 score are used to measure the model’s performance on the cloud type inversion task.
[0134] Accuracy is the ratio of the number of correctly predicted samples to the total number of samples. The calculation formula is:
[0135]
[0136] Among them, TP is true positive, the model correctly predicts positive samples as positive; TN is true negative, the model correctly predicts negative samples as negative; FP is false positive, the model incorrectly predicts negative samples as positive; FN is false negative, the model incorrectly predicts positive samples as negative.
[0137] Recall is the ratio of the number of samples correctly predicted as positive to the number of samples actually positive. The calculation formula is:
[0138]
[0139] Precision is the ratio of the number of samples correctly predicted as positive to the number of samples predicted as positive. The calculation formula is:
[0140]
[0141] The F1 score is the harmonic mean of precision and recall, and is used to comprehensively consider the balance between precision and recall. The calculation formula is:
[0142]
[0143] According to the above steps, an effective cloud type inversion network is constructed based on the hole window attention, which can achieve good cloud type inversion effect and make up for the shortcomings of traditional methods in dealing with multi-scale cloud structures and insufficient feature extraction and fusion.
[0144] Example 2:
[0145] This embodiment discloses a cloud type inversion system based on hole window attention, which specifically includes:
[0146] Data preprocessing module: According to the research area and time range, select the appropriate remote sensing satellite, download the image data through the official website, perform preprocessing operations such as normalization, data labeling, and data enhancement, and divide the data into training and test sets.
[0147] Feature encoding module: Performs preprocessing operations such as normalization and denoising on multidimensional spectral images, extracts image features from the images, obtains the latitude and longitude information of each remote sensing image using a geographic coordinate system, and uses sine and cosine functions for position encoding. Before concatenating the spectral and geographic features, the spectral features are reduced in dimension using methods such as principal component analysis (PCA), or the geographic features are increased in dimension using methods such as linear transformation to maintain consistency between the two dimensions. A fully connected layer (or linear transformation) is used to map the spectral and geographic feature vectors to this unified dimension, resulting in adjusted feature vectors. These adjusted spectral and geographic feature vectors are then concatenated along the channel dimension to obtain the final fused features. A three-layer convolutional network is constructed, and the fused features are input into the convolutional layer for downsampling, extracting features such as image edges and textures.
[0148] Dilated Window Attention Module: Multiple dilated window attention (WDWA) modules are stacked. In each WDWA module, the input feature first passes through a dilated window partitioning module (ODWP) operation, which partitions the input feature into local windows with a fixed window size and dilation rate r. The dilated windows slide in four directions (upper left, upper right, lower left, and lower right) to cover the entire region of the input feature. If the dilated window exceeds the bounds of the input feature, reflection padding is applied to ensure that all local windows have the same size. Each local window is embedded into the query (Q), key (K), and value (V) through a linear layer. Q and K are used to calculate the attention weights in the multi-head self-attention, while V is used in the subsequent weighted summation operation. The output of the ODWP operation is passed to a weighted attention restoration module (WDAR) operation, which first performs an inverse operation on each misaligned local attention window. Specifically, zero padding is applied to each local attention window in the opposite direction of the dilated window sliding. Then, the padded local attention windows in the same group are concatenated in height and width to restore them to the same shape as the input feature. Through this inverse operation, the four misaligned dilated attention window groups are aligned. Due to the presence of zero-padded regions, attention is fused using a weighted summation approach rather than a simple averaging operation. Weights are calculated based on the number of overlaps between each region to balance the attention between overlapping and non-overlapping regions. To this end, a base local weight window with the same shape as the local attention window is created, with all elements set to 1. The number of overlaps for each region in the inverse attention is calculated using the same inverse operation as the attention window. Finally, the inverse attention is divided by the inverse weight to balance the attention between different regions, ensuring that the attention values of overlapping and non-overlapping regions have equal weights, building long-range relationships, and enhancing the model's global perception. The output of each WDWA module serves as the input to the next WDWA module, obtaining a multi-scale feature representation that enhances the understanding and expression of the input features. After processing by multiple WDWA modules in the encoder, the final output features contain rich multi-scale information and long-range dependencies, making them suitable for use as input to the decoder.
[0149] Decoder module: The low-resolution features output by the encoder are gradually upsampled to the same spatial resolution as the input image through bilinear interpolation. After each upsampling, the resolution of the feature image increases, but the semantic information of the features may be lost. It is necessary to ensure that the upsampled features are consistent with the features of the encoder in spatial resolution and perform feature splicing in the channel dimension. The fused features are convolved using convolution kernels to extract higher-level feature representations. The output features of the convolution layer will serve as the input for the next upsampling. Finally, the feature map is mapped to the category space through the convolution kernel, and the output features are mapped to the classification probability distribution using the softmax activation function. The output features of each pixel are normalized to a probability value, and the category label can be determined by taking the category with the highest probability value.
[0150] Model Training and Validation Module: The network model is trained using preprocessed training data, using cross entropy as the loss function and backpropagation to optimize network parameters. The trained model is then trained on a test dataset, which is then fed into the model. Predictions are obtained through forward propagation, and the model's performance on cloud type inversion is measured using metrics such as accuracy, recall, precision, and F1 score.
[0151] The above describes the embodiments of the present invention in detail with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. After knowing the contents described in the present invention, ordinary technicians in this technical field can make several equivalent changes and substitutions without departing from the principles of the present invention. These equivalent changes and substitutions should also be regarded as falling within the scope of protection of the present invention.
Claims
1. A cloud type inversion method based on hole window attention, characterized in that: The following steps are involved: Step (1), data preprocessing, obtaining Landsat 8 and FY-4A satellite remote sensing image data, performing normalization, data annotation, data enhancement operations, and dividing the training set and test set; Step (2), feature encoding and fusion, extracts multidimensional spectral image features from the preprocessed data, and converts the image latitude and longitude information into spatial features through the position encoding module. The two are combined to obtain fused features, which are then input into the convolution layer to extract image edge and texture basic features; Step (3), hole window attention module processing, stacking multiple hole window attention modules, each hole window attention module includes a hole window partitioning module and a weighted attention recovery module operation, introducing adjacent local window information, balancing attention distribution, and obtaining multi-scale feature representation; Step (4), decoder classification, the decoder part gradually upsamples the low-resolution features output by the encoder to the same spatial resolution as the input image through feature upsampling, and finally maps the features to the category space through the convolution layer, outputting the classification probability of each pixel; Step (5), model training and verification, use the training set to train the network model, select a suitable optimizer, adopt the learning rate decay strategy, calculate the loss function through forward propagation, and then update the network parameters through back propagation, and use the test data set to learn and train the model.
2. The cloud type inversion method based on hole window attention according to claim 1 is characterized in that: The step (1) comprises the following steps: Step (1.1): Select appropriate Landsat 8 and FY-4A satellite remote sensing images according to the study area and time range, and download the remote sensing image data through the official website; In step (1.2), pixel values of remote sensing image data are normalized, scaling the original image pixel intensity values to between 0 and 1. After normalization, cloud areas in the image are accurately identified and categorized, clearly distinguishing between cloud and non-cloud areas. For identified cloud bodies, different category labels are assigned based on their morphology and texture characteristics, thus forming high-quality training labels. In step (1.3), data augmentation operations are performed on the labeled training data, including random rotation, flipping, cropping, and scaling. This artificially increases the diversity of the data and simulates various possible image changes. Random rotation simulates the effects of images at different shooting angles. Flipping enhances the model's ability to adapt to left-right or top-bottom symmetry. Cropping focuses on local areas in the image to better capture detailed features. Scaling helps the model adapt to image data at different resolutions. In step (1.4), the dataset is divided into 70% training set and 30% test set according to the size of the dataset.
3. The cloud type inversion method based on hole window attention according to claim 1 is characterized in that: The step (2) includes the following steps: In step (2.1), the multidimensional spectral image is normalized and denoised, and image features are extracted from the image. The normalization formula is as follows: Where x is the original data, is the normalized data, and are the minimum and maximum values of the data respectively; In step (2.2), the latitude and longitude information of each remote sensing image is obtained through the geographic coordinate system, and the sine and cosine functions are used for position encoding to provide a unique vector representation for each position in the sequence. The specific sine and cosine encoding formula is as follows: For each dimension 2i, the positional encoding PE is defined as: For each dimension 2i+1, the positional encoding PE is defined as: Among them, t represents the position index, i represents the dimension index, and d model Represents the dimension of the model embedding vector, the sine function is applied to all even dimensions, indexed as 2i; and the cosine function is applied to all odd dimensions, indexed as 2i+1; In step (2.3), use the principal component analysis method to reduce the dimensionality of the spectral features, or use the linear transformation method to increase the dimensionality of the geographic features, ensuring that the dimensions of the spectral features and geographic features remain consistent before splicing them; In step (2.4), the spectral feature vector and the geographic feature vector are mapped to the unified dimension using a fully connected layer or linear transformation to obtain the adjusted feature vector. In step (2.5), the adjusted spectral feature vector and geographic feature vector are concatenated along the channel dimension to obtain the final fusion feature; In step (2.6), a three-layer convolutional network is constructed, and the fused features are input into the convolutional layer to achieve downsampling and extract image edge and texture features.
4. The cloud type inversion method based on hole window attention according to claim 1 is characterized in that: The step (3) includes the following steps: In step (3.1), multiple hole window attention modules are stacked. In each hole window attention module, the input features are first processed by the hole window partitioning module to split the input features into local windows with a fixed window size and expansion rate r. The number of local windows is: Where H and W are the width and height of the feature map, h and w are the height and width of the local window, h ≥ 1 and w ≥ 1, ensuring a minimum 1 × 1 pixel window. and is an integer, otherwise the window cannot be evenly divided, resulting in incomplete coverage of the feature map; the expansion rate r is a positive integer. When r=0, it is equivalent to no expansion; when r=1, the expansion window degenerates into a regular sliding window with a step size of 1; the expansion window slides in four directions, namely, the upper left, upper right, lower left, and lower right, to cover all areas of the input feature; if the expansion window exceeds the boundary of the input feature, reflection padding is used to ensure that all local windows have the same size; Step (3.2), each local window Embedded into query Q, key K and value V through the linear layer, the expression is as follows: in, 、 and is the weight matrix, 、 and is the bias term; Q and K are used to calculate the attention weights in multi-head self-attention: in, is the dimension of K, The function normalizes the attention weights into a probability distribution; the attention weights are used to perform a weighted summation on V: in, is the output feature of the i-th local window; the output features of all local windows are concatenated into the final output feature map: The output features provide richer feature representations for subsequent analysis tasks; In step (3.3), the output of the hole window partitioning module is passed to the weighted attention restoration module, which performs an inverse operation on the output misaligned local attention window to restore it to the same shape as the input feature, aligns the misaligned window group, and fuses the attention using a weighted summation method, where the weight is calculated based on the number of overlaps, balancing the attention of overlapping and non-overlapping areas, building long-distance relationships, and enhancing the model's global perception ability; In step (3.4), the output of each hole window attention module is used as the input of the next hole window attention module to obtain multi-scale feature representation and enhance the understanding and expression ability of the input features; In step (3.5), after being processed by multiple hole window attention modules, the final output features contain rich multi-scale information and long-distance dependencies, which serve as subsequent input.
5. The cloud type inversion method based on hole window attention according to claim 1 is characterized in that: The step (4) comprises the following steps: In step (4.1), the decoder uses bilinear interpolation to upsample, generating new pixel values by calculating the linear combination of adjacent pixels, smoothly enlarging the image, and gradually upsampling the low-resolution features output by the encoder to the same spatial resolution as the input image; In step (4.2), after each upsampling, the resolution of the feature image will increase, but the feature semantic information may be lost. It is necessary to ensure that the upsampled features are consistent with the features of the encoder in spatial resolution and perform feature splicing in the channel dimension; In step (4.3), a 3×3 convolution kernel is used to perform a convolution operation on the fused features to extract a higher-level feature representation. The output features of the convolution layer will be used as the input for the next upsampling. In step (4.4), a 1×1 convolution kernel is used to reduce the number of channels of the feature map to the number of categories, and the feature map is mapped to the category space. The output features are mapped to the classification probability distribution using the softmax activation function. The output features of each pixel are normalized to a probability value, indicating the probability that the pixel belongs to each category. The category label of each pixel is determined by taking the category with the highest probability value.
6. The cloud type inversion method based on hole window attention according to claim 1 is characterized in that: The step (5) comprises the following steps: Step (5.1), model training: Use the preprocessed training set data to train the network model; set the initial learning rate to 0.002, the batch size to 32, the number of training rounds to 100, and use the Adam optimizer, which usually performs well due to its adaptive learning rate and momentum properties; use cross entropy as the loss function and optimize the network parameters through the backpropagation algorithm; Step (5.2), model validation: the model is trained on the test dataset, the test dataset is input into the model, the model prediction results are obtained through forward propagation, and the accuracy, recall, precision and F1 score evaluation indicators are used to measure the performance of the model in cloud type inversion.
7. A cloud type inversion system based on hole window attention, used to implement a cloud type inversion method based on hole window attention according to any one of claims 1 to 6, characterized in that: include: Data preprocessing module: responsible for satellite data download, normalization, cloud category labeling, data enhancement and data set division; Feature encoding module: includes spectral feature extraction, longitude and latitude sine and cosine position encoding, feature dimension unification, and a three-layer convolutional network to achieve spectral and spatial feature fusion and basic feature extraction; Hole Window Attention Module: It consists of a hole window partitioning module and a weighted attention recovery module. The hole window partitioning module implements multi-directional hole window partitioning and multi-head self-attention calculation. The weighted attention recovery module balances regional attention through inverse operations and weight calculation. Decoder module: achieves resolution restoration and pixel classification through bilinear interpolation upsampling, feature splicing, 3×3 convolution and 1×1 convolution and softmax activation function; Model training and verification module: Integrates the Adam optimizer, cross entropy loss function, and multi-indicator evaluation system to complete model training and performance verification.
8. The cloud type inversion system based on hole window attention according to claim 7, characterized in that: In the hole window attention module, the weighted attention recovery module calculates the weight by the number of overlaps, and the formula is: reverse attention value = attention value / number of overlaps, so as to balance the attention of different areas.
9. The cloud type inversion system based on hole window attention according to claim 7, characterized in that: In the feature encoding module, the position encoding dimension d model It is kept consistent with the spectral feature dimension through linear transformation to ensure the effectiveness of feature splicing.
Citation Information
Patent Citations
Remote sensing image cloud detection method based on cascade color and texture feature attention
CN113887472A
Remote sensing image cloud detection method and device, computer device and storage medium
CN115359370A
Cloud detection method and system based on coding and decoding attention interaction, medium and equipment
CN117292276A
Satellite remote sensing image cloud classification method based on space-time coding guidance
CN119495032A
Cited By
Image detection method, electronic equipment and storage medium
CN121280707A
Meteorological satellite real-time rainfall inversion method based on deep learning geographical attention mechanism
CN122176559A
Meteorological satellite real-time precipitation retrieval method based on deep learning geographical attention mechanism
CN122176559B