A cloud segmentation method based on a two-stage attention residual fusion network

CN122618249APending Publication Date: 2026-08-21SHANGHAI INSTITUTE OF TECHNICAL PHYSICS CHINESE ACADEMY OF SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611117454.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-27
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

[0004]然而,现有深度学习方法在细粒度云分割任务中仍存在以下不足:(1)多数方法采用预定义的跨层连接或固定融合方式,难以根据不同云类别对语义层级的差异化需求实现自适应特征筛选;(2)薄云、云阴影与背景之间在光谱和纹理上高度相似,现有方法对弱纹理区域和复杂边界区域的建模能力不足,容易产生漏检、误检及边界模糊问题;(3)现有方法通常将多尺度特征直接拼接或简单加权,缺乏对高层语义补充信息与低层细节修正信息的渐进式建模机制,难以兼顾全局语义一致性与局部边界完整性

Benefits of technology

[0016] (1) This invention replaces the fixed cross-layer feature fusion with a two-stage attention residual fusion mechanism, which can adaptively select information for different cloud categories and feature levels, thereby improving the ability to express semantic and detailed features in a coordinated manner.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122618249A_ABST
    Figure CN122618249A_ABST
Patent Text Reader

Abstract

The application discloses a cloud segmentation method based on a two-stage attention residual fusion network and belongs to the technical field of remote sensing image processing, and comprises the following steps: constructing a cloud fine segmentation training and test dataset; inputting training data into a backbone network to obtain an adapted feature set; taking the highest layer feature as a target feature, the adapted context feature as a guide feature, and the adapted feature set as a candidate feature set; obtaining a semantic enhancement feature; taking the lowest layer feature as the target feature, the semantic enhancement feature as the guide feature, and the adapted feature set as the candidate feature set; obtaining a detail enhancement feature; inputting the semantic enhancement feature and the detail enhancement feature into a decoder to output a segmentation result; setting main segmentation supervision, auxiliary semantic supervision and boundary supervision, constructing a joint loss function and training to obtain a two-stage attention residual fusion network model; and inputting test data into the model to output a segmentation result. The application can improve the precision of cloud detection results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of remote sensing image processing technology, specifically relating to a cloud segmentation method based on a two-stage attention residual fusion network. Background Technology

[0002] Cloud detection is a crucial preliminary step in remote sensing image quality control, surface parameter inversion, and atmospheric correction; its accuracy directly impacts the reliability and usability of subsequent quantitative remote sensing products. Compared to traditional cloud / non-cloud binary classification tasks, practical remote sensing applications typically require further differentiation into finer-grained categories such as thick clouds, thin clouds, cloud shadows, and background, as these different categories have varying effects on image quality, radiative transfer processes, and surface information recovery. Therefore, achieving fine-grained cloud segmentation in remote sensing images has become a significant technical challenge in the field of remote sensing information processing.

[0003] Existing cloud detection methods include threshold-based, physical model-based, and deep learning-based approaches. Traditional methods mainly rely on spectral thresholds, texture rules, or manually designed features to identify cloud regions. However, in complex scenes and when cloud types are blurred, it is difficult to balance detection accuracy and adaptability. With the development of deep learning technology, UNet (U-shaped encoder-decoder network), DeepLabv3+ (a deep semantic segmentation network based on dilated convolution and encoder-decoder structure), and their improved models have been widely used for cloud segmentation tasks. These models improve segmentation performance through encoder-decoder structures, multi-scale context modeling, and skip connection mechanisms.

[0004] However, existing deep learning methods still have the following shortcomings in fine-grained cloud segmentation tasks: (1) Most methods adopt predefined cross-layer connections or fixed fusion methods, which makes it difficult to achieve adaptive feature selection according to the different semantic level requirements of different cloud categories; (2) Thin clouds, cloud shadows and backgrounds are highly similar in spectrum and texture. Existing methods are not capable of modeling weak texture regions and complex boundary regions, which easily leads to missed detections, false detections and boundary ambiguity problems; (3) Existing methods usually directly splice or simply weight multi-scale features, lacking a progressive modeling mechanism for supplementary information of high-level semantics and correction information of low-level details, making it difficult to take into account both global semantic consistency and local boundary integrity.

[0005] Therefore, there is an urgent need to propose a new fine cloud segmentation technology for remote sensing images to improve the ability to distinguish between thick clouds, thin clouds, cloud shadows and complex backgrounds, and to enhance the model's ability to express fine-grained structure and boundary information, thereby improving the accuracy of fine cloud segmentation in remote sensing images. Summary of the Invention

[0006] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:

[0007] A cloud segmentation method based on a two-stage attention residual fusion network includes:

[0008] S1. Obtain the remote sensing scene image dataset and construct the cloud fine segmentation training dataset and test dataset;

[0009] S2 inputs the cloud fine segmentation training dataset into the backbone network and obtains the adapted feature set through the hollow spatial pyramid pooling module and feature adapter;

[0010] S3, Construct the first-stage semantic attention residual fusion module: Use the highest-level features as target features, the adapted context features as guiding features, and the adapted feature set as candidate feature set; obtain semantically enhanced features through the attention residual update mechanism;

[0011] S4, construct the second-stage detail attention residual fusion module: use the adapted lowest-level features as the target features, use the semantically enhanced features as the guiding features, and use the adapted feature set as the candidate feature set; adopt the same attention residual update mechanism as S3 to obtain the detail-enhanced features;

[0012] S5 inputs semantic enhancement features and detail enhancement features into the decoder, and outputs the segmentation result through low-level jumpers and upsampling;

[0013] S6. Set the main segmentation supervision at the final segmentation output, set the auxiliary semantic supervision at the semantic enhancement feature, set the boundary supervision at the detail enhancement feature, construct the joint loss function, perform end-to-end training, and obtain the two-stage attention residual fusion network model.

[0014] S7 inputs the test dataset into the trained two-stage attention residual fusion network model and outputs the segmentation result.

[0015] The present invention has the following beneficial effects:

[0016] (1) This invention replaces the fixed cross-layer feature fusion with a two-stage attention residual fusion mechanism, which can adaptively select information for different cloud categories and feature levels, thereby improving the ability to express semantic and detailed features in a coordinated manner.

[0017] (2) This invention enhances the network’s ability to distinguish edges, transition zones and weak texture regions by jointly constraining the intermediate layer feature learning process through auxiliary semantic supervision and boundary supervision, thereby reducing missed detections and boundary breaks.

[0018] (3) By focusing on high-level semantic enhancement in the first stage and low-level detail correction in the second stage, the present invention forms a two-stage progressive feature optimization path, which can effectively improve the segmentation accuracy of thin clouds, cloud shadows and complex boundary regions.

[0019] (4) This invention is compatible with multispectral remote sensing data and has good versatility and scalability. On the L1C (Level 1 C processing product) and L2A (Level 2 A processing product) processing levels of the CloudSEN12 (Sentinel-2 satellite cloud and cloud shadow semantic understanding dataset), the average intersection-union ratio (mIoU) reaches 71.72% and 71.77%, respectively, and on the GF1_WFV (Gaofen-1 wide-field camera dataset), the mIoU reaches 80.27%. The results show that the overall performance of this invention for cloud fine segmentation tasks is better than that of existing mainstream methods, and it has high application value. Attached Figure Description

[0020] Figure 1 The flowchart shows the remote sensing image cloud fine segmentation method based on a two-stage attention residual fusion network according to the present invention.

[0021] Figure 2 This is a schematic diagram showing the comparison of cloud segmentation visualization results of different methods on the CloudSEN12 L1C dataset in the example. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0023] To address the shortcomings of existing technologies, such as insufficient cross-layer feature fusion capabilities and low accuracy in identifying thin clouds and cloud shadows, this invention provides a cloud segmentation method based on a two-stage attention residual fusion network. Based on the DeepLabv3+ encoder-decoder framework, a two-stage attention residual fusion mechanism (S3 and S4) is introduced in the multi-scale feature interaction stage. Through a progressive feature optimization approach—first semantic enhancement, then detail enhancement—and the introduction of auxiliary semantic supervision and boundary supervision, high-precision and fine-grained segmentation of complex cloud types is achieved. The cloud segmentation method based on the two-stage attention residual fusion network of this invention includes:

[0024] S1. Obtain the remote sensing scene image dataset and construct the cloud fine segmentation training dataset and test dataset;

[0025] S2 inputs the cloud fine segmentation training dataset into the backbone network and obtains the adapted feature set through the hollow spatial pyramid pooling module and feature adapter;

[0026] S3, Construct the first-stage semantic attention residual fusion module: Use the highest-level features as target features, the adapted context features as guiding features, and the adapted feature set as candidate feature set; obtain semantically enhanced features through the attention residual update mechanism;

[0027] S4, construct the second-stage detail attention residual fusion module: use the adapted lowest-level features as the target features, use the semantically enhanced features as the guiding features, and use the adapted feature set as the candidate feature set; adopt the same attention residual update mechanism as S3 to obtain the detail-enhanced features;

[0028] S5 inputs semantic enhancement features and detail enhancement features into the decoder, and outputs the segmentation result through low-level jumpers and upsampling;

[0029] S6. Set the main segmentation supervision at the final segmentation output, set the auxiliary semantic supervision at the semantic enhancement feature, set the boundary supervision at the detail enhancement feature, construct the joint loss function, perform end-to-end training, and obtain the two-stage attention residual fusion network model.

[0030] S7 inputs the test dataset into the trained two-stage attention residual fusion network model and outputs the segmentation result.

[0031] The present invention will now be described in detail with reference to the accompanying drawings and embodiments:

[0032] like Figure 1 As shown, a cloud segmentation method based on a two-stage attention residual fusion network includes the following steps:

[0033] S1, Data Construction, Preprocessing, and Enhancement; specifically including:

[0034] Acquire a remote sensing scene image dataset and construct training and testing datasets for fine-grained cloud segmentation. The remote sensing scene image dataset can be CloudSEN12, GF1_WFV, or other multispectral remote sensing scene image datasets with pixel-level cloud annotations.

[0035] The original scene image is cropped into fixed-size image patches. In this embodiment, the fixed cropping size can be 512×512 pixels, and boundary image patches smaller than 512×512 pixels are filled with invalid pixels. During the training process, online data augmentation is performed on the image patches. The online data augmentation includes one or more of random horizontal flipping, random vertical flipping, random rotation, and Gaussian noise injection to improve the adaptability of the two-stage attention residual fusion network model (hereinafter referred to as the model) to imaging and scene changes.

[0036] Pixel-level classification labels are constructed, with categories including thick clouds, thin clouds, cloud shadows, and background. For samples with invalid pixels, invalid pixels can be set as ignored regions and not included in loss calculation and accuracy evaluation.

[0037] S2, multi-scale feature extraction and feature adaptation; specifically including:

[0038] Remote sensing scene image data is input into the backbone network of a two-stage attention residual fusion network model. In this embodiment, taking an input image patch size of 512×512 pixels as an example, the backbone network uses a pre-trained ResNet-34 (a 34-layer deep residual network) to output a four-level multi-scale feature set {C2, C3, C4, C5}, where C2, C3, C4, and C5 are multi-scale features, with C5 being the highest-level feature and C2 being the lowest-level feature. The number of channels and spatial dimensions of each multi-scale feature are as follows: C2: 64×128×128, C3: 128×64×64, C4: 256×32×32, C5: 512×32×32, where the first dimension represents the number of channels and the last two dimensions represent the feature map spatial dimensions (the scale data in this invention all have this meaning). The highest-level feature C5 is input into a hollow spatial pyramid pooling module to obtain the context feature F. aspp The dimensions are 256×32×32. The hollow spatial pyramid pooling module extracts multi-scale contextual information through convolutional and pooling branches with different dilation rates to improve the model's global receptive field and multi-scale semantic representation capabilities.

[0039] To support subsequent cross-layer fusion, the multi-scale feature set {C2, C3, C4, C5} and the context feature F are combined. aspp The adapted feature set is obtained by mapping the feature set to a 256-channel space via a feature adapter. , The adapted multi-scale features and adapted contextual features, among which, The highest-level feature after adaptation. The lowest layer feature after adaptation has the following channel count and spatial dimensions: : 256×128×128, : 256×64×64, : 256×32×32, : 256×32×32, 256×32×32. The feature adapter can be composed of a 1×1 convolutional layer, a normalization layer, and an activation layer. The connection relationship is as follows: the input features first enter the 1×1 convolutional layer, and the number of feature channels is adjusted through a pixel-by-pixel channel linear transformation to achieve dimensional alignment of features at different levels; then, the output features of the 1×1 convolutional layer are input to the normalization layer to standardize the feature distribution of each channel; finally, the normalized features are input to the activation layer, and the adapted output features are obtained through the GeLU (Gaussian Error Linear Unit) function.

[0040] S3, constructing the first-stage Semantic Attention Residual Fusion Module (SARF); specifically including:

[0041] S3.1, the first-stage semantic attention residual fusion module uses As the target feature X t With adapted context features As a guiding feature X g and adapt the multi-scale feature set As a set of candidate features { | },in For the first One candidate feature, The index is the candidate feature number.

[0042] In this embodiment, target feature X t With guiding feature X g The dimensions of all are 256×32×32. The dimensions are as follows: : 256×128×128, : 256×64×64, : 256×32×32, The model is 256×32×32, where the first dimension represents the number of channels and the last two dimensions represent the feature map space size. First, a size alignment operator is used... Guide feature X g and each candidate feature A unified mapping is performed to the target feature space resolution. Subsequently, the target feature X is... t With guiding feature X g Perform channel concatenation and mapping to obtain the query representation. :

[0043] ;

[0044] in, This indicates that the parameters of the query can learn a linear mapping.

[0045] No. Candidate features The key represents Sum value representation They are defined as follows:

[0046] ;

[0047] in, and The parameters representing the keys and values ​​can learn linear mappings.

[0048] S3.2, Calculate the query representation AND key representation Correlation score between them:

[0049] ;

[0050] in, Indicates the first The relevance score between the query representation and the key representation corresponding to each candidate feature. This represents the dot product operation along the channel dimension. This is a temperature coefficient used to adjust the smoothness of the attention weight distribution. This represents the number of feature channels. The relevance scores of all candidate features are normalized to obtain the attention weights. :

[0051] ;

[0052] in, Indicates the first The relevance score between the query representation and the key representation corresponding to each candidate feature; For the summation operator, The summation index takes values ​​from 1 to 4; It is a natural exponential function.

[0053] S3.3, based on attention weights The candidate features are weighted and aggregated to obtain cross-layer semantic consensus features. :

[0054] ;

[0055] in, This indicates element-wise multiplication.

[0056] To avoid directly replacing the target feature X t This results in the loss of original structural information; the residual representation is calculated first. :

[0057] ;

[0058] Among them, the residual representation R represents the cross-layer semantic consensus feature. With target feature X t The differences between them.

[0059] Finally, semantically enhanced features are obtained. (First-stage semantic attention residual fusion module):

[0060] ;

[0061] in, This represents the learnable channel scaling parameter. The parameters representing the residual output can be learned linearly.

[0062] Through the aforementioned attention residual update mechanism, effective information from different levels can be adaptively filtered and aggregated in the high-level semantic space. While maintaining the stability of the target feature's main structure, cross-layer semantic consistency is enhanced, thereby improving the overall discrimination ability of thin clouds and cloud shadows.

[0063] S4, Construct the second-stage Detail Attention Residual Fusion Module (DARF); specifically including:

[0064] The second-stage detail attention residual fusion module uses the adapted lowest-level features. As the target feature, the semantic enhancement feature F output by S3 is used. sem As a guiding feature, the adapted feature set will continue to be used. As a set of candidate features. In this embodiment, the target features The dimensions are 256×128×128, and the guiding feature F sem The original dimensions are 256×32×32. Similarly, in the size data, the first dimension represents the number of channels, and the last two dimensions represent the feature map space size.

[0065] This stage uses the same attention residual update mechanism as S3, providing high-level semantic priors through guiding features. This guides the network to focus on enhancing its ability to recognize cloud boundaries, thin cloud translucent textures, and cloud shadow regions in low-level spatial details. Specifically, the guiding feature F is first... sem Align to a size of 256×128×128, then... The splicing forms a joint feature; at the same time, Align each to a size of 256×128×128, and keep the original size. They jointly participate in attention aggregation and residual update, ultimately obtaining the detail-enhanced feature F. det In this embodiment, F det The output size is 256×128×128.

[0066] S5, based on a Deeplab-style decoder (which fuses high-level semantic features with shallow spatial detail features, recovering target boundaries and spatial resolution through a small amount of convolution and upsampling; the decoder structure described in S5.1 and S5.2 is the Deeplab-style decoder), performs decoding fusion and segmentation output; this includes inputting semantic enhancement features and detail enhancement features into the decoder, and outputting the segmentation result through low-level jumpers and upsampling. Specifically:

[0067] S5.1, semantic enhancement feature F sem Input the main branch of the decoder and upsample it; then use the detail enhancement feature F det As a low-level jumper feature, it is concatenated and fused with the upsampled semantic enhancement feature; after convolutional restoration, the output is a pixel-level segmentation result with the same size as the input image.

[0068] In this embodiment, the semantic enhancement feature F sem First, the 256×32×32 image is upsampled by a factor of 4 to obtain a high-level semantic feature of 256×128×128; detail enhancement feature F det The detail jumper features are kept at 256×128×128 and first compressed to 64×128×128 by 1×1 convolution.

[0069] S5.2, the 256×128×128 high-level semantic features and the 64×128×128 detail jump-connected features are concatenated in the channel dimension to obtain the 320×128×128 fused features; then convolved by 256 3×3 convolution kernels to output the 256×128×128 decoded features.

[0070] S5.3, the decoded features are mapped into a segmentation prediction map through the final prediction convolutional layer.

[0071] The predictive convolutional layer can output a 4×128×128 4-channel prediction map, which is further upsampled by 4 times to restore the output result to 4×512×512. For multi-class segmentation scenes of thick clouds, thin clouds, cloud shadows and background, the number of output channels of the predictive convolutional layer can be set to the number of classes, and the final class label of each pixel is determined according to the prediction results of each channel.

[0072] S6, multi-level supervised training and model optimization; specifically including:

[0073] A primary segmentation supervision is set at the final segmentation output, and semantic enhancement features F are applied. semAuxiliary semantic supervision is set at the location to enhance the details of feature F. det Set boundary supervision and construct a joint loss function. :

[0074] ;

[0075] in, , and The main segmentation loss is respectively Auxiliary semantic loss and detailed boundary supervision loss The weighting coefficients. In this embodiment, , , The values ​​can be 1.0, 0.01, and 0.01 respectively.

[0076] Main segmentation loss Auxiliary semantic loss and detailed boundary supervision loss Both can be achieved using a weighted combination of BCE (Binary Cross-Entropy) loss and Dice loss, specifically:

[0077] ;

[0078] ;

[0079] ;

[0080] in, , and These represent the real labels corresponding to the main segmentation branch, the auxiliary semantic branch, and the detail boundary branch, respectively. , and These represent the predicted probability graphs of the three branches mentioned above. , and These represent the weighting coefficients of the binary cross-entropy loss in each branch; , and These represent the weighting coefficients of the Dice loss in each branch; This represents the binary cross-entropy loss; This indicates Dice's loss.

[0081] The ground truth of the boundaries required for detail boundary supervision can be obtained by performing edge extraction on the original segmentation labels, for example, by using the Canny operator to extract the boundaries.

[0082] During training, the AdamW (Adaptive Moment Estimator with Decoupled Weight Decay) optimizer was used to update the model parameters, and a cosine preheating learning rate scheduling strategy was combined to complete end-to-end training. The initial learning rate and the final learning rate were 1×10⁻⁶. -4 and 1×10 -6 The image batch size is 4, and the training lasts for 50 epochs. After training, a two-stage attention residual fusion network model is obtained.

[0083] S7, Model Inference and Result Output; Specifically:

[0084] The dataset of remote sensing scene images to be processed is input into a trained two-stage attention residual fusion network model, which outputs pixel-level segmentation results for thick clouds, thin clouds, cloud shadows, and background. These results can be used for remote sensing image preprocessing, cloud mask generation, image selection, and subsequent quantitative remote sensing analysis.

[0085] Through the above implementation process, this invention enhances high-level semantic representation capabilities through a semantic attention residual fusion module and strengthens detailed information such as cloud boundaries, thin clouds, and cloud shadows through a detail attention residual fusion module, achieving effective fusion of multi-scale features. Compared with existing methods, this invention improves the recognition accuracy of thick clouds, thin clouds, and cloud shadows, and has good versatility and application value.

[0086] The cloud segmentation method based on a two-stage attention residual fusion network provided in this embodiment can be evaluated by calculating overall accuracy (OA), precision (Pr), recall (Re), F1 score, and intersection over union (IoU). Assuming that for a pixel of class c, where c represents the class, the number of true positive pixels obtained across all test images is TP. c The number of false positive pixels is FP c The number of false negative pixels is FN c Then the accuracy Pr of category c c Recall rate c F1 value F1 c And intersection and comparison of IoU c They are defined as follows:

[0087] ;

[0088] .

[0089] Overall accuracy (OA) is defined as:

[0090] ;

[0091] Where N represents the total number of pixels in the test set that participated in the evaluation. For the summation operator, For each category, the precision, recall, F1 score, and intersection-over-union ratio are averaged to obtain the average precision (mPr), average recall (mRe), average F1 score (mF1), and average intersection-over-union ratio (mIoU):

[0092] ;

[0093] Where K represents the total number of categories. All the above indicators are obtained based on statistics of all valid pixels in all test images.

[0094] In this embodiment, the method of the present invention, referred to as AttnRes-DeepLab (a two-stage attention residual fusion network model), is validated using the CloudSEN12 and GF1_WFV datasets, and compared with representative methods in the field such as UNet (U-shaped encoding and decoding with skip connections), DeepLabv3+ (dilated convolution and multi-scale fusion model), GCDB-UNet (global context dense fusion model), CDNetV2 (adaptive multi-layer feature fusion model), DABNet (deformable convolution and boundary weighting model), BoundaryNets (boundary-guided cloud detection model), AMCD-Net (attention-assisted multi-level cloud detection model), TransGA-Net (gradient-aware Transformer fusion model), and SCTNet (shallow convolution and Transformer fusion model).

[0095] As shown in Tables 1 and 2, on the CloudSEN12 dataset L1C processing level data, the method of this invention achieves 90.71% OA, 82.83% mPr, 82.07% mRe, 82.44% mF1, and 71.72% mIoU; on the CloudSEN12 dataset L2A processing level data, it achieves 90.70% OA, 82.51% mPr, 82.37% mRe, 82.44% mF1, and 71.77% mIoU; and on the GF1_WFV dataset, it achieves 95.45% OA, 89.76% mPr, 87.07% mRe, 88.33% mF1, and 80.27% mIoU. Experimental results demonstrate that, compared with existing mainstream methods, this invention achieves superior segmentation performance on different datasets and processing levels.

[0096] Table 1

[0097]

[0098] Table 2

[0099]

[0100] To verify the ability of this invention to identify different cloud categories, the IoU results of pixels in each category were statistically analyzed on the CloudSEN12 L1C dataset. As shown in Table 3, this invention achieves optimal results in the background, cloud shadow, and thin cloud categories, with IoUs of 90.70%, 53.22%, and 58.06% for the background, cloud shadow, and thin cloud categories, respectively; the IoU for thick clouds reaches 85.06%, which is basically equivalent to the optimal result of TransGA-Net. In particular, in the thin cloud category, this invention achieves an IoU of 58.06%, which is 3.77 percentage points higher than the best result of 54.29% in the existing methods. This indicates that this invention can more effectively utilize high-level semantic information and low-level detail information to accurately identify thin cloud regions with blurred boundaries, complex textures, and easy confusion with the background. At the same time, in the cloud shadow category, this invention achieves an IoU of 53.22%, further verifying that the proposed two-stage attention residual fusion mechanism has good discriminative ability for complex cloud structures and cloud shadow regions.

[0101] Table 3

[0102]

[0103] like Figure 2 As shown, by visually comparing the cloud segmentation results of different methods, it can be seen that the segmentation results obtained by this invention are comparable to those obtained by manual annotation. Figure 2 The model exhibits higher consistency with the actual annotations in the cloud. In thin cloud regions, cloud shadow regions, and cloud-background boundary regions (visually estimated and marked with red boxes), the present invention can more accurately identify target boundaries, effectively reducing missed detections, false detections, and boundary blurring, further verifying the effectiveness of the two-stage attention residual fusion network model in cloud fine segmentation tasks.

[0104] The above description is merely an embodiment of the present invention and does not limit the scope of the invention. Any equivalent structural or procedural transformations made based on the description and drawings of this invention, or direct or indirect applications in other related system fields, are similarly included within the protection scope of this invention. Contents not described in detail in this specification are prior art known to those skilled in the art.

Claims

1. A cloud segmentation method based on a two-stage attention residual fusion network, characterized in that, include: S1. Obtain the remote sensing scene image dataset and construct the cloud fine segmentation training dataset and test dataset; S2 inputs the cloud fine segmentation training dataset into the backbone network and obtains the adapted feature set through the hollow spatial pyramid pooling module and feature adapter; S3, Construct the first-stage semantic attention residual fusion module: Use the highest-level features as target features, the adapted context features as guiding features, and the adapted feature set as candidate feature set; obtain semantically enhanced features through the attention residual update mechanism; S4, construct the second-stage detail attention residual fusion module: use the adapted lowest-level features as the target features, use the semantically enhanced features as the guiding features, and use the adapted feature set as the candidate feature set; adopt the same attention residual update mechanism as S3 to obtain the detail-enhanced features; S5 inputs semantic enhancement features and detail enhancement features into the decoder, and outputs the segmentation result through low-level jumpers and upsampling; S6. Set the main segmentation supervision at the final segmentation output, set the auxiliary semantic supervision at the semantic enhancement feature, set the boundary supervision at the detail enhancement feature, construct the joint loss function, perform end-to-end training, and obtain the two-stage attention residual fusion network model. S7 inputs the test dataset into the trained two-stage attention residual fusion network model and outputs the segmentation result.

2. The cloud segmentation method based on a two-stage attention residual fusion network according to claim 1, characterized in that, S1 includes cropping the original scene image into image patches of a fixed crop size, and filling image patches that are smaller than the fixed crop size with invalid pixels; for the training process, online data augmentation is performed on the image patches, and the online data augmentation includes one or more of random horizontal flipping, random vertical flipping, random rotation, and Gaussian noise injection.

3. The cloud segmentation method based on a two-stage attention residual fusion network according to claim 1, characterized in that, In S2, the backbone network uses a pre-trained ResNet-34 to output a four-level multi-scale feature set.

4. The cloud segmentation method based on a two-stage attention residual fusion network according to claim 1, characterized in that, In S2, the feature adapter consists of a 1×1 convolutional layer, a normalization layer, and an activation layer. The input features first enter the 1×1 convolutional layer, where the number of feature channels is adjusted through a pixel-by-pixel linear channel transformation to achieve dimensional alignment of features at different levels. Subsequently, the output features of the 1×1 convolutional layer are input into the normalization layer to standardize the feature distribution of each channel. Finally, the normalized features are input into the activation layer to obtain the adapted output features.

5. The cloud segmentation method based on a two-stage attention residual fusion network according to claim 1, characterized in that, In S3, the attention residual update mechanism includes: mapping the guiding feature and each candidate feature to the target feature space resolution using a size alignment operator; concatenating the target feature and guiding feature channels and mapping them to obtain the query representation; calculating the relevance score between the query representation and the key representation of each candidate feature; normalizing the relevance scores of all candidate features to obtain attention weights; weighted aggregation of each candidate feature based on the attention weights to obtain cross-layer semantic consensus features; calculating the residual representation between the cross-layer semantic consensus features and the target feature, and obtaining semantically enhanced features by learning linear mappings through learnable channel scaling parameters and parameters of the residual output.

6. The cloud segmentation method based on a two-stage attention residual fusion network according to claim 1, characterized in that, In S4, the second-stage detail attention residual fusion module uses the adapted lowest-level features as the target features and the semantic enhancement features output from S3 as the guiding features. The guiding features provide high-level semantic priors, guiding the network to focus on enhancing the recognition capabilities of cloud boundaries, thin cloud semi-transparent textures, and cloud shadow areas in low-level spatial details.

7. The cloud segmentation method based on a two-stage attention residual fusion network according to claim 1, characterized in that, S5 includes: first upsampling the semantic enhancement features by 4 times to obtain high-level semantic features; first compressing the detail enhancement features into detail jumper features through 1×1 convolution; concatenating the high-level semantic features and detail jumper features in the channel dimension to obtain fused features; then convolving the fused features to output decoded features; and finally mapping the decoded features into a segmentation prediction map through the final prediction convolutional layer.

8. The cloud segmentation method based on a two-stage attention residual fusion network according to claim 1, characterized in that, In S6, the joint loss function is a weighted combination of the main segmentation loss, the auxiliary semantic loss, and the detail boundary supervision loss; the main segmentation loss, the auxiliary semantic loss, and the detail boundary supervision loss are all weighted combinations of the binary cross-entropy loss and the Dice loss.

9. The cloud segmentation method based on a two-stage attention residual fusion network according to claim 1, characterized in that, In S6, the AdamW optimizer is used to update the model parameters, and a cosine warm-up learning rate scheduling strategy is combined to complete end-to-end training.

10. The cloud segmentation method based on a two-stage attention residual fusion network according to claim 2, characterized in that, In S1, for samples with invalid pixels, the invalid pixels are set as ignored regions and are not included in loss calculation and accuracy evaluation.