Cloud detection method based on double-branch fusion feature enhancement
By building a cloud detection network model based on the E-TransA module and the dual-branch attention mechanism, the problem of insufficient accuracy of cloud boundary and thin cloud detection in remote sensing images is solved, and more accurate cloud structure analysis and detection is achieved.
Patent Information
- Application Number
- CN202510330386.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-18
AI Technical Summary
The existing remote sensing image cloud detection methods are not robust enough to handle the fine definition of cloud boundaries and the accurate identification of thin clouds, and it is difficult to effectively improve the detection accuracy.
Using a cloud detection method based on dual-branch fusion feature enhancement, a cloud detection network model based on the E-TransA module and the dual-branch attention mechanism is constructed, combining the multi-head self-attention mechanism and the residual fusion convolution module, global and local features are extracted, and feature fusion and decoding are performed to generate high-resolution cloud detection results.
It improves the cloud boundary detail extraction and thin cloud positioning capabilities, enhances the analytical ability of complex cloud structures, can accurately capture the detailed characteristics and spatial dependencies of cloud layers, and improves the accuracy and robustness of detection.
Smart Images

Figure CN120339792A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and particularly to a cloud detection method based on dual-branch fusion feature enhancement. Background Art
[0002] Deep neural networks have powerful feature extraction and representation learning capabilities, and can accurately identify different regions in images and extract objects of interest. Among them, the convolutional neural network (CNN), as a representative model of deep learning in the field of image processing, is good at automatically learning and refining efficient feature representations from massive image data, and thus has achieved remarkable achievements in many fields such as remote sensing image cloud detection. However, limited by the local receptive field characteristics of the convolutional kernel, CNN has deficiencies in capturing long-range dependence relationships. In this context, the Transformer architecture has attracted much attention because it can efficiently capture long-range dependence relationships in natural language sequences. In order to extend this advantage of Transformer to the visual field, researchers have innovatively proposed the vision Transformer (ViT). With its excellent global context modeling ability, ViT can efficiently learn and represent image features, thus showing great potential in the field of image processing. Given the expertise of CNN in local feature extraction and the advantage of Transformer in global context representation, using both of them for image processing has become a very promising strategy.
[0003] Although deep learning has shown great application potential in the field of remote sensing image cloud detection, many existing methods are still not robust enough in dealing with the fine definition of cloud boundaries and the accurate identification of thin clouds. Therefore, exploring new cloud detection schemes to improve the detection accuracy of cloud boundaries and thin cloud regions has become a key challenge to be solved urgently. Summary of the Invention
[0004] In view of this, in response to the challenges faced by cloud detection, the present invention proposes a cloud detection method based on dual-branch fusion feature enhancement, which can improve the ability to extract cloud boundary details and locate thin clouds. The specific technical solutions of the invention are as follows:
[0005] To solve the above problems, the present invention discloses a cloud detection method based on dual-branch fusion feature enhancement, including the following steps:
[0006] Step 1: Obtain a high-quality cloud detection data set and perform necessary preprocessing on the data set;
[0007] Step 2: Construct a cloud detection network model based on dual-branch fusion feature enhancement;
[0008] Step 3: Use the preprocessed data set for training the cloud detection model and save the best model parameters;
[0009] Step 4: Predict the images in the test set using the model parameters saved in Step 3, and quantify the model performance through evaluation metrics;
[0010] Furthermore, Step 1 involves the acquisition and preprocessing of the dataset. Specifically, three high-quality datasets are selected, which are sourced from different sensor types, cover diverse atmospheric conditions, and have different resolution sizes. After selecting the datasets, the necessary preprocessing operations are as follows:
[0011] Step 1.1: Perform the merging process of the RGB three channels for the acquired images;
[0012] Step 1.2: Randomly crop the images to a size of 256×256 pixels;
[0013] Step 1.3: Divide the entire dataset into two parts, where 80% of the images are used as the training set, and the remaining 20% are used as the test set and the validation set;
[0014] Furthermore, Step 2 is used to perform the following steps:
[0015] Step 2.1: Use the E-TransA module as the backbone branch to extract global features. This module captures four-level multi-level features T = {T1, T2, T3, T4} from the input feature map in stages;
[0016] Step 2.2: After extracting the features at each stage in Step 2.1, apply the residual fusion convolution module for feature enhancement processing, aiming to obtain a feature representation containing more detailed information. These enhanced four-level features are relabeled as R = {R1, R2, R3, R4};
[0017] Step 2.3: Use the dual-branch attention mechanism as the core to extract local features, and obtain four-stage multi-level features A = {A1, A2, A3, A4} from the input feature map. This mechanism cleverly combines the advantages of convolution operations and the attention mechanism. While extracting the local spatial features of the image, it can dynamically highlight and strengthen the feature representation of the key regions, thereby more accurately capturing the detailed features of the clouds and the spatial dependence relationship, and generating a hierarchical local feature representation;
[0018] Step 2.4: Each level of the global feature T extracted in Step 2.1 will first adjust the size of its feature map through linear interpolation to ensure that it matches the corresponding level of the local feature A in size in Step 2.3. Subsequently, adopt the cascade fusion strategy to fuse the global feature T and the local feature A to promote a comprehensive understanding of the cloud structure and generate more refined detection results. Finally, a synthetic feature set C = {C1, C2, C3, C4} is obtained through this process;
[0019] Step 2.5: Use a multi-layer perceptron to adjust the number of channels of the enhanced features R at all levels obtained in Step 2.2. Then, use linear interpolation technology to further adjust the sizes of these feature maps to ensure that they meet the conditions required for channel splicing;
[0020] Step 2.6: Input the highest-level synthetic feature C4 obtained in Step 2.4 into the dual convolutional module of the decoder. After processing, it is spliced with the fused features obtained in Step 2.5 to achieve further feature fusion;
[0021] Step 2.7: After obtaining the new synthetic feature in Step 2.6, send it to the last three-level decoding module for step-by-step decoding to generate the final output feature map. During the decoding process, the output features of the first two-level decoders are respectively spliced with C1 and C2 in Step 2.4 to make full use of low-level details and high-level semantics, improving the accuracy and robustness of detection;
[0022] Furthermore, in Step 2.1, a branch with the E-TransA module as the core architecture is used to extract features. The E-TransA module innovatively integrates the multi-head self-attention mechanism, which consists of two branches operating in parallel, focusing on extracting global and local information respectively. The local branch mainly uses two convolutional layers with kernel sizes of 1×3 and 3×1 to capture feature information in the horizontal and vertical directions. Subsequently, the extracted feature information is integrated to synthesize local features. The global branch first uses a 1×1 convolution to perform weighted summation on all channels of each pixel point, achieving information fusion between channels. Then, a 3×3 convolutional layer is used to generate query (Q), key (K), and value (V) matrices. To reduce the complexity of the algorithm, after the generation of the K and V matrices, they are respectively processed by a convolution with a stride of r and a kernel size of r, thereby achieving downsampling of the feature map. It should be noted that we have made a targeted setting for the value of r according to the specific level of the E-TransA module. Specifically, in the 4 levels of the encoder, the values of r are 8, 4, 2, and 1 respectively. Finally, the outputs of the local branch and the global branch are added together, and the feature map is output after passing through a depthwise separable convolution with a kernel size of 3×3. Generally speaking, the detailed process of the multi-head self-attention mechanism is defined as follows:
[0023] C = Conv 3×1 (X) + Conv 1×3 (X), Conv = Conv 1×1 (X)
[0024] Q, K, V = Conv 3×3 (Conv), K1 = Conv r×r (K), V1 = Conv r×r(V)
[0025] K1,
[0026]
[0027] where X is the input feature map of the multi-head self-attention mechanism.
[0028] Furthermore, in step 2.2, a residual fusion convolutional module is applied for feature enhancement processing. First, a 1×1 convolution is used to learn the weights between different channels, thereby more flexibly adjusting the importance of features. Then, a 3×3 convolution is used to capture local spatial features and enhance the sensitivity to image detail information. Subsequently, the output results of the above two convolutions are multiplied element by element to achieve deep fusion of features. Finally, the fused features are finely optimized through a 1×1 convolution to generate the output feature map C out . To retain the key information in the input feature map while enhancing the features, we add C out and the original input feature map element by element.
[0029] C out = Conv 1×1 {Conv 3×3 [Conv 1×1 (Y)]×Conv 1×1 (Y)}
[0030] Z out = C out + Y
[0031] where Y is the output of the E-TransA module.
[0032] Furthermore, the dual-branch attention mechanism proposed in step 2.3 performs precise local feature extraction on the input image to achieve more detailed cloud detection. First, the input feature map passes through a convolution sequence that contains i 3×3 convolutions for feature extraction, where the first convolution adjusts the channels and the size of the feature map. Through experimental verification, when i is set to 4, the best recognition and processing accuracy can be achieved. Then, the attention mechanism is used to improve the ability to capture key information, and the steps are as follows:
[0033] Step 2.3.1: The channel attention mechanism is adopted to dynamically adjust the importance weights of each channel in the output feature map of the convolutional sequence, so as to enhance the feature expression ability. First, the features are compressed in the spatial dimension through global average pooling, and each two-dimensional feature channel becomes a real number with a global receptive field. Secondly, after obtaining the global information, the generation and selection of feature channel weights are realized in a parameterized way. Finally, the adjusted importance weights are multiplied element-wise to the original feature map, so as to effectively recalibrate the original features in the channel dimension, further enhancing the feature expression ability and pertinence. The specific expression is as follows:
[0034]
[0035] Z out = W2[σ(W1C CA )],
[0036] W CC = F out × Sigmoid(Z out )
[0037] where F out is the output of the feature map after passing through the convolutional sequence.
[0038] Step 2.3.2: The parallel spatial attention mechanism is used to capture the spatial dependence relationships in W CC from different perspectives, effectively improving the model's ability to analyze the complex structure of clouds. First, the input feature map W CC is compressed in parallel in the channel dimension to generate two pairs of global average information and global significant information respectively. Then, these two types of global information are concatenated to provide richer input for the subsequent generation of spatial attention weights. Next, a 7×7 convolutional operation is performed on the concatenated feature map, and normalization is carried out through the Sigmoid function to learn the relationships between spatial positions and generate the importance weights of each spatial position. Subsequently, the generated spatial attention weights are multiplied element-wise with the input feature map W CC to obtain the weighted feature map. Finally, the weighted feature maps of the two parallel branches are added together to generate diverse feature representations. The specific expressions are as follows:
[0039] S1avg = AvgPool(W CC ), S1max = MaxPool(W CC )
[0040] S2avg = AvgPool(W CC ), S2max = MaxPool(W CC )
[0041] F S1 = W CC × Sigmoid{Conv 7×7 [Cat(S1avg, S1max)]}
[0042] F S2 = W CC × Sigmoid{Conv 7×7 [Cat(S2avg, S2max)]}
[0043] W SC = F S1 + F S2
[0044] Step 2.3.3: Extract the local details and multi-scale context information of the input feature map with the help of the convolutional branch to obtain the output F Cout , and the specific expression is as follows:
[0045] F CB = Conv 1×1 (F out )
[0046] E C1 = Conv 1×1 (F CB ), E C2 = Conv 3×3 [Conv 3×3 (F CB )]
[0047] F Cout = DWConv 3×3 (E C1 × E C2 )
[0048] Step 2.3.4: Add the output results of Step 2.3.2 and Step 2.3.3 element-wise to fuse the features of the two branches and achieve the complementarity and enhancement of the detailed information. Subsequently, further extract the fused features through a 1×1 convolution to output the final result, and the specific expression is as follows:
[0049] W CAout = Conv 1×1 (F Cout + W SC )
[0050] Further, in step 2.4, first, the output of the E-TransA module is used to adjust the size of the feature map through linear interpolation to make it the same as the size of the feature map output by the dual-branch attention mechanism; subsequently, the adjusted feature map is concatenated with the feature map output by the dual-branch attention mechanism to fuse the global context information and local detail information, further enhancing the feature expression ability.
[0051] F resize = Interpolate[Y, size=(H, W)]
[0052] Cat = Concat[W CAout , F resize
[0053] Where Y is the output after adjustment of the E-TransA module.
[0054] Further, in step 2.5, first, a multi-layer perceptron is used to adjust the number of channels of the encoded features output by the four-level residual fusion convolution module to 256. At the same time, each level of enhanced feature R is adjusted to the same size as R2 through linear interpolation. Secondly, the adjusted four-level features are concatenated in the channel dimension to fuse multi-scale information. Finally, a 1×1 convolution is used to adjust the number of channels of the fused features to 256, and the specific expression is as follows:
[0055] R C1 , R C2 , R C3 , R C4 = MLP(R1, R2, R3, R4)
[0056] R re1 , R re2 , R re3 , R re4 = Interpolate[R C1 , R C2 , R C3 , R C4 , size = R C2
[0057] R Cat = Concat[R re1 , R re2 , R re3 , R re4
[0058] R out = Conv 1×1 (R Cat )
[0059] Furthermore, in step 2.6, the highest-level fused feature C4 obtained in the encoder is used as the input of the decoder. The convolutional module design of the decoder includes two 3×3 convolutional layers. The main function of these convolutional layers is to adjust the number of channels of the feature map and effectively restore the detailed information in the image. During the decoding process, the fused feature C4 first undergoes feature extraction through the convolutional module, and then the image resolution is enhanced through bilinear interpolation operation. Next, the feature with adjusted resolution is concatenated with R out along the channel dimension. This design ensures that during the process of gradually restoring the spatial resolution of the image, the image details can be effectively and finely reconstructed, thereby providing a more comprehensive feature representation for the model and further improving the accuracy and refinement of the segmentation decision. The specific expression is as follows:
[0060] D1 = Conv 3×3 [Conv 3×3 (C4)]
[0061] D re = Interpolate[D1, size = R out
[0062] D cat = Concat[D re , R out
[0063] Furthermore, in step 2.6, D cat is used as the input of the last three levels of the decoder. Specifically, after the output of the first-level decoding module, the size of the feature map is increased through bilinear interpolation and concatenated with the feature map C2 in the encoder along the channel dimension; after the output of the second-level decoding module, the size of the feature map is also increased through bilinear interpolation and then concatenated with the feature map C1 in the encoder, aiming to simultaneously utilize the low-level details and high-level semantics to improve the comprehensiveness of the feature representation. Finally, after being processed by the third-level decoding module, the output feature map is normalized to generate a feature map with the same size as the original input feature map, thereby completing the decoding process and restoring the high-resolution features. The specific expression is as follows:
[0064] D2 = Conv 3×3 [Conv 3×3 (D cat )], D re2 = Interpolate[D2, size = C2]
[0065] D cat2 = Concat[D re2 , C2]
[0066] D3 = Conv 3×3 [Conv3×3 (D cat2 )],D re3 = Interpolate[D3, size = C1]
[0067] D cat3 = Concat[D re3 , C1]
[0068] D4 = Conv 3×3 [Conv 3×3 (D cat3 )]
[0069] D out = Sigmoid[Conv 1×1 (D4)]。
[0070] Compared with other methods of the prior art, the prominent advantages of the present invention are as follows:
[0071] The present invention can improve the ability to extract cloud boundary details and locate thin clouds, effectively enhance the model's ability to analyze the complex structure of clouds, skillfully combines the advantages of convolutional operations and attention mechanisms, while extracting local spatial features of the image, can dynamically highlight and strengthen the feature representation of key regions, so as to more accurately capture the detailed features and spatial dependence relationships of clouds, ensuring that during the process of gradually restoring the spatial resolution of the image, the image details can be effectively and finely reconstructed. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 is a schematic diagram of the overall method flow of the present invention;
[0073] Figure 2 is a schematic diagram of the process of extracting the input feature map by the dual-branch attention mechanism;
[0074] Figure 3 is a schematic diagram of the process of extracting features by the branch with the E-TransA module as the core architecture;
[0075] Figure 4 is a visualization experimental result diagram under different scenarios. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0076] The present invention will be described in detail below through specific embodiments, but the uses and purposes of these exemplary embodiments are only used to illustrate the present invention, and do not constitute any form of limitation to the actual protection scope of the present invention, nor limit the protection scope of the present invention thereto.
[0077] A cloud detection method based on dual-branch fusion feature enhancement in this embodiment,
[0078] such as Figure 1As shown in the figure, the cloud detection method based on dual-branch fusion feature enhancement includes the following steps:
[0079] Step 1: Obtain a high-quality cloud detection dataset and perform necessary preprocessing on the dataset;
[0080] The obtaining of the cloud detection dataset and the preprocessing are as follows. Specifically, three high-quality datasets from different sensor types, covering diverse atmospheric conditions, and having different resolution sizes are selected. After the datasets are selected, the necessary preprocessing operations include:
[0081] Step 1.1: Perform RGB three-channel merging processing on the obtained images;
[0082] Step 1.2: Randomly crop the images to a size of 256×256 pixels;
[0083] Step 1.3: Divide the entire dataset into two parts, 80% of the images are used as the training set, and the remaining 20% are used as the test set and the validation set;
[0084] Step 2: Build a cloud detection network model based on dual-branch fusion feature enhancement;
[0085] As Figure 2 shown in the figure, the detailed steps for building a cloud detection network model based on dual-branch fusion feature enhancement are as follows:
[0086] Step 2.1: Use the E-TransA module as the main branch to extract global features, and obtain four-stage features T = {T1, T2, T3, T4} from the input feature map;
[0087] As Figure 3As shown, the branch with the E-TransA module as the core architecture is adopted to extract features. Specifically, the E-TransA module innovatively integrates the multi-head self-attention mechanism, which consists of two branches operating in parallel, focusing on extracting global and local information respectively. The local branch mainly uses two convolutional layers with kernel sizes of 1×3 and 3×1 to capture feature information in the horizontal and vertical directions. Subsequently, the extracted feature information is integrated to synthesize local features. The global branch first uses a 1×1 convolution to perform weighted summation on all channels of each pixel point, achieving information fusion between channels. Then, a 3×3 convolutional layer is used to generate query (Q), key (K), and value (V) matrices. To reduce the complexity of the algorithm, the K and V matrices are respectively processed by a convolution with a stride of r and a kernel size of r after generation, thus realizing downsampling of the feature map. It should be noted that we have specifically set the value of r according to the specific level of the E-TransA module. Specifically, in the 4 levels of the encoder, the values of r are 8, 4, 2, and 1 respectively. Finally, the outputs of the local branch and the global branch are added together, and a depthwise separable convolution with a kernel size of 3×3 is used to output the feature map. Generally speaking, the detailed process of the multi-head self-attention mechanism is defined as follows:
[0088] C = Conv 3×1 (X) + Conv 1×3 (X), Conv = Conv 1×1 (X)
[0089] Q, K, V = Conv 3×3 (Conv), K1 = Conv r×r (K), V1 = Conv r×r (V)
[0090] K1,
[0091]
[0092] where X is the input feature map of the multi-head self-attention mechanism.
[0093] Step 2.2: After extracting the features at each stage in Step 2.1, apply the residual fusion convolution module for feature enhancement processing, aiming to obtain a feature representation containing more detailed information. These enhanced four-level features are re-labeled as R = {R1, R2, R3, R4};
[0094] As Figure 2As shown in the figure, a residual fusion convolution module is adopted to enhance the features output by the E-TransA module. This module is designed as follows: First, a 1×1 convolution is used to learn the weights between different channels, so as to more flexibly adjust the importance of features. Then, a 3×3 convolution is used to capture local spatial features and enhance the sensitivity to image detail information. Subsequently, the output results of the above two convolutions are multiplied element by element to achieve deep fusion of features. Finally, the fused features are finely optimized through a 1×1 convolution to generate the output feature map C out . To retain the key information in the input feature map while enhancing the features, we add C out and the original input feature map element by element.
[0095] C out = Conv 1×1 {Conv 3×3 [Conv 1×1 (Y)] × Conv 1×1 (Y)}
[0096] Z out = C out + Y
[0097] where Y is the output of the E-TransA module.
[0098] Step 2.3: Use a dual-branch attention mechanism as the core to extract local features and obtain four-stage features A = {A1, A2, A3, A4} from the input feature map. This mechanism cleverly combines the advantages of convolution operations and attention mechanisms. While extracting local spatial features of the image, it can dynamically highlight and strengthen the feature representation of key regions, so as to more accurately capture the detail features and spatial dependence relationships of clouds and generate a hierarchical local feature representation;
[0099] As Figure 2 shown, a dual-branch attention mechanism is used to extract local features of the input feature map. The dual-branch attention mechanism proposed in this design performs precise local feature extraction on the input image to achieve more detailed cloud detection. First, the input feature map passes through a convolution sequence, which contains i 3×3 convolutions for feature extraction, and the first convolution adjusts the channels and the size of the feature map. Through experimental verification, when i is set to 5, the best recognition and processing accuracy can be achieved. Then, the attention mechanism is used to improve the ability to capture key information, and the steps are as follows:
[0100] Step 2.3.1: The channel attention mechanism is adopted to dynamically adjust the importance weights of each channel in the output feature map of the convolutional sequence, so as to enhance the feature expression ability. First, the features are compressed in the spatial dimension through global average pooling, and each two-dimensional feature channel is transformed into a real number with a global receptive field. Secondly, after obtaining the global information, the generation and selection of the feature channel weights are realized in a parameterized way. Finally, the adjusted importance weights are multiplied element-wise to the original feature map, so as to effectively recalibrate the original features in the channel dimension, further enhancing the feature expression ability and pertinence. The specific expression is as follows:
[0101] Z out = W2[σ(W1C CA )],
[0102] W CC = F out × Sigmoid(Z out )
[0103] where F out is the output of the feature map after the convolutional sequence.
[0104] Step 2.3.2: The parallel spatial attention mechanism is used to capture the spatial dependence relationships in W CC from different perspectives, effectively improving the model's ability to analyze the complex structure of clouds. First, the input feature map W CC is compressed in parallel in the channel dimension to generate two pairs of global average information and global significant information respectively. Then, these two kinds of global information are concatenated to provide richer input for the subsequent generation of spatial attention weights. Then, a 7×7 convolution operation is performed on the concatenated feature map, and normalization processing is carried out through the Sigmoid function to learn the relationships between spatial positions and generate the importance weights of each spatial position. Subsequently, the generated spatial attention weights are multiplied element-wise to the input feature map W CC to obtain the weighted feature map. Finally, the weighted feature maps of the two parallel branches are added together to generate diverse feature representations. The specific expression is as follows:
[0105] S1avg = AvgPool(W CC ), S1max = MaxPool(W CC )
[0106] S2avg = AvgPool(W CC ), S2max = MaxPool(W CC )
[0107] F S1 = WCC × Sigmoid{Conv 7×7 [Cat(S1avg, S1max)]}
[0108] F S2 = W CC × Sigmoid{Conv 7×7 [Cat(S2avg, S2max)]}
[0109] W SC = F S1 + F S2
[0110] Step 2.3.3: Extract local details and multi-scale context information of the input feature map with the help of the convolutional branch to obtain the output F Cout , and the specific expression is as follows:
[0111] F CB = Conv 1×1 (F out )
[0112] E C1 = Conv 1×1 (F CB ), E C2 = Conv 3×3 [Conv 3×3 (F CB )]
[0113] F Cout = DWConv 3×3 (E C1 × E C2 )
[0114] Step 2.3.4: Add the output results of Step 2.3.2 and Step 2.3.3 element-wise to fuse the features of the two branches and achieve complementary and enhanced detailed information. Subsequently, further extract the fused features through 1×1 convolution to output the final result, and the specific expression is as follows:
[0115] W CAout = Conv 1×1 (F Cout + W SC )
[0116] Step 2.4: Each level of global feature T extracted in Step 2.1 will first adjust its feature map size through linear interpolation to ensure that it matches the size of the corresponding level of local feature A in Step 2.3. Subsequently, a cascaded fusion strategy is adopted to fuse the global feature T with the local feature A to promote a comprehensive understanding of the cloud structure and generate more refined detection results. Finally, a synthetic feature set C = {C1, C2, C3, C4} is obtained through this process;
[0117] As Figure 2 shown, first, the output of the E-TransA module is adjusted in size through linear interpolation to make its feature map size consistent with that of the feature map output by the dual-branch attention mechanism; subsequently, the adjusted feature map is concatenated with the feature map output by the dual-branch attention mechanism to fuse global context information and local detail information and further enhance the feature expression ability.
[0118] F resize = Interpolate[Y, size=(H, W)]
[0119] Cat = Concat[W CAout , F resize
[0120] where Y is the output after adjustment of the E-TransA module.
[0121] Step 2.5: A multi-layer perceptron is used to adjust the number of channels of each level of enhanced feature R obtained in Step 2.2. Then, linear interpolation technology is used to further adjust the size of these feature maps to ensure that they meet the conditions required for channel concatenation;
[0122] As Figure 2 shown, first, a multi-layer perceptron is used to adjust the number of channels of the encoded features output by the four-level residual fusion convolution module to 256. At the same time, each level of enhanced feature R is adjusted to the same size as R2 through linear interpolation. Secondly, the adjusted four-level features are concatenated in the channel dimension to fuse multi-scale information. Finally, a 1×1 convolution is used to adjust the number of channels of the fused features to 256, and the specific expression is as follows:
[0123] R C1 , R C2 , R C3 , R C4 = MLP(R1, R2, R3, R4)
[0124] R re1 , R re2 , R re3 , R re4 = Interpolate[R C1 , RC2 , R C3 , R C4 , size = R C2
[0125] R Cat = Concat[R re1 , R re2 , R re3 , R re4
[0126] R out = Conv 1×1 (R Cat )
[0127] Step 2.6: Input the highest-level synthetic feature C4 obtained in Step 2.4 into the convolutional module of the decoder. After processing, it is concatenated with the fused feature obtained in Step 2.5 to achieve further fusion of features;
[0128] As Figure 2 shown, take the highest-level fused feature C4 obtained at the decoding end as the input of the decoder. The dual convolutional module design of the decoder consists of two 3×3 convolutional layers. The main function of the two convolutional layers is to adjust the number of channels of the feature map and effectively restore the detailed information in the image. During the decoding process, the fused feature C4 first undergoes feature extraction through a dual convolutional module, and then the image resolution is enhanced through linear interpolation. Next, the feature with adjusted resolution is concatenated with R out in the channel dimension. This design ensures that during the gradual restoration of the image spatial resolution, the image details can be effectively and finely reconstructed, thereby providing a more comprehensive feature representation for the model and further improving the accuracy and refinement degree of the segmentation decision. The specific expression is as follows:
[0129] D1 = Conv 3×3 [Conv 3×3 (C4)]
[0130] D re = Interpolate[D1, size = R out
[0131] D cat = Concat[D re , R out
[0132] Step 2.7: After obtaining the new synthetic features in Step 2.6, they are fed into the three-level decoder for step-by-step decoding to generate the final output feature map. During the decoding process, the output features of the first two levels of the decoder are respectively concatenated with C1 and C2 in Step 2.4 to make full use of the low-level details and high-level semantics, improving the accuracy and robustness of detection;
[0133] As Figure 2 shown, take D cat as the input of the last three levels of the decoder. Specifically, after the output of the first-level decoding module, the size of the feature map is increased through linear interpolation and concatenated with the feature map C2 in the encoder in terms of channels; after the output of the second-level decoding module, the size of the feature map is also increased through linear interpolation and then concatenated with the feature map C1 in the encoder, aiming to utilize both low-level details and high-level semantics simultaneously to improve the comprehensiveness of feature representation. Finally, after being processed by the third-level decoding module, the output feature map is normalized to generate a feature map with the same size as the original input feature map, thus completing the decoding process and restoring the high-resolution features. The specific expressions are as follows:
[0134] D2 = Conv 3×3 [Conv 3×3 (D cat )], D re2 = Interpolate[D2, size = C2]
[0135] D cat2 = Concat[D re2 , C2]
[0136] D3 = Conv 3×3 [Conv 3×3 (D cat2 )], D re3 = Interpolate[D3, size = C1]
[0137] D cat3 = Concat[D re3 , C1]
[0138] D4 = Conv 3×3 [Conv 3×3 (D cat3 )]
[0139] D out = Sigmoid[Conv 1×1 (D4)]
[0140] Step 3: Use the preprocessed dataset for the training of the cloud detection model and save the best model parameters;
[0141] In step 3, after comprehensively considering the diversity of the dataset and the unique attributes of cloud pixels, the present invention adopts the cross-entropy loss function as the training criterion for the network model and introduces the Adam optimizer to efficiently optimize this loss function. During the training process, the system continuously monitors and saves the best-performing model for subsequent accurate prediction tasks.
[0142] Step 4: Predict the images in the test set from the model parameters saved in step 3 and quantify the model performance through evaluation metrics;
[0143] In step 4, we input the test set data into the previously saved best model to obtain the specific values of various evaluation metrics for measuring the model's performance. At the same time, the validation set data is also input into this model to obtain the corresponding prediction results, thereby further verifying the accuracy and generalization ability of the model.
[0144] The present invention has conducted multiple rigorous experiments on three publicly available datasets (95-Cloud, CHLandsat8, HRC_WHU), and the experimental data results are summarized and shown in Table 1.
[0145] Table 1: Evaluation of the performance metrics of the method of the present invention on three datasets.
[0146]
[0147] Furthermore, Figure 4 Intuitively shows the visualization experimental results of the model of the present invention under different scenarios. These selected images focus on cases where the cloud boundary is complex and the detection of thin cloud regions is extremely difficult. By observing the experimental results, it can be seen that the cloud detection method proposed by the present invention has achieved remarkable results in reducing false detections and missed detections.
[0148] The series of detailed descriptions listed above are only specific descriptions of the feasible embodiments of the present invention, and they are not intended to limit the protection scope of the present invention. Any equivalent embodiments or changes made without departing from the technical spirit of the present invention should be included within the protection scope of the present invention.
Claims
1. A cloud detection method based on dual-branch fusion feature enhancement, characterized in that It includes the following steps: Step 1: Obtain a cloud detection dataset and preprocess the dataset; Step 2: Construct a cloud detection network model based on dual-branch fusion feature enhancement; Step 3: Use the preprocessed dataset for training the cloud detection model and save the model parameters; Step 4: Predict the images in the test set from the model parameters saved in Step 3 and quantify the model performance through evaluation metrics.
2. The cloud detection method based on dual-branch fusion feature enhancement according to claim 1, wherein, In Step 1, three datasets from different sensor types, covering diverse atmospheric conditions, and having different resolution sizes are selected. After selecting the datasets, preprocessing is performed, specifically including: Step 1.1: Perform RGB three-channel merging processing on the acquired images; Step 1.2: Randomly crop the images to a size of 256×256 pixels; Step 1.3: Divide the entire dataset into two parts. 80% of the images are used as the training set, and the remaining 20% are used as the test set and validation set.
3. A cloud detection method based on dual-branch fusion feature enhancement according to claim 1, characterized in that, Step 2 is used to perform the following steps: Step 2.1: Use the E-TransA module as the backbone branch to extract global features. This module captures four-level multi-level features T = {T1, T2, T3, T4} from the input feature map in stages; Step 2.2: After extracting each stage of features in Step 2.1, apply the residual fusion convolution module for feature enhancement processing to obtain a feature representation containing more detailed information. The enhanced four-level features are relabeled as R = {R1, R2, R3, R4}; Step 2.3: Use the dual-branch attention mechanism as the core to extract local features and obtain four-stage multi-level features A = {A1, A2, A3, A4} from the input feature map. This mechanism cleverly combines the advantages of convolution operations and attention mechanisms. While extracting local spatial features of the image, it can dynamically highlight and strengthen the feature performance of key regions, thereby more accurately capturing the detailed features and spatial dependence relationships of the clouds and generating a hierarchical local feature representation; Step 2.4: Each level of global feature T extracted in Step 2.1 will first adjust its feature map size through linear interpolation to ensure that it matches the corresponding level of local feature A in size in Step 2.3, and use the cascade fusion strategy to fuse the global feature T and the local feature A to promote a comprehensive understanding of the cloud structure and generate a more refined detection result. Finally, a synthetic feature set C = {C1, C2, C3, C4} is obtained through this process; Step 2.5: Use a multi-layer perceptron to adjust the number of channels of each level of enhanced feature R obtained in Step 2.2, and then use linear interpolation technology to further adjust the size of these feature maps to ensure that they meet the conditions required for channel splicing; Step 2.6: Input the highest-level synthetic feature C4 obtained in Step 2.4 into the dual convolution module of the decoder. After processing, it is spliced with the fusion feature obtained in Step 2.5 to achieve further feature fusion; Step 2.7: After obtaining the new synthetic features in Step 2.6, send them into the last three-level decoding module for progressive decoding to generate the final output feature map. During the decoding process, the output features of the first two levels of decoders are respectively concatenated with C1 and C2 in Step 2.4 to make full use of low-level details and high-level semantics, improving the accuracy and robustness of detection.
4. A cloud detection method based on dual-branch fusion feature enhancement according to claim 3, characterized in that In Step 2.1, a branch with the E-TransA module as the core architecture is adopted to extract features. It consists of two branches operating in parallel, which respectively focus on the extraction of global and local information. Among them, the local branch mainly uses two convolutional layers with kernel sizes of 1×3 and 3×1 to capture feature information in the horizontal and vertical directions, and then integrates the extracted feature information to synthesize local features; the global branch first uses a 1×1 convolution to perform weighted summation on all channels of each pixel point, realizing information fusion between channels, and then uses a 3×3 convolutional layer to generate query (Q), key (K), and value (V) matrices. To reduce the complexity of the algorithm, after the generation of the K and V matrices, they are respectively processed by a convolution with a stride of r and a kernel size of r, thus realizing downsampling of the feature map. In the four levels of the encoder, the values of r are 8, 4, 2, and 1 respectively. Finally, the outputs of the local branch and the global branch are added together, and the output feature map is obtained after passing through a depthwise separable convolution with a kernel size of 3×3. The detailed process of the multi-head self-attention mechanism is defined as follows: C = Conv 3×1 (X) + Conv 1×3 (X), Conv = Conv 1×1 (X) Q, K, V = Conv 3×3 (Conv), K1 = Conv r×r (K), V1 = Conv r×r (V) where X is the input feature map of the multi-head self-attention mechanism.
5. A cloud detection method based on dual-branch fusion feature enhancement according to claim 3, characterized in that, In step 2.2, a residual fusion convolution module is applied for feature enhancement. A 1×1 convolution is used to learn the weights between different channels, so as to more flexibly adjust the importance of features. Local spatial features are captured through a 3×3 convolution to enhance the sensitivity to image detail information. The output results of the above two convolutions are multiplied element by element to achieve deep fusion of features. The fused features are finely optimized through a 1×1 convolution to generate the output feature map C out ; In order to retain the key information in the input feature map while enhancing the features, we take C out and perform an element-by-element addition operation with the original input feature map C out = Conv 1×1 {Conv 3×3 [Conv 1×1 (Y)] × Conv 1×1 (Y)} Z out = C out + Y where Y is the output of the E-TransA module.
6. The cloud detection method based on dual-branch fusion feature enhancement according to claim 3, wherein The dual-branch attention mechanism proposed in Step 2.3 performs precise local feature extraction on the input image to achieve more detailed cloud detection. First, the input feature map passes through a convolution sequence, which contains i 3×3 convolutions for feature extraction, and the first convolution adjusts the channels and the size of the feature map. Through experimental verification, when i is set to 5, the best recognition and processing accuracy can be achieved. Then, the attention mechanism is used to improve the ability to capture key information. The steps are as follows: Step 2.3.1: Adopt the channel attention mechanism to dynamically adjust the importance weights of each channel in the output feature map of the convolution sequence to enhance the feature expression ability. First, compress the features in the spatial dimension through global average pooling, turning each two-dimensional feature channel into a real number with a global receptive field; Secondly, after obtaining the global information, generate and select the weights of the feature channels in a parameterized manner; finally, multiply the adjusted importance weights element-wise to the original feature map, thus realizing effective recalibration of the original features in the channel dimension, further enhancing the feature expression ability and pertinence. The specific expression is as follows: W CC = F out × Sigmoid(Z out ) Among them, F out is the output after the feature map passes through the convolution sequence; Step 2.3.2: Use the parallel spatial attention mechanism to capture the spatial dependencies in W from different perspectives. First, compress the input feature map W in parallel along the channel dimension to generate two pairs of global average information and global saliency information respectively. Then, concatenate these two types of global information to provide richer input for subsequent generation of spatial attention weights. Next, perform a 7×7 convolution operation on the concatenated feature map and normalize it through the Sigmoid function to learn the relationships between spatial positions and generate the importance weights for each spatial position. CC in, first, CC compress it in parallel along the channel dimension to generate two pairs of global average information and global saliency information respectively; then, concatenate these two types of global information to provide richer input for subsequent generation of spatial attention weights; next, perform a 7×7 convolution operation on the concatenated feature map and normalize it through the Sigmoid function to learn the relationships between spatial positions and generate the importance weights for each spatial position. Subsequently, the generated spatial attention weights are multiplied element-wise with the input feature map W CC to obtain the weighted feature map; finally, the weighted feature maps of the two parallel branches are added together to generate a diverse feature representation. The specific expression is as follows: S1avg = AvgPool(W CC ), S1max = MaxPool(W CC ) S2avg = AvgPool(W CC ), S2max = MaxPool(W CC ) F S1 = W CC × Sigmoid{Conv 7×7 [Cat(S1avg, S1 max)]} F S2 = W CC × Sigmoid{Conv 7×7 [Cat(S2avg, S2 max)]} W SC = F S1 + F S2 Step 2.3.3: Extract the local details and multi-scale context information of the input feature map with the help of the convolutional branch, and the output is F Cout , and the specific expression is as follows: F CB = Conv 1×1 (F out ) E C1 = Conv 1×1 (F CB ), E C2 = Conv 3×3 [Conv 3×3 (F CB )] F Cout = DWConv 3×3 (E C1 × E C2 ); Step 2.3.4: Add the output results of Step 2.3.2 and Step 2.3.3 element-wise to fuse the features of the two branches, realizing the complementarity and enhancement of detailed information; subsequently, further extract the fused features through a 1×1 convolution, and output the final result. The specific expression is as follows: W CAout = Conv 1×1 (F Cout + W SC )。 7. A cloud detection method based on dual-branch fusion feature enhancement according to claim 3, characterized in that In step 2.4, the output of the E-TransA module is used to adjust the size of the feature map through linear interpolation to make it the same as the size of the feature map output by the dual-branch attention mechanism; subsequently, the adjusted feature map is concatenated with the feature map output by the dual-branch attention mechanism to fuse global context information and local detail information, further enhancing the feature expression ability. F resize = Interpolate[Y, size=(H, W)] Cat=Concat[W CAout ,F resize Where Y is the output after adjustment of the E-TransA module.
8. A cloud detection method based on dual-branch fusion feature enhancement according to claim 3, wherein In step 2.5, first, a multi-layer perceptron is used to adjust the number of channels of the encoded features output by the four-level residual fusion convolutional module to 256; at the same time, each level of enhanced feature R is adjusted to the same size as R2 through linear interpolation; secondly, the adjusted four-level features are concatenated in the channel dimension to fuse multi-scale information; finally, a 1×1 convolution is used to adjust the number of channels of the fused features to 256, and the specific expression is as follows: R C1 ,R C2 ,R C3 ,R C4 = MLP(R1, R2, R3, R4) R re1 ,R re2 ,R re3 ,R re4 = Interpolate[R C1 ,R C2 ,R C3 ,R C4 , size = R C2 R Cat = Concat[R re1 , R re2 , R re3 , R re4 R out = Conv 1×1 (R Cat )。 9. A cloud detection method based on dual-branch fusion feature enhancement according to claim 3, characterized in that, In step 2.6, the highest-level fused feature C4 obtained from the encoder is used as the input of the decoder. The convolutional module design of the decoder includes two 3×3 convolutional layers. The main function of these convolutional layers is to adjust the number of channels of the feature map and effectively restore the detailed information in the image. During the decoding process, the fused feature C4 first undergoes feature extraction through the convolutional module, and then the image resolution is enhanced through bilinear interpolation operation. Next, the feature with adjusted resolution is concatenated with R out along the channel dimension, and its specific expression is as follows: D1 = Conv 3×3 [Conv 3×3 (C4)] D re = Interpolate[D1, size = R out D cat = Concat[D re , R out .
10. A cloud detection method based on dual-branch fusion feature enhancement according to claim 3, characterized in that, In step 2.6, D cat is used as the input of the last three levels of the decoder. After the first-level decoding module outputs, the size of the feature map is increased by linear interpolation and channel concatenation is performed with the feature map C2 in the encoder; after the second-level decoding module also increases the size of the feature map by linear interpolation, channel concatenation is performed with the feature map C1 in the encoder, aiming to utilize both low-level details and high-level semantics simultaneously. Finally, after being processed by the third-level decoding module, normalization operation is performed on the output feature map to generate a feature map with the same size as the original input feature map, thus completing the decoding process and restoring the high-resolution features. The specific expression is as follows: D2 = Conv 3×3 [Conv 3×3 (D cat )], D re2 = Interpolate[D2, size = C2] D cat2 = Concat[D re2 , C2] D3 = Conv 3×3 [Conv 3×3 (D cat2 )], D re3 = Interpolate[D3, size = C1] D cat3 = Concat[D re3 , C1] D4 = Conv 3×3 [Conv 3×3 (D cat3 )]D out = Sigmoid[Conv 1×1 (D4)]。
Citation Information
Cited By
Crane with boom structure stress analysis capability
CN121787235A
Frequency modulation and wavelet sub-band guided double-domain cooperative Transform X-ray image denoising method
CN121981912A