A cloud image segmentation method

By improving the U-Net model and combining the self-attention mechanism of the transformer architecture, the problem that existing cloud detection technology is difficult to accurately identify cloud clusters and cloud shadow areas is solved, and higher cloud map segmentation accuracy and telemetry efficiency are achieved.

CN114898227BActive Publication Date: 2025-05-09WUXI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210643793.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-08
Publication Date
2025-05-09
Estimated Expiration
2042-06-08

AI Technical Summary

Technical Problem

Existing cloud detection technology is difficult to accurately identify cloud clusters and their cloud shadow areas. Especially in high-resolution remote sensing satellite images, cloud clusters and their projected shadows are easily confused with the characteristics of the land, affecting the classification, segmentation, change detection and other processes of remote sensing images.

Method used

The improved U-Net model is adopted to optimize the cloud map segmentation method by changing the convolution method, adding efficient channel attention mechanism, modifying the long jump connection method, and using the GeLU activation function, combined with the self-attention mechanism of the transformer architecture.

Benefits of technology

Effectively distinguish dark features such as land, surface shadows, and water bodies, reduce detection error rates, improve the accuracy of cloud telemetry image analysis and calculation, make cloud map prediction more accurate and stable, and improve telemetry efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114898227B_ABST
    Figure CN114898227B_ABST
Patent Text Reader

Abstract

The present invention discloses a cloud image segmentation method, comprising the following steps: S1, pre-processing the image of the visible band of the Sentinel-2 satellite to obtain a data set; S2, constructing an improved U-Net model by changing the convolution mode, adding efficient channel attention, modifying the long jump connection mode and modifying the activation function; S3, inputting the data set obtained in step S1 into the improved U-Net model for training and testing, and performing a cloud image segmentation experiment comparison with other segmentation networks to obtain a comparative output preview image; S4, optimizing the comparative output preview image in step S3 through a transformer architecture to obtain a final output effect image. The present invention uses the introduction of transformers and regression models in the U-Net model to significantly improve the calculation accuracy of telemetry image analysis of clouds, making the prediction of cloud images more accurate and stable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to cloud image detection, and in particular to a cloud image segmentation method. Background Art

[0002] With the development of remote sensing image processing technology, cloud detection, as an important step in remote sensing image preprocessing, has gradually become an issue of concern. The spectral information of clouds is determined by factors such as particle size, water vapor, height, and optical thickness. The spectral characteristics of clouds on images have various manifestations. The brightness, transparency, and texture shape of clouds themselves have different manifestations. Cloud shadows are easily confused with darker features such as land, surface shadows, and water bodies. In high-resolution remote sensing satellite images, clouds and their projected shadows are inevitable. Some areas in the remote sensing images will be contaminated by clouds or even completely covered, which will affect the classification, segmentation, change detection, and image matching of remote sensing images.

[0003] A lot of research has been done on cloud detection technology based on convolutional neural networks at home and abroad. For example, Wu Lifang et al. proposed a cloud image segmentation method based on FCN to achieve pixel-level segmentation. SegNet cleverly uses the encoding-decoding structure to optimize on the basis of FCN, but the advantage is not obvious and it is impossible to fully restore the information. Zhao et al. proposed PSPNet to aggregate more contextual information to achieve high-quality pixel-level scene analysis, but it is slow and takes a long time to train on remote sensing image datasets. Ronneberger et al. proposed U-Net for image segmentation. Its uniqueness is that it uses mirror folding to extrapolate missing contextual information, supplement the semantic information of the input image, and directly splice the feature maps in the codec through jump connections, effectively integrating deep detail information and shallow semantic information. However, this method will equally distribute the information on all spatial positions and channels on the feature tensor, thereby generating a large amount of computational redundancy, resulting in slower model training and lower segmentation accuracy. Summary of the invention

[0004] Purpose of the invention: The purpose of the present invention is to provide a cloud image segmentation method that can accurately identify cloud clusters and their cloud shadow areas.

[0005] Technical solution: The cloud image segmentation method of the present invention comprises the following steps:

[0006] S1, preprocessing the images of the visible band of the Sentinel-2 satellite to obtain a data set;

[0007] S2, an improved U-Net model is constructed by changing the convolution mode, adding efficient channel attention, modifying the long skip connection mode and modifying the activation function;

[0008] S3, inputting the data set obtained in step S1 into the improved U-Net model for training and testing, and performing a cloud image segmentation experiment comparison with the existing segmentation network to obtain a comparison output preview image;

[0009] S4, optimizing the comparison output preview image obtained in step S3 through the transformer architecture to obtain the final output effect image.

[0010] Further, the specific process of step S1 is as follows:

[0011] S11, obtain the images of band 2, band 3, and band 4 of the Sentinel-2 satellite, divide the large image into small blocks, manually annotate the small blocks with the annotation tool Labelme, obtain the corresponding label images, and use them to generate a data set with a size of 224×224×3;

[0012] S12, the data augmentation method is used on the data set to expand the data set to twice its original size, and the augmented data is divided into a training set, a validation set, and a test set.

[0013] Further, the specific process of step S2 is as follows:

[0014] S21, based on the U-Net segmentation model, replace the first convolution block of each layer in the encoding part with a variable convolution block to build an improved U-Net model;

[0015] S22, adding an efficient channel attention mechanism to the splicing operation of the decoding network and the splicing operation of the feature map, respectively. After the feature map output by the encoding part generates a one-dimensional attention vector through the efficient channel attention mechanism, the corresponding elements are multiplied with the original feature map to obtain a weighted feature map. The size of the feature map remains unchanged and the splicing operation is directly performed with the feature map of the decoding part;

[0016] S23, batch normalization is added between the convolution layer and the activation layer of the U-Net network, the original ReLU activation function is replaced by the GeLU activation function, each semantic segmentation category is trained separately by training two categories, and each binary classification training model is merged to obtain an improved U-Net model;

[0017] S24, jump-connect each layer of the decoding part with the feature map of the same layer of the encoding part and the feature map of the adjacent lower layer, ensuring that each layer of the decoding part has three input information streams; the last layer of the decoding part corresponds to the first layer of the same layer of the encoding part, the input information stream of the last layer of the decoding part remains unchanged, and the number of feature map channels after the splicing operation becomes 896, 448, 224, and 96.

[0018] Further, the specific process of step S3 is as follows:

[0019] S31, inputting 80% of the data set in step S1 as a training set into the improved U-Net model for training, and fine-tuning the entire network parameters using the gradient descent algorithm through labeled data supervised learning to obtain the optimal parameter model;

[0020] S32, inputting 10% of the data set in step S1 as a test set into the optimal parameter model in S31 for testing, and outputting a preliminary prediction effect diagram;

[0021] S33, compare the prediction effect graph in S32 with the label graph to obtain a comparison output result of the improved U-Net model.

[0022] Furthermore, in step S4, the comparison output image of the improved U-Net model in step S3 is subjected to Patch-Embedding using a convolutional layer convblock; Flatten is then performed to output a feature vector, and then cosine position encoding Position-Emdedding and a layer of dropout random inactivation are added to the feature vector; the input vector is placed in three different fully connected layers, and a query vector Query, a key vector Key and a value vector Value are output; the specific steps are as follows:

[0023] S41, use dot product to calculate the similarity between Q and K vectors:

[0024] f(Q,K i )=Q T K i

[0025] Among them, f(Q,K i ) is the similarity corresponding to each set of data, i = 1, 2, 3...m, Q is the query vector Query, K i For each key vector Key, Q T is the transpose of Q;

[0026] S42, normalize the similarity through the softmax function:

[0027]

[0028] Where i = 1, 2, 3...m, α i is the normalized similarity;

[0029] S43, perform weighted summation on all values ​​to obtain the Attention vector:

[0030]

[0031] Among them, Vi For each values.

[0032] Compared with the prior art, the present invention has the following significant effects:

[0033] 1. The present invention uses the self-attention mechanism of transformer. By introducing transformer and regression model into the U-Net model, the detection of cloud shadow pollution areas on the edge of cloud images is strengthened. It can effectively distinguish dark features of land, surface shadows, water bodies, etc., reduce the detection error rate, and significantly improve the calculation accuracy of remote sensing image analysis of clouds, making the prediction of cloud images more accurate and stable, and improving the remote sensing efficiency.

[0034] 2. The present invention adopts the U-Net model, which can effectively integrate deep detail information and shallow semantic information, improve the accuracy of remote sensing images, and provide a cloud map segmentation method. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a flow chart of cloud image segmentation of the present invention;

[0036] Figure 2 It is a structural diagram of the U-Net model of the present invention;

[0037] Figure 3 It is a variable convolution structure diagram of the present invention;

[0038] Figure 4 This is a structural diagram of the efficient channel attention mechanism of the present invention;

[0039] Figure 5 The U-shaped cloud image segmentation model based on efficient channel attention of the present invention;

[0040] Figure 6 This is a diagram of a long jump connection method of the present invention;

[0041] Figure 7 The transformer architecture diagram of the present invention;

[0042] Figure 8 It is a comparison chart of the generalization experiment of the present invention. DETAILED DESCRIPTION

[0043] The present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0044] like Figure 1 As shown, it is a cloud image segmentation flow chart of the present invention, which includes the following steps:

[0045] Step 1. The data set used in the present invention comes from the Sentinel-2 satellite, and uses images of the three visible bands of the Sentinel-2 satellite, Band 2 (red), Band 3 (green), and Band 4 (blue). The large image is divided into small blocks, and the small blocks are manually annotated with the labeling tool Labelme. Then, image enhancement methods such as random pruning, translation transformation, and noise perturbation are used to expand the data set to twice the original size, thereby expanding the diversity of existing data.

[0046] Step 2, such as Figure 2 The U-Net model structure diagram is shown below. Figure 3 It is a variable convolution structure diagram. The variable convolution is mainly composed of offset convolution and standard convolution. The standard convolution kernel size used in the present invention is 3×3. For an input feature map, in order to learn the offset offset, another offset convolution kernel size of 3×3 is defined. The output is the same size as the original feature map, and the number of channels is 2N. The variable convolution performs bilinear interpolation based on the offset offset, and then performs the standard convolution. The formula is as follows:

[0047]

[0048] Among them, p 0 is a pixel point in the feature map, y(p 0 ) is the convolution output, x is the set of input pixels, p n is any pixel point on the feature map, w(p n ) is the pixel p n The weight of {Δp n | n=1,2,...,N}(N=|R|) is the offset, R={(-1,-1),(-1,0),...,(0,1),(1,1)}, which defines the size and expansion of the receptive field.

[0049] like Figure 4 As shown, for a feature map U of size W×W×C, U=[x 1 ,x 2 ,...,x c ], perform one-dimensional operation on the feature map U to obtain the one-dimensional feature map Z. One-dimensional operation means to calculate the average value of each feature channel independently, compress each feature channel into a real number, which can characterize the global distribution on the feature channel. The formula is:

[0050]

[0051] Among them, z i ∈Z=[z 1 ,z 2 ,...,z c ], x i∈U=[x 1 ,x 2 ,...,x c ], F GAP (·) means that the feature map in the feature channel c is converted into a real number through linear operation, x i represents the i-th feature map in feature channel c, x i (m,n) represents the pixel value at the i-th feature map position (m,n), w represents the size of the feature map in the feature channel c, i = 1, 2, ..., c.

[0052] After completing the above operations, the feature map of the input feature W×W×C becomes 1×1×C. After that, the weight matrix is ​​constructed using each channel and its k nearest neighbors, that is, for the first channel, its 1st to kth items are non-zero items, and the other items are all zero. In the second channel, the 2nd to k+1th items are non-zero items, and the other items are all zero, and so on. The weight matrix is ​​used to capture the cross-channel interactions between feature maps, where k represents the coverage of local cross-channel interactions, that is, how many close neighbors participate in the attention prediction of a channel. The weight matrix is ​​expressed as follows:

[0053]

[0054] Among them, w c,c-k+1 represents the value of the first cross-channel interaction in feature channel c, w c,c represents the value of the kth cross-channel interaction in feature channel c. Therefore, the attention weight corresponding to the cth channel feature map in feature map U can be expressed as follows:

[0055]

[0056] Among them, w c represents the attention weight corresponding to the cth feature map, and W c =[w 1 ,w 2 ,...,w c ],w c j represents the weight matrix corresponding to the feature map, Ω c k Indicates z c The corresponding set of k adjacent feature channels, For the set Ω c k Further, in order to reduce the parameters to make it lightweight, while ensuring that the weights of each channel and its k neighboring channels can be optimized simultaneously, so that all feature channels share weight information, the above formula is updated to

[0057]

[0058] At this time, the number of parameters of the lightweight adaptive attention mechanism becomes k. For the above updated formula, it can be implemented through one-dimensional convolution. Therefore, in the lightweight adaptive attention mechanism, the information interaction between feature channels is finally completed through one-dimensional convolution with a convolution kernel size of k. The formula can be written as:

[0059] w" c =C1D k (z) (6)

[0060] Among them, C1D represents one-dimensional convolution. After obtaining the attention weight corresponding to the feature map, first use the Sigmoid gate to obtain the normalized weight between 0 and 1. Then use the final weight to weight the feature map U to obtain the optimized feature map. The formula is as follows:

[0061] U'=σ(w c )·U (7)

[0062] Among them, U′ is the optimized feature map of the cth feature channel, σ(w c ) is the weight after normalization using the Sigmoid gate. Through the above operations, the weight is suppressed or enhanced, that is, the significant feature map is enhanced, and the non-significant feature map is suppressed accordingly. The feature map after feature recalibration then enters the following network for learning.

[0063] like Figure 5 As shown, the main difference between the U-shaped cloud image segmentation model based on efficient channel attention and U-Net is whether the result obtained from the encoding part is directly used for decoding. The improved U-Net network can extract richer and more accurate feature information, making the segmentation results and generalization effects more accurate. At the same time, the present invention adds batch normalization between the convolution layer and the activation layer of the U-Net network, replaces the original ReLU activation function with the GeLU activation function, and adopts the method of training binary classification to train each semantic segmentation category separately, and merges the models trained for each binary classification.

[0064] The ReLU function is defined as:

[0065]

[0066] Here, x represents the input quantity.

[0067] The GeLU function is defined as:

[0068]

[0069] The activation function improves the nonlinear modeling ability of the network and defines the mapping relationship between input and output. When x≤0, the output of the ReLU function is 0, which will cause the death of neurons; the GeLU function effectively solves the problem of neuron death and improves the anti-noise performance of the activation function.

[0070] like Figure 6 As shown, En1 to En5 refer to each layer of the network coding part, and De4 to De1 refer to each layer of the network decoding part. In order to make the contour of the cloud map segmentation closer to the real label, the present invention jumps each layer of the improved U-Net decoding part with the feature map of the same layer of the encoding part and the feature map of the adjacent lower layer. Therefore, each layer of the decoding part has three input information streams. In addition to the input information of the next layer and the input information of the same layer of the corresponding encoding part, the low-level input information of the previous layer of the encoding part is also added. Since the output feature map size of the previous layer of the encoding part is twice the scale of the feature map of the current layer, the output feature map of the previous layer of the encoder is first subjected to the maximum pooling operation so that the feature map size is the same as the current feature map size. Since the last layer of the decoding part corresponds to the first layer of the same layer of the encoding part, and there is no previous layer, De1 has two input information streams as before.

[0071] Step 3: Input the training data into the model for training. Through supervised learning with labeled data, use the gradient descent algorithm to fine-tune the entire network parameters. Test the best trained model weights with test data and directly output the final prediction effect graph. Figure 7As shown in the figure, the data set size of the input network is 224×224×3. The encoding part has five layers. The first four layers are composed of convolution blocks, variable convolution blocks and maximum pooling modules. The convolution blocks include 3×3 convolution kernels, batch normalization bn and activation function Gelu. The variable convolution blocks include offset convolution kernels and the same convolution blocks as the same layer. There is no maximum pooling layer in the fifth layer; the decoding part has four layers, all of which are composed of upsampling modules, splicing operations, and two convolution blocks. A 1×1 convolution kernel is added at the end of the fourth layer to classify the cloud map. The 224×224×3 feature map is input to the first layer of the encoding part, and the convolution block conv11 outputs a 224×224×32 feature map, and the variable convolution block deform_conv11 outputs a 224×224×32 feature map, and the pooling layer Down1 outputs a 112×112×32 feature map; the 112×112×32 feature map is input to the second layer of the encoding part, and the convolution block conv12 outputs a 112×112×64 feature map, and the variable convolution block deform_conv12 outputs a 112×112×64 feature map. The feature map of 56×56×64 is output through the pooling layer Down2; the 56×56×64 feature map is input to the third layer of the encoding part, and the feature map of 56×56×128 is output through the convolution block conv13, and the feature map of 56×56×128 is output through the variable convolution block deform_conv13, and the feature map of 28×28×128 is output through the pooling layer Down3; the feature map of 28×28×128 is input to the fourth layer of the encoding part, and the feature map of 28×28×256 is output through the convolution block conv14, and the feature map of 28×28×256 is output through the variable convolution block deform_conv13. form_conv14 outputs a 28×28×256 feature map, and after the pooling layer Down4, it outputs a 14×14×256 feature map; the 14×14×256 feature map is input to the fifth layer of the encoding part, and after the convolution block conv15, it outputs a 14×14×512 feature map, and after the variable convolution block deform_conv15, it outputs a 14×14×512 feature map; the 14×14×512 feature map is input to the first layer of the decoding part, and after upsampling Up4, it outputs a 28×28×512 feature map, and after the concatenation operation Concat4, it is connected The feature maps output by Up4, deform_conv14 and Down3 are connected to obtain a 28×28×896 feature map, which is output by two convolution blocks conv24 to 28×28×256; the 28×28×256 feature map is input to the second layer of the decoding part, and the upsampling Up3 outputs a 56×56×256 feature map, which is connected by the concatenation operation Concat3 to obtain a 56×56×448 feature map, which is output by two convolution blocks conv23 to 56×56×128;The 56×56×128 feature map is input to the third layer of the decoding part, and the feature map of 112×112×128 is output after upsampling Up2. The feature map output by Up2, deform_conv12 and Down1 is connected through the concatenation operation Concat2 to obtain a feature map of 112×112×224. After two convolution blocks conv22, the output is 112×112×64; the 112×112×64 feature map is input to the fourth layer of the decoding part, and the feature map of 224×224×64 is output after upsampling Up1. After the concatenation operation Concat1, the feature map output by Up1 and deform_conv11 is connected to obtain a feature map of 224×224×96. After two convolution blocks conv21, the output is 224×224×32. Finally, after 1×1 convolution, the segmentation result feature map is 224×224×3. ;

[0072] like Figure 7 As shown, the improved U-Net comparison output image 224×224×3 is patch-embedded using a convolutional layer convblock, where the convblock consists of 16 standard convolution kernels with a step size of 1, a padding of 16, and a size of 16×16. The Flatten expansion is then performed to output a 196×768 feature vector. Subsequently, the cosine position encoding Position-Emdedding and a layer of dropout random inactivation are added to the 196×768 feature vector, and the output is a 197×768 vector. The input 197×768 is divided into 49 (2, 2, 768) vectors and put into three different fully connected layers, outputting Q, K, and V vectors (i.e., query vector Query, key vector Key, and value vector Value). The vector sizes are all (2, 2, 256) and multiplied by three weight matrices. The specific steps of the transformer formula are as follows:

[0073] Step 31, use dot product to calculate the similarity between Q and K vectors:

[0074] f(Q,K i )=Q T K i (10)

[0075] Among them, f(Q,K i ) is the similarity corresponding to each set of data, i = 1, 2, 3...m, Q is the query vector Query, K i For each key vector Key, Q T is the transpose of Q;

[0076] Step 32, normalize the similarity using the softmax function:

[0077]

[0078] Where i = 1, 2, 3...m, α i It is the similarity.

[0079] Step 33: Perform weighted summation on all values ​​to obtain the Attention vector:

[0080]

[0081] Among them, V i That is, for each value.

[0082] Finally, the output is (2, 2, 768). The output (49, 2, 2, 768) is concatenated into a feature vector of (196, 768), and then reshaped into a feature map of (224, 224, 3). Finally, a convolution layer is added, which consists of 3 standard convolution kernels with a step size of 1, a padding of 0, and a size of 1×1, and the resulting image is output.

[0083] like Figure 8 As shown in Figure 2, the improved U-Net segmentation model is compared with other segmentation networks in cloud segmentation experiments. Figure 8 It can be seen that the experiment selected four images with different distributions of clouds and cloud shadows in the data set. In experiment 1, clouds are mostly distributed below the cloud shadow, and the background area is small; in experiment 2, clouds are mostly distributed to the right of the cloud shadow, and the background area is small; in experiment 3, clouds are mostly distributed below the cloud shadow, and the background area is large; in experiment 4, clouds are mostly distributed to the upper right of the cloud shadow, and the background area is large. By comparing the segmentation of four remote sensing images with different distributions, it can be seen that the improved U-Net has the best generalization effect. The details and edges in the cloud image are clearer than the generalization effects of other models, and it can better complete the cloud and cloud shadow segmentation task.

Claims

1. A cloud image segmentation method, characterized in that: The steps include: S1, preprocess the images of the visible band of the Sentinel-2 satellite to obtain a data set; S2, an improved U-Net model is constructed by changing the convolution mode, adding efficient channel attention, modifying the long skip connection mode and modifying the activation function; S3, inputting the data set obtained in step S1 into the improved U-Net model for training and testing, and performing a cloud image segmentation experiment comparison with the existing segmentation network to obtain a comparison output preview image; The specific process is as follows: S31, inputting 80% of the data set in step S1 as a training set into the improved U-Net model for training, and fine-tuning the entire network parameters using the gradient descent algorithm through labeled data supervised learning to obtain the optimal parameter model; S32, inputting 10% of the data set in step S1 as a test set into the optimal parameter model in S31 for testing, and outputting a preliminary prediction effect diagram; S33, comparing the prediction effect graph in S32 with the label graph to obtain a comparison output result of the improved U-Net model; S4, use a convolutional layer convblock to complete Patch-Embedding of the comparison output image of the improved U-Net model in step S3; then Flatten to expand the output feature vector, then add cosine position encoding Position-Emdedding and a layer of dropout random inactivation to the feature vector; put the input vector into three different fully connected layers, output the query vector Query, key vector Key and value vector Value, and obtain the final output effect diagram; the specific steps are as follows: S41, use dot product to calculate the similarity between Q and K vectors: f(Q,K i )=Q T K i Among them, f(Q,K i ) is the similarity corresponding to each set of data, i = 1, 2, 3…m, Q is the query vector Query, K i For each key vector Key, Q T is the transpose of Q; S42, normalize the similarity through the softmax function: Where i = 1, 2, 3…m, α i is the normalized similarity; S43, perform weighted summation on all values ​​to obtain the Attention vector: Among them, V i For each values.

2. The cloud image segmentation method according to claim 1, characterized in that: The specific process of step S1 is as follows: S11, obtain the images of band 2, band 3, and band 4 of the Sentinel-2 satellite, divide the large image into small blocks, manually annotate the small blocks with the annotation tool Labelme, obtain the corresponding label images, and use them to generate a data set with a size of 224×224×3; S12, a data augmentation method is used on the data set to expand the data set to twice its original size, and the augmented data is divided into a training set, a validation set, and a test set.

3. The cloud image segmentation method according to claim 1, characterized in that: The specific process of step S2 is as follows: S21, based on the U-Net segmentation model, replace the first convolution block of each layer in the encoding part with a variable convolution block to build an improved U-Net model; S22, adding an efficient channel attention mechanism to the splicing operation of the decoding network and the splicing operation of the feature map, respectively. After the feature map output by the encoding part generates a one-dimensional attention vector through the efficient channel attention mechanism, the corresponding elements are multiplied with the original feature map to obtain a weighted feature map. The size of the feature map remains unchanged and the splicing operation is directly performed with the feature map of the decoding part; S23, batch normalization is added between the convolution layer and the activation layer of the U-Net network, the original ReLU activation function is replaced by the GeLU activation function, each semantic segmentation category is trained separately by training two categories, and each binary classification training model is merged to obtain an improved U-Net model; S24, jump-connect each layer of the decoding part with the feature map of the same layer of the encoding part and the feature map of the adjacent lower layer, ensuring that each layer of the decoding part has three input information streams; the last layer of the decoding part corresponds to the first layer of the same layer of the encoding part, the input information stream of the last layer of the decoding part remains unchanged, and the number of feature map channels after the splicing operation becomes 896, 448, 224, and 96.

Citation Information

Patent Citations

  • Remote sensing image cloud layer detection method based on deep convolutional neural network

    CN112749621A

  • Lightweight remote sensing image cloud detection method

    CN114120036A