Unet-based cloud and cloud shadow improved semantic segmentation method
By introducing an improved ASPP spatial pyramid pooling module and attention mechanism into the Unet model, combined with the design of edge detection subnet, the problems of missed detection, false detection and boundary detection in complex backgrounds by traditional cloud and cloud shadow detection methods are solved, and the cloud image recognition accuracy and edge information extraction capabilities are achieved.
Patent Information
- Application Number
- CN202510372456.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-03-27
AI Technical Summary
When traditional cloud and cloud shadow detection methods face different types of clouds and cloud shadows, there are detection challenges. For example, the threshold method is prone to missed or misdetection in complex backgrounds, high light reflection interference, high algorithm complexity and calculation amount, and rough boundary detection, and low edge information extraction ability.
The improved semantic segmentation method of cloud and cloud shadow based on Unet is adopted, and the downsampling of the Unet model is replaced as an improved ASPP spatial pyramid pooling module, and the attention mechanism module is embedded in the jump connection to build a semantic segmentation subnet; at the same time, an edge detection subnet is built, and the edge detection feature map is obtained through four convolution stages and deconvolution operations. Finally, the semantic segmentation feature map and edge detection feature map are weighted and fusion is performed to output the cloud and cloud shadow segmentation results.
It improves the accuracy of cloud map recognition, enhances the complementarity of edge features and semantic features, reduces isolated misdetection in edge detection, reduces the rate of misjudgment and misjudgment, and is especially more complementarity to edges and semantics in complex backgrounds.
Smart Images

Figure CN119942306A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to an improved semantic segmentation method for clouds and cloud shadows based on Unet, which is particularly suitable for application scenarios such as surface feature extraction, climate detection, and atmospheric correction. Background Art
[0002] The detection and segmentation of clouds and cloud shadows are of great significance in remote sensing image processing, and are of great help in surface feature extraction, climate detection, etc. The detection of clouds and cloud shadows can effectively reduce their interference in the identification of objects, especially in the fields of vegetation cover, land use change and water resources monitoring, and can help predict the trend of climate change. At the same time, observing the changes in cloud shadows can better evaluate the temporal and spatial distribution of solar energy resources, and provide a scientific basis for the site selection of solar power stations and production activities such as agriculture and industry.
[0003] Traditional cloud and cloud shadow detection methods are mainly divided into three categories: one is based on statistics, one is the statistical equation method, which uses data samples to build a mathematical model, detects cloud layers by calculating parameters such as brightness and reflectivity, and identifies cloud shadows based on geometric relationships. Another is the cluster analysis method, which clusters and groups remote sensing image pixels to identify and distinguish cloud layers; the second is based on spectral thresholds, which uses the spectral characteristics of single-phase or multi-phase images to set threshold detection, compares the grayscale value or reflectivity of each pixel in the image with a predefined threshold, and classifies the pixels as cloud shadows or non-cloud shadows based on the results; the third is based on morphological and texture features, and detects and segments based on the typical features of clouds and cloud shadows. However, traditional methods face detection challenges for different types of clouds and cloud shadows. For example, the threshold method is prone to missed detection or false detection under complex backgrounds, and light reflection can also cause greater interference, requiring additional discrimination rules, increasing algorithm complexity and computational complexity. Affected by factors such as noise interference, the accuracy of marking and detection efficiency are relatively low, and the detection of cloud and cloud shadow boundaries is very rough, and the ability to extract edge information is not high.
[0004] With the continuous advancement of deep learning technology, semantic segmentation networks have been introduced into remote sensing image processing. The network model based on CNN has performed well in image classification tasks, and has paved the way for the development of pixel-level classification tasks, namely semantic segmentation. Some researchers first introduced the concept of fully convolutional neural network FCN, and achieved the classification of image pixels by improving the CNN network, thus achieving the effect of image segmentation. Subsequently, the encoder-decoder structure was introduced, and pixel-level classification was achieved through maximum pooling index transfer. By designing a U-shaped network (Unet) based on encoder and decoder, and by establishing a jump connection between the encoder and decoder, the detailed information of the image is effectively retained. UNet++ further optimized and improved the UNet structure in medical image segmentation, and effectively enhanced the feature expression ability of the network with the help of nested and densely connected modules. Some researchers have also proposed an image segmentation method called DeepLab, which uses a deep convolutional neural network (DCNN) structure, uses a variety of different network layers and atrous convolution to expand the receptive field of each layer of the network, and uses the fully connected random field CRF to refine the feature information to achieve accurate semantic segmentation of the image. However, these networks often ignore the interaction between local and global features, making it impossible to accurately understand information when complex features or interfering noise exist.
[0005] The invention with patent publication number CN114943876A discloses a cloud and cloud shadow detection method, device and storage medium with multi-level semantic fusion. It uses residual network (ResNet) as the backbone network and combines the encoder-decoder structure to propose a cloud image segmentation method including a multi-branch residual context semantic module, a multi-scale convolution sub-channel attention module and a feature fusion upsampling module. However, the invention focuses on cloud edge and thin cloud detection. When there are many types or shapes of clouds and cloud shadows in the cloud image, the edge detection of the patent depends on the feature extraction of the main network and lacks independent branches, which will destroy the integrity of different features. In addition, the boundary segmentation is not sharp enough, and the cloud shadow edge will appear jagged or discontinuous, thereby affecting the effect of cloud shadow segmentation, and the misjudgment and missed judgment rate is high. At the same time, the fusion method adopted by the patent will also lead to information loss in the aforementioned complex background. When high-level features are fused, its low-contrast texture may be smoothed by the pooling operation, resulting in the loss of texture information in the thin cloud area. If the fusion does not fully separate the semantics and edge channels, the model may misjudge the shadow of the ground object as a cloud shadow. Summary of the invention
[0006] The purpose of the present invention is to provide an improved semantic segmentation method for clouds and cloud shadows based on Unet. For complex cloud images with many types or shapes of clouds and cloud shadows, a new edge detection and feature fusion method is proposed, which improves the accuracy of cloud image recognition by enhancing the complementarity of edge features and semantic features.
[0007] In order to achieve the above technical objectives, the technical solution adopted by the present invention is: The present invention discloses an improved semantic segmentation method for clouds and cloud shadows based on Unet, the method comprising the following steps: S1, replace the downsampling of the Unet model with the improved ASPP spatial pyramid pooling module, and embed the attention mechanism module in the skip connection to construct a semantic segmentation subnetwork; use the semantic segmentation subnetwork to process the input cloud map, and obtain the semantic segmentation feature map through upsampling; S2, connect the four convolutional layers in sequence to form four convolutional stages. After adding a 1×1 convolution operation in each convolutional stage, use the deconvolution operation to restore the feature maps of the last three convolutional stages to the same scale as the input image. Finally, fuse the feature maps obtained in the four convolutional stages through a 1×1 convolution operation to construct an edge detection subnetwork. Use the edge detection subnetwork to process the input cloud map to obtain an edge detection feature map. S3 performs weighted fusion on the input cloud map, semantic segmentation feature map and edge detection feature map, and then processes the fused features through the Sigmoid activation function layer to output the cloud and cloud shadow segmentation results.
[0008] Further, in step S2, the improved ASPP spatial pyramid pooling module includes a first semantic convolution layer, a second semantic convolution layer, a third semantic convolution layer, a fourth semantic convolution layer, an average pooling layer, a maximum pooling layer, a semantic feature fusion layer and a 1×1 convolution layer; The first semantic convolution layer, the second semantic convolution layer, the third semantic convolution layer, and the fourth semantic convolution layer respectively extract features of the input cloud image; the maximum pooling layer performs a maximum value operation on the local area of the input cloud image to extract local detail features of the input cloud image including the edge and texture of the cloud shadow; the average pooling layer performs an average operation on the local area of the input cloud image to extract global context features of the input cloud image including the overall shape of the cloud and the distribution of the cloud and cloud shadows.
[0009] Furthermore, the first semantic convolution layer uses a 1×1 convolution kernel to perform feature extraction on the input cloud map and extracts preliminary features of the cloud map; the second semantic convolution layer uses a 3×3 convolution kernel and sets the hole rate to 6 to capture a larger range of contextual information in the cloud map; the third semantic convolution layer uses a 3×3 convolution kernel and the hole rate is increased to 12 to capture a larger range of contextual information in the cloud map and extract large-scale features between clouds and cloud shadows in the cloud map; the fourth semantic convolution layer uses a 3×3 convolution kernel and the hole rate is increased to 18 to capture the largest range of contextual information in the cloud map as the global features between clouds and cloud shadows in the cloud map; the semantic feature fusion layer concats the feature maps from different convolution layers and pooling layers, integrates local details and global contextual information, and generates a richer feature representation; the 1×1 convolution layer performs further convolution operations on the feature maps fused by the semantic feature fusion layer, adjusts the number of channels of the feature maps, and finally generates output feature maps for subsequent tasks.
[0010] Further, in step S2, the edge detection subnetwork includes a first edge convolution layer, a second edge convolution layer, a third edge convolution layer, a fourth edge convolution layer and an edge feature fusion layer connected in sequence, and four 1×1 convolution layers connected to the first convolution layer, the second convolution layer, the third convolution layer and the fourth convolution layer respectively; A pooling layer is inserted into each of the first edge convolution layer, the second edge convolution layer, the third edge convolution layer, and the fourth edge convolution layer. The first edge convolution layer, the second edge convolution layer, the third edge convolution layer, and the fourth edge convolution layer sequentially extract features of the input cloud image to obtain cloud shadow edge features of different scales, and reduce the resolution of the feature image through the corresponding pooling layers. The output features of the first edge convolution layer are subjected to a 1×1 convolution operation to obtain a feature map of the first convolution stage; the output features of the second edge convolution layer, the third edge convolution layer, and the fourth edge convolution layer are subjected to a 1×1 convolution operation respectively, and then the feature maps of the corresponding convolution stages are restored to the same scale as the input cloud map by using a deconvolution operation, thereby obtaining a feature map of the second convolution stage, a feature map of the third convolution stage, and a feature map of the fourth convolution stage respectively; The edge feature fusion layer is a 1×1 convolution layer, which fuses the feature map of the first convolution stage, the feature map of the second convolution stage, the feature map of the third convolution stage, and the feature map of the fourth convolution stage to obtain an edge detection feature map that simultaneously contains edge information of cloud shadow areas in different ranges.
[0011] Furthermore, the first edge convolution layer, the second edge convolution layer, the third edge convolution layer, and the fourth edge convolution layer all adopt 3×3 convolution layers, and the number of channels are 64, 128, 256, and 512, respectively.
[0012] Step S3 further comprises: Input cloud map , semantic segmentation feature map and edge detection feature map Perform splicing and fusion to obtain the fused feature map : ; The fused feature map The number of feature channels included is the input cloud map , semantic segmentation feature map and edge detection feature map The sum of the number of channels; The fused feature map The average pooling and maximum pooling operations are performed on each feature channel to obtain the feature maps The global average and maximum value for each feature channel: ; ; Among them, i represents the index of the feature channel, ; Weights are assigned to each feature channel through multiplication calculation to extract effective feature information; the multiplication calculation formula is as follows: ; in It represents the eigenvalue of the i-th channel on the feature map after weighted fusion; Represented as the eigenvalue of the kth channel on the feature map; It is represented as the weight coefficient of the i-th channel; N is represented as the number of feature channels.
[0013] Compared with the prior art, the present invention has the following beneficial effects: The improved semantic segmentation method of cloud and cloud shadow based on Unet of the present invention proposes a new edge detection and feature fusion method for complex cloud images with many types or shapes of clouds and cloud shadows, effectively splices the original image, semantic segmentation features and edge detection features, and uses concat fusion to splice semantic features and edge features in the channel dimension, completely retaining the original information of the two, retaining the integrity of the original features, and reducing isolated false detections in edge detection. The present invention deploys an attention mechanism and an improved ASPP void space pyramid pooling module in the semantic segmentation subnetwork, so that the model can adaptively adjust the feature expression of cloud and cloud shadow feature maps, and improve the model's ability to capture key information; at the same time, it captures rich multi-scale information at the local feature level, and effectively reduces the number of parameters compared to ordinary convolution operations; the fusion of edge features and semantic features effectively improves the model's ability to extract edge features, and obtains clearer and more complete image target edge information. The present invention intuitively combines the advantages of different subnetworks, and also avoids the information loss that may be caused by early fusion, especially in complex backgrounds, the complementarity of edges and semantics is stronger, and the misjudgment and missed judgment rates are low. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a schematic diagram of the improved semantic segmentation of clouds and cloud shadows based on Unet designed by the present invention; Figure 2 It is a schematic diagram of the CBAM designed by the present invention; Figure 3 It is a schematic diagram of the improved ASPP designed by the present invention; Figure 4 It is a schematic diagram of the edge detection subnetwork designed by the present invention; Figure 5 It is a schematic diagram of the feature fusion module designed in the present invention.
[0015] Figure 6 Schematic diagram of the recognition results of the present invention, wherein (a) is a cluster cloud image over a large area of vegetation-covered land in the remote sensing image data set, and (e) is the experimental result image corresponding to (a); (b) is a cluster cloud image over the coast in the remote sensing image data set, and (f) is the experimental result image corresponding to (b); (c) is a cluster cloud image over a barren mountain in the remote sensing image data set, and (g) is the experimental result image corresponding to (c); (d) is a fragmented and scattered cloud image over the ocean in the remote sensing image data, and (h) is the experimental result image corresponding to (d). DETAILED DESCRIPTION
[0016] The embodiments of the present invention are further described in detail below in conjunction with the accompanying drawings.
[0017] The present invention discloses an improved semantic segmentation method for clouds and cloud shadows based on Unet, the method comprising the following steps: S1, replace the downsampling of the Unet model with the improved ASPP spatial pyramid pooling module, and embed the attention mechanism module in the skip connection to construct a semantic segmentation subnetwork; use the semantic segmentation subnetwork to process the input cloud map, and obtain the semantic segmentation feature map through upsampling; S2, connect the four convolutional layers in sequence to form four convolutional stages. After adding a 1×1 convolution operation in each convolutional stage, use the deconvolution operation to restore the feature maps of the last three convolutional stages to the same scale as the input image. Finally, fuse the feature maps obtained in the four convolutional stages through a 1×1 convolution operation to construct an edge detection subnetwork. Use the edge detection subnetwork to process the input cloud map to obtain an edge detection feature map. S3 performs weighted fusion on the input cloud map, semantic segmentation feature map and edge detection feature map, and then processes the fused features through the Sigmoid activation function layer to output the cloud and cloud shadow segmentation results.
[0018] like Figure 1 As shown in the figure, the input of the model is divided into three parts: semantic segmentation subnetwork, edge detection subnetwork and feature fusion module. The implementation steps mainly include: changing the downsampling of the semantic segmentation subnetwork to the improved ASPP spatial pyramid pooling module, and embedding the attention mechanism module (CBAM) in the jump connection to obtain the semantic segmentation feature map through upsampling; in the edge detection subnetwork, the contour feature information of the input image is obtained through four consecutive stages of convolutional layers to obtain the edge detection feature map; the input feature map, semantic segmentation feature map and edge detection feature map are weightedly fused through the feature fusion module, and finally the image is output through Sigmoid.
[0019] like Figure 2 The CBAM shown is an adaptive feature importance learning mechanism widely used in convolutional neural networks, which can improve network performance. When a feature map is given, the module will autonomously infer the attention weight values along two independent dimensions in sequence, and then multiply the generated attention weight map by the input feature map for adaptive refinement. The mechanism consists of two parts: the channel attention module and the spatial attention module: the channel attention module learns and highlights the importance of each channel in feature extraction through weighted averaging; the spatial attention module learns and emphasizes the importance of each spatial position in feature extraction through weighted averaging, effectively highlighting the target area and suppressing irrelevant information. The formula is as follows: ; ; ; ; in, A feature map representing the input; Represents the Sigmoid function; Represents the feature map after channel attention; Input feature map representing spatial attention; Represents the feature map after spatial attention; Represents the weight feature vector output after the CBAM fusion attention operation.
[0020] The Unet network can capture features at different levels through the encoder-decoder structure, but its ability to extract multi-scale information is relatively limited. In the encoder stage, it mainly focuses on extracting local information, and in the decoder stage, although upsampling and feature fusion are performed, it may still not be able to effectively integrate local information with global information, resulting in insufficient utilization of the receptive field, affecting the understanding of the overall structure and target of the image.
[0021] As a means to effectively expand the receptive field of ordinary convolution kernels, dilated convolution has always been favored by researchers in the field of segmentation. Its ability lies in: significantly expanding the receptive field of the convolution kernel without increasing the number of model parameters and computational complexity. Adding intervals to ordinary convolution kernels to form dilated convolutions allows the convolution kernel to cover a larger range without changing the size or number of layers of the convolution kernel. The relationship between the convolution kernel range of dilated convolution and its dilation rate is shown in the formula: ; Where r is the void ratio; is the range of the initial convolution kernel; is the actual receptive field size of the dilated convolution.
[0022] However, the original ASPP lacks balanced processing of local and overall features, resulting in insufficient extraction of some information. Therefore, this paper adds an Avgpool and a Maxpool to the ASPP module: a) AvgPool (average pooling): By averaging the local areas of the feature map, more global context information can be retained, which helps the network capture the overall structure of the image.
[0023] b) MaxPool: By performing the maximum value operation on the local area of the feature map, more local detail information can be retained, which helps the network capture the edge, texture and other detailed features of the image.
[0024] By combining AvgPool and MaxPool, the module can retain both global context information and local detail information, thus achieving a better balance in the feature extraction process. Figure 3As shown in the figure, the improved ASPP spatial pyramid pooling module includes a first semantic convolution layer, a second semantic convolution layer, a third semantic convolution layer, a fourth semantic convolution layer, an average pooling layer, a maximum pooling layer, a semantic feature fusion layer and a 1×1 convolution layer; the first semantic convolution layer, the second semantic convolution layer, the third semantic convolution layer and the fourth semantic convolution layer respectively extract features of the input cloud image; the maximum pooling layer performs a maximum value operation on the local area of the input cloud image to extract local detail features of the input cloud image including the edge and texture of the cloud shadow; the average pooling layer performs an average operation on the local area of the input cloud image to extract global context features of the input cloud image including the overall shape of the cloud and the distribution of the cloud and the cloud shadow. Specifically, the first semantic convolution layer uses a 1×1 convolution kernel to extract features from the input cloud map and extracts preliminary features of the cloud map; the second semantic convolution layer uses a 3×3 convolution kernel and sets the hole rate to 6 to capture a larger range of contextual information in the cloud map; the third semantic convolution layer uses a 3×3 convolution kernel and the hole rate is increased to 12 to capture a larger range of contextual information in the cloud map and extract large-scale features between clouds and cloud shadows in the cloud map; the fourth semantic convolution layer uses a 3×3 convolution kernel and the hole rate is increased to 18 to capture the largest range of contextual information in the cloud map as the global features between clouds and cloud shadows in the cloud map; the semantic feature fusion layer concats feature maps from different convolution layers and pooling layers, integrates local details and global contextual information, and generates a richer feature representation; the 1×1 convolution layer performs further convolution operations on the feature maps fused by the semantic feature fusion layer, adjusts the number of channels of the feature maps, and finally generates output feature maps for subsequent tasks. This combination can enhance the network's ability to extract multi-scale features, especially when dealing with complex segmentation tasks, and can better cope with targets of different scales.
[0025] In the edge detection subnetwork, the convolution layer is divided into four stages. A pooling layer is inserted in the middle of each of the four convolution stages to reduce the dimension of the information extracted by the convolution. After adding a 1×1 convolution operation in each convolution stage, the deconvolution operation is used to restore the feature map to the same scale as the input image because it is necessary to ensure the same scale as the input image. Finally, the feature maps obtained in the four stages are fused through a 1×1 convolution operation. This can effectively obtain the contour of the image and obtain clearer and more complete image target edge information, such as Figure 4As shown in the figure, the edge detection subnetwork includes a first edge convolution layer, a second edge convolution layer, a third edge convolution layer, a fourth edge convolution layer and an edge feature fusion layer connected in sequence, and four 1×1 convolution layers connected to the first convolution layer, the second convolution layer, the third convolution layer and the fourth convolution layer respectively; a pooling layer is inserted in each of the first edge convolution layer, the second edge convolution layer, the third edge convolution layer and the fourth edge convolution layer, and the first edge convolution layer, the second edge convolution layer, the third edge convolution layer and the fourth edge convolution layer extract features of the input cloud image in sequence to obtain cloud shadow edge features of different scales, and reduce the resolution of the feature map through their corresponding pooling layers; the output features of the first edge convolution layer are subjected to a 1×1 convolution operation to obtain the first convolution stage feature map; the output features of the second edge convolution layer, the third edge convolution layer and the fourth edge convolution layer are respectively subjected to a 1×1 convolution operation to obtain the first convolution stage feature map. After the convolution operation, the deconvolution operation is used to restore the feature map of the corresponding convolution stage to the same scale as the input cloud map, and the feature map of the second convolution stage, the feature map of the third convolution stage and the feature map of the fourth convolution stage are obtained respectively; the edge feature fusion layer is a 1×1 convolution layer, which fuses the feature map of the first convolution stage, the feature map of the second convolution stage, the feature map of the third convolution stage and the feature map of the fourth convolution stage to obtain an edge detection feature map that contains edge information of cloud shadow areas of different ranges at the same time. Preferably, the first edge convolution layer, the second edge convolution layer, the third edge convolution layer and the fourth edge convolution layer all use 3×3 convolution layers, and the number of channels are 64, 128, 256 and 512 respectively.
[0026] 1) Input feature map The input feature map is denoted as I, and its size is H×W×C, where H is the height, W is the width, and C is the number of channels.
[0027] 2) The network is divided into four convolution stages, each of which contains convolution operations and pooling operations. The calculation formula is as follows: ; in, is the result of convolution operation at each stage, Indicates the number of network layers, ; Indicates Layer convolution kernel size; Indicates The number of convolution output channels of the layer; Indicates the pooling kernel size; 3) After each convolution stage, a 1×1 convolution operation is added to further extract features and reduce the number of channels. The formula is as follows: ; in, For the stage The result of convolution after pooling; 4) In order to restore the feature map to the same scale as the input image, the deconvolution operation is used to upsample the feature map at each stage. The formula is as follows: ; in, is the result of deconvolution, is the deconvolution kernel size; 5) The feature maps of the four stages are transformed through 1×1 convolution operation Fusion is performed to obtain the final feature map: ; in, It is the final output feature map of the edge detection sub-network; After the semantic segmentation subnetwork and the edge detection subnetwork extract semantic features and edge features respectively, the two features are fused. The concat fusion method can completely retain the original information of each feature map involved in the fusion without losing features. Therefore, this paper designs a feature fusion module based on the concat fusion method to perform weighted fusion of the two features. Figure 5 shown.
[0028] 1) Feature Map The feature map input by this module is divided into three parts, namely the original image , edge detection subnetwork output feature map And the output feature map of the semantic segmentation sub-network is .
[0029] 2) Concatenate the input feature map, semantic feature map and edge feature map to obtain the fused feature map : .
[0030] 3) Perform average pooling and maximum pooling operations on the fused feature map to obtain the global average and maximum value of each feature channel respectively: ; ; Among them, i represents the index of the feature channel, In semantic segmentation, feature channels refer to the number of channels in the feature map output by each layer of the convolutional neural network (CNN). Each channel represents a specific feature that is extracted from the input image through the convolution operation. In this step, feature channels refer to the fused feature map. When performing average pooling and maximum pooling operations, each channel is performed separately to calculate the global mean and maximum value of each channel. It is formed by concatenating the input cloud map, semantic segmentation feature map and edge detection feature map. The number of its channels is the sum of the three concats. Each channel corresponds to a specific feature, which may be edges, textures or more advanced features.
[0031] 4) Assign weights to each feature channel through multiplication calculation to highlight important features and suppress unimportant features, thereby extracting effective feature information. The multiplication calculation formula is as follows: ; in, It represents the weighted fusion feature map. The characteristic value of the channel; Represented as the first The characteristic value of the channel; Expressed as The weight coefficient of the channel; Expressed as the number of feature channels. By assigning weights to each feature channel, the result can be made more robust.
[0032] The present invention conducts experiments on the GF1_WHU dataset, which contains different backgrounds such as natural water bodies, artificial objects, water bodies and mixed surfaces, and covers a variety of typical cloud types such as thick clouds, thin clouds and broken clouds. The experiments show that the model has good segmentation effects on different backgrounds and different cloud types. Some cloud images are as follows Figure 6 shown.
[0033] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0034] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. An improved semantic segmentation method for cloud and cloud shadow based on Unet, characterized in that: The method comprises the following steps: S1, replace the downsampling of the Unet model with the improved ASPP spatial pyramid pooling module, and embed the attention mechanism module in the skip connection to construct a semantic segmentation subnetwork; use the semantic segmentation subnetwork to process the input cloud map, and obtain the semantic segmentation feature map through upsampling; S2, connect the four convolutional layers in sequence to form four convolutional stages. After adding a 1×1 convolution operation in each convolutional stage, use the deconvolution operation to restore the feature maps of the last three convolutional stages to the same scale as the input image. Finally, fuse the feature maps obtained in the four convolutional stages through a 1×1 convolution operation to construct an edge detection subnetwork. Use the edge detection subnetwork to process the input cloud map to obtain an edge detection feature map. S3 performs weighted fusion on the input cloud map, semantic segmentation feature map and edge detection feature map, and then processes the fused features through the Sigmoid activation function layer to output the cloud and cloud shadow segmentation results.
2. The improved cloud and cloud shadow semantic segmentation method based on Unet according to claim 1, characterized in that: In step S2, the improved ASPP spatial pyramid pooling module includes a first semantic convolution layer, a second semantic convolution layer, a third semantic convolution layer, a fourth semantic convolution layer, an average pooling layer, a maximum pooling layer, a semantic feature fusion layer and a 1×1 convolution layer; The first semantic convolution layer, the second semantic convolution layer, the third semantic convolution layer, and the fourth semantic convolution layer respectively extract features of the input cloud image; the maximum pooling layer performs a maximum value operation on the local area of the input cloud image to extract local detail features of the input cloud image including the edge and texture of the cloud shadow; the average pooling layer performs an average operation on the local area of the input cloud image to extract global context features of the input cloud image including the overall shape of the cloud and the distribution of the cloud and cloud shadows.
3. The improved cloud and cloud shadow semantic segmentation method based on Unet according to claim 2, characterized in that: The first semantic convolution layer uses a 1×1 convolution kernel to perform feature extraction on the input cloud image and extracts preliminary features of the cloud image; the second semantic convolution layer uses a 3×3 convolution kernel and sets the hole rate to 6 to capture a larger range of context information in the cloud image; the third semantic convolution layer uses a 3×3 convolution kernel and increases the hole rate to 12 to capture a larger range of context information in the cloud image and extract large-scale features between clouds and cloud shadows in the cloud image; The fourth semantic convolution layer uses a 3×3 convolution kernel, and the void rate is increased to 18, so as to capture the maximum range of contextual information in the cloud map as the global feature between clouds and cloud shadows in the cloud map; the semantic feature fusion layer concats the feature maps from different convolutional layers and pooling layers, integrates local details and global contextual information, and generates a richer feature representation; the 1×1 convolution layer performs further convolution operations on the feature maps fused by the semantic feature fusion layer, adjusts the number of channels of the feature maps, and finally generates an output feature map for subsequent tasks.
4. The improved cloud and cloud shadow semantic segmentation method based on Unet according to claim 1, characterized in that: In step S2, the edge detection subnetwork includes a first edge convolution layer, a second edge convolution layer, a third edge convolution layer, a fourth edge convolution layer and an edge feature fusion layer connected in sequence, and four 1×1 convolution layers connected to the first convolution layer, the second convolution layer, the third convolution layer and the fourth convolution layer respectively; A pooling layer is inserted into each of the first edge convolution layer, the second edge convolution layer, the third edge convolution layer, and the fourth edge convolution layer. The first edge convolution layer, the second edge convolution layer, the third edge convolution layer, and the fourth edge convolution layer sequentially extract features of the input cloud image to obtain cloud shadow edge features of different scales, and reduce the resolution of the feature image through the corresponding pooling layers. The output features of the first edge convolution layer are subjected to a 1×1 convolution operation to obtain a first convolution stage feature map; the output features of the second edge convolution layer, the third edge convolution layer, and the fourth edge convolution layer are subjected to a 1×1 convolution operation respectively, and then the deconvolution operation is used to restore the feature maps of the corresponding convolution stages to the same scale as the input cloud map, to obtain the second convolution stage feature map, the third convolution stage feature map, and the fourth convolution stage feature map respectively; The edge feature fusion layer is a 1×1 convolution layer, which fuses the feature map of the first convolution stage, the feature map of the second convolution stage, the feature map of the third convolution stage, and the feature map of the fourth convolution stage to obtain an edge detection feature map that simultaneously contains edge information of cloud shadow areas in different ranges.
5. The improved cloud and cloud shadow semantic segmentation method based on Unet according to claim 4, characterized in that: The first edge convolution layer, the second edge convolution layer, the third edge convolution layer, and the fourth edge convolution layer all use 3×3 convolution layers, and the number of channels are 64, 128, 256, and 512, respectively.
6. The improved cloud and cloud shadow semantic segmentation method based on Unet according to claim 1, characterized in that: Step S3 further comprises: Input cloud map , semantic segmentation feature map and edge detection feature map Perform splicing and fusion to obtain the fused feature map : ; The fused feature map The number of feature channels included is the input cloud map , semantic segmentation feature map and edge detection feature map The sum of the number of channels; The fused feature map The average pooling and maximum pooling operations are performed on each feature channel to obtain the feature maps The global average and maximum value for each feature channel: ; ;; Among them, i represents the index of the feature channel, ; Weights are assigned to each feature channel through multiplication calculation to extract effective feature information; the multiplication calculation formula is as follows:
7. Among them It represents the eigenvalue of the i-th channel on the feature map after weighted fusion; Represented as the eigenvalue of the kth channel on the feature map; It is represented as the weight coefficient of the i-th channel; N is represented as the number of feature channels.
Citation Information
Patent Citations
Multilevel semantic fusion cloud and cloud shadow detection method and device, and storage medium
CN114943876A
Remote sensing building image extraction method based on improved U-Net model
CN115841625A
Stacked object recognition method, apparatus and device, and computer storage medium
WO2023047167A1