A lightweight semantic segmentation method for remote sensing cloud images
By building a lightweight residual module, a spatial pyramid pooling module and a structural reparameterization module, the problem of difficulty in distinguishing clouds from surface objects in satellite images is solved, and more efficient cloud map detection and segmentation performance is achieved, which is suitable for various remote sensing image processing.
Patent Information
- Application Number
- CN202410322088.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-20
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-03-20
AI Technical Summary
In the prior art, when using satellite images to analyze ground objects, clouds serve as interference signals, making it difficult to distinguish between clouds and surface objects. Especially when the spectral band is limited, it is more difficult to automatically segment and extract clouds.
A lightweight remote sensing cloud image semantic segmentation method is proposed, including building a lightweight residual module, a spatial pyramid pooling module and a structural re-parameterization module. Through grouping convolution, multi-scale feature extraction and fusion of convolution layers and batch normalization layers, the calculation amount and parameter amount are reduced, while improving the accuracy and generalization ability of the model.
In terms of comprehensive consideration of parameter quantity, calculation quantity and segmentation accuracy, the performance is better, reducing the parameter quantity of traditional models. At the same time, in the segmentation performance of cloud image detection, it can better avoid missing alarms and false alarms, and is suitable for remote sensing image processing in environments such as ice and snow.
Smart Images

Figure CN118115741B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing cloud detection, and specifically, is a lightweight remote sensing cloud image semantic segmentation method for mobile edge computing. Background Art
[0002] Remote sensing technology is a means of obtaining information about the earth's surface through long-distance sensing technologies such as satellites, aircraft, and drones. With the advancement of remote sensing technology and satellite technology, remote sensing images have become crucial in studying information about the earth's surface, especially the prominent position of satellite remote sensing images in this field.
[0003] The continuous development of science and technology and the emergence of various remote sensing image analysis software tools have reduced the cost of acquiring and studying satellite remote sensing images, and optical remote sensing technology has been widely used in Earth system science. According to the International Satellite Cloud Climate Program (ISCCP), the global average annual cloud coverage is about 68%. Among them, the cloud cover on land is about 55%, while the cloud cover on the ocean is about 72%, highlighting the large proportion of cloud pixels in remote sensing images. Therefore, cloud detection has become a crucial research task in satellite remote sensing image processing. Although observing cloud cover is of great significance to the study of the earth's water cycle and climate change, when using satellite images for ground object analysis, clouds may become an interference signal and pose an obstacle to the accurate acquisition of surface information. Therefore, accurately detecting cloud-covered areas and removing clouds are crucial to obtaining reliable surface information.
[0004] In cloud detection, when clouds and other surface objects with similar reflective properties coexist in remote sensing images, it may be difficult to distinguish them. In order to improve recognition accuracy, some researchers have proposed methods to utilize information from other relevant bands. However, due to the limited number and coverage of spectral bands in some satellite data, it is more difficult to automatically segment and extract clouds when the channels only include red, green, blue and near-infrared bands. On the other hand, the rapid development of deep learning algorithms in recent years has made semantic segmentation algorithms based on convolutional neural networks (CNNs) possible. These algorithms segment remote sensing images by combining spatial and spectral patterns, and show superior performance compared to early hand-crafted feature extraction methods. Summary of the invention
[0005] The purpose of the present invention is to overcome the deficiencies of the above-mentioned prior art and to provide a lightweight remote sensing cloud image semantic segmentation method.
[0006] A lightweight remote sensing cloud image semantic segmentation method includes the following steps:
[0007] Acquire remote sensing cloud images;
[0008] Construct a lightweight residual module, perform group convolution on the features of the received remote sensing cloud image to form multiple groups, perform independent convolution operations on each group, and finally merge the outputs of each group to form a total feature data;
[0009] Construct a spatial pyramid pooling module ASPP, divide the total feature data into N equal parts through channels, extract feature information of different scales from the N equal parts, thus generating receptive fields of different scales, and obtaining context information of different scales. Finally, fuse and splice the N equal parts of features together;
[0010] Construct a structural reparameterization module to perform zero padding and fusion processing on the fused and spliced feature data to form a data set;
[0011] Train the network, use the Adam optimizer and linear decay of the learning rate, mark the cloud parts in the data set as white, and the rest as black, train the network, and output the segmentation map.
[0012] In the lightweight residual module, first, the module uses 1×1 group convolution to perform group convolution operation on the input features to divide them into multiple groups;
[0013] Then use 3×3 depth-separable convolution to extract input features;
[0014] Then use a 1×1 grouped convolution kernel to reduce the dimension of the features;
[0015] The result of dimensionality reduction through the 1×1 grouped convolutional layer is added to the residual connection and then input into the SE attention module. The output result of the SE module is rearranged in channels to further strengthen the interaction between channels.
[0016] When processing in the spatial pyramid pooling module, the multi-scale feature extraction module is used to divide the total feature data into four equal parts, and the dimension of each part is one quarter of the number of channels of the total input feature data;
[0017] Then, three dilated convolutions with dilation rates of 1, 3, and 5 and a 1×1 convolution layer are used to extract feature information of different scales from the four equal parts of features, thereby generating receptive fields of different scales and obtaining contextual information of different scales.
[0018] Finally, the four equal features after convolution processing are fused and spliced together, and the number of channels is adjusted through a 1x1 convolution layer to form a single channel, thereby reducing the number of channels of the feature map.
[0019] The structure reparameterization module specifically includes the following contents:
[0020] Convert the 1×1 convolution layer to a 3×3 convolution layer to keep the convolution kernel size consistent. Padding is used here for zero filling.
[0021] The identity mapping is converted to a 3×3 convolutional layer, that is, the separate 1×1 convolutional layer of each channel is converted into a 3×3 convolutional layer in the depth-separable convolution, and padding is used for zero filling;
[0022] After each convolution layer is processed, the corresponding batch normalization layer is processed, and the convolution layer and the batch normalization layer are fused, that is, the two operators of the convolution layer and the batch normalization layer are fused into RepConv, thereby reducing the amount of calculation and memory usage. The fusion of operators is performed after the network training is completed.
[0023] The specific process of fusion of the convolution layer and the batch normalization layer is:
[0024] Assume that the remote sensing cloud image is x, the convolution kernel size is w, and the expression of the convolution layer is as follows:
[0025] Conv(χ)=w*χ+b
[0026] Among them, b represents the bias of the convolutional layer, and the subsequent input value batch normalization layer BN is as follows:
[0027]
[0028] Among them, μ represents the mean of the BN layer data, σ represents the variance of the BN layer data, ε represents a very small constant used to prevent the divisor from being 0, γ and β represent the scale factor and translation factor of the BN layer respectively, which are used to adjust the data to the standard normal distribution. Substitute the operation of the convolution layer into the BN layer. The formula is as follows:
[0029]
[0030] After simplification, we can get:
[0031]
[0032] From the formula, we can see that and They are used as the convolution kernel and bias of the convolution layer respectively to achieve the fusion of the convolution layer and the BN layer.
[0033] When the SE attention module rearranges channels, the original channels are output in different connection modes.
[0034] The present invention has the following beneficial technical effects:
[0035] In terms of comprehensive consideration of the number of parameters, the amount of calculation, and the segmentation accuracy, the present invention performs better, which is specifically manifested in that the number of residuals of the present invention is greatly reduced compared to the traditional MobileNet, and is between the semantic segmentation model DFANet and the lightweight segmentation model LEDNet, but the required amount of calculation is close to that of DFANet and less than that of LEDNet. Secondly, in terms of the segmentation performance of cloud image detection, this solution can pay attention to the edge structure of clouds of different shapes through an improved pyramid multi-scale feature extraction module, which can objectively better avoid missed alarms and false alarms, and can be applied to remote sensing image processing in ice and snow environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 Schematic diagram of the processing process of the lightweight residual module of the present invention;
[0037] Figure 2 Schematic diagram of the processing process of the spatial pyramid pooling module of the present invention;
[0038] Figure 3 Schematic diagram of the processing process of the structure reparameterization module of the present invention. DETAILED DESCRIPTION
[0039] In order to further understand the features, technical means, specific objectives and functions of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0040] A lightweight remote sensing cloud image semantic segmentation method is proposed. Based on Yolov3, the invention introduces a lightweight residual module, which can significantly reduce model parameters and computational complexity while maintaining a high accuracy, thereby achieving a faster reasoning speed. Then, an improved ASPP module and a structural reparameterization module are introduced, which can effectively extract feature information from the image and have good support for features of various scales, thereby achieving a higher model accuracy. More specifically, the following steps are included:
[0041] Remote sensing cloud images can be obtained from various environments, such as ice and snow.
[0042] A lightweight residual module is constructed to perform group convolution on the features of the received remote sensing cloud images to form multiple groups, perform independent convolution operations on each group, and finally merge the outputs of each group to form a total feature data.
[0043] A spatial pyramid pooling module ASPP is constructed to divide the total feature data into N equal features through channels. Feature information of different scales is extracted from the N equal features to generate receptive fields of different scales, thereby obtaining contextual information of different scales. Finally, the N equal features are fused and spliced together.
[0044] A structural reparameterization module is constructed to perform zero padding and fusion processing on the fused and spliced feature data to form a data set.
[0045] Train the network, use the Adam optimizer and linear decay of the learning rate, mark the cloud parts in the data set as white, and the rest as black, train the network, and output the segmentation map.
[0046] refer to Figure 1 As shown in the figure, the lightweight residual module first uses 1×1 grouped convolution to adjust the input features. Grouped convolution is a convolution operation that divides the input feature map into multiple groups, performs independent convolution operations on each group, and finally merges the outputs of each group. Compared with traditional convolution operations, grouped convolution can reduce the amount of calculation and parameters, thereby accelerating calculation and reducing the risk of overfitting. Then use 3×3 depth-separable convolution to extract the input features, and then use 1×1 grouped convolution kernel to reduce the dimension of the features to reduce the amount of calculation of the model. Then the result of 1×1 convolution is added to the residual connection and input into the SE attention module to further strengthen the interaction between channels, so that the generalization ability and accuracy of the model are improved. Since grouped convolution may cause some features to be isolated in different groups and cannot be effectively interactively fused, thereby limiting the expression ability of the network, the output results of the SE attention module are subjected to channel shuffle operation, which can strengthen the information exchange between different channels, strengthen the interaction and information transmission between feature channels, and thus improve the effect of feature extraction. Channel rearrangement is to output the original channels in different connection modes, so that some features isolated in different groups can also be output.
[0047] When processing in the spatial pyramid pooling module, the multi-scale feature extraction module is used to divide the total feature data into four equal parts, and the dimension of each part is one quarter of the number of channels of the total input feature data;
[0048] Then, three dilated convolutions with dilation rates of 1, 3, and 5 and a 1×1 convolution layer are used to extract feature information of different scales from the four equal parts of features, thereby generating receptive fields of different scales and obtaining contextual information of different scales.
[0049] Finally, the four equal features after convolution processing are fused and spliced together, and the number of channels is adjusted through a 1x1 convolution layer to form a single channel, thereby reducing the number of channels of the feature map.
[0050] The structure reparameterization module specifically includes the following contents:
[0051] Convert the 1×1 convolution layer to a 3×3 convolution layer to keep the convolution kernel size consistent. Padding is used here for zero filling.
[0052] The identity mapping is converted to a 3×3 convolutional layer, that is, the separate 1×1 convolutional layer of each channel is converted into a 3×3 convolutional layer in the depth-separable convolution, and padding is used for zero filling;
[0053] After each convolution layer is processed, the corresponding batch normalization layer is processed, and the convolution layer and the batch normalization layer are fused, that is, the two operators of the convolution layer and the batch normalization layer are fused into RepConv, thereby reducing the amount of calculation and memory usage. The fusion of operators is performed after the network training is completed.
[0054] The specific process of fusion of the convolution layer and the batch normalization layer is:
[0055] Assume that the remote sensing cloud image is x, the convolution kernel size is w, and the expression of the convolution layer is as follows:
[0056] Conv(χ)=w*χ+b
[0057] Among them, b represents the bias of the convolutional layer, and the subsequent input value batch normalization layer BN is as follows:
[0058]
[0059] Among them, μ represents the mean of the BN layer data, σ represents the variance of the BN layer data, ε represents a very small constant used to prevent the divisor from being 0, γ and β represent the scale factor and translation factor of the BN layer respectively, which are used to adjust the data to the standard normal distribution. Substitute the operation of the convolution layer into the BN layer. The formula is as follows:
[0060]
[0061] After simplification, we can get:
[0062]
[0063] From the formula, we can see that and They are used as the convolution kernel and bias of the convolution layer respectively to achieve the fusion of the convolution layer and the BN layer, which can be replaced with a form with fewer parameters while maintaining higher performance.
[0064] It should be noted that the above description is not a limitation of the present invention. Without departing from the creative concept of the present invention, any obvious replacement is within the protection scope of the present invention.
Claims
1. A lightweight remote sensing cloud image semantic segmentation method, characterized in that: The following steps are involved: Acquire remote sensing cloud images; Construct a lightweight residual module, perform group convolution on the features of the received remote sensing cloud image to form multiple groups, perform independent convolution operations on each group, and finally merge the outputs of each group to form a total feature data; Construct a spatial pyramid pooling module ASPP, divide the total feature data into N equal parts through channels, extract feature information of different scales from the N equal parts, thus generating receptive fields of different scales, and obtaining context information of different scales. Finally, fuse and splice the N equal parts of features together; Construct a structural reparameterization module to perform zero padding and fusion processing on the fused and spliced feature data to form a data set; Train the network, use the Adam optimizer and linear decay of the learning rate, mark the cloud parts of the data set as white, and the rest as black, train the network, and output the segmentation map; The structure reparameterization module specifically includes the following contents: Convert the 1×1 convolution layer to a 3×3 convolution layer to keep the convolution kernel size consistent. Padding is used here for zero filling. The identity mapping is converted to a 3×3 convolutional layer, that is, the separate 1×1 convolutional layer of each channel is converted into a 3×3 convolutional layer in the depth-separable convolution, and padding is used for zero filling; After each convolution layer is processed, the corresponding batch normalization layer is processed, and the convolution layer and the batch normalization layer are fused, that is, the two operators of the convolution layer and the batch normalization layer are fused into RepConv, thereby reducing the amount of calculation and memory usage. The fusion of operators is performed after the network training is completed.
2. The lightweight remote sensing cloud image semantic segmentation method according to claim 1, characterized in that: In the lightweight residual module, first, the module uses 1×1 group convolution to perform group convolution operation on the input features to divide them into multiple groups; Then use 3×3 depth-separable convolution to extract input features; Then use a 1×1 grouped convolution kernel to reduce the dimension of the features; The result of dimensionality reduction through the 1×1 grouped convolutional layer is added to the residual connection and then input into the SE attention module. The output result of the SE module is rearranged in channels to further strengthen the interaction between channels.
3. The lightweight remote sensing cloud image semantic segmentation method according to claim 2, characterized in that: When processing in the spatial pyramid pooling module, the multi-scale feature extraction module is used to divide the total feature data into four equal parts, and the dimension of each part is one quarter of the number of channels of the total input feature data; Then, three dilated convolutions with dilation rates of 1, 3, and 5 and a 1×1 convolution layer are used to extract feature information of different scales from the four equal parts of features, thereby generating receptive fields of different scales and obtaining contextual information of different scales. Finally, the four equal features after convolution processing are fused and spliced together, and the number of channels is adjusted through a 1x1 convolution layer to form a single channel, thereby reducing the number of channels of the feature map.
4. The lightweight remote sensing cloud image semantic segmentation method according to claim 3, characterized in that: The specific process of fusion of the convolution layer and the batch normalization layer is: Assume that the remote sensing cloud image is x, the convolution kernel size is w, and the expression of the convolution layer is as follows: Conv(χ)=w*χ+b Among them, b represents the bias of the convolutional layer, and the subsequent input value batch normalization layer BN is as follows: Among them, μ represents the mean of the BN layer data, σ represents the variance of the BN layer data, ε represents a very small constant used to prevent the divisor from being 0, γ and β represent the scale factor and translation factor of the BN layer respectively, which are used to adjust the data to the standard normal distribution. Substitute the operation of the convolution layer into the BN layer. The formula is as follows: After simplification, we can get: From the formula, we can see that and They are used as the convolution kernel and bias of the convolution layer respectively to achieve the fusion of the convolution layer and the BN layer.
5. The lightweight remote sensing cloud image semantic segmentation method according to claim 4, characterized in that: When the SE attention module rearranges channels, the original channels are output in different connection modes.
Citation Information
Patent Citations
Remote sensing image road segmentation method based on contextual information and multi-scale feature fusion
CN113850825A
Multilevel semantic fusion cloud and cloud shadow detection method and device, and storage medium
CN114943876A