Remote sensing image cloud detection method, computer readable storage medium and computer device
By replacing the DeepLabV3+ encoder with the RegNetY-016 encoder and combining it with the ASPP module and DICE loss function, the problem of high computing power requirements in remote sensing image cloud detection was solved, achieving high-precision and high-speed cloud extraction.
Patent Information
- Application Number
- CN202310183338.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-01
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2043-03-01
AI Technical Summary
The existing DeepLabV3+ model has high computational requirements in remote sensing image cloud detection, which cannot meet the high accuracy and high speed requirements of massive remote sensing images.
The DeepLabV3+ encoder was replaced with the RegNetY-016 encoder. Combined with the ASPP module and DICE loss function, the model weights were updated through the training optimizer to perform cloud detection in remote sensing images.
It improves the accuracy and speed of cloud detection in remote sensing images, meets the needs of algorithms for massive amounts of remote sensing images, and achieves more accurate cloud extraction.
Smart Images

Figure CN116416215B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of optical image processing, in particular to a remote sensing image cloud detection method, a computer readable storage medium and a computer device. BACKGROUND
[0002] At present, the segmentation of the region of interest in the remote sensing image based on the convolutional neural network has been widely applied, and the commonly used segmentation models include U-Net, DeepLab and the like. Among them, DeepLabV3+ is a deep learning model used for image segmentation. Image segmentation refers to dividing an image into different regions and assigning a label to each region. This technology can be used to automatically extract the contours of objects in an image or separate different parts of an image.
[0003] DeepLabV3+ is based on a residual network (ResNet) and an atrous convolution (Atrous Convolution). It uses a novel structure called "multi-scale residual module" (MSRA) to process features in different receptive fields (ranges seen in an image). This allows DeepLabV3+ to identify different objects in different parts of an image.
[0004] DeepLabV3+ also uses an ASPP (Atrous Convolution Parallel Pooling Layer) module to extract global context information. In this way, even if some parts of the image are missing or have a lot of noise, DeepLabV3+ can still accurately segment the objects in the image.
[0005] However, the current DeepLabV3+ model encoder uses Xception, which requires a large amount of computing power. Cloud detection is a necessary step in batch processing of massive remote sensing images, and the algorithm needs to improve the detection speed while maintaining high accuracy. Therefore, the current DeepLabV3+ model encoder using Xception cannot meet the needs of existing massive remote sensing image algorithms.
[0006] It should be noted that the information disclosed in this BACKGROUND section is intended only to increase an understanding of the general context of the present application, and should not be taken as an acknowledgement or implication that this information constitutes prior art that is already widely known in the art. SUMMARY
[0007] An embodiment of the present application provides a remote sensing image cloud detection method, which comprises the following steps:
[0008] S1, acquiring a data set;
[0009] S2, constructing a segmentation model;
[0010] S2.1, constructing a RegNetY-016 encoder based on a RegNet model;
[0011] The RegNetY-016 encoder is composed of a convolutional layer and a trunk layer, the convolutional layer is composed of a 3*3 convolutional kernel, a batch normalization layer and a ReLU activation function, the trunk layer is composed of a first stage, a second stage, a third stage and a fourth stage, the first stage, the second stage, the third stage and the fourth stage will reduce the height and width of the input feature matrix to half of the original, and each is composed of multiple modules, in the first module of each stage, there is a group convolution with a stride of 2 on the main branch and a normal convolution with a stride of 2 on the shortcut branch, and the convolution stride in the remaining modules is 1;
[0012] S2.2, modifying DeepLabV3+ based on the RegNetY-016 encoder; connecting the RegNetY-016 encoder to the ASPP module in DeepLabV3+; taking the output result of the first module block1 in the RegNetY-016 encoder as input to a 1*1 convolutional layer, a batch normalization layer and a ReLU activation function, and then connecting it with the output of the ASPP module after twice up-sampling;
[0013] S3, constructing a loss function, the loss function is as follows:
[0014]
[0015] Wherein, X is a predicted image, Y is a real image, wc is a weight, M is a class number, yc is a one-hot vector value, which is 0 or 1, pc is the probability of the predicted sample belonging to c;
[0016] S4, obtaining a use model through training;
[0017] After the images of the data set in S1 are input into the segmentation model constructed in S2 after data augmentation, the prediction result is obtained, the prediction result and the label of the data set are input into the loss function in S3, the loss is calculated, and then the model weight is updated by the optimizer. After one round of training, the model obtained by training is used to predict the validation set data, and the prediction result is evaluated using the evaluation index, and the model obtained in the round of training with the optimal evaluation index on the validation set is saved as the final use model.
[0018] In some embodiments, the S1 data set acquisition step includes: constructing a remote sensing image cloud detection semantic segmentation data set, wherein the target to be recognized in the image is labeled with a pixel level, and then the samples are divided into a training set and a validation set in a ratio of 4:1.
[0019] In some embodiments, the RegNetY-016 encoder does not include a global average pooling layer and a fully connected layer.
[0020] In some embodiments, the data enhancement in S4 includes brightness transformation of pixel value change in the form of flipping, rotation or scaling of morphological change.
[0021] In some embodiments, the optimizer in S4 selects an AdamW optimizer.
[0022] In some embodiments, after S4, there is further S5, prediction result post-processing; the image is divided for prediction, the overlapped division is used in the process of division, and the neural network is used for prediction, and finally only the middle part of the divided prediction image is taken to splice to obtain the final prediction image.
[0023] An embodiment of the present application also provides a computer readable storage medium, which stores computer instructions, and the computer is executed by a processor to implement the remote sensing image cloud detection method in any of the above embodiments.
[0024] An embodiment of the present application also provides a computer device, which comprises at least one processor and a memory connected with the processor in communication, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to make the processor execute the remote sensing image cloud detection method in any of the above embodiments.
[0025] In summary, an embodiment of the present application provides a remote sensing image cloud detection method, a computer readable storage medium and a computer device, a full convolutional neural network with an encoder-decoder structure is used to perform semantic segmentation on high-resolution remote sensing images to extract cloud layer pixel points in the remote sensing images, and then a morphological method is used for post-processing to obtain the final extraction result; and by replacing the encoder of DeepLabV3+ with the RegNetY-016 encoder, DeepLabV3+ can maintain high accuracy and greatly improve the detection speed when facing massive remote sensing image cloud detection, which can meet the needs of existing massive remote sensing image algorithms, and the remote sensing image cloud detection effect is more accurate.
[0026] Other features and advantages of the present application will be illustrated in the following description, and some features and advantages can be obviously obtained from the description, or can be understood by implementing the present application. The purposes and other advantages of the present application can be achieved and obtained by the structures specifically indicated in the description. BRIEF DESCRIPTION OF DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, part of the drawings described below are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0028] Figure 1 is a flowchart of the remote sensing image cloud detection method provided by an embodiment of the present application;
[0029] Figure 2 is a structural diagram of RegNet series model;
[0030] Figure 3 is a branch structure diagram of block structure;
[0031] Figure 4 is a structural diagram of RegNetY-016 encoder;
[0032] Figure 5 is a diagram of each block structure in RegNetY-016 encoder;
[0033] Figure 6 is a structural diagram of DeepLabV3+ improved by adding an encoder and a decoder;
[0034] Figure 7 is a comparison diagram of the to-be-detected image and the detection result in coastal areas;
[0035] Figure 8 is a comparison diagram of the to-be-detected image and the detection result in mountainous areas;
[0036] Figure 9 is a comparison diagram of the to-be-detected image and the detection result in snow-covered areas. DETAILED DESCRIPTION
[0037] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the following will combine the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, rather than all the embodiments. The technical features designed in different embodiments of the present application can be combined with each other as long as they do not conflict with each other. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.
[0038] In the description of the present application, it needs to be understood that the terms "center", "transverse", "upper", "lower", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, which is only for the convenience of describing the present application and simplifying the description, and does not indicate or imply that the device or component referred to must have a particular orientation or be constructed and operated in a particular orientation, therefore it cannot be understood as a limitation on the present application. In addition, the terms "first", "second" are only for description purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the technical features indicated. Therefore, the features defined as "first", "second" can be explicitly or implicitly included one or more of the features. In the description of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more. In addition, the term "comprising" and any variations thereof mean "at least including".
[0039] In the description of the present application, it needs to be understood that unless otherwise explicitly specified and limited, the terms "mounting", "connecting", "connection" should be understood in a broad sense, for example, it can be fixed connection, or detachable connection, or integral connection; it can be mechanical connection, or electrical connection; it can be direct connection, or indirect connection through intermediate medium, or internal communication of two components. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0040] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the example embodiments. Unless the context clearly indicates otherwise, as used herein, the singular forms "a", "an", and "the" are intended to include the plural forms as well. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0041] Please refer to Figures 1 to 6 To achieve at least one of the above-mentioned advantages or other advantages, an embodiment of the present application provides a remote sensing image cloud detection method based on convolutional neural network. As shown in the figure, the remote sensing image cloud detection method comprises the following steps:
[0042] S1: Obtain a data set. A remote sensing image cloud detection semantic segmentation data set is constructed, and the target to be identified in the image needs to be labeled with pixel level. Then, the samples are divided into a training set and a verification set according to a ratio of 4:1 for subsequent training and verification.
[0043] S2: Constructing the segmentation model. A fully convolutional neural network with an encoder-decoder structure is used for image interpretation. The overall structure of the model uses DeepLabV3+, and the encoder part is replaced with RegNetY-016.
[0044] S2.1: Constructing RegNetY-016 encoder based on RegNet model. RegNet is a deep neural network architecture for computer vision tasks. It is designed as a lightweight model that can efficiently run on resource-constrained devices. RegNet adopts the structure of "residual module" to improve the accuracy of the model through residual connection. Unlike other deep neural network architectures, the design parameters of RegNet (such as the depth and width of each residual module) are data-driven, which allows it to be fine-tuned according to the needs of specific tasks.
[0045] In addition, RegNet also uses the "adaptive span" technology to make the model better handle features of different scales. These technologies make RegNetY perform well in various computer vision tasks and achieve high accuracy with small model size.
[0046] The structure of RegNet series model is shown in Figure 2 The network is composed of convolutional layer stem, trunk layer body and head. The stem is composed of a 3*3 convolution kernel convolution layer, a batch normalization layer and a ReLU activation function. The body is composed of 4 stages. Each stage will reduce the height and width of the input feature matrix to half of the original. And each stage is composed of a series of blocks. In the first block, there is a group convolution with a stride of 2 (on the main branch) and a normal convolution (on the shortcut branch), and the convolution stride in the remaining blocks is 1, similar to ResNet.
[0047] As shown in Figure 3As shown, the block structure contains a main branch and a shortcut branch. The main branch typically consists of a 1x1 convolutional layer, a 3x3 group convolutional layer, and a 1x1 convolutional layer, and improves network performance by adding batch normalization and activation functions (such as ReLU) between each layer. The shortcut branch performs different processing depending on the stride. When stride = 1, this branch usually does no processing and directly outputs the input data, skipping subsequent layers; when stride = 2, this branch usually downsamples the input data through a 1x1 convolutional layer to reduce the size of the input data. r represents the resolution, i.e., the height and width of the feature matrix. When stride s = 1, the input and output r remain unchanged; when s = 2, the output r is half that of the input. w represents the channel of the feature matrix. Note that when s = 2, the input is wi-1, and the output is wi, so the channel changes. g represents the group width of each group in the group convolution, and b represents the bottleneck ratio, which means that the channels of the output feature matrix will be reduced to 1 / b of the channels of the input feature matrix.
[0048] The head is a common classifier in classification networks, consisting of a global average pooling layer and a fully connected layer.
[0049] There are two architectures in the RegNet series models: RegNetX and RegNetY. The main difference is that RegNetY connects an SE (Squeeze-and-Excitation) module after the Group Conv in the block. Figure 3 As shown, the SE module typically consists of a global average pooling layer and two fully connected layers. In RegNet, the number of nodes in fully connected layer 1 (FC1) is equal to one-quarter of the input feature matrix channel of the block (not one-quarter of the GroupConv output feature matrix channel), and the activation function is ReLU. The number of nodes in fully connected layer 2 (FC2) is equal to the group Conv output feature matrix channel, and the activation function is Sigmoid.
[0050] Specifically, in this embodiment, the algorithm selects RegNetY-016 as the encoder, and its structure is as follows: Figure 4 As shown, its four stages (stage one, stage two, stage three, and stage four) contain 2, 6, 17, and 2 module blocks, respectively.
[0051] wherein each block structure is as shown in Figure 5 For the first block of each stage, the stride of the 3*3 convolution on the main branch is 2, there is a 1*1 convolution on the shortcut branch with a stride of 2, and the stride of the remaining convolution layers is 1. In the remaining blocks, the stride of all convolution layers is 1, and the shortcut branch is directly output to the next block. It should be noted that RegNetY-016 is used as an encoder to output a feature map, so a global average pooling layer and a fully connected layer (FC 1000) are not needed, and thus these two layers can be deleted, i.e., the RegNetY-016 encoder does not include a global average pooling layer and a fully connected layer.
[0052] S2.2: Modify DeepLabV3+ based on RegNetY-016 encoder.
[0053] DeepLabV3+ uses an Atrous Spatial Pyramid Pooling (ASPP) module to capture rich contextual information through pooling operations on features at different resolutions, and adds an encoder-decoder structure to recover spatial information to capture clear target boundaries based on DeepLabV3. After the standard DeepLabV3+ encoder uses the Xception output feature map to input into the ASPP module of the decoder, DeepLabV3+ uses the encoding-decoding structure, so only the part after the ASPP module needs to be replaced.
[0054] As shown in Figure 6 In this embodiment, the global average pooling layer and the fully connected layer are deleted in RegNetY-016 and then connected to the ASPP module of the DeepLabV3+ model. Considering that the output of the ASPP in DeepLabV3+ still needs to be upsampled by two times and then concatenated with the feature map from the init block in Xception, the output of the first block of RegNetY-016 is taken as the input to a 1*1 convolution layer, a batch normalization layer, and a ReLU activation function, and then concatenated with the output of the ASPP module after being upsampled by two times.
[0055] S3, construct a loss function.
[0056] The loss function uses a weighted cross-entropy loss combined with a DICE loss. The reason for using the DICE loss is that there may be a class imbalance problem in the training data, and the target occupies a small proportion of pixels, which leads the model to learn to tend to the background class, and the DICE loss can alleviate this situation. The DICE loss formula is: Wherein X is a predicted image, Y is a real image.
[0057] Considering that only using DICE loss may lead to unstable training process, a weighted cross-entropy loss is introduced, and the formula is: Wherein wc is a weight, M is the number of categories, yc is a one-hot vector with values of 0 and 1 only, and if the category is the same as the category of the sample, it takes 1, otherwise it takes 0, and pc is the probability of the predicted sample belonging to c.
[0058] The weight formula is: Wherein N represents the total number of pixels, and Nc represents the total number of pixels of category C in the real situation.
[0059] In the specific implementation, a larger weight is applied to the target edge pixel, so that the training process tends to the edge pixel, thereby improving the segmentation effect of the edge. Therefore, the complete loss function is:
[0060]
[0061] Wherein X is a predicted image, Y is a real image, wc is a weight, M is the number of categories, yc is a one-hot vector with values of 0 or 1, and pc is the probability of the predicted sample belonging to c. In addition to the above method, OHEM (Online Hard Example Mining) is also used to make the model training tend to be more difficult samples.
[0062] S4: Obtain a using model through training.
[0063] The images of the data set in S1 are input into the segmentation model constructed in S2 after data augmentation to obtain a prediction result, and the prediction result and the label of the data set are input into the loss function in S3. After calculating the loss, the optimizer updates the model weight. After one round of training, the model obtained by training is used to predict the validation set data, and the prediction result is evaluated using the evaluation index. The model obtained by the round of training with the optimal evaluation index on the validation set is saved as the final using model.
[0064] The data augmentation in S4 includes: using morphological changes such as flipping, rotating or scaling to realize brightness transformation of pixel value changes.
[0065] The optimizer in S4 selects AdamW optimizer, which is Adam optimizer plus L2 regularization. The learning rate adjustment uses cosine annealing learning rate adjustment with warm-up, thereby improving the stability of the model in the early stage of training and avoiding the model from falling into a local minimum point in the whole training process.
[0066] After S4, the following steps can also be included:
[0067] S5, post-processing of the prediction result. Considering that the remote sensing image is large in size, the image is divided into blocks for prediction. In order to avoid the relatively poor accuracy of the edge part of the divided image due to the lack of context information compared with the central region, overlapping block division is adopted in the process of division, and the neural network is used for prediction, and finally only the middle part of the divided prediction image is taken to splice to obtain the final prediction image.
[0068] An embodiment of the present application also provides a computer readable storage medium, which stores computer instructions, and the computer is executed by a processor to implement the remote sensing image cloud detection method in any of the above embodiments.
[0069] An embodiment of the present application also provides a computer device, which comprises at least one processor and a memory connected with the processor in communication, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to make the processor execute the remote sensing image cloud detection method in any of the above embodiments.
[0070] Please refer to Figures 7 to 9 , which are respectively the comparison schematic diagrams of the to-be-detected images and the detection results in the coastal area, the mountainous area and the snow-covered area. As shown in the figure, by using the remote sensing image cloud detection method of the present application, when facing the cloud detection of a large amount of remote sensing images, the detection speed can be greatly improved while maintaining high accuracy, which can meet the demand of the existing large amount of remote sensing image algorithm, and the remote sensing image cloud detection extraction effect is more accurate.
[0071] In summary, an embodiment of the present application provides a remote sensing image cloud detection method, a computer readable storage medium and a computer device, a full convolutional neural network with an encoder-decoder structure is used to perform semantic segmentation on high-resolution remote sensing images to extract cloud layer pixel points in the remote sensing images, and then a morphological method is used for post-processing to obtain the final extraction result; and by replacing the encoder of DeepLabV3+ with RegNetY-016 encoder, DeepLabV3+ can maintain high accuracy while greatly improving the detection speed when facing the cloud detection of a large amount of remote sensing images, which can meet the demand of the existing large amount of remote sensing image algorithm, and the remote sensing image cloud detection effect is more accurate.
[0072] In addition, those skilled in the art should understand that although there are many problems in the prior art, each embodiment or technical solution of the present application can only improve in one or several aspects, and it is not necessary to solve all the technical problems listed in the prior art or background art at the same time. Those skilled in the art should understand that what is not mentioned in a claim should not be regarded as a limitation of the claim.
[0073] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions recorded in the above embodiments can be modified, or some or all of the technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for cloud detection in remote sensing images, characterized in that: The remote sensing image cloud detection method comprises the following steps: S1, acquiring a data set; S2, constructing a segmentation model; S2.1, constructing a RegNetY-016 encoder based on a RegNet model; The RegNetY-016 encoder is composed of a convolutional layer and a trunk layer, the convolutional layer is composed of a 3*3 convolutional kernel, a batch normalization layer and a ReLU activation function, the trunk layer is composed of a first stage, a second stage, a third stage and a fourth stage, the first stage, the second stage, the third stage and the fourth stage all reduce the height and width of the input feature matrix to half of the original, and are all composed of multiple modules, in the first module of each stage, there is group convolution with a stride of 2 on the main branch and ordinary convolution with a stride of 2 on the shortcut branch, and the convolution stride in the remaining modules is 1; S2.2, modifying DeepLabV3+ based on the RegNetY-016 encoder; connecting the RegNetY-016 encoder to the ASPP module in DeepLabV3+; inputting the output result of the first module block1 in the RegNetY-016 encoder to a 1*1 convolutional layer, a batch normalization layer and a ReLU activation function, and then connecting with the output of the ASPP module after two times of upsampling; S3, constructing a loss function, the loss function is as follows: where X is the predicted image, Y is the real image, w c is the weight, M is the number of categories, y c is the one-hot vector value, which is 0 or 1, p c is the probability that the predicted sample belongs to c; S4, obtaining a use model through training; After the data set images in S1 are input into the segmentation model constructed in S2 after data augmentation, the prediction result is obtained, the prediction result and the label of the data set are input into the loss function in S3, the loss is calculated, the model weight is updated by the optimizer, the model obtained by training is used to predict the validation set data after one round of training, and the prediction result is evaluated using the evaluation index, and the model obtained by one round of training with the optimal evaluation index on the validation set is saved as the final use model.
2. The remote sensing image cloud detection method of claim 1, wherein: The S1 acquiring a data set step comprises: constructing a remote sensing image cloud detection semantic segmentation data set, wherein the target to be recognized in the image is labeled with a pixel level, and then the samples are divided into a training set and a validation set in a ratio of 4:
1. 3.The remote sensing image cloud detection method of claim 1, wherein: The RegNetY-016 encoder does not contain a global average pooling layer and a fully connected layer. 4.The method of claim 1, wherein: The data augmentation in S4 comprises: adopting a flipping, rotating or scaling mode of morphological change to realize a brightness transformation of pixel value change.
5. The remote sensing image cloud detection method of claim 1, wherein: The optimizer in S4 is an AdamW optimizer.
6. The remote sensing image cloud detection method of claim 1, wherein: After S4, there is also: S5, post-processing of the prediction result; the image is predicted in blocks during prediction, overlapping blocks are used in the blocking process, and the neural network is input for prediction, and finally only the middle part of the block prediction image is taken to obtain the final prediction image.
7. A computer-readable storage medium, characterized in that: The computer readable storage medium stores computer instructions, and the computer is executed by the processor to realize the remote sensing image cloud detection method in any one of claims 1-6.
8. A computer device, comprising: A computer program product comprising a computer readable medium, the computer readable medium having stored thereon the program element of a computer program comprising program code means adapted to perform the method of any one of claims 1-6 when the program is executed on a computer or computer network. A computer program comprising program code means adapted to perform the method of any one of claims 1-6 when the program is executed on a computer or computer network. A computer program product comprising a computer readable medium, the computer readable medium having stored thereon the program element of a computer program comprising program code means adapted to perform the method of any one of claims 1-6 when the program is executed on a computer or computer network. A computer program comprising program code means adapted to perform the method of
Citation Information
Patent Citations
Digital pathological full-slice image rapid analysis method based on JPEG compressed coding
CN112767503A
Neural network design and optimization method based on software and hardware joint learning
CN113902099A