Multispectral remote sensing image cloud detection method based on multilevel feature fusion U-Net network

By using a multi-level feature fusion U-Net network and attention mechanism, the problem of insufficient feature saliency in thin cloud detection is solved, improving cloud detection accuracy and robustness, especially cloud detection performance in complex terrain backgrounds.

CN120877088APending Publication Date: 2025-10-31CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510744438.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing deep learning network models suffer from problems in thin cloud detection, such as difficulty in recognizing the shape of thin clouds, their scattered distribution, and high transparency, making it difficult to distinguish them from the background, resulting in insufficient cloud detection accuracy.

Method used

We employ a U-Net network based on multi-level feature fusion, combined with a coordinate attention module, encoder, decoder, and attention gating module. Through multi-level feature fusion and skip connections, we improve the feature saliency and robustness of cloud detection.

Benefits of technology

It effectively alleviates the difficulty of thin cloud detection and improves the accuracy and robustness of cloud detection in multispectral remote sensing images, especially the cloud detection performance against complex surface backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877088A_ABST
    Figure CN120877088A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-spectral remote sensing image cloud detection method, system and product based on a multi-level feature fusion U-Net network, and the method comprises the steps: inputting a to-be-detected multi-spectral remote sensing image into the multi-level feature fusion U-Net network, and outputting a multi-spectral remote sensing image cloud detection result; the neural network comprises a coordinate attention module CA, an encoder, a decoder and an attention gating module AG; according to the invention, a multi-level feature fusion module is introduced into an encoder, and low-level structure information and high-level semantic information are fused to enhance thin cloud sensing capability; a coordinate attention module CA is introduced at the front end of the network, direction perception and position sensitive features are extracted, high-albedo interferents are effectively inhibited, and the target positioning precision is improved; an attention gating module AG is embedded in jump connection between an encoder and a decoder, salient region response is enhanced, irrelevant feature interference is restrained, and therefore fine prediction of the cloud is achieved. According to the method, the precision and robustness of cloud detection in the multispectral remote sensing image under the complex earth surface background are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of remote sensing image processing technology, and relates to a method, system and product for detecting cloud in multispectral remote sensing images, and particularly to a method, system and product for detecting cloud in multispectral remote sensing images based on a multi-level feature fusion U-Net network. Background Technology

[0002] Remote sensing, as a non-contact technology for acquiring information about ground features, obtains information about ground features by receiving electromagnetic waves reflected, radiated, or scattered by targets through sensors. It boasts advantages such as high timeliness, wide coverage, abundant data, and ease of acquisition, playing a crucial role in geospatial exploration. With the continuous advancement of Earth observation by various countries, numerous remote sensing satellites have been launched and put into use, providing rich data for exploring the Earth's surface and being widely applied in many fields such as agricultural production, environmental monitoring, geographic mapping, water conservancy and transportation, and military defense. However, according to long-term observations from the International Satellite Cloud Climatology Project, approximately 66.7% of the Earth's surface is covered by clouds. Optical satellite remote sensing imagery, as an important data source for Earth observation, is inevitably affected by cloud obstruction due to its imaging mechanism, leading to problems such as loss of underlying surface information and degraded data radiometric quality, severely restricting the application level of optical remote sensing.

[0003] Cloud detection is a crucial step in the preprocessing of optical satellite remote sensing images. Traditional cloud detection methods rely on setting thresholds, expert knowledge, and manually designing shallow features. These methods typically have limited robustness and discriminative power, making them ill-suited for complex real-world applications. In contrast, deep learning-based cloud detection methods can deeply mine high-level semantic information from images through a data-driven approach. By offering superior robustness and discriminative power, they provide a new solution and technical framework for cloud detection, significantly improving its accuracy, efficiency, and reliability. This has become a cutting-edge research direction in the field of cloud detection.

[0004] In recent years, deep learning, with its powerful feature extraction capabilities, has provided new solutions for cloud detection. Numerous deep learning network models have been proposed, greatly promoting the rapid development of the cloud detection field. However, due to the heterogeneity of clouds and the diversity of underlying surfaces, as well as strong interference from similar land features with high albedo (such as snow / ice, desert / rock, etc.), existing cloud detection networks still have certain shortcomings. This is especially true for thin clouds that are difficult to identify in shape, scattered in distribution, and highly transparent; their visual features are weak, and they are highly similar to non-cloud pixels, making them difficult to distinguish from the background, which to some extent limits the accuracy of cloud detection. Summary of the Invention

[0005] To address the technical challenges of thin cloud detection and insufficient feature saliency in existing deep learning network models, this invention proposes a cloud detection method, system, and product for multispectral remote sensing images based on a Multilevel Feature fusion U-Net (MFU-Net) network.

[0006] The technical solution adopted by the present invention is: a multispectral remote sensing image cloud detection method based on a multi-level feature fusion U-Net network, comprising the following steps: Step 1: Acquire the multispectral remote sensing image to be detected; Step 2: Input the multispectral remote sensing image to be detected into the U-Net network based on multi-level feature fusion, and output the multispectral remote sensing image cloud detection results; The U-Net network based on multi-level feature fusion includes a coordinate attention module (CA), an encoder, a decoder, and an attention gating module (AG). The coordinate attention module (CA) is deployed at the network front end to capture orientation-aware and position-sensitive features. Each layer of the encoder is equipped with a multi-level feature fusion module (MFF) to simultaneously transmit the input and output of the current layer to the next layer and fuse them with the input of the first layer through feature concatenation, thereby inheriting the original features of the input data and providing a more useful feature representation for the next layer. The encoder and decoder use the attention gating module (AG) to achieve a skip connection.

[0007] Preferably, the encoder consists of five deep convolutional layers, including four downsampling layers and one convolutional layer. The first downsampling layer consists of three 3×3 convolutions and one 2×2 max pooling, the other downsampling layers consist of two 3×3 convolutions and one 2×2 max pooling, and the fifth convolutional layer generates the final encoded features through a regular 3×3 convolution. For an input ,in x and y Where d is the spatial dimension and d is the number of channels, the convolution kernel of each deep convolutional layer generates a feature map with a specific pattern; ; in, Indicates the first Feature mapping of each convolutional kernel Represents kernel parameters, Indicates the number of cores. This indicates a non-linear activation operation.

[0008] Preferably, the multi-level feature fusion module (MFF) uses 2×2 strided convolution to perform cross-layer feature downsampling, thereby effectively reducing feature size and information loss.

[0009] Preferably, the attention gating module AG is derived from the first... Layer coding features and from the Layer upsampling decoding features As input, generate optimized features based on the attention mechanism. First, regarding and Performing a 1×1 convolution yields features of the same size, i.e. and ;Then, and Feature fusion is performed by element-wise addition, followed by ReLU activation, 1×1 convolution, and Sigmoid activation in sequence to obtain the attention function. ,in, Each element is considered a pixel. Spatial weight This indicates the importance of the current position; among which, , , , , and Represents the parameters of the convolution operation. This indicates an element-wise addition operation. and For ReLU and Sigmoid functions; finally, the attention of resampling. Multiplying with the encoded features yields the optimized features. .

[0010] Preferably, the decoder outputs the optimized features from the attention gating module AG. With upsampling decoding features Feature stitching is performed, followed by two 3×3 convolutions and one upsampling operation. After four upsampling layers with the same operation, a feature map of the same size as the original image patch is obtained. Finally, the final cloud probability map is obtained through 1×1 convolution and Sigmoid operation.

[0011] Preferably, the coordinate attention module (CA) consists of two parts: coordinate information embedding and coordinate attention generation. For input data First, use a 1D pooling kernel ( h , 1) or (1, w Feature aggregation is performed along the horizontal or vertical direction, effectively embedding coordinate information in a specific direction into the feature representation to obtain aggregated features. and ;in, This represents the channel features aggregated along the height direction, with dimension [dimensional value missing]. h × c , The channel features are aggregated along the horizontal direction, with dimension 1. w × c ,I( h , i ) indicates that the input feature map is at position ( h , i The value of ) h This indicates a fixed row index (vertical direction). i It is a column index (horizontal direction), I( j , w ) indicates that the input feature map is at position ( j , w The value of ) j It is a row index (in the height direction). w For fixed column indexes (horizontal direction) h and w These represent the height and width of the feature map, respectively. c Indicate the number of channels; then aggregate the features. and Concatenate along spatial direction and input convolution function ,in, To share a 1×1 convolution, Indicates nonlinear activation. It is an intermediate feature that contains coordinate information of two spatial directions. This is the downsampling ratio; then... Divide into two tensors along two spatial directions and These two tensors are fed into two separate convolution operations. In, to generate and input Having the same channel characteristics, ., , It is the attention weight in the vertical direction. The attention weights in the horizontal direction, It is the Sigmoid function; then, using the obtained and Highlight target features ,in, It is a scalar multiplication operation. As a target indication, For the first c The first channel i Vertical attention coefficient of the row, For the first c The first channelj The horizontal attention coefficient of the column.

[0012] Preferably, for the cloud prediction probability image output by the decoder, pixels in the cloud prediction probability image that are greater than a threshold are determined to be clouds, and pixels that are less than a threshold are determined to be non-clouds. Based on this, the cloud prediction probability image is thresholded to obtain the final binarized cloud detection result, and the image blocks are stitched together to form a whole-scene image prediction cloud mask result to obtain the cloud detection result.

[0013] Preferably, the U-Net network based on multi-level feature fusion is a pre-trained network. First, an initial learning rate is set, and a learning rate decay strategy is selected to adjust the training parameters and optimizer. Accuracy verification is performed simultaneously during training. The accuracy of the model is evaluated for each training session using a validation sample dataset. The network model parameters are adjusted based on the evaluation accuracy and loss value of the validation set, and the accuracy and loss value of the network model are recorded for each training session to obtain the optimal network parameter model. The cloud prediction probability of the output training samples is calculated, the network loss function is calculated, and backpropagation is performed. The loss function is: ; in, and These represent the generated cloud mask and the actual cloud label, respectively. h and w These represent the height and width of the input image, respectively. i , j Indicates the pixel index position in the image. Represents the true label image L Middle position ( i , j The pixel value at position ) takes a value of 0 or 1, representing either a non-cloud pixel or a cloud pixel, respectively; cloud mask. This indicates that the predicted cloud mask output by the model is located at ( i , j The pixel value at position ) ranges from 0 to 1. The larger the value, the greater the probability that the corresponding pixel belongs to the cloud.

[0014] The technical solution adopted by the system of this invention is: a multispectral remote sensing image cloud detection system based on a multi-level feature fusion U-Net network, comprising: One or more processors; A storage device is provided for storing one or more programs, which, when executed by one or more processors, enable the one or more processors to implement the multispectral remote sensing image cloud detection method based on a multi-level feature fusion U-Net network.

[0015] The technical solution adopted by the product of the present invention is: a multispectral remote sensing image cloud detection product based on a multi-level feature fusion U-Net network, including computer program instructions, which, when the computer program instructions are run on a computer, cause the computer to execute the multispectral remote sensing image cloud detection method based on a multi-level feature fusion U-Net network.

[0016] Compared with the prior art, the beneficial effects of the present invention include: (1) The present invention introduces a multilevel feature fusion module (MFF) in the encoder to fuse low-level structural information and high-level semantic information, thereby enhancing the ability to perceive thin clouds; (2) The present invention introduces a coordinate attention module (CA) at the network front end to extract direction perception and position sensitivity features, effectively suppressing high albedo interference and improving target positioning accuracy; (3) The present invention embeds an attention gate (AG) module in the skip connection between the encoder and the decoder to enhance the response of salient regions and suppress interference from irrelevant features, thereby achieving fine prediction of clouds; (4) This invention effectively alleviates the problems of thin cloud detection difficulty and insufficient feature saliency by constructing and training the MFU-Net model, and improves the accuracy and robustness of cloud detection in multispectral remote sensing images under complex surface background. Attached Figure Description The technical solutions of the present invention will be further illustrated below using embodiments and specific implementation methods. In addition, some accompanying drawings are used in the description of the technical solutions. Those skilled in the art can obtain other drawings and the intent of the present invention from these drawings without any creative effort.

[0017] Figure 1 This is a diagram of the U-Net network structure framework based on multi-level feature fusion provided in this embodiment of the invention; Figure 2 This is a schematic diagram of the coordinate attention module structure provided in the embodiments of the invention; Figure 3 This is a schematic diagram of the attention gating module structure provided in an embodiment of the present invention; Figure 4 This is a flowchart of the training process of the U-Net network based on multi-level feature fusion according to an embodiment of the present invention; Figure 5 This is a structural diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0018] To facilitate understanding and implementation of the present invention by those skilled in the art, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0019] Please see Figure 1 This embodiment provides a method for cloud detection in multispectral remote sensing images based on a multi-level feature fusion U-Net network, which includes the following steps: Step 1: Acquire the multispectral remote sensing image to be detected; Step 2: Input the multispectral remote sensing image to be detected into the U-Net network based on multi-level feature fusion, and output the multispectral remote sensing image cloud detection results; The U-Net network based on multi-level feature fusion includes a coordinate attention module (CA), an encoder, a decoder, and an attention gating module (AG). The coordinate attention module (CA) is deployed at the network front end to capture orientation-aware and position-sensitive features. Each layer of the encoder is equipped with a multi-level feature fusion module (MFF) to simultaneously transmit the input and output of the current layer to the next layer and fuse them with the input of the first layer through feature concatenation, thereby inheriting the original features of the input data and providing a more useful feature representation for the next layer. The encoder and decoder use the attention gating module (AG) to achieve a skip connection.

[0020] In one implementation, the encoder consists of five deep convolutional layers, including four downsampling layers and one convolutional layer. To increase the richness of the encoded features, 32, 64, 128, 256, and 512 convolutional kernels are used in the five convolutions, respectively. 3×3 convolutions are used to extract local contextual information, and padding operations are employed to maintain the size of the output feature map. Furthermore, exponential linear units (ELU) are used for non-linear feature activation, and 2×2 max pooling is used for feature downsampling. After four downsampling layers, the feature size becomes 1 / 16 of the original image patch.

[0021] Convolution is the core operation of an encoder. Within a specific receptive field, it uses a set of learnable convolutional kernels to extract local contextual information from a cloud, generating rich feature representations. For example, for an input... ,in x and y Let d be the spatial dimension and d be the number of channels. Each convolutional kernel can generate a feature map with a specific pattern, as shown in the following formula: ; in, express Feature mapping of each kernel, Represents kernel parameters, Indicates the number of cores. This represents a non-linear activation operation. Based on this, through continuous convolutional learning in the encoder, we can delve deeper into the abstract high-level semantic representation of the cloud, thereby approximating its intrinsic properties.

[0022] Considering the significant differences in target description between different levels of features, with lower-level features focusing more on target structural information and higher-level features focusing more on target semantic representation, a multilevel feature fusion (MFF) module containing four downsampling layers was constructed in the encoder to fully integrate complementary information from multiple levels of features, promote accurate cloud identification, and improve thin cloud detection capabilities.

[0023] Specifically, the first downsampling layer consists of three 3×3 convolutions and one 2×2 max pooling operation, while the other downsampling layers each consist of two 3×3 convolutions and one 2×2 max pooling operation. Furthermore, in the MFF module, the input and output of the current layer are simultaneously passed to the next layer and fused with the input of the first layer through feature concatenation, thus inheriting the original features of the input data and providing a more beneficial feature representation for the next layer. Considering the feature scale differences caused by the max pooling operation, the MFF module uses 2×2 strided convolutions for cross-layer feature downsampling, effectively reducing feature size and information loss. In this way, multi-level feature representations with different scales and receptive fields can be fully fused, further improving the model's feature representation and discriminative capabilities. Finally, the features are passed to the last convolutional layer, where two regular 3×3 convolutions generate the final encoded features, which contain rich semantic information and form the basis of the decoder.

[0024] In this embodiment, a coordinate attention module is introduced at the network front end to extract cloud indicative features from the raw data, thereby promoting better learning of the network.

[0025] In one implementation, please see Figure 2 The coordinate attention module consists of two parts: coordinate information embedding and coordinate attention generation. Channel attention often uses two-dimensional global pooling to globally encode spatial information into channel descriptors, discarding positional information and hindering target localization. To alleviate this problem, the coordinate attention module decomposes channel attention into two one-dimensional feature encoding blocks and uses one-dimensional global pooling to aggregate information along two different directions, effectively promoting direction-aware and position-sensitive feature learning, as well as inter-channel correlation mining. Then, the aggregated features are encoded into two complementary feature maps and applied to the input data to emphasize the feature representation of the object of interest.

[0026] Specifically, for input data Using 1D pooling kernels ( h , 1) or (1, w Features are aggregated along the horizontal or vertical direction, as shown in the following formula: ; ; in, This represents the channel features aggregated along the height direction, with dimension [dimensional value missing]. h × c , The channel features are aggregated along the horizontal direction, with dimension 1. w × c ,I( h , i ) indicates that the input feature map is at position ( h , i The value of ) h This indicates a fixed row index (vertical direction). i It is a column index (horizontal direction), I( j , w ) indicates that the input feature map is at position ( j , w The value of ) j It is a row index (in the height direction). w For fixed column indexes (horizontal direction) h and w These represent the height and width of the feature map, respectively. c Indicates the number of channels.

[0027] In this way, coordinate information in a specific direction is effectively embedded into the feature representation. While learning long-range dependencies in one direction, accurate positional information in another direction can be preserved, thus helping the network to locate targets more accurately.

[0028] The learned features are then further processed to generate coordinate attention, fully utilizing location information and channel correlation. The aggregated features are then... and The convolution function is input along the spatial direction, and the formula is as follows: ; in, To share a 1×1 convolution, Indicates nonlinear activation. It is an intermediate feature that contains coordinate information of two spatial directions. The downsampling ratio is adjusted for the module size; here it's set to 32. Next, Divide into two tensors along two spatial directions and These two tensors are fed into two separate convolution operations. To generate and input Having the same channel characteristics, the formula is as follows: ; ; in, It's the Sigmoid function. Then, using the obtained... and The formula for highlighting target features is as follows: ; in, It is a scalar multiplication operation. As a target indication, For the first c The first channel i Vertical attention coefficient of the row, For the first c The first channel j The horizontal attention coefficient of the column.

[0029] In one implementation, please see Figure 3 To extract effective information to assist feature decoding and more accurately recover the spatial distribution of clouds, this embodiment introduces a lightweight attention gate (AG) module to optimize low-level encoded features. It suppresses irrelevant regions and emphasizes salient regions through implicit learning.

[0030] The attention gating module AG, which is derived from the first Layer coding features and from the Layer upsampling decoding features As input, generate optimized features based on the attention mechanism. First, regarding and Performing a 1×1 convolution yields features of the same size, i.e. and As shown in equations 3.2 and 3.3. Then, and Feature fusion is performed by element-wise addition, followed by ReLU activation, 1×1 convolution, and Sigmoid activation in sequence to obtain the attention function. As shown in Equation 3.4. Wherein, Each element can be considered as a pixel. Spatial weight This indicates the importance of the current position. The formula is as follows: ; ; ; Where W1,b1,W2,b2,W and b represent the convolution operation parameters. This indicates an element-wise addition operation. and For ReLU and Sigmoid functions.

[0031] Finally, the attention of resampling Multiplying with the encoded features yields the optimized features. This feature focuses only on salient regions and adaptively adjusts the contribution of different positions in the low-level encoded features, which helps improve cloud prediction accuracy. The formula is as follows: ; In one implementation, the decoder optimizes features during the network decoding process. With upsampling decoding features Feature concatenation is performed, followed by two 3×3 convolutions and one upsampling operation. After four identical upsampling layers, a feature map of the same size as the original image patch is obtained. Finally, a 1×1 convolution and a sigmoid operation are applied to obtain the final cloud probability map.

[0032] In one implementation, for the cloud prediction probability image output by the decoder, pixels in the cloud prediction probability image that are greater than a threshold are determined to be clouds, and pixels that are less than the threshold are determined to be non-clouds. Based on this, the cloud prediction probability image is thresholded to obtain the final binarized cloud detection result, and the image blocks are stitched together to form a whole-scene image prediction cloud mask result to obtain the cloud detection result.

[0033] In this embodiment, the cloud segmentation threshold is set to 0.5. In the binarized threshold segmentation, pixels with a cloud probability value greater than 0.5 in the test results are segmented as clouds and assigned a value of 1, while pixels with a probability value less than 0.5 are segmented as non-clouds and assigned a value of 0, ultimately obtaining a binarized cloud mask image. The cloud mask image blocks are stitched together in sequence to output the cloud detection results corresponding to the entire scene image.

[0034] In one implementation, please see Figure 4 The U-Net network based on multi-level feature fusion is a pre-trained network; the training process specifically includes the following steps: (1) Collect and acquire multispectral remote sensing cloud detection datasets as data sources for model training and testing; Step 1.1: Select a multispectral remote sensing cloud detection dataset; In this embodiment, two typical public datasets for multispectral remote sensing cloud detection, namely the 38-Cloud dataset and the SPARCS dataset, are selected as data sources for model training and validation. These datasets contain various cloud types and complex underlying surface scenarios, which are suitable for evaluating the performance of cloud detection algorithms. Step 1.2: Divide the 38-Cloud cloud detection dataset; In this embodiment, 38 Landsat-8 OLI images selected from the 38-Cloud dataset are divided into 18 training images and 20 test images according to the conventional strategy. All images contain four bands: red, green, blue and near-infrared. The cloud cover ranges from 0% to 100%, and the underlying surface types include barren land, grassland / farmland, snow / ice and water bodies. Step 1.3: Preprocess the remote sensing images from the 38-Cloud dataset; In this embodiment, each image in the 38-Cloud dataset is cropped into image patches of size 384×384, generating 8400 training image patches and 9201 test image patches; 20% of the image patches in the training set are randomly divided as a validation set, and finally a data subset containing 6720 training samples and 1680 validation samples is constructed.

[0035] Step 1.4: Divide the SPARCS cloud detection dataset; In this embodiment, 80 Landsat-8 OLI images from the SPARCS dataset are processed. Each image has an original size of 1000×1000 pixels and includes various cloud types and typical underlying surface types. The "cloud" in the image label is used as a separate detection category, and other categories are uniformly classified as background. Step 1.5: Preprocess remote sensing images from the SPARCS dataset; In this embodiment, each image in the SPARCS dataset is cropped into 384×384 image patches to generate 540 training image patches and 180 test image patches. At the same time, 20% of the training image patches are divided into a validation set to construct a stable and representative training and validation dataset.

[0036] (2) Set the loss function for the MFU-Net model, use the training subset in the public multispectral remote sensing cloud detection dataset for model training, optimize the network parameters through multiple rounds of iteration, and use the validation set to monitor the training effect until the model converges and obtain the optimal cloud detection model after training. In this embodiment, the cloud detection training set of multispectral remote sensing images is input into the U-Net network model MFU-Net based on multi-level feature fusion, the network weights are initialized in [-1, 1], the cloud prediction probability of the training samples is output, the cloud detection network loss function is calculated and backpropagation is performed.

[0037] The entire network framework is trained using the simple and efficient binary cross-entropy (BCE) loss function. The BCE loss function is widely used in various tasks such as classification, segmentation, and change detection. It allows the network to focus on learning discriminative information, making it very suitable for binary classification and cloud detection tasks. It can be expressed as the following formula: ; in, and These represent the generated cloud mask and the actual cloud label, respectively. h and w These represent the height and width of the input image, respectively. i , j Indicates the pixel index position in the image. Represents the true label image L Middle position ( i , j The pixel value at position ) takes a value of 0 or 1, representing either a non-cloud pixel or a cloud pixel, respectively; cloud mask. This indicates that the predicted cloud mask output by the model is located at ( i , j The pixel value at position ) ranges from 0 to 1. The larger the value, the greater the probability that the corresponding pixel belongs to the cloud.

[0038] The Adam optimizer is used to solve the above optimization problems and to train the network.

[0039] Model training was implemented in Keras (2.4.3) and TensorFlow (2.4.0). During training, the learning rate and batch size for the 38-Cloud and SPARCS datasets were set to 0.0001 and 24, respectively, and the number of iterations were set to 200 and 400, respectively. Batch normalization was performed on each convolutional layer in the encoder, with a momentum of 0.7. 20% of the training samples were allocated as validation samples, and the accuracy of the network model at each training iteration was evaluated using these validation samples. The network model parameters were adjusted based on the validation accuracy and performance to obtain the optimal model.

[0040] (3) Input the test samples in batches into the best model obtained from training to predict cloud probability, and adjust its size to 384×384. By selecting an appropriate threshold, the final cloud mask is generated.

[0041] (4) Finally, the training results are evaluated by calculating the preset indicators; the preset indicators include: overall accuracy, recall, precision, F1 index and Jaccard index.

[0042] This embodiment also provides a multispectral remote sensing image cloud detection system based on a multi-level feature fusion U-Net network, including: One or more processors; A storage device is provided for storing one or more programs, which, when executed by one or more processors, enable the one or more processors to implement the multispectral remote sensing image cloud detection method based on a multi-level feature fusion U-Net network.

[0043] This embodiment also provides a multispectral remote sensing image cloud detection product based on a multi-level feature fusion U-Net network, including computer program instructions. When the computer program instructions are run on a computer, the computer executes the multispectral remote sensing image cloud detection method based on a multi-level feature fusion U-Net network.

[0044] In one implementation, please see Figure 5 The embodiments also provide an electronic device 10. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0045] Electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0046] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0047] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the methods and processes described above.

[0048] In some embodiments, the U-Net network multispectral image cloud detection method based on multi-level feature fusion can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the U-Net network multispectral image cloud detection method based on multi-level feature fusion described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the multispectral remote sensing image cloud detection method by any other suitable means (e.g., by means of firmware).

[0049] Various embodiments of the systems and techniques described above in this invention can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0050] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0051] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0052] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0053] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0054] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0055] The invention will be further illustrated below through specific experiments.

[0056] The cloud detection experiments were conducted on the multispectral cloud detection datasets 38-Cloud and SPARCS, and RS-Net and Cloud-Net models were selected as comparison methods. The quantitative evaluation of their cloud detection performance is shown in Table 1.

[0057] Table 1

[0058] As shown in Table 1, the method proposed in this invention achieves the best results on this dataset, and significantly improves upon other methods in most metrics. Specifically, the recall rate reaches 88.14%, the overall precision reaches 96.61%, the F1-Score is 88.62%, and the Jaccard score is 81.83%. Compared with the suboptimal method RS-Net, it shows significant improvements of 1.38% and 1.96% in F1-Score and Jaccard score, respectively, demonstrating the effectiveness and superiority of the proposed method.

[0059] This invention collects a multispectral remote sensing cloud detection dataset as a data source for model training and testing; constructs a U-Net network model MFU-Net based on multi-level feature fusion, including an encoder, decoder, and skip connection paths; trains the initial multispectral cloud detection model using the training set; sets a loss function for the MFU-Net model, trains the model using a training subset from the publicly available multispectral remote sensing cloud detection dataset, optimizes the network parameters through multiple iterations, and monitors the training effect using a validation set until the model converges, obtaining the optimal cloud detection model after training; inputs the test set into the trained MFU-Net model to obtain cloud probability images as prediction results; takes a threshold and performs binarized threshold segmentation on the test results accordingly, stitches the segmentation results into a complete image, obtains the final predicted cloud mask, and obtains the final cloud detection result.

[0060] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0061] It should be understood that the embodiments described above are only some, not all, of the embodiments of the present invention. Furthermore, the technical features of the various embodiments or individual embodiments provided by the present invention can be arbitrarily combined to form feasible technical solutions. Such combinations are not constrained by the order of steps and / or structural composition patterns, but must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.

[0062] It should be understood that the above description of the preferred embodiments is quite detailed, but it should not be considered as a limitation on the scope of protection of this invention. Those skilled in the art, under the guidance of this invention, can make substitutions or modifications without departing from the scope of protection of the claims of this invention, and all such substitutions or modifications fall within the scope of protection of this invention. The scope of protection of this invention should be determined by the appended claims.

Claims

1. A cloud detection method for multispectral remote sensing images based on a multi-level feature fusion U-Net network, characterized in that, Includes the following steps: Step 1: Acquire the multispectral remote sensing image to be detected; Step 2: Input the multispectral remote sensing image to be detected into the U-Net network based on multi-level feature fusion, and output the multispectral remote sensing image cloud detection results; The U-Net network based on multi-level feature fusion includes a coordinate attention module (CA), an encoder, a decoder, and an attention gating module (AG). The coordinate attention module (CA) is deployed at the network front end to capture orientation-aware and position-sensitive features. Each layer of the encoder is equipped with a multi-level feature fusion module (MFF) to simultaneously transmit the input and output of the current layer to the next layer and fuse them with the input of the first layer through feature concatenation, thereby inheriting the original features of the input data and providing a more useful feature representation for the next layer. The encoder and decoder use the attention gating module (AG) to achieve a skip connection.

2. The multispectral remote sensing image cloud detection method based on a multi-level feature fusion U-Net network according to claim 1, characterized in that: The encoder consists of five deep convolutional layers, including four downsampling layers and one convolutional layer. The first downsampling layer consists of three 3×3 convolutions and one 2×2 max pooling, and the other downsampling layers consist of two 3×3 convolutions and one 2×2 max pooling. The fifth convolutional layer generates the final encoded features through a regular 3×3 convolution. For an input ,in x and y Where d is the spatial dimension and d is the number of channels, the convolution kernel of each deep convolutional layer generates a feature map with a specific pattern; ; in, Indicates the first Feature mapping of each convolutional kernel Represents kernel parameters, Indicates the number of cores. This indicates a non-linear activation operation.

3. The multispectral remote sensing image cloud detection method based on a multi-level feature fusion U-Net network according to claim 1, characterized in that: The multi-level feature fusion module (MFF) uses 2×2 strided convolution to perform cross-layer feature downsampling, thereby effectively reducing feature size and information loss.

4. The multispectral remote sensing image cloud detection method based on a multi-level feature fusion U-Net network according to claim 1, characterized in that: The attention gating module AG, from the first Layer coding features and from the Layer upsampling decoding features As input, generate optimized features based on the attention mechanism. First, regarding and Performing a 1×1 convolution yields features of the same size, i.e. and ;Then, and Feature fusion is performed by element-wise addition, followed by ReLU activation, 1×1 convolution, and Sigmoid activation in sequence to obtain the attention function. ,in, Each element is considered a pixel. Spatial weight , indicating the importance of the current position; where W1, b1, W2, b2, W and b represent the convolution operation parameters, This indicates an element-wise addition operation. and For ReLU and Sigmoid functions; finally, the attention of resampling. Multiplying with the encoded features yields the optimized features. .

5. The multispectral remote sensing image cloud detection method based on a multi-level feature fusion U-Net network according to claim 1, characterized in that: The decoder will output the optimized features from the attention-gated module AG. With upsampling decoding features Feature stitching is performed, followed by two 3×3 convolutions and one upsampling operation. After four upsampling layers with the same operation, a feature map of the same size as the original image patch is obtained. Finally, the final cloud probability map is obtained through 1×1 convolution and Sigmoid operation.

6. The multispectral remote sensing image cloud detection method based on a multi-level feature fusion U-Net network according to claim 1, characterized in that: The coordinate attention module (CA) consists of two parts: coordinate information embedding and coordinate attention generation. For input data First, use a 1D pooling kernel ( h , 1) or (1, w Feature aggregation is performed along the horizontal or vertical direction, effectively embedding coordinate information in a specific direction into the feature representation to obtain aggregated features. and ;in, This represents the channel features aggregated along the height direction, with dimension [dimensional value missing]. h × c ; The channel features are aggregated along the horizontal direction, with dimension 1. w × c ;I( h , i ) indicates that the input feature map is at position ( h , i The value of ) h This indicates a fixed row index. i It is a column index; I( j , w ) indicates that the input feature map is at position ( j , w The value of ) j It is a row index. w For fixed column indexes; h and w These represent the height and width of the feature map, respectively. c Indicate the number of channels; then aggregate the features. and Concatenate along spatial direction and input convolution function ,in, To share a 1×1 convolution, Indicates nonlinear activation. It is an intermediate feature that contains coordinate information of two spatial directions. This is the downsampling ratio; then... Divide into two tensors along two spatial directions and These two tensors are fed into two separate convolution operations. In, to generate and input Having the same channel characteristics, ., , It is the attention weight in the vertical direction. The attention weights in the horizontal direction, It is the Sigmoid function; then, using the obtained and Highlight target features ,in, It is a scalar multiplication operation. As a target indication, For the first c The first channel i Vertical attention coefficient of the row, For the first c The first channel j The horizontal attention coefficient of the column.

7. The multispectral remote sensing image cloud detection method based on a multi-level feature fusion U-Net network according to claim 1, characterized in that: For the cloud prediction probability image output by the decoder, pixels with values ​​greater than a threshold are determined to be clouds, and pixels with values ​​less than a threshold are determined to be non-clouds. Based on this, the cloud prediction probability image is segmented by threshold to obtain the final binarized cloud detection result. The image blocks are then stitched together to form a whole-scene image cloud prediction mask result to obtain the cloud detection result.

8. The multispectral remote sensing image cloud detection method based on a multi-level feature fusion U-Net network according to any one of claims 1-7, characterized in that: The U-Net network based on multi-level feature fusion is a pre-trained network. First, an initial learning rate is set, and a learning rate decay strategy is selected to adjust the training parameters and optimizer. Accuracy verification is performed simultaneously during training. The accuracy of the model is evaluated for each training session using the validation sample dataset. The network model parameters are adjusted based on the evaluation accuracy and loss value of the validation set, and the accuracy and loss value of the network model are recorded for each training session to obtain the optimal network parameter model. The cloud prediction probability of the output training samples is calculated, the network loss function is calculated, and backpropagation is performed. The loss function is: ; in, and These represent the generated cloud mask and the actual cloud label, respectively. h and w These represent the height and width of the input image, respectively. i , j Indicates the pixel index position in the image. Represents the true label image L Middle position ( i , j The pixel value at position ) takes a value of 0 or 1, representing either a non-cloud pixel or a cloud pixel, respectively; cloud mask. This indicates that the predicted cloud mask output by the model is located at ( i , j The pixel value at position ) ranges from 0 to 1. The larger the value, the greater the probability that the corresponding pixel belongs to the cloud.

9. A multispectral remote sensing image cloud detection system based on a multi-level feature fusion U-Net network, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the multispectral remote sensing image cloud detection method based on a multi-level feature fusion U-Net network as described in any one of claims 1 to 8.

10. A multispectral remote sensing image cloud detection product based on a multi-level feature fusion U-Net network, comprising computer program instructions, characterized in that: When the computer program instructions are executed on a computer, the computer performs the multispectral remote sensing image cloud detection method based on a multi-level feature fusion U-Net network as described in any one of claims 1 to 8.