Remote sensing image cloud removal method and system based on attention mechanism and diffusion model
By introducing a high grayscale attention mechanism and a block-based denoising diffusion model in the remote sensing image cloud removal method, the problems of inaccurate positioning and long inference time are solved, and a more efficient and accurate remote sensing image declouding effect is achieved.
Patent Information
- Application Number
- CN202510135463.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-07
AI Technical Summary
The existing remote sensing image cloud removal methods have problems such as inaccurate cloud area positioning, long inference time, and poor recovery effect in thick cloud area.
Using an attention mechanism and diffusion model method, the remote sensing image is weighted through a high grayscale attention mechanism, and combined with a block-based denoising diffusion model to perform denoising processing, and different attention weights are allocated to improve the cloud removal effect.
It effectively improves the cloud removal effect of remote sensing images, improves the accuracy and efficiency of cloud removal, solves the problems of weak adaptability and long inference time, and performs more stable and accurate in the recovery effect of thick cloud areas.
Smart Images

Figure CN120070218A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and in particular, to a method and system for removing clouds from remote sensing images based on an attention mechanism and a diffusion model. Background Art
[0002] Remote sensing image information plays a crucial role in earth observation tasks due to its wide coverage, high - efficiency data acquisition ability, and diverse data types. However, in practical applications, natural factors such as cloud cover and atmospheric conditions often hinder the extraction of effective information from remote sensing images, leading to image data distortion and affecting subsequent analysis and applications.
[0003] Traditional methods for removing clouds from remote sensing images, such as physical - model - based cloud - removal methods, interpolation - based methods, and deep - learning - based methods, although certain achievements have been made, still face many challenges. For example, physical - model methods are computationally complex and difficult to adapt to complex and variable cloud conditions; interpolation methods are prone to losing image details and affecting image quality; while deep - learning - based methods, such as generative adversarial networks (GANs) and variational autoencoders (VAEs), although having excellent performance, still face problems such as insufficient preservation of generated image details, large data requirements, and poor training stability when dealing with complex cloud types such as thick clouds and cloud shadows. It can be seen that there are defects in the existing technology and urgent solutions are needed. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method and system for removing clouds from remote sensing images based on an attention mechanism and a diffusion model. By using a block - based denoising diffusion model combined with a high - gray - value attention mechanism, different attention weights are assigned to regions with different cloud contents in the remote sensing image, avoiding problems such as inaccurate cloud region positioning and long inference time, and improving the cloud - removal effect of remote sensing images.
[0005] To solve the above - mentioned technical problems, the first aspect of the present invention discloses a method for removing clouds from remote sensing images based on an attention mechanism and a diffusion model, and the method includes:
[0006] Obtain a remote sensing image, and perform gray - scale conversion on the remote sensing image to obtain a gray - scale remote sensing image;
[0007] Perform weighted processing on the remote sensing image through a high - gray - value attention mechanism and the gray - scale remote sensing image to obtain a weighted remote sensing image;
[0008] Input the weighted remote sensing image into a diffusion model for denoising processing to obtain a plurality of denoised image blocks;
[0009] Perform merging processing on all the denoised image blocks to obtain a cloud - removed remote sensing image.
[0010] As an alternative implementation, in the first aspect of the present invention, the weighted remote sensing image obtained by performing weighted processing on the remote sensing image through the high grayscale value attention mechanism and the grayscale remote sensing image includes:
[0011] Extracting a high grayscale value region from the grayscale remote sensing image using a learnable grayscale threshold to obtain a high grayscale value region mask;
[0012] Generating an attention weight matrix using the high grayscale value region mask and the remote sensing image;
[0013] Performing weighted processing on the remote sensing image using the attention weight matrix to obtain a weighted remote sensing image.
[0014] As an alternative implementation, in the first aspect of the present invention, the generating an attention weight matrix using the high grayscale value region mask and the remote sensing image includes:
[0015] Multiplying the high grayscale value region mask by the remote sensing image to obtain a high grayscale value feature;
[0016] After passing the high grayscale value feature and the remote sensing image through a convolutional layer and a batch normalization layer respectively, performing element-wise merging and using a first activation function for non-linear function mapping to obtain a merged feature after non-linear mapping;
[0017] Performing a convolutional normalization operation on the merged feature after non-linear mapping and using a second activation function to map the data to obtain an attention weight matrix.
[0018] As an alternative implementation, in the first aspect of the present invention, the performing weighted processing on the remote sensing image using the attention weight matrix to obtain a weighted remote sensing image includes:
[0019] Multiplying the attention weight matrix by the remote sensing image to obtain a weighted remote sensing image.
[0020] As an alternative implementation, in the first aspect of the present invention, the inputting the weighted remote sensing image into a diffusion model for denoising processing to obtain a plurality of denoised image blocks includes:
[0021] Randomly sampling from a standard normal distribution to generate an initial noise image, the size of the initial noise image being the same as the size of the weighted remote sensing image;
[0022] Cropping the initial noise image and the weighted remote sensing image according to a preset position dictionary to obtain a plurality of noise image blocks and a plurality of degraded image blocks, wherein there is an overlapping region between two adjacent noise image blocks;
[0023] Use a conditional diffusion model to estimate the noise of the noise image block and the degraded image block to obtain the estimated noise;
[0024] Calculate the noise mean according to the estimated noise and the cumulative weight to obtain the average estimated noise;
[0025] Update multiple noise image blocks according to the average estimated noise to obtain multiple denoised image blocks.
[0026] As an optional implementation manner, in the first aspect of the present invention, the calculating the noise mean according to the estimated noise and the cumulative weight to obtain the average estimated noise includes:
[0027] Perform element-wise division of the estimated noise by the cumulative weight to obtain the average estimated noise.
[0028] As an optional implementation manner, in the first aspect of the present invention, the merging all the denoised image blocks to obtain a cloud-removed remote sensing image includes:
[0029] Merge all the denoised image blocks, and use the weighted average method for the estimated noise in the overlapping area to ensure smooth transition between the merged denoised image blocks.
[0030] The second aspect of the present invention discloses a remote sensing image cloud removal system based on an attention mechanism and a diffusion model, and the system includes:
[0031] A grayscale image acquisition module, which is used to acquire a remote sensing image and perform grayscale conversion on the remote sensing image to obtain a grayscale remote sensing image;
[0032] A weighted image acquisition module, which is used to perform weighted processing on the remote sensing image through a high grayscale value attention mechanism and the grayscale remote sensing image to obtain a weighted remote sensing image;
[0033] A denoising module, which is used to input the weighted remote sensing image into a diffusion model for denoising processing to obtain multiple denoised image blocks;
[0034] A merging module, which is used to merge all the denoised image blocks to obtain a cloud-removed remote sensing image.
[0035] As an optional implementation manner, in the second aspect of the present invention, the weighted image acquisition module performs weighted processing on the remote sensing image through a high grayscale value attention mechanism and the grayscale remote sensing image to obtain a weighted remote sensing image, including:
[0036] Extract the high gray - value region from the gray - scale remote - sensing image using a learnable gray - scale threshold to obtain a high gray - value region mask;
[0037] Generate an attention weight matrix using the high gray - value region mask and the remote - sensing image;
[0038] Perform weighted processing on the remote - sensing image using the attention weight matrix to obtain a weighted remote - sensing image.
[0039] As an alternative implementation, in the second aspect of the present invention, the weighted image acquisition module generating an attention weight matrix using the high gray - value region mask and the remote - sensing image includes:
[0040] Multiply the high gray - value region mask by the remote - sensing image to obtain high gray - value features;
[0041] After passing the high gray - value features and the remote - sensing image through a convolutional layer and a batch normalization layer respectively, perform element - by - element merging and use a first activation function for non - linear function mapping to obtain a merged feature after non - linear mapping;
[0042] Perform convolutional normalization on the merged feature after non - linear mapping and use a second activation function to map the data to obtain an attention weight matrix.
[0043] As an alternative implementation, in the second aspect of the present invention, the weighted image acquisition module performing weighted processing on the remote - sensing image using the attention weight matrix to obtain a weighted remote - sensing image includes:
[0044] Multiply the attention weight matrix by the remote - sensing image to obtain a weighted remote - sensing image.
[0045] As an alternative implementation, in the second aspect of the present invention, the denoising module inputting the weighted remote - sensing image into a diffusion model for denoising processing to obtain multiple denoised image blocks includes:
[0046] Randomly sample from a standard normal distribution to generate an initial noise image, the size of the initial noise image being the same as that of the weighted remote - sensing image;
[0047] Crop the initial noise image and the weighted remote - sensing image according to a preset position dictionary to obtain multiple noise image blocks and multiple degraded image blocks, where there is an overlapping region between two adjacent noise image blocks;
[0048] Use a conditional diffusion model to estimate the noise of the noise image blocks and the degraded image blocks to obtain an estimated noise;
[0049] Calculate the noise mean according to the estimated noise and the accumulated weight to obtain the average estimated noise;
[0050] Update the multiple noise image blocks according to the average estimated noise to obtain multiple denoised image blocks.
[0051] As an alternative implementation, in the second aspect of the present invention, the denoising module calculates the noise mean according to the estimated noise and the accumulated weight to obtain the average estimated noise, including:
[0052] Perform element-wise division of the estimated noise by the accumulated weight to obtain the average estimated noise.
[0053] As an alternative implementation, in the second aspect of the present invention, the merging module merges all the denoised image blocks to obtain a cloud-removed remote sensing image, including:
[0054] Merge all the denoised image blocks, and use the weighted average method for the estimated noise in the overlapping regions to ensure smooth transition between the merged denoised image blocks.
[0055] The third aspect of the present invention discloses a remote sensing image cloud removal device based on an attention mechanism and a diffusion model, the device includes:
[0056] A memory storing executable program code;
[0057] A processor coupled to the memory;
[0058] The processor calls the executable program code stored in the memory and executes some or all of the steps of a remote sensing image cloud removal method disclosed in the first aspect of the embodiments of the present invention.
[0059] The fourth aspect of the embodiments of the present invention discloses a computer storage medium, the computer storage medium stores computer instructions, and when the computer instructions are called, they are used to execute some or all of the steps of a remote sensing image cloud removal method disclosed in the first aspect of the embodiments of the present invention.
[0060] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0061] By introducing a high gray - value attention mechanism, the attention to the high - gray - value regions in the image can be effectively enhanced. By suppressing or filtering out the low - gray - value regions, the model is allowed to automatically identify thick cloud regions, better learn the characteristic information of clouds, allocate more attention weights during the restoration and reconstruction process of thick cloud regions, and prioritize their reconstruction. Secondly, on the above basis, combined with a block - based denoising diffusion model, by using a guided denoising process for smooth noise estimation of overlapping image patches during the inference process, image restoration of unknown size is achieved. The image reconstruction process is more stable than other generative models, and the restoration of detail texture and color accuracy is more precise, effectively solving the problems of weak adaptability, long inference time, and poor restoration effect of thick cloud regions in the existing remote - sensing image cloud - removal methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0063] Figure 1 is a schematic flowchart of a method for removing clouds from remote - sensing images based on an attention mechanism and a diffusion model disclosed in an embodiment of the present invention;
[0064] Figure 2 is a structural block diagram of a system for removing clouds from remote - sensing images based on an attention mechanism and a diffusion model disclosed in an embodiment of the present invention;
[0065] Figure 3 is a schematic structural diagram of a device for removing clouds from remote - sensing images based on an attention mechanism and a diffusion model disclosed in an embodiment of the present invention;
[0066] Figure 4 is a network structure diagram of a block - based denoising probability diffusion model in an embodiment of the present invention;
[0067] Figure 5 is a schematic diagram of the network architecture of the high - gray - value attention mechanism in an embodiment of the present invention;
[0068] Figure 6 is a comparison chart of the cloud - removal effect of a cloud - covered image sample 1 by a method for removing clouds from remote - sensing images based on an attention mechanism and a diffusion model disclosed in an embodiment of the present invention and the effects of other latest cloud - removal models;
[0069] Figure 7 is a comparison chart of the cloud - removal effect of a cloud - covered image sample 2 by a method for removing clouds from remote - sensing images based on an attention mechanism and a diffusion model disclosed in an embodiment of the present invention and the effects of other latest cloud - removal models. Detailed implementation manners
[0070] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.
[0071] The terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or terminal including a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products or terminals.
[0072] Referring to "embodiment" herein means that a specific feature, structure or characteristic described in conjunction with the embodiment may be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.
[0073] The present invention discloses a method and system for removing clouds from remote sensing images based on an attention mechanism and a diffusion model. By using a block-based denoising diffusion model combined with a high gray value attention mechanism, different attention weights are assigned to regions with different cloud contents in the remote sensing image, avoiding problems such as inaccurate cloud region positioning and long inference time, and improving the cloud removal effect of the remote sensing image. The following will be described in detail respectively.
[0074] Embodiment 1
[0075] Please refer to Figure 1 , Figure 1 which is a schematic flow chart of a method for removing clouds from remote sensing images based on an attention mechanism and a diffusion model disclosed in an embodiment of the present invention. Among them, Figure 1 The described method is applied to a device for removing clouds from remote sensing images based on an attention mechanism and a diffusion model. The cloud removal device may be a corresponding cloud removal terminal, cloud removal device or server, and the server may be a local server or a cloud server, which is not limited in the embodiments of the present invention. As Figure 1As shown, the method for removing clouds from remote sensing images based on the attention mechanism and diffusion model may include the following operations:
[0076] 101. Obtain a remote sensing image, perform gray-scale conversion on the remote sensing image, and obtain a gray-scale remote sensing image.
[0077] In an embodiment of the present invention, first, the optical remote sensing image I is converted into a gray-scale remote sensing image. For an RGB color image, it includes three channels: red (R), green (G), and blue (B). When converting this color image into a gray-scale image, a weighted average formula is used, so that different color channels contribute differently to the gray-scale value.
[0078] The calculation formula for the gray-scale value is as follows:
[0079] Gray = 0.2989×R + 0.5870×G + 0.1140×B
[0080] Where: R represents the value of the red channel; G represents the value of the green channel; B represents the value of the blue channel.
[0081] These weights come from the perception model of human vision. The human eye is more sensitive to different colors, most sensitive to green, followed by red, and finally blue. Therefore, the weight of the green channel is the largest, followed by red, and the smallest for blue.
[0082] 102. Perform weighted processing on the remote sensing image through the high gray-scale value attention mechanism and the gray-scale remote sensing image to obtain a weighted remote sensing image.
[0083] 103. Input the weighted remote sensing image into a diffusion model for denoising processing to obtain multiple denoised image blocks.
[0084] 104. Perform merging processing on all the denoised image blocks to obtain a cloud-free remote sensing image.
[0085] In an optional embodiment, the performing weighted processing on the remote sensing image through the high gray-scale value attention mechanism and the gray-scale remote sensing image in step 102 to obtain a weighted remote sensing image includes:
[0086] Extract a high gray-scale value region from the gray-scale remote sensing image using a learnable gray-scale threshold to obtain a high gray-scale value region mask;
[0087] Generate an attention weight matrix using the high gray-scale value region mask and the remote sensing image;
[0088] Perform weighted processing on the remote sensing image using the attention weight matrix to obtain a weighted remote sensing image.
[0089] In an alternative embodiment, the step of generating the attention weight matrix using the high gray value region mask and the remote sensing image includes:
[0090] Multiply the high gray value region mask by the remote sensing image to obtain high gray value features;
[0091] After passing the high gray value features and the remote sensing image through a convolutional layer and a batch normalization layer respectively, perform element-wise merging, and use a first activation function for non-linear function mapping to obtain the merged features after non-linear mapping;
[0092] Perform convolutional normalization on the merged features after non-linear mapping, and use a second activation function to map the data to obtain the attention weight matrix.
[0093] In an alternative embodiment, the step of using the attention weight matrix to perform weighted processing on the remote sensing image to obtain the weighted remote sensing image includes:
[0094] Multiply the attention weight matrix by the remote sensing image to obtain the weighted remote sensing image.
[0095] In the embodiment of the present invention, a learnable gray threshold is used to extract a high gray value region from the gray remote sensing image to generate a high gray value region mask M 1 , which is mainly used to identify cloud-covered areas. Among them, the gray threshold is a learnable parameter, and the initial value is set to 0.5, allowing the model to dynamically adjust the threshold according to different features of the input remote sensing image. This adaptive mechanism enables the attention module to focus more accurately on cloud regions, thereby effectively improving the feature expression ability during the cloud removal process. Multiply the high gray value region mask M 1 by the input remote sensing image I to generate high gray value features F t = M 1 × I; then, after passing the input remote sensing image I and the high gray value features through a convolutional layer (W1) and a batch normalization layer (W2) respectively, perform element-wise merging, and then use a first activation function for non-linear function mapping to obtain the merged features after non-linear mapping. It should be noted that the ReLU activation function is used as the first activation function here; perform convolutional normalization on the merged features after non-linear mapping (W3), and map the data to the interval (0,1) through a second activation function, thereby obtaining the attention weight matrix HGA(I). It should be noted that the Sigmoid activation function is used as the second activation function here, and the expression of HGA(I) is as follows:
[0096] HGA(I) = σ Sigmoid (W 2 (σ Relu(W 0 (I)+W 1 (F t ))))
[0097] Finally, multiply HGA(I) by the input remote sensing image to obtain the weighted remote sensing image F′, and the expression of F′ is as follows:
[0098] F′ = I × HGA(I)
[0099] It can be seen that by using a learnable gray threshold to extract high-gray-value regions from the gray-scale remote sensing image, a high-gray-value region mask is generated. This mask can accurately reflect the position and range of clouds in the image, providing key information for generating the attention weight matrix subsequently. Then, use the high-gray-value region mask and the remote sensing image to generate the attention weight matrix, and perform weighted processing on the remote sensing image. This weighted processing can enhance the feature representation of the high-gray-value regions in the image, enabling the model to more accurately identify and process clouds in subsequent denoising processing, improving the accuracy and efficiency of cloud removal. At the same time, due to the use of a convolutional neural network for feature extraction, this method also has a certain degree of adaptability and robustness, and can handle the cloud removal tasks of remote sensing images under different scenarios and conditions. Therefore, this step improves the accuracy and efficiency of cloud removal, laying a foundation for generating high-quality cloud-free remote sensing images.
[0100] In an optional embodiment, the step of inputting the weighted remote sensing image into the diffusion model for denoising processing to obtain multiple denoised image blocks includes:
[0101] Randomly sample from the standard normal distribution to generate an initial noise image, and the size of the initial noise image is the same as that of the weighted remote sensing image;
[0102] Crop the initial noise image and the weighted remote sensing image according to a preset position dictionary to obtain multiple noise image blocks and multiple degraded image blocks, where there is an overlapping area between two adjacent noise image blocks;
[0103] Use the conditional diffusion model to estimate the noise of the noise image blocks and the degraded image blocks to obtain the estimated noise;
[0104] Calculate the noise mean according to the estimated noise and the cumulative weight to obtain the average estimated noise;
[0105] Update the multiple noise image blocks according to the average estimated noise to obtain multiple denoised image blocks.
[0106] In an optional embodiment, the step of calculating the noise mean according to the estimated noise and the cumulative weight to obtain the average estimated noise in the above steps includes:
[0107] Perform an element-wise division of the estimated noise by the cumulative weight to obtain the average estimated noise.
[0108] In the embodiments of the present invention, first, initial noise sampling is performed, and initial noise images X are randomly sampled from the standard normal distribution t , and the cumulative estimated noise is initialized and the cumulative weight M = 0 is initialized. According to a preset dictionary containing D overlapping image block positions, the block position P d is used to crop the current noise image X t and the weighted remote sensing image respectively, to obtain the noise image block and the degraded image block where P d is a binary mask matrix of the same dimension as X 0 and , representing the position of the d-th p×p block in the image, and d ∈ D.
[0109] Randomly sample a time step t from the uniform distribution set {1, 2,..., T} to determine a specific stage in the diffusion process of the model, and perform step-by-step sampling according to the preset number of implicit sampling steps S. The following operations are performed for each sampling step i until i = 1 to obtain multiple denoised image blocks:
[0110] Calculate the time step t according to the current step i, and the calculation expression is: t = (i - 1)·T / S + 1. If i > 1, then calculate the next time step t next , and the calculation expression is: t next = (i - 2)·T / S + 1. If i = 1, then t next = 0;
[0111] Use the conditional diffusion model to update the estimated noise and the cumulative weight for each image block position d. The update expression for the estimated noise is:
[0112] Perform an element-wise division of the estimated noise by the cumulative weight to obtain the average estimated noise of the block, and its calculation expression is: where, represents element-wise division;
[0113] Update the noise image block through the following formula:
[0114] Among them, the training process of the diffusion model includes the following steps:
[0115] During the training process, an unknown cloud-free image of any size is defined as X0 ∈R C×H×W , define the cloudy observation image as Starting element by element, P i is the same dimension as X 0 and a binary mask matrix of the same dimension as X, representing the position of the i-th p×p block in the image.
[0116] Training step S1: Input a pair of images, including a cloudy occluded picture and a cloudless clear image, where: X 0 is the clear image; is the image affected by cloud cover.
[0117] Training step S2: Randomly sample the binary mask P i :
[0118] Randomly generate a binary mask P i , which is used to select a part of the image area.
[0119] Training step S3: Randomly crop the image area:
[0120] Use the binary mask P i to select a part of the image area for cropping, which specifically includes the following steps: For the clear image X 0 , use the mask to crop to obtain the cropped image patch i.e., For the image affected by weather also use the mask to crop to obtain the cropped degraded image patch
[0121] Training step S4: Sample the time step t:
[0122] Randomly sample a time step t from the uniform distribution set {1, 2,..., T}, which is used to determine a specific stage in the diffusion process of the model.
[0123] Training step S5: Noise sampling t :
[0124] Sample a noise from the normal distribution t , which is used to add randomness to the model input so that the model can be trained to process and remove noise.
[0125] Training step S6: Gradient descent optimization:
[0126] Perform one gradient descent step to optimize the model parameter θ to minimize the following objective function:
[0127]
[0128] Among them, the objective function is used to measure the error between the model output and the image after adding noise, aiming to enable the model to learn to restore clear visual effects in degraded images of different degrees.
[0129] Training step S7: Repeat training steps S2 to S6 until the loss function of the model reaches a preset convergence threshold. Finally, output the trained model parameters θ to complete the model training.
[0130] In an optional embodiment, the merging and processing of all the denoised image patches in step 104 to obtain a cloud-removed remote sensing image includes:
[0131] Merge all the denoised image patches, and for the estimated noise in the overlapping regions, use the method of weighted average to ensure a smooth transition between the merged denoised image patches.
[0132] In the embodiment of the present invention, for the overlapping regions, the average update formula for estimating noise is:
[0133]
[0134] Where, represents the average estimated noise of the overlapping pixel region. d represents different sampling directions or blocks in the overlapping region (for example, 1 to 4 represent four directions). represents the noise estimation of the model for the d-th overlapping region at time step t, where: represents the current noise image patch of the d-th direction or sampling block; represents the corresponding degraded image patch; t represents the current time step, which controls the stage of the model in the diffusion process.
[0135] It can be seen that by introducing the high gray value attention mechanism, the attention to the high gray value regions in the image can be effectively enhanced. By suppressing or filtering out the low gray value regions, the model is allowed to automatically identify the thick cloud regions, better learn the characteristic information of the clouds, assign more attention weights to the thick cloud regions during the restoration and reconstruction process, and prioritize their reconstruction; secondly, on the above basis, combined with the block-based denoising diffusion model, by using a guided denoising process for smooth noise estimation of the overlapping image patches during the inference process, image restoration of unknown size is achieved, and the image reconstruction process is more stable than other generative models, and the detail texture and color accuracy restoration processing are more accurate, effectively solving the problems of weak adaptability, long inference time, and poor restoration effect of thick cloud regions in the existing remote sensing image cloud removal methods.
[0136] Embodiment 2
[0137] Please refer to Figure 2 , Figure 2It is a schematic structural diagram of a remote sensing image cloud removal device based on an attention mechanism and a diffusion model disclosed in an embodiment of the present invention. Among them, Figure 3 The described device can be applied to corresponding cloud removal terminals, cloud removal devices or servers, and the server can be a local server or a cloud server, which is not limited in the embodiments of the present invention. As Figure 3 shown, the device may include:
[0138] A grayscale image acquisition module 201, which is configured to acquire a remote sensing image and perform grayscale conversion on the remote sensing image to obtain a grayscale remote sensing image.
[0139] In an embodiment of the present invention, first, an optical remote sensing image I is converted into a grayscale remote sensing image. For an RGB color image, it contains three channels: red (R), green (G), and blue (B). When converting this color image into a grayscale image, a weighted average formula is used, so that different color channels contribute differently to the grayscale value.
[0140] The calculation formula of the grayscale value is as follows:
[0141] Gray = 0.2989×R + 0.5870×G + 0.1140×B
[0142] Where: R represents the value of the red channel; G represents the value of the green channel; B represents the value of the blue channel.
[0143] These weights come from the perception model of human vision. The human eye is more sensitive to different colors, most sensitive to green, followed by red, and finally blue. Therefore, the weight of the green channel is the largest, followed by red, and the smallest for blue.
[0144] A weighted image acquisition module 202, which is configured to perform weighted processing on the remote sensing image through a high grayscale value attention mechanism and the grayscale remote sensing image to obtain a weighted remote sensing image.
[0145] A denoising module 203, which is configured to input the weighted remote sensing image into a diffusion model for denoising processing to obtain a plurality of denoised image blocks.
[0146] A merging module 204, which is configured to merge all the denoised image blocks to obtain a cloud-free remote sensing image.
[0147] In an optional embodiment, the weighted image acquisition module 202 performs weighted processing on the remote sensing image through a high grayscale value attention mechanism and the grayscale remote sensing image to obtain a weighted remote sensing image, including:
[0148] Extract a high gray - value region from the grayscale remote - sensing image using a learnable gray - scale threshold to obtain a high gray - value region mask;
[0149] Generate an attention weight matrix using the high gray - value region mask and the remote - sensing image;
[0150] Perform weighted processing on the remote - sensing image using the attention weight matrix to obtain a weighted remote - sensing image.
[0151] In an alternative embodiment, the weighted image acquisition module 202 generating an attention weight matrix using the high gray - value region mask and the remote - sensing image includes:
[0152] Multiply the high gray - value region mask by the remote - sensing image to obtain high gray - value features;
[0153] After passing the high gray - value features and the remote - sensing image through a convolutional layer and a batch normalization layer respectively, perform element - by - element merging, and perform non - linear function mapping using a first activation function to obtain merged features after non - linear mapping;
[0154] Perform a convolutional normalization operation on the merged features after non - linear mapping, and map the data using a second activation function to obtain an attention weight matrix.
[0155] In an alternative embodiment, the weighted image acquisition module 202 performing weighted processing on the remote - sensing image using the attention weight matrix to obtain a weighted remote - sensing image includes:
[0156] Multiply the attention weight matrix by the remote - sensing image to obtain a weighted remote - sensing image.
[0157] In the embodiments of the present invention, a learnable gray - scale threshold is used to extract a high gray - value region from the grayscale remote - sensing image to generate a high gray - value region mask M 1 , which is mainly used to identify cloud - covered regions. Among them, the gray - scale threshold is a learnable parameter, and the initial value is set to 0.5, allowing the model to dynamically adjust the threshold according to different features of the input remote - sensing image. This adaptive mechanism enables the attention module to more accurately focus on cloud regions, thereby effectively improving the feature expression ability in the cloud removal process. Multiply the high gray - value region mask M 1 by the input remote - sensing image I to generate high gray - value features F t = M 1×I; Then, the input remote sensing image I and the high - gray - value features are respectively passed through a convolutional layer (W1) and a batch normalization layer (W2), and then merged element - by - element. Then, a non - linear function mapping is performed using the first activation function to obtain the merged features after non - linear mapping. It should be noted that the ReLU activation function is used as the first activation function here; perform a convolutional normalization operation (W3) on the merged features after non - linear mapping, and map the data to the interval (0, 1) through the second activation function, thereby obtaining the attention weight matrix HGA(I). It should be noted that the Sigmoid activation function is used as the second activation function here, and the expression of HGA(I) is as follows:
[0158] HGA(I)=σ Sigmoid (W 2 (σ Relu (W 0 (I)+W 1 (F t ))))
[0159] Finally, multiply HGA(I) by the input remote sensing image to obtain the weighted remote sensing image F′, and the expression of F′ is as follows:
[0160] F′ = I×HGA(I)
[0161] It can be seen that by using a learnable gray - scale threshold to extract the high - gray - value region from the gray - scale remote sensing image, a high - gray - value region mask is generated. This mask can accurately reflect the position and range of clouds in the image, providing key information for generating the attention weight matrix subsequently. Then, use the high - gray - value region mask and the remote sensing image to generate the attention weight matrix and perform weighted processing on the remote sensing image. This weighted processing can enhance the feature representation of the high - gray - value region in the image, enabling the model to more accurately identify and process clouds in subsequent denoising processing, improving the accuracy and efficiency of cloud removal. At the same time, due to using a convolutional neural network for feature extraction, this method also has a certain degree of adaptability and robustness, and can handle the cloud - removal tasks of remote sensing images under different scenarios and conditions. Therefore, this step improves the accuracy and efficiency of cloud removal, laying a foundation for generating high - quality cloud - free remote sensing images.
[0162] In an optional embodiment, the denoising module 203 inputs the weighted remote sensing image into a diffusion model for denoising processing to obtain multiple denoised image blocks, including:
[0163] Randomly sample from the standard normal distribution to generate an initial noise image, and the size of the initial noise image is the same as that of the weighted remote sensing image;
[0164] Crop the initial noise image and the weighted remote sensing image according to a preset position dictionary to obtain a plurality of noise image patches and a plurality of degraded image patches, wherein there is an overlapping area between two adjacent noise image patches;
[0165] Use a conditional diffusion model to estimate the noise of the noise image patches and the degraded image patches to obtain an estimated noise;
[0166] Calculate the noise mean according to the estimated noise and the cumulative weight to obtain an average estimated noise;
[0167] Update the plurality of noise image patches according to the average estimated noise to obtain a plurality of denoised image patches.
[0168] In an alternative embodiment, the denoising module 203 calculates the noise mean according to the estimated noise and the cumulative weight to obtain an average estimated noise, including:
[0169] Perform an element-wise division of the estimated noise by the cumulative weight to obtain the average estimated noise.
[0170] In the embodiment of the present invention, first, initial noise sampling is performed, and initial noise image X is randomly sampled from a standard normal distribution t , and the cumulative estimated noise is initialized, and the cumulative weight M = 0 is initialized. According to a preset dictionary containing the positions of D overlapping image patches, the block position P d is used to crop the current noise image X t and the weighted remote sensing image respectively, to obtain the noise image patches and the degraded image patches where P d is a binary mask matrix of the same dimension as X 0 and , representing the position of the d-th p×p block in the image, and d ∈ D.
[0171] Randomly sample a time step t from the uniform distribution set {1, 2,..., T} to determine a specific stage in the diffusion process of the model, and perform step-by-step sampling according to the preset implicit sampling step number S. The following operations are performed for each sampling step i until i = 1 to obtain a plurality of denoised image patches:
[0172] Calculate the time step t according to the current step i, and the calculation formula is: t = (i - 1)·T / S + 1. If i > 1, then calculate the next time step t next , and the calculation formula is: t next = (i - 2)·T / S + 1. If i = 1, then tnext = 0;
[0173] Using the conditional diffusion model, estimate noise update and cumulative weight are performed for each image patch position d, and the update expression of the estimated noise is:
[0174] Perform element-wise division of the estimated noise by the cumulative weight to obtain the average estimated noise of the patch, and its calculation expression is: Where, represents element-wise division;
[0175] Update the noisy image patch through the following formula:
[0176] In an optional embodiment, the merging module 204 merges all the denoised image patches to obtain a cloud-removed remote sensing image, including:
[0177] Merge all the denoised image patches, and for the estimated noise in the overlapping region, use the weighted average method to ensure a smooth transition between the merged denoised image patches.
[0178] In the embodiment of the present invention, for the overlapping region, the average update formula of the estimated noise is:
[0179]
[0180] Where, represents the average estimated noise of the overlapping pixel region. d represents different sampling directions or patches in the overlapping region (for example, 1 to 4 represent four directions). represents the noise estimation of the model for the d-th overlapping region at the time step t, where: represents the current noisy image patch of the d-th direction or sampling patch; represents the corresponding degraded image patch; t represents the current time step, controlling the stage of the model in the diffusion process.
[0181] It can be seen that by introducing the high gray value attention mechanism, the attention to the high gray value region in the image can be effectively enhanced. By suppressing or filtering out the low gray value region, the model is allowed to automatically identify the thick cloud region, better learn the feature information of the cloud, allocate more attention weights during the restoration and reconstruction process of the thick cloud region, and prioritize its reconstruction; secondly, on the above basis, combined with the block-based denoising diffusion model, by using the guided denoising process of smooth noise estimation for overlapping image patches during the inference process, image restoration of unknown size is achieved, and the image reconstruction process is more stable than other generative models, and the detail texture and color accuracy restoration processing are more accurate, effectively solving the problems of weak adaptive ability, long inference time, and poor restoration effect of thick cloud regions in the existing remote sensing image cloud removal methods.
[0182] Embodiment III
[0183] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of another remote sensing image cloud removal device based on the attention mechanism and diffusion model disclosed in the embodiments of the present invention. As Figure 3 shown, the device may include:
[0184] A memory 301 storing executable program code;
[0185] A processor 302 coupled to the memory 301;
[0186] The processor 302 calls the executable program code stored in the memory 301 and executes some or all of the steps in a remote sensing image cloud removal method disclosed in Embodiment I of the present invention.
[0187] Embodiment IV
[0188] The embodiments of the present invention disclose a computer storage medium storing computer instructions, which are used to execute some or all of the steps in a remote sensing image cloud removal method disclosed in Embodiment I of the present invention when the computer instructions are called.
[0189] The device embodiments described above are only illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0190] Through the specific descriptions of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, and the storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other medium that can be used to carry or store data and is computer-readable.
[0191] Finally, it should be noted that: The remote sensing image cloud removal method and system based on the attention mechanism and diffusion model disclosed in the embodiments of the present invention only disclose the preferred embodiments of the present invention, and are only used to illustrate the technical solutions of the present invention, rather than limiting them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A remote sensing image cloud removal method based on attention mechanism and diffusion model, characterized in that: The method comprises: Acquire a remote sensing image, and perform grayscale conversion on the remote sensing image to obtain a grayscale remote sensing image; Performing weighted processing on the remote sensing image through a high grayscale value attention mechanism and the grayscale remote sensing image to obtain a weighted remote sensing image; Inputting the weighted remote sensing image into a diffusion model for denoising to obtain a plurality of denoised image blocks; All the denoised image blocks are merged to obtain a cloud-free remote sensing image.
2. The remote sensing image cloud removal method based on attention mechanism and diffusion model according to claim 1 is characterized in that: The step of performing weighted processing on the remote sensing image through the high grayscale value attention mechanism and the grayscale remote sensing image to obtain the weighted remote sensing image comprises: A learnable grayscale threshold is used to extract a high grayscale value region from the grayscale remote sensing image to obtain a high grayscale value region mask; Generate an attention weight matrix using the high gray value region mask and the remote sensing image; The remote sensing image is weighted using the attention weight matrix to obtain a weighted remote sensing image.
3. The remote sensing image cloud removal method based on attention mechanism and diffusion model according to claim 2 is characterized in that: The generating of the attention weight matrix by using the high gray value area mask and the remote sensing image comprises: Multiplying the high gray value area mask with the remote sensing image to obtain high gray value features; The high gray value features and the remote sensing image are respectively passed through a convolution layer and a batch normalization layer, and then merged element by element, and a nonlinear function mapping is performed using a first activation function to obtain a merged feature after nonlinear mapping; A convolution normalization operation is performed on the merged features after the nonlinear mapping, and the data is mapped using a second activation function to obtain an attention weight matrix.
4. The remote sensing image cloud removal method based on attention mechanism and diffusion model according to claim 2 is characterized in that: The step of using the attention weight matrix to perform weighted processing on the remote sensing image to obtain a weighted remote sensing image comprises: The attention weight matrix is multiplied by the remote sensing image to obtain a weighted remote sensing image.
5. The remote sensing image cloud removal method based on attention mechanism and diffusion model according to claim 1, characterized in that: The step of inputting the weighted remote sensing image into a diffusion model for denoising to obtain a plurality of denoised image blocks comprises: Generating an initial noise image by random sampling from a standard normal distribution, wherein the size of the initial noise image is the same as the size of the weighted remote sensing image; The initial noise image and the weighted remote sensing image are cropped according to a preset position dictionary to obtain a plurality of noise image blocks and a plurality of degraded image blocks, wherein there is an overlapping area between two adjacent noise image blocks; Using a conditional diffusion model, the noise image block and the degraded image block are subjected to noise estimation to obtain estimated noise; Calculate the noise mean according to the estimated noise and the accumulated weight to obtain the average estimated noise; The multiple noisy image blocks are updated according to the average estimated noise to obtain multiple denoised image blocks.
6. The remote sensing image cloud removal method based on attention mechanism and diffusion model according to claim 5, characterized in that: The calculating the noise mean according to the estimated noise and the accumulated weight to obtain the average estimated noise comprises: The estimated noise is divided by the accumulated weight element-wise to obtain an average estimated noise.
7. The remote sensing image cloud removal method based on attention mechanism and diffusion model according to claim 5, characterized in that: The merging of all the denoised image blocks to obtain a cloud-free remote sensing image includes: All denoised image blocks are merged, and the estimated noise in the overlapping area is weighted averaged to ensure a smooth transition between the merged denoised image blocks.
8. A remote sensing image cloud removal system based on attention mechanism and diffusion model, characterized in that: The system comprises: A grayscale image acquisition module, wherein the grayscale image acquisition module is used to acquire a remote sensing image and perform grayscale conversion on the remote sensing image to obtain a grayscale remote sensing image; A weighted image acquisition module, wherein the weighted image acquisition module is used to perform weighted processing on the remote sensing image through a high gray value attention mechanism and the gray remote sensing image to obtain a weighted remote sensing image; A denoising module, wherein the denoising module is used to input the weighted remote sensing image into a diffusion model for denoising processing to obtain a plurality of denoised image blocks; A merging module is used to merge all the denoised image blocks to obtain a cloud-free remote sensing image.
9. A remote sensing image cloud removal device based on attention mechanism and diffusion model, characterized in that: The device comprises: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the remote sensing image cloud removal method based on the attention mechanism and diffusion model as described in any one of claims 1-7.
10. A computer storage medium, characterized in that: The computer storage medium stores computer instructions, which, when called, are used to execute the remote sensing image cloud removal method based on the attention mechanism and diffusion model as described in any one of claims 1-7.
Citation Information
Patent Citations
Thick cloud remote sensing removal method and device
CN103729831A
Method for removing thick cloud of high-resolution optical image based on fully convolutional network
CN107590782A
Remote sensing image cloud removal method based on double-branch channel and feature enhancement mechanism
CN113935908A
Optical remote sensing image cloud removal method and system based on diffusion model
CN116777764A
Remote sensing image cloud removal method and system
CN116823664A