A method for processing cloud occlusion in water area remote sensing images

By synchronously collecting water remote sensing images and synthesized aperture radar images, combining cloud segmentation model and diffusion processing technology, the problem of cloud occlusion in water remote sensing images is solved, and high-quality cloudless water remote sensing images are generated, improving the spatial structure and semantic information of the image.

CN120014483BActive Publication Date: 2025-06-24SHANDONG YUANMINGQING TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510486962.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-06-24
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

Cloud occlusion in water remote sensing images causes spatial structure, texture and context information to be occluded, affecting the pixel distribution of image scenes. The prior art is not effective when covered by thick clouds.

Method used

By synchronously collecting water remote sensing images and synthetic aperture radar images, the cloud segmentation model and synthetic aperture radar image recognition model extract the target object semantics of the cloud cover area and input them into the pre-trained cloud processing diffusion model. The diffusion process and the Unet model are used to process the cloud cover area to generate cloud-free water remote sensing images.

Benefits of technology

Effectively remove cloud occlusion and generate high-quality remote sensing images of cloudless waters, improving the accuracy of the spatial structure and semantic information of the image, especially in the case of thick cloud coverage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014483B_ABST
    Figure CN120014483B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for processing cloud occlusion in water area remote sensing images. The present invention synchronously acquires water area remote sensing images and water area synthetic aperture radar images of a target water area; screens the target water area remote sensing images that need cloud occlusion processing; aligns the target water area remote sensing images with the associated water area synthetic aperture radar images; after alignment, uses a synthetic aperture radar image recognition model to extract the semantic information of the target objects in the cloud-covered area in the water area synthetic aperture radar image domain; inputs the aligned target water area remote sensing images, the associated water area synthetic aperture radar images, and the target object semantics into a pre-trained cloud processing diffusion model, and the cloud processing diffusion model processes the cloud-covered area of the target water area remote sensing images according to the guidance of the water area synthetic aperture radar images and the target object semantics to generate cloud-free target water area remote sensing images. This application generates images of the cloud-covered area in the water area remote sensing image domain guided by the synchronous water area synthetic aperture radar images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud processing in remote sensing images, and particularly to a method for processing cloud occlusion in water area remote sensing images. Background Art

[0002] The water area remote sensing images taken by drones are an important means for water area inspection. However, there are clouds in many water area remote sensing images. The clouds obscure the spatial structure, texture, and context of the water area remote sensing images, cover the effective information in the water area remote sensing images, and significantly affect the pixel distribution of the water area remote sensing image scenes.

[0003] The current solutions for processing cloud occlusion are as follows: migrating the trained cloud-free optical image model to the cloud-covered optical image processed by the cloud removal method. For example, reconstructing the cloud-covered area by propagating the texture structure of the cloud-free background, and further optimizing the result by enriching the feature space and search range of the background information. This method is easily affected by edge effects, and when there is thick cloud cover, the generation effect is very poor. For the serious information loss caused by thick cloud cover, in the prior art, a unified spatio-temporal-spectral model is constructed to remove thick clouds. This model can make full use of the information from multi-source and multi-temporal data to accurately reconstruct the cloud-damaged area, but historical unoccluded images are required, and the reconstruction is not based on the real-time scene. Summary of the Invention

[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present invention provides a method for processing cloud occlusion in water area remote sensing images.

[0005] The present invention provides a method for processing cloud occlusion in water area remote sensing images, including:

[0006] Synchronously collecting the water area remote sensing image and the water area synthetic aperture radar image of the target water area by using a remote sensing imaging device and a synthetic aperture radar, and associating the two;

[0007] Identifying the target water area remote sensing image that needs to be processed for cloud occlusion by using a cloud segmentation model that supports cloud recognition;

[0008] Performing alignment processing on the target water area remote sensing image and the associated water area synthetic aperture radar image;

[0009] After alignment, obtaining the cloud-covered area from the water area synthetic aperture radar image by using a cloud mask, and extracting the semantic information of the target objects in the cloud-covered area in the water area synthetic aperture radar image domain by using a synthetic aperture radar image recognition model;

[0010] Input the aligned remote sensing image of the target water area, the synthetic aperture radar image of the water area associated therewith, and the semantics of the target object into a pre-trained cloud processing diffusion model. The cloud processing diffusion model processes the cloud-covered area of the remote sensing image of the target water area according to the guidance of the synthetic aperture radar image of the water area and the semantics of the target object, and generates a cloud-free remote sensing image of the target water area.

[0011] Further, the process of aligning the remote sensing image of the target water area and the synthetic aperture radar image of the water area associated therewith includes: segmenting the non-cloud area in the remote sensing image of the target water area through the cloud segmentation model, and extracting the alignment reference target in the remote sensing image domain from the non-cloud area of the remote sensing image of the target water area;

[0012] Extract the alignment reference target in the synthetic aperture radar image domain from the synthetic aperture radar image of the water area associated with the remote sensing image of the target water area;

[0013] Match the alignment reference target in the remote sensing image domain within the range of the alignment reference target in the synthetic aperture radar image domain;

[0014] Align the remote sensing image of the water area and the synthetic aperture radar image of the water area according to the rotation and translation relationship between the matching alignment reference targets in the two domains.

[0015] Further, the cloud processing diffusion model includes: a VQVAE encoder, a VQVAE decoder, a Unet model, a radar feature encoder for the object under the cloud, a shared temporal encoder, and a shared text encoder for the object under the cloud;

[0016] Among them, the pre-trained VQVAE encoder converts the input remote sensing image of the target water area and the synthetic aperture radar image of the water area associated therewith into the latent space;

[0017] Add Gaussian noise iteratively to the latent space representation of the remote sensing image of the target water area by using the diffusion process;

[0018] The diffusion result of the latent space representation of the remote sensing image of the target water area is input into the Unet model for iterative denoising processing;

[0019] The radar feature encoder for the object under the cloud and the Unet model are set in parallel, and are used to provide the radar features of the object under the cloud extracted from the synthetic aperture radar image of the water area to the intermediate module and the decoding module of the Unet model during each iterative denoising process. The radar features of the object under the cloud guide the Unet model to generate a representation in the remote sensing image domain of the water area based on the object under the cloud in the corresponding area under the cloud;

[0020] The shared sub-cloud object text encoder encodes the semantics of the objects extracted from the cloud-covered area in the water synthetic aperture radar image domain by using the synthetic aperture radar image recognition model, and provides the object semantic encoding to the intermediate module and the decoding module of the Unet model during the noise reduction process of each iteration. The object semantic encoding guides the Unet model to generate a representation of the water remote sensing image domain based on the sub-cloud objects in the corresponding sub-cloud area;

[0021] The shared time encoder provides the iteration time step encoding for the sub-cloud object radar feature encoder and the Unet model during the noise reduction process of each iteration;

[0022] The pre-trained VQVAE decoder decodes the processed target water remote sensing image based on the noise reduction result generated by the Unet model.

[0023] Furthermore, the Unet model includes: a number of cascaded encoding modules, an intermediate module, and decoding modules corresponding to the number of encoding modules. Among them, skip connections are set between the corresponding encoding modules and decoding modules; for any level of decoding module, the output of the decoding module of the previous level and the output of the encoder of the same level are combined and input into this decoding module.

[0024] Furthermore, the sub-cloud object radar feature encoder includes: a number of cascaded encoding modules and an intermediate module copied from the Unet model. Each encoding module and intermediate module are connected to a 1×1 convolutional layer. The 1×1 convolutional layer connected to the intermediate module is connected to the intermediate module of the Unet model, and the 1×1 convolutional layer connected to the encoding module is connected to the corresponding decoding module in the Unet model.

[0025] Furthermore, the training process of the cloud processing diffusion model includes:

[0026] Training the parameters of the Unet model so that the Unet model can generally restore the diffused water remote sensing image. The training method includes: inputting the water remote sensing images in the training set into the cloud processing diffusion model, and the pre-trained VQVAE encoder converts the water remote sensing images into the latent space; then, Gaussian noise is iteratively added to the latent space representation of the water remote sensing images by using the diffusion process; multiple arbitrary samplings are performed during the diffusion process to obtain the iteration time steps, noises, and latent space characterizations before and after adding noise of all samplings; the shared time encoder encodes the iteration time steps of the samplings, and the latent space characterization after adding noise and the time step encoding are combined and input into the Unet model, expecting the latent space characterization after adding noise to be restored to the latent space characterization before adding noise according to the noise predicted by the Unet; finally, the pre-trained VQVAE encoder will restore the water remote sensing image based on the final noise reduction result.

[0027] Further, the training process of the cloud processing diffusion model includes: after training the Unet model, copying a number of cascaded encoding modules and intermediate modules of the Unet model and constructing a radar feature encoder for the object under the cloud; training the radar feature encoder for the object under the cloud and the shared text encoder for the object under the cloud to provide optimization conditions for the Unet, and after using the optimization conditions to control the denoising of the noise predicted by the Unet model, a latent space representation of a cloud-free water area image can be generated.

[0028] Further, the method for training the radar feature encoder for the object under the cloud and the shared text encoder for the object under the cloud includes:

[0029] For any synthetic aperture radar image of a water area in the training set, the cloud-covered area is extracted by using the corresponding cloud mask; the semantic information of the object in the cloud-covered area in the synthetic aperture radar image domain of the water area is extracted by using the synthetic aperture radar image recognition model;

[0030] The water area remote sensing image and the synthetic aperture radar image of the water area are input into the pre-trained VQVAE encoder of the cloud processing diffusion model; the pre-trained VQVAE encoder converts the water area remote sensing image and the synthetic aperture radar image of the water area into the latent space;

[0031] Then, consistent Gaussian noise is synchronously added to the latent space representations of the water area remote sensing image and the cloud-free water area remote sensing image by using the diffusion process;

[0032] The shared text encoder for the object under the cloud encodes the semantic information of the object in the cloud-covered area in the synthetic aperture radar image domain of the water area to obtain the object semantic encoding and provides it to the Unet model. The object semantic encoding provides the semantic information of the object under the cloud for the Unet model to strengthen the semantic representation;

[0033] Multiple arbitrary samplings are performed during the diffusion processes of the water area remote sensing image and the cloud-free water area remote sensing image to obtain the iteration time steps, noise, and latent space representations before and after noise addition of the two samplings;

[0034] The shared time encoder encodes the iteration time steps of the sampling;

[0035] The latent space representation of the sampled synthetic aperture radar image of the water area is processed by a 1×1 convolutional layer and then combined with the latent space representation of the cloud water area remote sensing image after noise addition and input into the radar feature encoder for the object under the cloud. At the same time, the encoded sampling time steps and the object semantic encoding are combined and input into the radar feature encoder for the object under the cloud. The radar feature encoder for the object under the cloud provides the spatial information of the object under the cloud for the Unet model;

[0036] The latent space representation of the sampled and noisy cloud water area remote sensing image is input into the Unet model. At the same time, the sampled time step encoding and the target object semantic encoding are input into the Unet model;

[0037] The Unet model predicts the diffused noise under the additional conditions provided by the radar feature encoder of the object under the cloud and the object semantic encoding, so that the latent space representation of the cloud water area remote sensing image after denoising according to the predicted noise has the smallest distance from the latent space representation of the cloudless water area remote sensing image before noise addition at the same sampling stage; Finally, the pre-trained VQVAE encoder will generate a cloudless water area remote sensing image based on the final denoising result.

[0038] Furthermore, the training set includes: associated water area remote sensing images, cloud masks of water area remote sensing images, water area synthetic aperture radar images, cloudless water area remote sensing images, where the associated cloudless water area remote sensing images have the same scene as the water area remote sensing images but are not blocked by clouds.

[0039] Furthermore, during the generation process of the cloud processing diffusion model, the target water area remote sensing image and the target water area synthetic aperture radar image are provided;

[0040] The target water area remote sensing image and the target water area synthetic aperture radar image are input into the pre-trained VQVAE encoder of the cloud processing diffusion model; the pre-trained VQVAE encoder converts the target water area remote sensing image and the target water area synthetic aperture radar image into the latent space;

[0041] Then, Gaussian noise is added to the latent space representation of the target water area remote sensing image using the diffusion process;

[0042] The shared text encoder of the object under the cloud encodes the semantics of the object in the cloud-covered area in the target water area synthetic aperture radar image domain to obtain the object semantic encoding and provides it to the Unet model;

[0043] The diffusion result of the target water area remote sensing image is denoised to obtain the diffusion iteration time step, noise, and latent space representations before and after noise addition;

[0044] The shared time encoder encodes the iteration time step corresponding to the denoising process;

[0045] The latent space representation of the target water area synthetic aperture radar image during the denoising process is processed by a 1×1 convolutional layer and combined with the latent space representation of the target cloud water area remote sensing image after noise addition and input into the radar feature encoder of the object under the cloud. At the same time, the time step encoding and the object semantic encoding are combined and input into the radar feature encoder of the object under the cloud;

[0046] The latent space representation of the target cloud water area remote sensing image after adding noise at each diffusion step is input into the Unet model. At the same time, the time step encoding and the target object semantic encoding are input into the Unet model;

[0047] The Unet model predicts the noise under the additional conditions provided by the radar feature encoder of the target object under the cloud and the target object semantic encoding, so that the latent space representation of the target cloud water area remote sensing image after adding noise is denoised according to the predicted noise; finally, the pre-trained VQVAE encoder will generate a cloud-free target water area remote sensing image based on the final denoising result.

[0048] The above technical solutions provided by the embodiments of the present invention have the following advantages compared with the prior art:

[0049] The present invention synchronously collects the water area remote sensing image and the water area synthetic aperture radar image of the target water area; screens the target water area remote sensing image that needs to be processed for cloud occlusion; aligns the target water area remote sensing image and the associated water area synthetic aperture radar image; after alignment, uses the synthetic aperture radar image recognition model to extract the target object semantics in the cloud-covered area in the water area synthetic aperture radar image domain; inputs the aligned target water area remote sensing image, the associated water area synthetic aperture radar image, and the target object semantics into the pre-trained cloud processing diffusion model, and the cloud processing diffusion model processes the cloud-covered area of the target water area remote sensing image according to the guidance of the water area synthetic aperture radar image and the target object semantics to generate a cloud-free target water area remote sensing image. This application generates an image of the cloud-covered area in the water area remote sensing image domain guided by the synchronous water area synthetic aperture radar image. This application uses the ability of the cloud processing diffusion model to learn the object distribution information to generatively generate the area covered by clouds in the water area remote sensing image. Since the synchronously collected water area synthetic aperture radar image is used as a guide, the semantic space information of the target object under the cloud can be well transferred from the water area synthetic aperture radar image to the water area remote sensing image domain, and the generated image of the cloud-covered area conforms to the current real situation. By enhancing the target object semantics in the cloud-covered area through the target object semantics, the spatial information of the target object in the water area synthetic aperture radar image is supplemented to ensure the correct generation of the target object. Using the synergistic and complementary effects of optics and synthetic aperture radar, the cloud processing diffusion model can naturally connect the target under the cloud area and the non-cloud area, and the quality of the generated cloud-free water area image is higher. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.

[0051] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0052] Figure 1 It is a flowchart of a method for processing cloud occlusion in water area remote sensing images provided by an embodiment of the present invention;

[0053] Figure 2 It is an architecture diagram of a cloud processing diffusion model, a cloud segmentation model, and a synthetic aperture radar image recognition model cooperating to achieve cloud occlusion processing in water area remote sensing images provided by an embodiment of the present invention;

[0054] Figure 3 It is a schematic diagram of each model in the cloud processing diffusion model provided by an embodiment of the present invention;

[0055] Figure 4 It is a schematic diagram of the Unet model and the radar feature encoder of the target object under the cloud provided by an embodiment of the present invention;

[0056] Figure 5 It is a training flowchart of the cloud processing diffusion model provided by an embodiment of the present invention;

[0057] Figure 6 It is a flowchart of training the Unet model provided by an embodiment of the present invention;

[0058] Figure 7 It is a flowchart of the optimization conditions provided by the training of the radar feature encoder of the target object under the cloud and the shared text encoder of the target object under the cloud for the Unet, and using the optimization conditions to control the denoising of the noise predicted by the Unet model to generate the latent space representation of the cloud-free water area image;

[0059] Figure 8 It is a schematic diagram of the device for processing cloud occlusion in water area remote sensing images provided by an embodiment of the present invention. Specific Embodiments

[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0061] It should be noted that in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or elements inherent to such a process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0062] Embodiment 1

[0063] Refer to Figure 1 As shown, the method for processing cloud occlusion in the water area remote sensing image provided by this application includes:

[0064] S100, synchronously collect the water area remote sensing image and the water area synthetic aperture radar image of the target water area by using a remote sensing imaging device and a synthetic aperture radar. Since the synthetic aperture radar wave has a longer wavelength, compared with the visible light wave, it has the ability to penetrate clouds, can penetrate the cloud cover, and sense the structural information of the target object on the water area under the clouds; therefore, in the process of removing clouds in the remote sensing image domain, the water area synthetic aperture radar image can provide guiding conditions for the generation of pixels in the cloud-covered area. Through subsequent processes, this application effectively converts the water area synthetic aperture radar image into guiding conditions that can guide the generation of target object pixels in the cloud-covered area.

[0065] S200, use a pre-trained cloud segmentation model that supports cloud recognition to identify and screen out the target water area remote sensing images that need to be processed for cloud occlusion; in the prior art, there are many semantic segmentation models that support cloud recognition that can be used as the cloud segmentation model of this application, such as: the semantic segmentation model with the Unet architecture. Here, the semantic segmentation model is not the same as the subsequent Unet model, and the architectures of the two models are both Unet. Use the cloud segmentation model to traverse and process each water area remote sensing image to generate a cloud mask. For any water area remote sensing image, when the cloud in the generated cloud mask is non-empty, select this water area remote sensing image as the target water area remote sensing image.

[0066] Although both the remote sensing imaging device and the synthetic aperture radar collect information on the target water area, due to the different installation positions of the two on the flight vehicle, the collected water area remote sensing image and the water area synthetic aperture radar image may not be aligned. Therefore, after screening out the target water area remote sensing image, S300, first align the target water area remote sensing image and the associated water area synthetic aperture radar image, as Figure 2 shown, the process includes:

[0067] Segment the non-cloud area in the remote sensing image of the target water area through the cloud segmentation model, and extract the remote sensing image domain alignment reference target from the non-cloud area of the remote sensing image of the target water area;

[0068] Extract the synthetic aperture radar image domain alignment reference target from the synthetic aperture radar image of the water area associated with the remote sensing image of the target water area;

[0069] Match the remote sensing image domain alignment reference target within the range of the alignment reference target in the synthetic aperture radar image domain;

[0070] Align the remote sensing image of the water area and the synthetic aperture radar image of the water area according to the rotation and translation relationship between the matching alignment reference targets in the two domains. In the specific implementation process, use the findFundamentalMat function in opencv to determine the rotation and translation relationship according to the positions of the matching alignment reference targets in the two domains, and use the warpPerspective function in opencv to align the remote sensing image of the target water area and the synthetic aperture radar image of the water area associated with it according to the rotation and translation relationship.

[0071] S400. After alignment, use the cloud mask to obtain the cloud-covered area from the synthetic aperture radar image of the water area, and use the synthetic aperture radar image recognition model to extract the semantic information of the target objects in the cloud-covered area in the synthetic aperture radar image domain. The synthetic aperture radar image recognition model uses the CLIP model trained for target object recognition in the synthetic aperture radar image.

[0072] S500. Input the aligned remote sensing image of the target water area, the synthetic aperture radar image of the water area associated with it, and the target object semantics into the pre-trained cloud processing diffusion model. The cloud processing diffusion model processes the cloud-covered area of the remote sensing image of the target water area according to the guidance of the synthetic aperture radar image and the target object semantics, and generates a cloud-free remote sensing image of the target water area.

[0073] During the generation process of the cloud processing diffusion model, the remote sensing image of the target water area and the synthetic aperture radar image of the target water area are input into the pre-trained VQVAE encoder of the cloud processing diffusion model; the pre-trained VQVAE encoder converts the remote sensing image of the target water area and the synthetic aperture radar image of the target water area into the latent space; then, Gaussian noise is added to the latent space representation of the remote sensing image of the target water area by using the diffusion process; the shared sub-cloud target text encoder encodes the semantics of the targets in the cloud-covered area in the synthetic aperture radar image domain of the target water area to obtain the target semantic encoding and provides it to the Unet model; the diffusion result of the remote sensing image of the target water area is denoised to obtain the diffusion iteration time steps, noise, and the latent space representations before and after adding noise; the shared time encoder encodes the iteration time steps corresponding to the denoising process; the latent space representation of the synthetic aperture radar image of the target water area in the denoising process is combined with the latent space representation of the remote sensing image of the target cloud water area after adding noise and input into the sub-cloud target radar feature encoder. At the same time, the time step encoding and the target semantic encoding are combined and input into the sub-cloud target radar feature encoder; the latent space representation of the remote sensing image of the target cloud water area after adding noise at each diffusion step is input into the Unet model. At the same time, the time step encoding and the target semantic encoding are input into the Unet model; the Unet model predicts the noise under the additional conditions provided by the sub-cloud target radar feature encoder and the target semantic encoding, so that the latent space representation of the remote sensing image of the target cloud water area after adding noise is denoised according to the predicted noise; finally, the pre-trained VQVAE encoder will generate a cloud-free remote sensing image of the target water area based on the final denoising result.

[0074] To achieve the above effects, it is necessary to construct and train the cloud processing diffusion model.

[0075] Such as Figure 3As shown in the figure, the cloud processing diffusion model includes: a VQVAE encoder, a VQVAE decoder, a Unet model, a radar feature encoder for targets under clouds, a shared time encoder, and a shared text encoder for targets under clouds; among them, the pre-trained VQVAE encoder converts the input remote sensing image of the target water area and the associated synthetic aperture radar image of the water area into the latent space; the cloud processing diffusion model iteratively adds Gaussian noise to the latent space representation of the remote sensing image of the target water area by using the diffusion process; the diffusion result of the latent space representation of the remote sensing image of the target water area is input into the Unet model for iterative denoising processing; the radar feature encoder for targets under clouds and the Unet model are set in parallel, and are used to provide the intermediate module and the decoding module of the Unet model with the radar features of targets under clouds extracted from the synthetic aperture radar image of the water area during each iterative denoising process, and the radar features of targets under clouds guide the Unet model to generate a domain representation of the remote sensing image of the water area based on the targets under clouds in the corresponding area under clouds; the shared text encoder for targets under clouds encodes the target semantic information extracted from the cloud-covered area in the synthetic aperture radar image domain by using the synthetic aperture radar image recognition model, and provides the intermediate module and the decoding module of the Unet model with the target semantic encoding during each iterative denoising process, and the target semantic encoding guides the Unet model to generate a domain representation of the remote sensing image of the water area based on the targets under clouds in the corresponding area under clouds. The shared time encoder provides the iterative time step encoding for the radar feature encoder for targets under clouds and the Unet model during each iterative denoising process; the pre-trained VQVAE decoder decodes the processed remote sensing image of the target water area based on the denoising result generated by the Unet model.

[0076] As Figure 4 shown in the figure, the Unet model includes: a number of cascaded encoding modules, an intermediate module, and decoding modules corresponding to the number of encoding modules, where skip connections are set between the corresponding encoding modules and decoding modules; for any level of decoding module, the output of the decoding module or the intermediate module of the previous level and the output of the encoder of the same level are combined and input into this decoding module.

[0077] The radar feature encoder for targets under clouds includes: a number of cascaded encoding modules and an intermediate module copied from the Unet model, each encoding module and intermediate module are connected to a 1×1 convolutional layer, the 1×1 convolutional layer connected to the intermediate module is connected to the intermediate module of the Unet model, and the 1×1 convolutional layer connected to the encoding module is connected to the corresponding decoding module in the Unet model. A 1×1 convolutional layer is set at the input of the radar feature encoder for targets under clouds, and the output of the 1×1 convolution is combined with the latent space representation input into the Unet model and then input into the radar feature encoder for targets under clouds.

[0078] The shared time encoder provides the iterative time step encoding of the iterative time steps to each encoding module, intermediate module, and decoding module of the Unet model, and the shared time encoder provides the iterative time step encoding of the iterative time steps to the encoding module and intermediate module of the radar feature encoder of the sub-cloud target object.

[0079] The shared text encoder of the sub-cloud target object provides the semantic encoding of the sub-cloud target object to each encoding module, intermediate module, and decoding module of the Unet model, and the shared text encoder of the sub-cloud target object provides the semantic encoding of the sub-cloud target object to the encoding module and intermediate module of the radar feature encoder of the sub-cloud target object.

[0080] As Figure 5 shown, the training process of the cloud layer processing diffusion model includes:

[0081] Construct a training set for training the cloud layer processing diffusion model, where the training set includes: associated and aligned water area remote sensing images, cloud masks of water area remote sensing images, water area synthetic aperture radar images, cloud-free water area remote sensing images, where the associated cloud-free water area remote sensing images have the same scene as the water area remote sensing images but are not blocked by clouds.

[0082] First, use the water area remote sensing images in the training set to train the parameters of the Unet model so that the Unet model can generally restore the diffused water area remote sensing images. As Figure 6 shown, it includes:

[0083] Input the water area remote sensing images in the training set into the cloud layer processing diffusion model, and the pre-trained VQVAE encoder converts the water area remote sensing images into the latent space;

[0084] Then, use the diffusion process to iteratively add Gaussian noise to the latent space representation of the water area remote sensing images;

[0085] Perform multiple arbitrary samplings from the diffusion process to obtain the iterative time steps, noise, and latent space representations before and after adding noise of all samplings;

[0086] The shared time encoder encodes the sampled iterative time steps;

[0087] The combined noisy latent space representation and time step encoding are input into the Unet model, expecting the noisy latent space representation to be restored to the latent space representation before denoising according to the noise predicted by the Unet. Finally, the pre-trained VQVAE encoder will restore the water area remote sensing image based on the final denoising result. During this training process, the loss function used includes the sum of the L2 distances between all sampled diffusion noises and the noises predicted by the Unet model and the cross-entropy loss between the restored water area remote sensing image and the input water area remote sensing image. The Unet model is trained with the goal of minimizing the loss function. During this training process, the Unet model actually learns the essential structures and distribution characteristics of all objects in the training set of water area remote sensing images through the two processes of diffusion and denoising. Through the training of the denoising process, the Unet model learns to gradually generate remote sensing images that conform to the water area scene distribution from pure noise.

[0088] After training the Unet model, train the cloud-covered object radar feature encoder and the shared cloud-covered object text encoder for the optimization conditions provided by the Unet. After using the optimization conditions to control the denoising of the noise predicted by the Unet model, a latent space representation capable of generating a cloud-free water area image can be generated. As Figure 7 shown, the process includes:

[0089] After training the Unet model, copy several cascaded encoding modules and intermediate modules of the Unet model and construct a cloud-covered object radar feature encoder.

[0090] For any synthetic aperture radar image of water area in the training set, use the corresponding cloud mask to extract the cloud-covered area; use the synthetic aperture radar image recognition model to extract the semantics of the objects in the cloud-covered area in the synthetic aperture radar image domain of the water area.

[0091] Input the water area remote sensing image, cloud-free water area remote sensing image, and synthetic aperture radar image of the water area into the pre-trained VQVAE encoder of the cloud processing diffusion model; the pre-trained VQVAE encoder converts the water area remote sensing image, cloud-free water area remote sensing image, and synthetic aperture radar image of the water area into the latent space.

[0092] Then, use the diffusion process to synchronously add consistent Gaussian noise to the latent space representations of the water area remote sensing image and the cloud-free water area remote sensing image.

[0093] The shared cloud-covered object text encoder encodes the semantics of the objects in the cloud-covered area in the synthetic aperture radar image domain of the water area to obtain the object semantic encoding and provides it to the Unet model. The object semantic encoding provides the cloud-covered object semantic information for the Unet model, strengthens the semantic representation, and ensures the generation of a representation that conforms to the cloud-covered object semantics in the remote sensing image pixel domain.

[0094] Perform multiple arbitrary samplings from the diffusion processes of water area remote sensing images and cloudless water area remote sensing images to obtain the iterative time steps, noise, and latent space representations before and after adding noise for the two samplings;

[0095] The shared time encoder encodes the iterative time steps of the samplings;

[0096] The latent space representation of the sampled synthetic aperture radar image of the water area is processed by a 1×1 convolutional layer and combined with the latent space representation of the cloud water area remote sensing image after adding noise and input into the radar feature encoder of the object under the cloud. At the same time, the encoded time steps of the sampling and the encoded object semantics are combined and input into the radar feature encoder of the object under the cloud. The radar feature encoder of the object under the cloud provides the spatial information of the object under the cloud for the Unet model;

[0097] The latent space representation of the sampled cloud water area remote sensing image after adding noise is input into the Unet model. At the same time, the encoded time steps of the sampling and the encoded object semantics are input into the Unet model;

[0098] The Unet model predicts the diffusion noise under the additional conditions provided by the radar feature encoder of the object under the cloud and the object semantics encoding, so that the latent space representation of the cloud water area remote sensing image after adding noise is denoised according to the predicted noise and has the minimum distance from the latent space representation of the cloudless water area remote sensing image before adding noise in the same sampling stage; The finally pre-trained VQVAE encoder will generate a cloudless water area remote sensing image based on the final denoising result.

[0099] The loss function in the training process includes: the cross-entropy loss between the cloudless water area remote sensing image in the training set and the generated cloudless water area remote sensing image, and the L2 loss accumulation between the latent space representation diffusion sampling of the cloudless water area remote sensing image and the denoising result of the Unet model based on the diffusion sampling of the water area remote sensing image.

[0100] After generating the cloudless water area remote sensing image, the quality of the cloudless water area remote sensing image is further optimized through super-resolution processing.

[0101] Embodiment 2

[0102] Refer to Figure 8As shown in the figure, an embodiment of the present invention provides a device for processing cloud occlusion in water area remote sensing images, including: at least one processing unit, which is connected to a storage unit and a collection unit through a bus unit. The storage unit, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the software programs, computer-executable programs, and modules corresponding to a method for processing cloud occlusion in water area remote sensing images in an embodiment of the present invention. The collection unit includes a synthetic aperture radar and a remote sensing imaging device. The processing unit realizes the above-mentioned method for processing cloud occlusion in water area remote sensing images by running the software programs, computer-executable programs, and modules stored in the storage unit, including:

[0103] Synchronously collect a water area remote sensing image and a water area synthetic aperture radar image of a target water area by using the remote sensing imaging device and the synthetic aperture radar, and associate the two;

[0104] Use a cloud segmentation model that supports cloud recognition to identify the target water area remote sensing image that needs to be processed for cloud occlusion;

[0105] Perform alignment processing on the target water area remote sensing image and the associated water area synthetic aperture radar image;

[0106] After alignment, use a cloud mask to obtain the cloud-covered area from the water area synthetic aperture radar image, and use a synthetic aperture radar image recognition model to extract the semantics of the target objects in the cloud-covered area of the water area synthetic aperture radar image domain;

[0107] Input the aligned target water area remote sensing image, the associated water area synthetic aperture radar image, and the target object semantics into a pre-trained cloud processing diffusion model. The cloud processing diffusion model processes the cloud-covered area of the target water area remote sensing image according to the guidance of the water area synthetic aperture radar image and the target object semantics to generate a cloud-free target water area remote sensing image.

[0108] Of course, the computer program stored in the storage unit of the device for processing cloud occlusion in water area remote sensing images provided by an embodiment of the present invention is not limited to the method operations described above, and can also execute the related operations in a method for processing cloud occlusion in water area remote sensing images provided by any embodiment of the present invention.

[0109] Embodiment 3

[0110] An embodiment of the present invention provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed, it realizes the method for processing cloud occlusion in water area remote sensing images, including:

[0111] Synchronously collect the water area remote sensing image and the water area synthetic aperture radar image of the target water area by using a remote sensing imaging device and a synthetic aperture radar, and associate the two;

[0112] Use a cloud segmentation model that supports cloud recognition to identify the target water area remote sensing image that needs to be processed for cloud occlusion;

[0113] Perform alignment processing on the target water area remote sensing image and the associated water area synthetic aperture radar image;

[0114] After alignment, use a cloud mask to obtain the cloud-covered area from the water area synthetic aperture radar image, and use a synthetic aperture radar image recognition model to extract the semantic information of the targets in the cloud-covered area of the water area synthetic aperture radar image domain;

[0115] Input the aligned target water area remote sensing image, the associated water area synthetic aperture radar image, and the target semantic information into a pre-trained cloud processing diffusion model. The cloud processing diffusion model processes the cloud-covered area of the target water area remote sensing image according to the guidance of the water area synthetic aperture radar image and the target semantic information to generate a cloud-free target water area remote sensing image.

[0116] A computer-readable storage medium provided by an embodiment of the present invention stores a computer program that is not limited to the method operations described above, and can also execute related operations in a method for processing cloud occlusion in a water area remote sensing image provided by any embodiment of the present invention.

[0117] In the embodiments provided by the present invention, it should be understood that the disclosed structures and methods can be implemented in other ways. For example, the structural embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces, and the indirect couplings or communication connections of structures or units can be in electrical, mechanical or other forms.

[0118] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0119] In addition, in each embodiment of the present invention, each functional unit may be integrated into one processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units.

[0120] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for processing cloud occlusion in water remote sensing images, characterized in that: include: Using remote sensing imaging equipment and synthetic aperture radar to synchronously collect water remote sensing images and water synthetic aperture radar images of the target water area, and associate the two; Using a cloud segmentation model that supports cloud recognition to identify target water area remote sensing images that need to be processed for cloud occlusion; Aligning the remote sensing image of the target water area with the synthetic aperture radar image of the water area associated with it; After alignment, the cloud mask is used to obtain the cloud coverage area from the water synthetic aperture radar image, and the synthetic aperture radar image recognition model is used to extract the semantics of the target object in the cloud coverage area in the water synthetic aperture radar image domain; The aligned target water area remote sensing image, the associated water area synthetic aperture radar image, and the target object semantics are input into a pre-trained cloud layer processing diffusion model, and the cloud layer processing diffusion model processes the cloud coverage area of ​​the target water area remote sensing image according to the guidance of the water area synthetic aperture radar image and the target object semantics to generate a cloud-free target water area remote sensing image; wherein the cloud layer processing diffusion model includes: a VQVAE encoder, a VQVAE decoder, a Unet model, a radar feature encoder for targets under clouds, a shared time encoder, and a shared text encoder for targets under clouds; The pre-trained VQVAE encoder converts the input target water area remote sensing image and the water area synthetic aperture radar image associated therewith into a latent space; The diffusion process is used to iteratively add Gaussian noise to the latent space representation of the target water area remote sensing image; The diffusion results of the latent space representation of the target water area remote sensing image are input into the Unet model for iterative denoising; The radar feature encoder of the target under the cloud and the Unet model are set in parallel to provide the intermediate module and the decoding module of the Unet model with the radar features of the target under the cloud extracted from the water synthetic aperture radar image in each iterative denoising process. The radar features of the target under the cloud guide the Unet model to generate a water remote sensing image domain representation based on the target under the cloud in the corresponding sub-cloud area; The shared sub-cloud object text encoder encodes the semantics of the object extracted from the cloud-covered area in the water area synthetic aperture radar image domain using the synthetic aperture radar image recognition model, and provides the object semantic encoding to the intermediate module and decoding module of the Unet model in each iterative denoising process. The object semantic encoding guides the Unet model to generate a water area remote sensing image domain representation based on the sub-cloud object in the corresponding sub-cloud area. The shared time encoder provides iterative time step encoding for the radar feature encoder and Unet model of sub-cloud targets in each iterative denoising process; The pre-trained VQVAE decoder decodes the processed target water area remote sensing image based on the denoising result generated by the Unet model.

2. The method for processing cloud occlusion in water area remote sensing images according to claim 1, characterized in that: The process of aligning the target water area remote sensing image and the water area synthetic aperture radar image associated therewith comprises: segmenting the non-cloud area in the target water area remote sensing image by the cloud segmentation model, and extracting the remote sensing image domain alignment reference target from the non-cloud area in the target water area remote sensing image; Extracting a synthetic aperture radar image domain alignment reference target from a water area synthetic aperture radar image associated with a target water area remote sensing image; Matching the remote sensing image domain alignment reference target within the range of the synthetic aperture radar image domain alignment reference target; The water area remote sensing image and water area synthetic aperture radar image are aligned according to the rotation and translation relationship between the matching alignment reference targets in the two domains.

3. The method for processing cloud occlusion in water area remote sensing images according to claim 1, characterized in that: The Unet model includes: a plurality of cascaded encoding modules, intermediate modules and decoding modules corresponding to the number of encoding modules, wherein a jump chain is set between the corresponding encoding modules and decoding modules; for a decoding module at any level, the output of the decoding module at the previous level and the output of the encoder at the same level are combined and input into the decoding module.

4. The method for processing cloud occlusion in water area remote sensing images according to claim 3, characterized in that: The radar feature encoder for targets under clouds includes: a number of encoding modules and intermediate modules copied in the cascade of the Unet model, each encoding module and intermediate module is connected to a 1×1 convolutional layer, the 1×1 convolutional layer connected to the intermediate module is connected to the intermediate module of the Unet model, and the 1×1 convolutional layer connected to the encoding module is connected to the corresponding decoding module in the Unet model.

5. The method for processing cloud occlusion in water area remote sensing images according to claim 1, characterized in that: The training process of the cloud layer processing diffusion model includes: The parameters of the Unet model are trained so that the Unet model can generalize and restore the diffused water remote sensing image. The training method includes: inputting the water remote sensing image in the training set into the cloud processing diffusion model, and the pre-trained VQVAE encoder converts the water remote sensing image into a latent space; then using the diffusion process to iteratively add Gaussian noise to the latent space representation of the water remote sensing image; performing multiple arbitrary samplings from the diffusion process to obtain the iterative time step, noise and latent space representation before and after noise addition of all samples; the shared time encoder encodes the iterative time step of the sampling, and the latent space representation after noise addition and the time step encoding are combined and input into the Unet model, and it is expected that the latent space representation after noise addition will restore the latent space representation before noise addition after denoising according to the Unet prediction noise; finally, the pre-trained VQVAE encoder will restore the water remote sensing image based on the final denoising result.

6. The method for processing cloud occlusion in water area remote sensing images according to claim 5, characterized in that: The training process of the cloud layer processing diffusion model includes: after training the Unet model, copying several encoding modules and intermediate modules of the cascade of the Unet model and constructing a radar feature encoder for targets under clouds; training the radar feature encoder for targets under clouds and the shared text encoder for targets under clouds to provide optimization conditions for Unet, and using the optimization conditions to control the noise denoising predicted by the Unet model, so as to generate a latent space representation of a cloud-free water image.

7. The method for processing cloud occlusion in water area remote sensing images according to claim 6, characterized in that: Methods for training the radar feature encoder for sub-cloud objects and the shared text encoder for sub-cloud objects include: For any water synthetic aperture radar image in the training set, the cloud coverage area is extracted using the corresponding cloud mask; the semantics of the target object in the cloud coverage area in the water synthetic aperture radar image domain is extracted using the synthetic aperture radar image recognition model; Inputting the water area remote sensing image and the water area synthetic aperture radar image into the pre-trained VQVAE encoder of the cloud layer processing diffusion model; the pre-trained VQVAE encoder converts the water area remote sensing image and the water area synthetic aperture radar image into a latent space; Then, the diffusion process is used to synchronously add consistent Gaussian noise to the latent space representations of the water remote sensing image and the cloud-free water remote sensing image; The shared under-cloud object text encoder encodes the semantics of the objects in the cloud-covered area in the water synthetic aperture radar image domain to obtain the object semantic encoding and provide it to the Unet model. The object semantic encoding provides the Unet model with the semantic information of the under-cloud objects and strengthens the semantic representation. Multiple random samplings are performed from the diffusion process of water remote sensing images and cloud-free water remote sensing images to obtain the iterative time step, noise, and latent space representations before and after noise addition of the two samplings. The shared temporal encoder encodes the sampled iterative time steps; The latent space representation of the sampled water synthetic aperture radar image is processed by a 1×1 convolution layer and combined with the latent space representation of the cloud water remote sensing image after noise addition, and then input into the radar feature encoder of the target under the cloud. At the same time, the sampled time step encoding and the semantic encoding of the target are combined and input into the radar feature encoder of the target under the cloud. The radar feature encoder of the target under the cloud provides the spatial information of the target under the cloud for the Unet model. The latent space representation of the sampled cloud-water area remote sensing image after noise addition is input into the Unet model. At the same time, the sampled time step encoding and target object semantic encoding are input into the Unet model. The Unet model predicts diffuse noise under the additional conditions provided by the radar feature encoder of sub-cloud targets and the semantic encoding of targets, so that the latent space representation of the cloudy water remote sensing image after noise addition is denoised according to the predicted noise and the distance between the latent space representation of the cloudless water remote sensing image before noise addition in the same sampling stage is minimized; finally, the pre-trained VQVAE encoder will generate a cloudless water remote sensing image based on the final denoising result.

8. The method for processing cloud occlusion in water area remote sensing images according to claim 5, characterized in that: The training set includes: associated water remote sensing images, cloud masks of water remote sensing images, water synthetic aperture radar images, and cloudless water remote sensing images, wherein the associated cloudless water remote sensing images have the same scene as the water remote sensing images but are not blocked by clouds.

9. The method for processing cloud occlusion in water area remote sensing images according to claim 1, characterized in that: In the process of generating the cloud processing diffusion model, a remote sensing image of the target water area and a synthetic aperture radar image of the target water area are provided; Inputting the target water area remote sensing image and the target water area synthetic aperture radar image into the pre-trained VQVAE encoder of the cloud processing diffusion model; the pre-trained VQVAE encoder converts the target water area remote sensing image and the target water area synthetic aperture radar image into a latent space; Then, the diffusion process is used to add Gaussian noise to the latent space representation of the remote sensing image of the target water area; The shared under-cloud object text encoder encodes the semantics of the object in the cloud-covered area of ​​the target water area in the synthetic aperture radar image domain to obtain the semantic encoding of the object and provide it to the Unet model; The diffusion results of the remote sensing image of the target water area are denoised to obtain the iterative time step of the diffusion, the noise, and the latent space representation before and after the noise addition; The shared temporal encoder encodes the iterative time steps corresponding to the denoising process; The latent space representation of the synthetic aperture radar image of the target water area in the denoising process is processed by a 1×1 convolutional layer and combined with the latent space representation of the target cloud water area remote sensing image after noise addition, and then input into the radar feature encoder of the target object under the cloud. At the same time, the time step encoding and the semantic encoding of the target object are combined and then input into the radar feature encoder of the target object under the cloud. The latent space representation of the target cloud water area remote sensing image after each diffusion step of noise is input into the Unet model. At the same time, the time step encoding and the semantic encoding of the target object are input into the Unet model. The Unet model predicts the noise under the additional conditions provided by the radar feature encoder of the target under the cloud and the semantic encoding of the target, so that the latent space representation of the noisy target cloud water area remote sensing image is denoised according to the predicted noise; finally, the pre-trained VQVAE encoder will generate a cloud-free target water area remote sensing image based on the final denoising result.

Citation Information

Patent Citations

  • High-resolution remote sensing image and laser radar point cloud fused individual tree segmentation method and system

    CN111462134A

  • Deep learning cloud removal method based on SAR-optical remote sensing image combination

    CN115809970A