Cloud layer shielding processing method in water area remote sensing image
By synchronously collecting water remote sensing images and synthesizing aperture radar images, using diffusion model and Unet model to process cloud cover areas, the problem of cloud occlusion in water remote sensing images is solved, and high-quality cloudless water remote sensing images are generated.
Patent Information
- Application Number
- CN202510486962.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-18
AI Technical Summary
Cloud occlusion in water remote sensing images causes spatial structure, texture and context information to be occluded, affecting the pixel distribution of image scenes. The prior art is not effective when covered by thick clouds.
By synchronously collecting water remote sensing images and synthetic aperture radar images, the cloud segmentation model and synthetic aperture radar image recognition model extract the target object semantics of the cloud cover area and input them into the pre-trained cloud processing diffusion model. The diffusion process and the Unet model are used to process the cloud cover area to generate cloud-free water remote sensing images.
Effectively remove the cloud cover area and generate higher quality remote sensing images in cloudless waters, solving the problem of missing information caused by thick cloud coverage, and the generated images are in line with the current real situation.
Smart Images

Figure CN120014483A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of cloud layer processing in remote sensing images, and in particular to a method for processing cloud layer occlusion in water area remote sensing images. Background Art
[0002] Water remote sensing images taken by drones are an important means of water inspection. However, many water remote sensing images contain clouds, which obscure the spatial structure, texture, and context of water remote sensing images, cover up the effective information in water remote sensing images, and significantly affect the pixel distribution of water remote sensing image scenes.
[0003] The current solution for dealing with cloud occlusion is to migrate the trained cloud-free optical image model to the cloud-covered optical image that has been processed by the cloud removal method. For example, the cloud-covered area is reconstructed by propagating the texture structure of the cloud-free background, and the result is further optimized by enriching the feature space and search range of the background information. This method is easily affected by edge effects, and the generation effect is very poor when there is thick cloud cover. In order to address the serious information loss caused by thick cloud cover, the existing technology constructs a unified space-time-spectral model to remove thick clouds. This model can make full use of information from multi-source and multi-time data to accurately reconstruct cloud-damaged areas, but it requires the use of historical unobstructed images, and the reconstruction is not based on real-time scenes. Summary of the invention
[0004] In order to solve the above technical problem or at least partially solve the above technical problem, the present invention provides a method for processing cloud occlusion in water area remote sensing images.
[0005] The present invention provides a method for processing cloud occlusion in a water area remote sensing image, comprising: Using remote sensing imaging equipment and synthetic aperture radar to synchronously collect water remote sensing images and water synthetic aperture radar images of the target water area, and associate the two; Using a cloud segmentation model that supports cloud recognition to identify target water area remote sensing images that need to be processed for cloud occlusion; Aligning the remote sensing image of the target water area with the synthetic aperture radar image of the water area associated with it; After alignment, the cloud mask is used to obtain the cloud coverage area from the water synthetic aperture radar image, and the synthetic aperture radar image recognition model is used to extract the semantics of the target object in the cloud coverage area in the water synthetic aperture radar image domain; The aligned target water area remote sensing image, the associated water area synthetic aperture radar image, and the target object semantics are input into a pre-trained cloud processing diffusion model. The cloud processing diffusion model processes the cloud coverage area of the target water area remote sensing image according to the guidance of the water area synthetic aperture radar image and the target object semantics to generate a cloud-free target water area remote sensing image.
[0006] Furthermore, the process of aligning the target water area remote sensing image and the water area synthetic aperture radar image associated therewith includes: segmenting the non-cloud area in the target water area remote sensing image by the cloud segmentation model, and extracting the remote sensing image domain alignment reference target from the non-cloud area in the target water area remote sensing image; Extracting a synthetic aperture radar image domain alignment reference target from a water area synthetic aperture radar image associated with a target water area remote sensing image; Matching the remote sensing image domain alignment reference target within the range of the synthetic aperture radar image domain alignment reference target; The water area remote sensing image and water area synthetic aperture radar image are aligned according to the rotation and translation relationship between the matching alignment reference targets in the two domains.
[0007] Furthermore, the cloud layer processing diffusion model includes: a VQVAE encoder, a VQVAE decoder, a Unet model, a radar feature encoder for sub-cloud targets, a shared time encoder, and a shared text encoder for sub-cloud targets; The pre-trained VQVAE encoder converts the input target water area remote sensing image and the water area synthetic aperture radar image associated therewith into a latent space; The diffusion process is used to iteratively add Gaussian noise to the latent space representation of the target water area remote sensing image; The diffusion results of the latent space representation of the target water area remote sensing image are input into the Unet model for iterative denoising; The radar feature encoder of the target under the cloud and the Unet model are set in parallel to provide the intermediate module and the decoding module of the Unet model with the radar features of the target under the cloud extracted from the water synthetic aperture radar image in each iterative denoising process. The radar features of the target under the cloud guide the Unet model to generate a water remote sensing image domain representation based on the target under the cloud in the corresponding sub-cloud area; The shared sub-cloud object text encoder encodes the semantics of the object extracted from the cloud-covered area in the water area synthetic aperture radar image domain using the synthetic aperture radar image recognition model, and provides the object semantic encoding to the intermediate module and decoding module of the Unet model in each iterative denoising process. The object semantic encoding guides the Unet model to generate a water area remote sensing image domain representation based on the sub-cloud object in the corresponding sub-cloud area. The shared time encoder provides iterative time step encoding for the radar feature encoder and Unet model of sub-cloud targets in each iterative denoising process; The pre-trained VQVAE decoder decodes the processed target water area remote sensing image based on the denoising result generated by the Unet model.
[0008] Furthermore, the Unet model includes: a plurality of cascaded encoding modules, intermediate modules and decoding modules corresponding to the number of encoding modules, wherein a jump chain is set between the corresponding encoding modules and decoding modules; for a decoding module at any level, the output of the decoding module at the previous level and the output of the encoder at the same level are combined and input into the decoding module.
[0009] Furthermore, the radar feature encoder for targets under clouds includes: a plurality of encoding modules and intermediate modules copied in the cascade of the Unet model, each encoding module and the intermediate module is connected to a 1×1 convolutional layer, the 1×1 convolutional layer connected to the intermediate module is connected to the intermediate module of the Unet model, and the 1×1 convolutional layer connected to the encoding module is connected to the corresponding decoding module in the Unet model.
[0010] Furthermore, the training process of the cloud layer processing diffusion model includes: The parameters of the Unet model are trained so that the Unet model can generalize and restore the diffused water remote sensing image. The training method includes: inputting the water remote sensing image in the training set into the cloud processing diffusion model, and the pre-trained VQVAE encoder converts the water remote sensing image into a latent space; then using the diffusion process to iteratively add Gaussian noise to the latent space representation of the water remote sensing image; performing multiple arbitrary samplings from the diffusion process to obtain the iterative time step, noise and latent space representation before and after noise addition of all samples; the shared time encoder encodes the iterative time step of the sampling, and the latent space representation after noise addition and the time step encoding are combined and input into the Unet model, and it is expected that the latent space representation after noise addition will restore the latent space representation before noise addition after denoising according to the Unet prediction noise; finally, the pre-trained VQVAE encoder will restore the water remote sensing image based on the final denoising result.
[0011] Furthermore, the training process of the cloud layer processing diffusion model includes: after training the Unet model, copying several encoding modules and intermediate modules of the cascade of the Unet model and constructing a radar feature encoder for targets under clouds; training the radar feature encoder for targets under clouds and the shared text encoder for targets under clouds to provide optimization conditions for Unet, and using the optimization conditions to control the noise denoising predicted by the Unet model, so as to generate a latent space representation of a cloud-free water image.
[0012] Furthermore, the method of training the radar feature encoder of the sub-cloud object and the shared text encoder of the sub-cloud object includes: For any water synthetic aperture radar image in the training set, the cloud coverage area is extracted using the corresponding cloud mask; the semantics of the target object in the cloud coverage area in the water synthetic aperture radar image domain is extracted using the synthetic aperture radar image recognition model; Inputting the water area remote sensing image and the water area synthetic aperture radar image into the pre-trained VQVAE encoder of the cloud layer processing diffusion model; the pre-trained VQVAE encoder converts the water area remote sensing image and the water area synthetic aperture radar image into a latent space; Then, the diffusion process is used to synchronously add consistent Gaussian noise to the latent space representations of the water remote sensing image and the cloud-free water remote sensing image; The shared under-cloud object text encoder encodes the semantics of the objects in the cloud-covered area in the water synthetic aperture radar image domain to obtain the object semantic encoding and provide it to the Unet model. The object semantic encoding provides the Unet model with the semantic information of the under-cloud objects and strengthens the semantic representation. Multiple random samplings are performed from the diffusion process of water remote sensing images and cloud-free water remote sensing images to obtain the iterative time step, noise, and latent space representations before and after noise addition of the two samplings. The shared temporal encoder encodes the sampled iterative time steps; The latent space representation of the sampled water synthetic aperture radar image is processed by a 1×1 convolution layer and combined with the latent space representation of the cloud water remote sensing image after noise addition, and then input into the radar feature encoder of the target under the cloud. At the same time, the sampled time step encoding and the semantic encoding of the target are combined and input into the radar feature encoder of the target under the cloud. The radar feature encoder of the target under the cloud provides the spatial information of the target under the cloud for the Unet model. The latent space representation of the sampled cloud-water area remote sensing image after noise addition is input into the Unet model. At the same time, the sampled time step encoding and target object semantic encoding are input into the Unet model. The Unet model predicts diffuse noise under the additional conditions provided by the radar feature encoder of sub-cloud targets and the semantic encoding of targets, so that the latent space representation of the cloudy water remote sensing image after noise addition is denoised according to the predicted noise and the distance between the latent space representation of the cloudless water remote sensing image before noise addition in the same sampling stage is minimized; finally, the pre-trained VQVAE encoder will generate a cloudless water remote sensing image based on the final denoising result.
[0013] Furthermore, the training set includes: associated water remote sensing images, cloud masks of water remote sensing images, water synthetic aperture radar images, and cloudless water remote sensing images, wherein the associated cloudless water remote sensing images have the same scene as the water remote sensing images but are not obstructed by clouds.
[0014] Furthermore, in the process of generating the cloud layer processing diffusion model, a remote sensing image of the target water area and a synthetic aperture radar image of the target water area are provided; Inputting the target water area remote sensing image and the target water area synthetic aperture radar image into the pre-trained VQVAE encoder of the cloud layer processing diffusion model; the pre-trained VQVAE encoder converts the target water area remote sensing image and the target water area synthetic aperture radar image into a latent space; Then, the diffusion process is used to add Gaussian noise to the latent space representation of the remote sensing image of the target water area; The shared under-cloud object text encoder encodes the semantics of the object in the cloud-covered area of the target water area in the synthetic aperture radar image domain to obtain the semantic encoding of the object and provide it to the Unet model; The diffusion results of the remote sensing image of the target water area are denoised to obtain the iterative time step of the diffusion, the noise, and the latent space representation before and after the noise addition; The shared temporal encoder encodes the iterative time steps corresponding to the denoising process; The latent space representation of the synthetic aperture radar image of the target water area in the denoising process is processed by a 1×1 convolutional layer and combined with the latent space representation of the target cloud water area remote sensing image after noise addition, and then input into the radar feature encoder of the target object under the cloud. At the same time, the time step encoding and the semantic encoding of the target object are combined and then input into the radar feature encoder of the target object under the cloud. The latent space representation of the target cloud water area remote sensing image after each diffusion step of noise is input into the Unet model. At the same time, the time step encoding and the semantic encoding of the target object are input into the Unet model. The Unet model predicts the noise under the additional conditions provided by the radar feature encoder of the target under the cloud and the semantic encoding of the target, so that the latent space representation of the noisy target cloud water area remote sensing image is denoised according to the predicted noise; finally, the pre-trained VQVAE encoder will generate a cloud-free target water area remote sensing image based on the final denoising result.
[0015] The above technical solution provided by the embodiment of the present invention has the following advantages compared with the prior art: The present invention synchronously collects water remote sensing images and water synthetic aperture radar images of the target water area; screens the target water remote sensing images that need to be processed for cloud occlusion; aligns the target water remote sensing images and the water synthetic aperture radar images associated therewith; after alignment, uses the synthetic aperture radar image recognition model to extract the semantics of the target object in the cloud-covered area in the water synthetic aperture radar image domain; inputs the aligned target water remote sensing images, the water synthetic aperture radar images associated therewith, and the semantics of the target object into the pre-trained cloud layer processing diffusion model, and the cloud layer processing diffusion model processes the cloud-covered area of the target water remote sensing image according to the guidance of the water synthetic aperture radar image and the semantics of the target object to generate a cloud-free target water remote sensing image. The present application generates an image of the cloud-covered area in the water remote sensing image domain guided by the synchronized water synthetic aperture radar image. The present application uses the ability of the cloud layer processing diffusion model to learn object distribution information to generatively generate the cloud-covered area in the water remote sensing image. Since the synchronously collected water synthetic aperture radar images are used as a guide, the semantic spatial information of the sub-cloud targets can be well transferred from the water synthetic aperture radar images to the water remote sensing image domain, and the image of the cloud-covered area formed conforms to the current real situation. The semantics of the targets in the cloud-covered area are enhanced by the semantics of the targets, and the spatial information of the targets in the water synthetic aperture radar images is supplemented to ensure the generation of correct targets. By using the synergistic and complementary effects of optics and synthetic aperture radars, the cloud processing diffusion model can naturally connect the sub-cloud area targets and non-sub-cloud areas, and the quality of the generated cloud-free water images is higher. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0018] Figure 1 A flow chart of a method for processing cloud occlusion in a water area remote sensing image provided by an embodiment of the present invention; Figure 2 An architecture diagram of the cloud processing diffusion model, cloud segmentation model and synthetic aperture radar image recognition model provided in an embodiment of the present invention cooperating to implement cloud occlusion processing in water remote sensing images; Figure 3 A schematic diagram of each model in the cloud layer processing diffusion model provided by an embodiment of the present invention; Figure 4A schematic diagram of a Unet model and a radar feature encoder for targets under clouds provided in an embodiment of the present invention; Figure 5 A training flow chart of a cloud layer processing diffusion model provided by an embodiment of the present invention; Figure 6 A flowchart of training a Unet model provided by an embodiment of the present invention; Figure 7 A flowchart of the optimization conditions provided by the embodiment of the present invention for training the radar feature encoder of the target under the cloud and the shared text encoder of the target under the cloud for Unet, and generating a latent space representation of the cloud-free water image after the noise denoising predicted by the Unet model is controlled by using the optimization conditions; Figure 8 A schematic diagram of a cloud occlusion processing device in a water area remote sensing image provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0021] Example 1 See also Figure 1 As shown, the cloud occlusion processing method in the water area remote sensing image provided by this application includes: S100, using remote sensing imaging equipment and synthetic aperture radar to synchronously collect water area remote sensing images and water area synthetic aperture radar images of the target water area. Due to the longer wavelength, synthetic aperture radar waves have the ability to penetrate clouds compared to visible light waves, and can penetrate the cover of clouds and perceive the structural information of targets on the water area under the clouds; therefore, in the cloud removal process in the remote sensing image domain, the water area synthetic aperture radar image can provide guiding conditions for the generation of pixels in the cloud-covered area. Through subsequent processes, the present application effectively converts the water area synthetic aperture radar image into guiding conditions that can guide the generation of target object pixels in the cloud-covered area.
[0022] S200, using a pre-trained cloud segmentation model that supports cloud recognition to identify and select the target water remote sensing images that need to be processed for cloud occlusion; in the prior art, there are many semantic segmentation models that support cloud recognition that can be used as cloud segmentation models for this application, such as: semantic segmentation models of the Unet architecture, the semantic segmentation model here is not consistent with the subsequent Unet model, and the architectures of both models are Unet. Use the cloud segmentation model to traverse and process each water remote sensing image to generate a cloud mask. For any water remote sensing image, when the cloud layer in the generated cloud mask is not empty, the water remote sensing image is selected as the target water remote sensing image.
[0023] Although both the remote sensing imaging device and the synthetic aperture radar collect information of the target water area, due to the different installation positions of the two on the flight vehicle, the collected water area remote sensing image and the water area synthetic aperture radar image may not be aligned. Therefore, after the target water area remote sensing image is screened out, S300 first aligns the target water area remote sensing image and the water area synthetic aperture radar image associated therewith, such as Figure 2 As shown, the process includes: Segmenting the non-cloud area in the remote sensing image of the target water area by using the cloud segmentation model, and extracting the remote sensing image domain alignment reference target from the non-cloud area in the remote sensing image of the target water area; Extracting a synthetic aperture radar image domain alignment reference target from a water area synthetic aperture radar image associated with a target water area remote sensing image; Matching the remote sensing image domain alignment reference target within the range of the synthetic aperture radar image domain alignment reference target; The water area remote sensing image and the water area synthetic aperture radar image are aligned according to the rotation and translation relationship between the matching alignment reference targets in the two domains. In the specific implementation process, the findFundamentalMat function in opencv is used to determine the rotation and translation relationship according to the positions of the matching alignment reference targets in the two domains, and the warpPerspective function in opencv is used to align the target water area remote sensing image and the water area synthetic aperture radar image associated with it according to the rotation and translation relationship.
[0024] S400, after alignment, using the cloud mask to obtain the cloud coverage area from the water synthetic aperture radar image, and using the synthetic aperture radar image recognition model to extract the semantics of the target object in the cloud coverage area in the water synthetic aperture radar image domain. The synthetic aperture radar image recognition model uses the CLIP model trained for target object recognition in synthetic aperture radar images.
[0025] S500, input the aligned target water area remote sensing image, the associated water area synthetic aperture radar image, and the target object semantics into a pre-trained cloud layer processing diffusion model, and the cloud layer processing diffusion model processes the cloud coverage area of the target water area remote sensing image according to the guidance of the water area synthetic aperture radar image and the target object semantics to generate a cloud-free target water area remote sensing image.
[0026] In the process of generating the cloud processing diffusion model, the target water area remote sensing image and the target water area synthetic aperture radar image are input into the pre-trained VQVAE encoder of the cloud processing diffusion model; the pre-trained VQVAE encoder converts the target water area remote sensing image and the target water area synthetic aperture radar image into latent space; then Gaussian noise is added to the latent space representation of the target water area remote sensing image by using the diffusion process; the shared under-cloud target object text encoder encodes the semantics of the target object in the cloud coverage area in the synthetic aperture radar image domain of the target water area to obtain the target object semantic encoding and provides it to the Unet model; the diffusion result of the target water area remote sensing image is denoised to obtain the diffusion iteration time step, noise, and latent space representation before and after noise addition; the shared time encoder encodes the iteration time step corresponding to the denoising process; The latent space representation of the target water synthetic aperture radar image after the noise process is processed by a 1×1 convolution layer and combined with the latent space representation of the target cloud water remote sensing image after noise is input into the radar feature encoder of the target object under the cloud, and the time step encoding and the target semantic encoding are combined and input into the radar feature encoder of the target object under the cloud; the latent space representation of the target cloud water remote sensing image after noise is added in each diffusion step is input into the Unet model, and the time step encoding and the target semantic encoding are input into the Unet model; the Unet model predicts the noise under the additional conditions provided by the radar feature encoder of the target object under the cloud and the target semantic encoding, so that the latent space representation of the target cloud water remote sensing image after noise is denoised according to the predicted noise; finally, the pre-trained VQVAE encoder generates a cloud-free target water remote sensing image based on the final denoising result.
[0027] In order to achieve the above effects, it is necessary to construct and train the cloud layer processing diffusion model.
[0028] like Figure 3As shown, the cloud layer processing diffusion model includes: a VQVAE encoder, a VQVAE decoder, a Unet model, a radar feature encoder for targets under clouds, a shared time encoder, and a shared text encoder for targets under clouds; wherein the pre-trained VQVAE encoder converts the input target water area remote sensing image and the water area synthetic aperture radar image associated therewith into a latent space; the cloud layer processing diffusion model uses a diffusion process to iteratively add Gaussian noise to the latent space representation of the target water area remote sensing image; the diffusion result of the latent space representation of the target water area remote sensing image is input into the Unet model for iterative denoising processing; the radar feature encoder for targets under clouds and the Unet model are set in parallel for each In the iterative denoising process, radar features of sub-cloud targets extracted from water synthetic aperture radar images are provided to the intermediate module and decoding module of the Unet model, and the radar features of sub-cloud targets guide the Unet model to generate a water remote sensing image domain representation based on sub-cloud targets in the corresponding sub-cloud area; the shared sub-cloud target text encoder encodes the semantics of the target extracted from the cloud-covered area in the water synthetic aperture radar image domain using the synthetic aperture radar image recognition model, and provides the target semantic encoding to the intermediate module and decoding module of the Unet model in each iterative denoising process, and the target semantic encoding guides the Unet model to generate a water remote sensing image domain representation based on sub-cloud targets in the corresponding sub-cloud area. The shared time encoder provides iterative time step encoding for the sub-cloud target radar feature encoder and the Unet model in each iterative denoising process; the pre-trained VQVAE decoder decodes the processed target water remote sensing image based on the denoising result generated by the Unet model.
[0029] like Figure 4 As shown, the Unet model includes: a plurality of cascaded encoding modules, intermediate modules and decoding modules corresponding to the number of encoding modules, wherein a jump chain is set between the corresponding encoding modules and decoding modules; for a decoding module at any level, the output of the decoding module or the intermediate module at the previous level and the output of the encoder at the same level are combined and input into the decoding module.
[0030] The radar feature encoder for targets under clouds includes: a plurality of encoding modules and intermediate modules copied to the cascade of the Unet model, each encoding module and intermediate module is connected to a 1×1 convolution layer, the 1×1 convolution layer connected to the intermediate module is connected to the intermediate module of the Unet model, and the 1×1 convolution layer connected to the encoding module is connected to the corresponding decoding module in the Unet model. The input of the radar feature encoder for targets under clouds is set to a 1×1 convolution layer, and the output of the 1×1 convolution is combined with the latent space representation of the input Unet model and then input to the radar feature encoder for targets under clouds.
[0031] The shared time encoder provides the iterative time step encoding of the iterative time step to each encoding module, intermediate module and decoding module of the Unet model, and the shared time encoder provides the iterative time step encoding of the iterative time step to the encoding module and intermediate module of the radar feature encoder of the target under the cloud.
[0032] The shared sub-cloud object text encoder provides the sub-cloud object semantic encoding to each encoding module, intermediate module and decoding module of the Unet model, and the shared sub-cloud object text encoder provides the sub-cloud object semantic encoding to the encoding module and intermediate module of the sub-cloud object radar feature encoder.
[0033] like Figure 5 As shown, the training process of the cloud layer processing diffusion model includes: A training set for training a cloud processing diffusion model is constructed, wherein the training set includes: associated and aligned water remote sensing images, cloud masks of water remote sensing images, water synthetic aperture radar images, and cloudless water remote sensing images, wherein the associated cloudless water remote sensing images have the same scene as the water remote sensing images but are not obstructed by clouds.
[0034] First, the parameters of the Unet model are trained using the water remote sensing images in the training set so that the Unet model can generalize and restore the diffused water remote sensing images. Figure 6 As shown, including: The water area remote sensing image in the training set is input into the cloud processing diffusion model, and the pre-trained VQVAE encoder converts the water area remote sensing image into a latent space; Then, the diffusion process is used to iteratively add Gaussian noise to the latent space representation of the water remote sensing image; Perform multiple arbitrary samplings from the diffusion process to obtain the iterative time steps, noise, and latent space representations before and after noise addition for all samples; The shared temporal encoder encodes the sampled iterative time steps; The latent space representation after noise addition and the time step encoding are combined and input into the Unet model, and it is expected that the latent space representation after noise addition will restore the latent space representation before noise addition after denoising according to the noise predicted by Unet; finally, the pre-trained VQVAE encoder will restore the water remote sensing image based on the final denoising result. During the training process, the loss function used includes the sum of the L2 distance accumulation between all sampled diffusion noise and the noise predicted by the Unet model and the cross entropy loss of the restored water remote sensing image and the input water remote sensing image. The Unet model is trained with the minimum loss function as the goal. During this training process, the Unet model actually learns the essential structure and distribution characteristics of all targets in the water remote sensing images of the training set by using the two processes of diffusion and denoising. Through the training of the denoising process, the Unet model learns to gradually generate remote sensing images that conform to the distribution of water scenes from pure noise.
[0035] After training the Unet model, the radar feature encoder for sub-cloud objects and the shared sub-cloud object text encoder are trained to provide optimization conditions for Unet. After denoising the noise predicted by the Unet model, the latent space representation of the cloud-free water image can be generated. Figure 7 As shown, the process includes: After training the Unet model, several encoding modules and intermediate modules of the cascade of the Unet model are copied and the radar feature encoder of the target under the cloud is constructed.
[0036] For any water synthetic aperture radar image in the training set, the cloud coverage area is extracted using the corresponding cloud mask; the synthetic aperture radar image recognition model is used to extract the semantics of the target object in the cloud coverage area in the water synthetic aperture radar image domain.
[0037] The water area remote sensing image, the cloud-free water area remote sensing image and the water area synthetic aperture radar image are input into the pre-trained VQVAE encoder of the cloud layer processing diffusion model; the pre-trained VQVAE encoder converts the water area remote sensing image, the cloud-free water area remote sensing image and the water area synthetic aperture radar image into latent space.
[0038] Then, the diffusion process is used to synchronously add consistent Gaussian noise to the latent space representations of the water remote sensing image and the cloud-free water remote sensing image.
[0039] The shared sub-cloud object text encoder encodes the semantics of objects in the cloud-covered area in the water synthetic aperture radar image domain to obtain the target semantic code and provide it to the Unet model. The target semantic code provides the Unet model with semantic information of sub-cloud objects, strengthens the semantic representation, and ensures that the representation that conforms to the semantics of sub-cloud objects is generated in the pixel domain of the remote sensing image.
[0040] Multiple random samplings are performed from the diffusion process of water remote sensing images and cloud-free water remote sensing images to obtain the iterative time step, noise, and latent space representations before and after noise addition of the two samplings. The shared temporal encoder encodes the sampled iterative time steps; The latent space representation of the sampled water synthetic aperture radar image is processed by a 1×1 convolution layer and combined with the latent space representation of the cloud water remote sensing image after noise addition, and then input into the radar feature encoder of the target under the cloud. At the same time, the sampled time step encoding and the semantic encoding of the target are combined and input into the radar feature encoder of the target under the cloud. The radar feature encoder of the target under the cloud provides the spatial information of the target under the cloud for the Unet model. The latent space representation of the sampled cloud-water area remote sensing image after noise addition is input into the Unet model. At the same time, the sampled time step encoding and target object semantic encoding are input into the Unet model. The Unet model predicts diffuse noise under the additional conditions provided by the radar feature encoder of sub-cloud targets and the semantic encoding of targets, so that the latent space representation of the cloudy water remote sensing image after noise addition is denoised according to the predicted noise and the distance between the latent space representation of the cloudless water remote sensing image before noise addition in the same sampling stage is minimized; finally, the pre-trained VQVAE encoder will generate a cloudless water remote sensing image based on the final denoising result.
[0041] The loss function of the training process includes: the cross entropy loss between the cloud-free water remote sensing images in the training set and the generated cloud-free water remote sensing images, the L2 loss accumulation between the diffusion sampling of the latent space representation of the cloud-free water remote sensing images and the denoising results of the diffusion sampling of the water remote sensing images based on the Unet model.
[0042] After generating cloud-free water remote sensing images, super-resolution processing is used to further optimize the quality of cloud-free water remote sensing images.
[0043] Example 2 See also Figure 8 As shown, an embodiment of the present invention provides a device for processing cloud occlusion in a water remote sensing image, comprising: at least one processing unit, the processing unit is connected to a storage unit and an acquisition unit via a bus unit, the storage unit is a computer-readable storage medium, and can be used to store software programs, computer executable programs, and modules, such as the software programs, computer executable programs, and modules corresponding to a method for processing cloud occlusion in a water remote sensing image in an embodiment of the present invention. The acquisition unit includes a synthetic aperture radar and a remote sensing imaging device, and the processing unit implements the above-mentioned method for processing cloud occlusion in a water remote sensing image by running the software programs, computer executable programs, and modules stored in the storage unit, comprising: Using remote sensing imaging equipment and synthetic aperture radar to synchronously collect water remote sensing images and water synthetic aperture radar images of the target water area, and associate the two; Using a cloud segmentation model that supports cloud recognition to identify target water area remote sensing images that need to be processed for cloud occlusion; Aligning the remote sensing image of the target water area with the synthetic aperture radar image of the water area associated with it; After alignment, the cloud mask is used to obtain the cloud coverage area from the water synthetic aperture radar image, and the synthetic aperture radar image recognition model is used to extract the semantics of the target object in the cloud coverage area in the water synthetic aperture radar image domain; The aligned target water area remote sensing image, the associated water area synthetic aperture radar image, and the target object semantics are input into a pre-trained cloud processing diffusion model. The cloud processing diffusion model processes the cloud coverage area of the target water area remote sensing image according to the guidance of the water area synthetic aperture radar image and the target object semantics to generate a cloud-free target water area remote sensing image.
[0044] Of course, the computer program stored in the storage unit of the cloud occlusion processing device for water remote sensing images provided by an embodiment of the present invention is not limited to the operation of the method described above, and can also execute related operations in the cloud occlusion processing method for water remote sensing images provided by any embodiment of the present invention.
[0045] Example 3 An embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program. When the computer program is executed, the method for processing cloud occlusion in a water area remote sensing image is implemented, including: Using remote sensing imaging equipment and synthetic aperture radar to synchronously collect water remote sensing images and water synthetic aperture radar images of the target water area, and associate the two; Using a cloud segmentation model that supports cloud recognition to identify target water area remote sensing images that need to be processed for cloud occlusion; Aligning the remote sensing image of the target water area with the synthetic aperture radar image of the water area associated with it; After alignment, the cloud mask is used to obtain the cloud coverage area from the water synthetic aperture radar image, and the synthetic aperture radar image recognition model is used to extract the semantics of the target object in the cloud coverage area in the water synthetic aperture radar image domain; The aligned target water area remote sensing image, the associated water area synthetic aperture radar image, and the target object semantics are input into a pre-trained cloud processing diffusion model. The cloud processing diffusion model processes the cloud coverage area of the target water area remote sensing image according to the guidance of the water area synthetic aperture radar image and the target object semantics to generate a cloud-free target water area remote sensing image.
[0046] A computer-readable storage medium provided in an embodiment of the present invention stores a computer program which is not limited to the method operations described above, but can also execute related operations in a method for processing cloud occlusion in a water area remote sensing image provided in any embodiment of the present invention.
[0047] In the embodiments provided by the present invention, it should be understood that the disclosed structures and methods can be implemented in other ways. For example, the structural embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, structures or units, which can be electrical, mechanical or other forms.
[0048] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0049] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0050] The foregoing is merely a specific embodiment of the present invention, which enables those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for processing cloud occlusion in water remote sensing images, characterized in that: include: Using remote sensing imaging equipment and synthetic aperture radar to synchronously collect water remote sensing images and water synthetic aperture radar images of the target water area, and associate the two; Using a cloud segmentation model that supports cloud recognition to identify target water area remote sensing images that need to be processed for cloud occlusion; Aligning the remote sensing image of the target water area with the synthetic aperture radar image of the water area associated with it; After alignment, the cloud mask is used to obtain the cloud coverage area from the water synthetic aperture radar image, and the synthetic aperture radar image recognition model is used to extract the semantics of the target object in the cloud coverage area in the water synthetic aperture radar image domain; The aligned target water area remote sensing image, the associated water area synthetic aperture radar image, and the target object semantics are input into a pre-trained cloud processing diffusion model. The cloud processing diffusion model processes the cloud coverage area of the target water area remote sensing image according to the guidance of the water area synthetic aperture radar image and the target object semantics to generate a cloud-free target water area remote sensing image.
2. The method for processing cloud occlusion in water area remote sensing images according to claim 1, characterized in that: The process of aligning the target water area remote sensing image and the water area synthetic aperture radar image associated therewith comprises: segmenting the non-cloud area in the target water area remote sensing image by the cloud segmentation model, and extracting the remote sensing image domain alignment reference target from the non-cloud area in the target water area remote sensing image; Extracting a synthetic aperture radar image domain alignment reference target from a water area synthetic aperture radar image associated with a target water area remote sensing image; Matching the remote sensing image domain alignment reference target within the range of the synthetic aperture radar image domain alignment reference target; The water area remote sensing image and water area synthetic aperture radar image are aligned according to the rotation and translation relationship between the matching alignment reference targets in the two domains.
3. The method for processing cloud occlusion in water area remote sensing images according to claim 1, characterized in that: The cloud layer processing diffusion model includes: a VQVAE encoder, a VQVAE decoder, a Unet model, a radar feature encoder for sub-cloud targets, a shared time encoder, and a shared text encoder for sub-cloud targets; The pre-trained VQVAE encoder converts the input target water area remote sensing image and the water area synthetic aperture radar image associated therewith into a latent space; The diffusion process is used to iteratively add Gaussian noise to the latent space representation of the target water area remote sensing image; The diffusion results of the latent space representation of the target water area remote sensing image are input into the Unet model for iterative denoising; The radar feature encoder of the target under the cloud and the Unet model are set in parallel to provide the intermediate module and the decoding module of the Unet model with the radar features of the target under the cloud extracted from the water synthetic aperture radar image in each iterative denoising process. The radar features of the target under the cloud guide the Unet model to generate a water remote sensing image domain representation based on the target under the cloud in the corresponding sub-cloud area; The shared sub-cloud object text encoder encodes the semantics of the object extracted from the cloud-covered area in the water area synthetic aperture radar image domain using the synthetic aperture radar image recognition model, and provides the object semantic encoding to the intermediate module and decoding module of the Unet model in each iterative denoising process. The object semantic encoding guides the Unet model to generate a water area remote sensing image domain representation based on the sub-cloud object in the corresponding sub-cloud area. The shared time encoder provides iterative time step encoding for the radar feature encoder and Unet model of sub-cloud targets in each iterative denoising process; The pre-trained VQVAE decoder decodes the processed target water area remote sensing image based on the denoising result generated by the Unet model.
4. The method for processing cloud occlusion in water area remote sensing images according to claim 3, characterized in that: The Unet model includes: a plurality of cascaded encoding modules, intermediate modules and decoding modules corresponding to the number of encoding modules, wherein a jump chain is set between the corresponding encoding modules and decoding modules; for a decoding module at any level, the output of the decoding module at the previous level and the output of the encoder at the same level are combined and input into the decoding module.
5. The method for processing cloud occlusion in water area remote sensing images according to claim 4, characterized in that: The radar feature encoder for targets under clouds includes: a number of encoding modules and intermediate modules copied in the cascade of the Unet model, each encoding module and intermediate module is connected to a 1×1 convolutional layer, the 1×1 convolutional layer connected to the intermediate module is connected to the intermediate module of the Unet model, and the 1×1 convolutional layer connected to the encoding module is connected to the corresponding decoding module in the Unet model.
6. The method for processing cloud occlusion in water area remote sensing images according to claim 3, characterized in that: The training process of the cloud layer processing diffusion model includes: The parameters of the Unet model are trained so that the Unet model can generalize and restore the diffused water remote sensing image. The training method includes: inputting the water remote sensing image in the training set into the cloud processing diffusion model, and the pre-trained VQVAE encoder converts the water remote sensing image into a latent space; then using the diffusion process to iteratively add Gaussian noise to the latent space representation of the water remote sensing image; performing multiple arbitrary samplings from the diffusion process to obtain the iterative time step, noise and latent space representation before and after noise addition of all samples; the shared time encoder encodes the iterative time step of the sampling, and the latent space representation after noise addition and the time step encoding are combined and input into the Unet model, and it is expected that the latent space representation after noise addition will restore the latent space representation before noise addition after denoising according to the Unet prediction noise; finally, the pre-trained VQVAE encoder will restore the water remote sensing image based on the final denoising result.
7. The method for processing cloud occlusion in water area remote sensing images according to claim 6, characterized in that: The training process of the cloud layer processing diffusion model includes: after training the Unet model, copying several encoding modules and intermediate modules of the cascade of the Unet model and constructing a radar feature encoder for targets under clouds; training the radar feature encoder for targets under clouds and the shared text encoder for targets under clouds to provide optimization conditions for Unet, and using the optimization conditions to control the noise denoising predicted by the Unet model, so as to generate a latent space representation of a cloud-free water image.
8. The method for processing cloud occlusion in water area remote sensing images according to claim 7, characterized in that: Methods for training the radar feature encoder for sub-cloud objects and the shared text encoder for sub-cloud objects include: For any water synthetic aperture radar image in the training set, the cloud coverage area is extracted using the corresponding cloud mask; the semantics of the target object in the cloud coverage area in the water synthetic aperture radar image domain is extracted using the synthetic aperture radar image recognition model; Inputting the water area remote sensing image and the water area synthetic aperture radar image into the pre-trained VQVAE encoder of the cloud layer processing diffusion model; the pre-trained VQVAE encoder converts the water area remote sensing image and the water area synthetic aperture radar image into a latent space; Then, the diffusion process is used to synchronously add consistent Gaussian noise to the latent space representations of the water remote sensing image and the cloud-free water remote sensing image; The shared under-cloud object text encoder encodes the semantics of the objects in the cloud-covered area in the water synthetic aperture radar image domain to obtain the object semantic encoding and provide it to the Unet model. The object semantic encoding provides the Unet model with the semantic information of the under-cloud objects and strengthens the semantic representation. Multiple random samplings are performed from the diffusion process of water remote sensing images and cloud-free water remote sensing images to obtain the iterative time step, noise, and latent space representations before and after noise addition of the two samplings. The shared temporal encoder encodes the sampled iterative time steps; The latent space representation of the sampled water synthetic aperture radar image is processed by a 1×1 convolution layer and combined with the latent space representation of the cloud water remote sensing image after noise addition, and then input into the radar feature encoder of the target under the cloud. At the same time, the sampled time step encoding and the semantic encoding of the target are combined and input into the radar feature encoder of the target under the cloud. The radar feature encoder of the target under the cloud provides the spatial information of the target under the cloud for the Unet model. The latent space representation of the sampled cloud-water area remote sensing image after noise addition is input into the Unet model. At the same time, the sampled time step encoding and target object semantic encoding are input into the Unet model. The Unet model predicts diffuse noise under the additional conditions provided by the radar feature encoder of sub-cloud targets and the semantic encoding of targets, so that the latent space representation of the cloudy water remote sensing image after noise addition is denoised according to the predicted noise and the distance between the latent space representation of the cloudless water remote sensing image before noise addition in the same sampling stage is minimized; finally, the pre-trained VQVAE encoder will generate a cloudless water remote sensing image based on the final denoising result.
9. The method for processing cloud occlusion in water area remote sensing images according to claim 6, characterized in that: The training set includes: associated water remote sensing images, cloud masks of water remote sensing images, water synthetic aperture radar images, and cloudless water remote sensing images, wherein the associated cloudless water remote sensing images have the same scene as the water remote sensing images but are not blocked by clouds.
10. The method for processing cloud occlusion in water area remote sensing images according to claim 3, characterized in that: In the process of generating the cloud processing diffusion model, a remote sensing image of the target water area and a synthetic aperture radar image of the target water area are provided; Inputting the target water area remote sensing image and the target water area synthetic aperture radar image into the pre-trained VQVAE encoder of the cloud layer processing diffusion model; the pre-trained VQVAE encoder converts the target water area remote sensing image and the target water area synthetic aperture radar image into a latent space; Then, the diffusion process is used to add Gaussian noise to the latent space representation of the remote sensing image of the target water area; The shared under-cloud object text encoder encodes the semantics of the object in the cloud-covered area of the target water area in the synthetic aperture radar image domain to obtain the semantic encoding of the object and provide it to the Unet model; The diffusion results of the remote sensing image of the target water area are denoised to obtain the iterative time step of the diffusion, the noise, and the latent space representation before and after the noise addition; The shared temporal encoder encodes the iterative time steps corresponding to the denoising process; The latent space representation of the synthetic aperture radar image of the target water area in the denoising process is processed by a 1×1 convolutional layer and combined with the latent space representation of the target cloud water area remote sensing image after noise addition, and then input into the radar feature encoder of the target object under the cloud. At the same time, the time step encoding and the semantic encoding of the target object are combined and then input into the radar feature encoder of the target object under the cloud. The latent space representation of the target cloud water area remote sensing image after each diffusion step of noise is input into the Unet model. At the same time, the time step encoding and the semantic encoding of the target object are input into the Unet model. The Unet model predicts the noise under the additional conditions provided by the radar feature encoder of the target under the cloud and the semantic encoding of the target, so that the latent space representation of the noisy target cloud water area remote sensing image is denoised according to the predicted noise; finally, the pre-trained VQVAE encoder will generate a cloud-free target water area remote sensing image based on the final denoising result.
Citation Information
Patent Citations
High-resolution remote sensing image and laser radar point cloud fused individual tree segmentation method and system
CN111462134A
Deep learning cloud removal method based on SAR-optical remote sensing image combination
CN115809970A
Remote sensing image cloud removal method based on synthetic aperture radar and visible light fusion
CN116993598A
Progressive repair frame cloud removal method for fusion of optical remote sensing image and SAR (Synthetic Aperture Radar) image
CN117058059A
Remote sensing optical time sequence image reconstruction method and device, terminal and storage medium
CN119672158A
Cited By
Full-time-sequence lake water hyacinth monitoring method and system based on optical and SAR image fusion
CN121837964A