SAR-to-optical image generation method and system based on remote sensing spectral feature guidance

By introducing multimodal training with spectral feature-guided in SAR to optical image generation, the problem of insufficient SAR image generation quality in the prior art is solved, high-quality optical image generation is achieved, and accurate recognition of complex geological areas is supported.

CN120355586APending Publication Date: 2025-07-22POWERCHINA FUJIAN ELECTRIC POWER SURVEY & DESIGN INST CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510419422.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art has high-frequency details loss, blurred texture, unstable training process, and lack of spectral feature guidance in SAR to optical image generation, resulting in spectral distortion of vegetation and water bodies, and semantically controllable image generation cannot be achieved.

Method used

Using a method based on remote sensing spectral feature guidance, a generation model is pre-trained and multi-modal training is carried out in combination with spectral feature data and prompt words, the first and second generative models are established to generate high-quality optical images.

Benefits of technology

It enhances the ability of generative models to reconstruct visual spectral information of different surfaces, alleviates the blur, confusion and loss of generated optical image information, and provides accurate identification data support for poor geological areas in complex mountainous areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355586A_ABST
    Figure CN120355586A_ABST
Patent Text Reader

Abstract

The invention relates to an SAR-to-optical image generation method and system based on remote sensing spectral feature guidance. The method comprises the following steps: acquiring an SAR image and an optical image, and preprocessing the SAR image and the optical image to form a sample set; establishing a first generation model, pre-training the first generation model through the sample set, and generating a pseudo optical image by the first generation model according to the input SAR image; selecting a spectral band of the target optical image and extracting spectral features, and constructing multi-modal training guide data according to spectral feature data and set cues; establishing a second generation model, and performing joint training on the first generation model and the second generation model through multi-modal training guide data and a sample SAR image; and generating an SAR-to-optical image through the trained first generation model and second generation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for generating SAR-to-optical images guided by remote sensing spectral features, and belongs to the technical field of remote sensing image processing. Background Art

[0002] Optical images are a type of remote sensing data that uses visible light or multispectral sensors to digitally record the reflection characteristics of the Earth's surface. It can directly reflect different spectral, texture, and spatial characteristics of the Earth's surface and is the most commonly used remote sensing data source. However, factors such as cloud cover or unstable satellite revisit cycles can interfere with the normal application of optical images. Especially in some special areas, there may even be a situation where no available optical images can be obtained for a long time, which has a negative impact on the spatio-temporal interpretation of the Earth's surface. In contrast, synthetic aperture radar (SAR), as an active sensor data source, relies on the penetration ability of microwaves and can overcome the influence of clouds and rain to achieve surface imaging under any weather conditions. With its stable revisit cycle and anti-interference ability, synthetic aperture radar has become an important auxiliary data source for remote sensing technology.

[0003] However, due to its unique imaging mechanism, the surface features presented by synthetic aperture radar often show obvious clutter and ambiguity, posing challenges to manual visual interpretation. Therefore, converting synthetic aperture radar into optical images has become an important method to reduce the complexity of visual interpretation and make up for the lack of optical image data. The essence of synthetic aperture radar-to-optical image conversion is the image domain conversion process between different modalities. In the past, pixel regression mapping methods and machine learning techniques were used for this conversion. However, these methods are difficult to apply to highly dynamic and complex surface scenes. In recent years, the research on generative models based on deep learning has grown explosively. Through deep training, generative neural networks can capture the high-level mapping relationships and distribution patterns between the source domain and the target domain, thereby generating pseudo-data consistent with the characteristics of the target domain. In the prior art, SAR-to-optical image mainly includes the following two types of methods:

[0004] 1. Conversion methods based on generative adversarial networks (GANs):

[0005] As proposed in the patent US20200193652A1, "SYSTEM AND METHOD FOR SAR TO OPTICAL IMAGE TRANSLATION USING GENERATIVE ADVERSARIAL NETWORKS", a cGAN architecture is used to achieve the conversion from SAR to optical images. However, this method has the following defects: serious loss of high-frequency details, blurred textures in the generated images (PSNR is generally lower than 30dB); unstable training process, prone to mode collapse; lack of spectral feature guidance, resulting in spectral distortion of ground objects such as vegetation and water bodies (SAD>15°);

[0006] 2. Improved solution based on diffusion model:

[0007] As proposed in the patent CN113592987A, "A Remote Sensing Image Generation Method Based on Diffusion Model", the DDPM framework is used to improve the quality of image generation. However, this technology has limitations: it does not consider multi-modal feature fusion and only uses a single SAR data source; there is a large deviation in the spectral distribution between the generated image and the real optical data (NDVI error>0.2); it cannot achieve semantically controllable image generation. Summary of the Invention

[0008] In order to solve the problems existing in the above-mentioned prior art, the present invention proposes a method and system for generating SAR-to-optical images guided by remote sensing spectral features. First, in the process of converting synthetic aperture radar to optical images, an end-to-end generation model is pre-trained, and the spectral feature data of the target optical image is calculated. Then, the spectral feature data and the corresponding prompt words are used as prior conditional information to guide the multi-modal training of the pre-trained model, so as to generate higher-quality optical images and provide data support for the accurate identification of bad geological areas in complex mountainous areas.

[0009] The technical solution of the present invention is as follows:

[0010] On the one hand, the present invention proposes a method for generating SAR-to-optical images guided by remote sensing spectral features, including the following steps:

[0011] Obtain SAR images and optical images and preprocess them to form a sample set;

[0012] Establish a first generation model, and pre-train the first generation model through the sample set. The first generation model generates pseudo-optical images according to the input SAR images;

[0013] Select the spectral bands of the target optical image and extract spectral features, and construct multi-modal training guidance data according to the spectral feature data and the set prompt words;

[0014] Build a second generation model, and jointly train the first generation model and the second generation model through multimodal training guidance data and sample SAR images;

[0015] Generate SAR-to-optical images through the trained first generation model and second generation model.

[0016] As a preferred embodiment, the first generation model adopts a DDPM denoising diffusion probability model;

[0017] In the forward process, the first generation model adds noise and concatenates channels to the VV and VH polarization information of the SAR image.

[0018] Input the SAR image and the time step into the first generation model, and the first generation model predicts a random noise and generates a corresponding pseudo-optical image.

[0019] As a preferred embodiment, in the step of selecting the spectral bands of the target optical image and extracting spectral features, the selected spectral bands include:

[0020] R red, G green, B blue bands, as well as the NIR near-infrared spectral band and the SWIR short-wave infrared spectral band;

[0021] The extracted spectral features include:

[0022] NDVI Normalized Difference Vegetation Index, NDWI Normalized Difference Water Index, NDBI Normalized Difference Building Index, and BSI Bare Soil Index.

[0023] As a preferred embodiment, the second generation model adopts a CLIP contrastive language-image pre-training model;

[0024] The multimodal training guidance data includes first guidance data obtained by concatenating channels of spectral feature data and SAR image data, and second guidance data composed of prompt words;

[0025] The first guidance data is input into the image encoder of the second generation model to extract image features, the second guidance data is input into the text encoder of the second generation model to obtain text features, and the second generation model is subjected to similarity learning by calculating the contrast loss through the image features and the text features;

[0026] The image decoder of the second generation model generates a feature optical image according to the image features.

[0027] As a preferred embodiment, in the step of jointly training the first generation model and the second generation model through multimodal training guidance data and sample SAR images, based on the second generation model, a conditional mapping constraint is imposed on the pre-trained first generation model, including:

[0028] The sample SAR image and spectral feature data are concatenated by channel concatenation as the input data of the first generation model;

[0029] The image features extracted by the image encoder of the second generation model are mapped to the latent space of the same dimension of the first generation model;

[0030] The feature optical image generated by the image decoder of the second generation model is mapped to the generation space of the same dimension of the first generation model, and the synthetic optical image is output from the generation space of the first generation model.

[0031] As a preferred embodiment, in the step of jointly training the first generation model and the second generation model by using the multi-modal training guidance data and the sample SAR image, joint training is performed by calculating the joint loss, and the joint loss includes:

[0032] The contrast loss, and the metric loss between the synthetic optical image and the target optical image.

[0033] As a preferred embodiment, the image encoder of the second generation model has the same structure as the image encoder of the first generation model.

[0034] On the other hand, the present invention also proposes a SAR-to-optical image generation system guided by remote sensing spectral features, including:

[0035] A data acquisition module, configured to acquire SAR images and optical images and perform preprocessing to form a sample set;

[0036] A pre-training module, configured to establish a first generation model, and pre-train the first generation model through the sample set, and the first generation model generates a pseudo-optical image according to the input SAR image;

[0037] A multi-modal data construction module, configured to select the spectral bands of the target optical image and extract spectral features, and construct multi-modal training guidance data according to the spectral feature data and the set prompt words;

[0038] A joint training module, configured to establish a second generation model, and jointly train the first generation model and the second generation model by using the multi-modal training guidance data and the sample SAR image;

[0039] A SAR-to-optical image generation module, configured to generate SAR-to-optical images by using the trained first generation model and second generation model.

[0040] On yet another aspect, the present invention also proposes an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the SAR-to-optical image generation method guided by remote sensing spectral features according to any embodiment of the present invention.

[0041] In another aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method for generating SAR-to-optical images guided by remote sensing spectral features as described in any embodiment of the present invention.

[0042] The beneficial effects of the present invention are as follows:

[0043] Integrating SAR images with multiple spectral feature data and corresponding prompt words enhances the ability of the generation model to reconstruct visual spectral information of different surfaces, and alleviates the problems of information blurring, confusion, and loss in the generated optical images.

[0044] The additional aspects and advantages of the present invention will be set forth in the following description, and some of them will be obvious from the description, or can be learned by practicing the present invention. In addition, the various aspects and advantages of the present invention can be realized and obtained by the method steps and combinations specifically pointed out in the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 It is a flowchart of Embodiment 1 of the present invention;

[0046] Figure 2 It is a specific flowchart for generating synthetic optical images in the embodiments of the present invention. DETAILED DESCRIPTION

[0047] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0048] It should be understood that the step numbers used in the text are only for convenient description and do not limit the execution order of the steps.

[0049] It should be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0050] The terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0051] The term "and / or" refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0052] Introduce some of the nouns applied in the embodiments:

[0053] Synthetic Aperture Radar (SAR): It is an active microwave imaging radar. By moving along a trajectory with a moving platform (such as an airplane, a satellite), it emits and receives microwave signals, and uses signal processing technology to synthesize a virtual large-aperture antenna, thereby obtaining high-resolution images. SAR has the imaging ability all day and all weather, and is widely used in terrain mapping, disaster monitoring, military reconnaissance and other fields.

[0054] DDPM (Denoising Diffusion Probabilistic Models): It is a generative model based on the diffusion process, used to generate high-quality data samples (such as images, audio, etc.). DDPM generates new samples consistent with the training data distribution by simulating the step-by-step denoising process of data from noise to the target distribution. It is one of the important advances in generative models in the field of deep learning in recent years, especially performing well in image generation tasks. The sample quality generated by DDPM is usually better than that of traditional Generative Adversarial Networks (GANs) and Variational Autoencoders (VAEs).

[0055] tanh activation function: The full name of the tanh activation function is the Hyperbolic Tangent Function. It is one of the commonly used activation functions in deep learning, mainly used to introduce non-linear characteristics, thereby enhancing the expression ability of neural networks.

[0056] Embodiment 1:

[0057] Refer to Figure 1 , this embodiment proposes a method for generating SAR-to-optical images guided by remote sensing spectral features to achieve an effective conversion of SAR images to high-precision optical images. First, pre-train a first generation model for converting synthetic aperture radar images to optical images; and calculate the spectral feature data of the target optical image; then perform multi-modal training, integrating synthetic aperture radar images, spectral feature data and prompts on the basis of the pre-trained model, enhancing the ability of the generation model to reconstruct different surface visual spectral information, and reducing problems such as information blur, confusion and loss in the generated optical images.

[0058] The method provided in this embodiment specifically includes the following steps:

[0059] S100. Obtain SAR images and optical images as basic samples, and perform preprocessing to form a sample set.

[0060] Specifically, the preprocessing steps of the SAR image include:

[0061] Radiometric calibration: Convert the original DN value to the backscattering coefficient.

[0062] Terrain correction: Use SRTM 30m DEM data and apply the Range-Doppler terrain correction algorithm.

[0063] Median synthesis: Perform temporal median synthesis on SAR images within the same area over 3 months to suppress speckle noise.

[0064] The preprocessing steps of the optical image include:

[0065] Atmospheric correction: Use the Sen2Cor algorithm to eliminate the atmospheric influence.

[0066] Geometric registration: Using Sentinel-2 as a reference, resample the SAR image to the same coordinate system (UTM / WGS84).

[0067] Data cropping: Taking pixels as units, set the sliding window step size for image cropping, and mirror fill the boundary area.

[0068] Through cropping based on a sliding window and manual filtering, a sample data set with a size of H×W pixels is obtained. And it is divided into a training data set and a test data set according to a certain ratio.

[0069] S200. Establish a first generation model, and pre-train the first generation model through the sample set. The first generation model generates a pseudo-optical image according to the input SAR image; in this step, pre-training on the sample set constructed in step S100 can enable the first generation model to independently learn the initial feature distribution of a single target, thereby reducing the training instability encountered in the multi-modal training process and being conducive to more effectively integrating information from different modalities.

[0070] S300. Select the spectral bands of the target optical image and extract spectral features, and construct multi-modal training guidance data according to the spectral feature data and the set prompt words, which serves as the basis for subsequent multi-modal training.

[0071] S400. Establish a second generation model, and jointly train the first generation model and the second generation model through the multi-modal training guidance data and the sample SAR image.

[0072] S500. Generate SAR-to-optical images through the trained first generation model and second generation model.

[0073] As a preferred implementation manner of this embodiment, the first generation model adopts the DDPM denoising diffusion probabilistic model;

[0074] During the pre-training process, in the forward process, the first generation model adds noise and concatenates channels to the VV and VH polarization information of the SAR image (VV represents the Vertical-Vertical polarization information; VH represents the Vertical-Horizontal polarization information).

[0075] The SAR image and the time step are input into the backbone U-Net network of the first generation model. The U-Net network processes and analyzes the input image data and time step information, predicts a random noise, and the image decoder of the first generation model generates a corresponding pseudo-optical image based on the random noise and the features of the SAR image output by the image encoder.

[0076] As a preferred implementation manner of this embodiment, in the step of selecting the spectral bands of the target optical image and extracting spectral features, the selected spectral bands include:

[0077] R red, G green, B blue bands, as well as the NIR near-infrared spectral band and the SWIR short-wave infrared spectral band;

[0078] The extracted spectral features include:

[0079] Normalized Difference Vegetation Index (NDVI), Normalized Difference Water Index (NDWI), Normalized Difference Built-Up Index (NDBI), and Bare Soil Index (BSI). Specifically, the calculation formulas for each index are as follows:

[0080] NDVI = (NIR - R) / (NIR + R);

[0081] NDWI = (G - NIR) / (G + NIR);

[0082] NDBI = (SWIR - NIR) / (SWIR + NIR);

[0083] BSI = ((SWIR + R) - (NIR + B)) / ((SWIR + R) + (NIR + B)).

[0084] As a preferred embodiment of this embodiment, the second generation model adopts the CLIP contrastive language-image pre-training model; in terms of multi-modal training, the pre-trained first generation model is retrained and combined with the calculated spectral feature data to achieve multi-modal joint training. Specifically:

[0085] In this embodiment, by using the prompt as the supervision information to drive the image encoder for similarity learning, the network's ability to distinguish and match different index attributes is enhanced. First, the SAR data and each spectral feature data are fused through channel splicing to construct the first guiding data S n :

[0086] S n = Cat(S, I n );

[0087] Among them, S represents the SAR image data, and I n represents the nth spectral feature data, and Cat represents the channel splicing operation.

[0088] Subsequently, the second guiding data composed of the first guiding data S n and the corresponding prompt is jointly input into the image-text encoder of the second generation model to obtain the corresponding deep features, that is, the image feature F n and the text feature F t :

[0089]

[0090] Among them, represents the image encoder, represents the text encoder, and T n is the second guiding data.

[0091] Similarity learning is performed on the second generation model by calculating the contrast loss through the image feature and the text feature:

[0092] L n = (L clip (F n , F t )) / 4;

[0093] Among them, L n represents the contrast loss calculation, and L clip represents the average value of the CLIP calculation results of the losses of each spectral feature data.

[0094] The overall architecture of the second generation model also includes an image decoder in addition to the image-text encoder. The image decoder generates the feature optical image O n according to the image feature F n :

[0095] On = D clip (F n );

[0096] Among them, D clip represents the image decoder of the second generation model.

[0097] As a preferred implementation of this embodiment, in the step of jointly training the first generation model and the second generation model through multimodal training guidance data and sample SAR images, based on the second generation model, a conditional mapping constraint is imposed on the pre-trained first generation model, including:

[0098] At the data input layer, the sample SAR image and spectral feature data are concatenated by channels as the input data S of the first generation model m :

[0099] I m = Cat(I ndvi , I ndwi , I ndbi , I bs i );

[0100] S m = Cat(S, I m );

[0101] Among them, I ndvi , I ndwi , I ndbi , I bsi respectively represent the Normalized Difference Vegetation Index NDVI, the Normalized Difference Water Index NDWI, the Normalized Difference Built-up Index NDBI, and the Bare Soil Index BSI;

[0102] At the deep inference layer, the image features F extracted by the image encoder of the second generation model n are mapped to the latent space of the same dimension of the first generation model to obtain the mapped features F m :

[0103] F m = Cat(E p (S m ), F n );

[0104] Among them, E p represents the image encoder of the first generation model.

[0105] At the decision output layer, the feature optical image generated by the image decoder of the second generation model is mapped to the generation space of the same dimension of the first generation model, and the generation space of the first generation model outputs the synthetic optical image O m , specifically:

[0106] O m = Tn(Conv(Cat(D p (F m ), O n )));

[0107] Among them, D p is the image decoder of the first generation model, Tn is the tanh activation function, and Conv is the convolution operation.

[0108] As a preferred implementation mode of this embodiment, in the step of jointly training the first generation model and the second generation model through multi-modal training guidance data and sample SAR images, joint training is performed by calculating the joint loss, and the joint loss includes:

[0109] Contrast loss L clip , and the metric loss between the synthesized optical image and the target optical image. In this embodiment, the metric loss between the synthesized optical image and the target optical image is calculated by the mean square error (MSE), that is:

[0110] L mse = |O m - O r | 2 ;

[0111] Among them, O t represents the target optical image.

[0112] The joint loss is specifically:

[0113] L j = αL mse + βL clip ;

[0114] Among them, L j is the joint loss, and α and β are respectively preset weight coefficients.

[0115] As a preferred implementation mode of this embodiment, the image encoder of the second generation model has the same structure as the image encoder of the first generation model.

[0116] Based on the method provided in this embodiment, specifically refer to Figure 2 , and the specific process of generating the synthesized optical image is:

[0117] First, obtain the sample set, pre-train the first generation model through the sample set, and the pre-trained first generation model can generate a pseudo-optical image according to the input SAR image.

[0118] Next, spectral feature calculation is performed, and spectral feature data: NVDI, NDWI, NDBI, and BSI are extracted based on the target optical image.

[0119] Finally, multi-modal training is carried out. The prompt words and spectral feature data are used as guiding data to jointly train the first generation model and the second generation model. The feature images output by the image decoders of the trained first generation model and the second generation model are synthesized to obtain the final synthesized image.

[0120] Embodiment 2:

[0121] This embodiment proposes a SAR-to-optical image generation system guided by remote sensing spectral features, including:

[0122] A data acquisition module for acquiring SAR images and optical images and performing preprocessing to form a sample set; this module is used to implement the function of step S100 in Embodiment 1 and will not be elaborated here;

[0123] A pre-training module for establishing a first generation model and pre-training the first generation model through the sample set. The first generation model generates a pseudo-optical image based on the input SAR image; this module is used to implement the function of step S200 in Embodiment 1 and will not be elaborated here;

[0124] A multi-modal data construction module for selecting the spectral bands of the target optical image and extracting spectral features, and constructing multi-modal training guiding data based on the spectral feature data and the set prompt words; this module is used to implement the function of step S300 in Embodiment 1 and will not be elaborated here;

[0125] A joint training module for establishing a second generation model and jointly training the first generation model and the second generation model through the multi-modal training guiding data and the sample SAR image; this module is used to implement the function of step S400 in Embodiment 1 and will not be elaborated here;

[0126] A SAR-to-optical image generation module for generating SAR-to-optical images through the trained first generation model and the second generation model; this module is used to implement the function of step S500 in Embodiment 1 and will not be elaborated here.

[0127] Embodiment 3:

[0128] This embodiment proposes an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the SAR-to-optical image generation method guided by remote sensing spectral features as described in any embodiment of the present invention.

[0129] Embodiment 4:

[0130] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for generating SAR-to-optical images guided by remote sensing spectral features as described in any embodiment of the present invention.

[0131] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent the cases of A existing alone, A and B existing simultaneously, and B existing alone. Where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one of the following" and its similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c can be single or multiple.

[0132] Those of ordinary skill in the art can realize that the units and algorithm steps described in the embodiments disclosed herein can be implemented by a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0133] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0134] In several embodiments provided by the present application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (hereinafter referred to as ROM), random access memories (hereinafter referred to as RAM), magnetic disks, or optical discs that can store program codes.

[0135] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.

Claims

1. A method for generating SAR-to-optical images guided by remote sensing spectral features, characterized in that It includes the following steps: Obtain SAR images and optical images and perform preprocessing to form a sample set; Establish a first generation model and pre-train the first generation model with the sample set. The first generation model generates pseudo-optical images based on the input SAR images; Select the spectral bands of the target optical image and extract spectral features, and construct multi-modal training guidance data according to the spectral feature data and the set prompt words; Establish a second generation model and jointly train the first generation model and the second generation model with the multi-modal training guidance data and sample SAR images; Generate SAR-to-optical image conversion through the trained first generation model and second generation model.

2. The method for generating SAR-to-optical images guided by remote sensing spectral features according to claim 1, wherein: The first generation model uses the DDPM denoising diffusion probability model; In the forward process, the first generation model adds noise and concatenates channels to the VV and VH polarization information of the SAR image; Input the SAR image and the time step into the first generation model, and the first generation model predicts a random noise and generates a corresponding pseudo-optical image.

3. A method for generating SAR-to-optical images guided by remote sensing spectral features according to claim 1, characterized in that In the step of selecting the spectral bands of the target optical image and extracting spectral features, the selected spectral bands include: R red, G green, B blue bands, as well as the NIR near-infrared spectral band and the SWIR short-wave infrared spectral band; The extracted spectral features include: NDVI normalized difference vegetation index, NDWI normalized difference water index, NDBI normalized difference building index, and BSI bare soil index.

4. The method for generating SAR-to-optical images guided by remote sensing spectral features according to claim 1, wherein: The second generation model uses the CLIP contrastive language-image pre-training model; The multi-modal training guidance data includes first guidance data obtained by concatenating channels of the spectral feature data and the SAR image data, and second guidance data composed of prompt words; The first guidance data is input into the image encoder of the second generation model to extract image features, the second guidance data is input into the text encoder of the second generation model to obtain text features, and the second generation model is subjected to similarity learning by calculating the contrastive loss through the image features and the text features; The image decoder of the second generation model generates a feature optical image according to the image features.

5. A method for generating SAR-to-optical images guided by remote sensing spectral features according to claim 4, characterized in that In the step of jointly training the first generation model and the second generation model with the multi-modal training guidance data and sample SAR images, based on the second generation model, a conditional mapping constraint is imposed on the pre-trained first generation model, including: Concatenate the sample SAR image and the spectral feature data through channel concatenation as the input data of the first generation model; Map the image features extracted by the image encoder of the second generation model to the latent space of the same dimension as the first generation model; Map the feature optical image generated by the image decoder of the second generation model to the generation space of the same dimension as the first generation model, and the generation space of the first generation model outputs a synthetic optical image.

6. The SAR-to-optical image generation method based on remote sensing spectral feature guidance according to claim 5, wherein In the step of jointly training the first generation model and the second generation model by using the multi-modal training guiding data and the sample SAR images, joint training is performed by calculating a joint loss, and the joint loss includes: a contrast loss, and a metric loss between the synthesized optical image and the target optical image.

7. A method for generating SAR-to-optical images guided by remote sensing spectral features according to claim 4, wherein: the image encoder of the second generation model has the same structure as the image encoder of the first generation model.

8. A SAR-to-optical image generation system guided by remote sensing spectral features, characterized in that, It includes: a data acquisition module, configured to acquire SAR images and optical images and perform preprocessing to form a sample set; a pre-training module, configured to establish a first generation model and pre-train the first generation model by using the sample set, and the first generation model generates a pseudo-optical image according to the input SAR image; a multi-modal data construction module, configured to select the spectral bands of the target optical image and extract spectral features, and construct multi-modal training guiding data according to the spectral feature data and the set prompt words; a joint training module, configured to establish a second generation model and jointly train the first generation model and the second generation model by using the multi-modal training guiding data and the sample SAR images; a SAR-to-optical image generation module, configured to generate SAR-to-optical images by using the trained first generation model and the second generation model.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the method for generating SAR-to-optical images guided by remote sensing spectral features according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the method for generating SAR-to-optical images guided by remote sensing spectral features according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • High visibility overlay systems and methods

    US20200193652A1