Synthetic aperture radar to optical image conversion method based on diffusion driving and wavelet enhancement

By using the diffusion-driven and wavelet-enhanced methods, modal mapping and back-diffusion networks are constructed to solve the problem of difference in feature distribution between SAR and optical images, improve the interpretability and translation accuracy of SAR images, and achieve high-quality optical image generation.

CN120807301APending Publication Date: 2025-10-17HARBIN INST OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510917437.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

The significant differences in feature distribution between SAR and optical images in existing methods make feature alignment difficult, affecting the accuracy of SAR to optical image translation and the quality of generated images.

Method used

A method based on diffusion drive and wavelet enhancement is adopted. By constructing a modal mapping network and a back diffusion network, the feature distributions of SAR and optical images are gradually aligned. The intermediate state in the forward diffusion process and wavelet transform processing are utilized to optimize the low-frequency alignment and high-frequency detail reconstruction of the image.

Benefits of technology

It has significantly improved the interpretability and application value of SAR images, improved the accuracy and visual quality of cross-modal image translation, and promoted the effective use and application expansion of remote sensing data in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807301A_ABST
    Figure CN120807301A_ABST
Patent Text Reader

Abstract

The invention discloses a synthetic aperture radar to optical image conversion method based on diffusion driving and wavelet enhancement, relates to a remote sensing image multi-modal image conversion method, and belongs to the technical field of multi-modal satellite remote sensing. The objective of the invention is to solve the problem that the accuracy of SAR-to-optical image translation and the quality of a generated image are affected due to difficult feature alignment between an SAR and an optical image caused by the significant difference between the SAR and the optical image in feature distribution in the prior art. The method comprises the following steps: 1, constructing a training sample pair; 2, obtaining x0 and y0; 3, respectively adding the same Gaussian noise to x0 and y0 to generate xt and yt; 4, obtaining training data of the back diffusion network; 6, obtaining a trained back diffusion network; and 7, obtaining a feature space of an optical image after the SAR remote sensing image to be detected is decoded.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a remote sensing image multi-modal image conversion method, belonging to the technical field of multi-modal satellite remote sensing. BACKGROUND

[0002] Synthetic Aperture Radar (SAR) has outstanding advantages in the field of remote sensing due to its unique imaging mechanism. It has all-weather, all-day imaging capability, and can stably obtain high-quality images in complex environmental conditions such as clouds, haze and night, especially in situations where optical sensors cannot work normally. Therefore, SAR has important application value in application scenarios such as disaster monitoring, land use classification and target detection. However, due to its characteristics of coherent imaging based on electromagnetic waves, SAR images are usually disturbed by significant speckle noise; at the same time, its imaging method also leads to blurred texture, low contrast and lack of color information. These characteristics make SAR images have great differences in visual perception from optical images, increase the difficulty of interpretation and understanding, and become an important factor restricting its in-depth application.

[0003] In contrast, optical images have rich texture information and strong interpretability, and are widely used in various remote sensing tasks such as feature identification and change detection. However, optical imaging is sensitive to weather conditions and lighting conditions, and often has imaging interruption or information loss problems in cloudy, low-light and other environments. In order to fully play the complementary advantages of the two types of sensors, the SAR-to-optical image conversion task has emerged, thereby improving the comprehensive utilization efficiency of remote sensing data. In recent years, this task has shown good application prospects in disaster assessment, land cover classification and other fields, and has gradually become one of the important directions of remote sensing research. SUMMARY

[0004] The purpose of the present application is to solve the problem that due to the significant difference in feature distribution between SAR and optical images in the prior art, the feature alignment between the two is difficult, thereby affecting the accuracy of SAR-to-optical image translation and the quality of the generated image, and a SAR-to-optical image conversion method based on diffusion driving and wavelet enhancement is proposed.

[0005] The specific process of the synthetic aperture radar-to-optical image conversion method based on diffusion driving and wavelet enhancement is as follows:

[0006] Step one, obtain SAR remote sensing and optical paired images of the same target area, and construct a training sample pair;

[0007] Step two, input the SAR remote sensing image in the training sample pair into the encoder of the pre-trained variational autoencoder, and the encoder outputs the feature vector x0 of the encoded SAR remote sensing image;

[0008] input the optical image in the training sample pair into the encoder of the pre-trained variational autoencoder, and the encoder outputs a feature vector y0 of the optical image after encoding;

[0009] Step three, respectively adding the same Gaussian noise to the feature vector x0 of the SAR remote sensing image after encoding and the feature vector y0 of the optical image after encoding output by step two, to generate a noisy SAR remote sensing image x t and a noisy optical image y t in the time step t noise domain;

[0010] Step four, constructing a modal mapping network;

[0011] inputting the noisy SAR remote sensing image x t generated in step three into the modal mapping network as input, inputting the noisy optical image y t generated in step three into the modal mapping network as output, training the modal mapping network until the loss function L G converges, and obtaining the trained modal mapping network;

[0012] Step five, taking the feature vector x0 of the SAR remote sensing image after encoding output by step two, the Gaussian noise added in the noisy optical image y j in the time step j noise domain, and the noise added in the noisy optical image y j in the time step j noise domain as training data of the backward diffusion network;

[0013] j=1,2,…,t;t=1,2,…,T;

[0014] T represents the total time step;

[0015] Step six, constructing a backward diffusion network;

[0016] inputting the feature vector x0 of the SAR remote sensing image after encoding output by step two and the noisy optical image y j in the time step j noise domain into the backward diffusion network as input, taking the noise added in the noisy optical image y j in the time step j noise domain as output of the backward diffusion network, training the backward diffusion network until the loss function L p converges, and obtaining the trained backward diffusion network;

[0017] Step seven, inputting the to-be-detected SAR remote sensing image into the encoder of the pre-trained variational autoencoder, and the encoder outputs a feature vector of the SAR remote sensing image after encoding;

[0018] Gaussian noise is added to the encoded feature vector of the output SAR remote sensing image to generate a noisy SAR remote sensing image x in the noise domain t ;

[0019] The noisy SAR remote sensing image x is input into the trained modal mapping network, and the modal mapping network outputs a noisy optical image t The trained modal mapping network is input into the trained modal mapping network, and the modal mapping network outputs a noisy optical image

[0020] The noisy optical image is input into the trained reverse diffusion network, and the trained reverse diffusion network outputs a noisy optical image The trained reverse diffusion network is input into the trained reverse diffusion network, and the trained reverse diffusion network outputs a noisy optical image Noise added in the current time step.

[0021] The real non-noisy optical image y0 is calculated based on the noise step by step,

[0022] The real non-noisy optical image y0 is input into the decoder of the pre-trained variational autoencoder to obtain the feature space of the decoded optical image.

[0023] The beneficial effects of the present application are:

[0024] The present application proposes a SAR-to-optical image translation method based on diffusion modeling and wavelet enhancement (Diffusion-Driven Wavelet-enhanced SAR-to-Optical Translation Framework, DDW-SOT), which aims to improve the interpretability and application value of SAR images, and to alleviate the low translation accuracy and poor generated image quality caused by the large difference in feature distribution between SAR and optical images, thereby promoting the effective use and application expansion of remote sensing data in complex environments. The technical innovation of the present application lies in constructing a generation framework combining diffusion process and frequency domain information guidance, which effectively improves the accuracy and visual quality of cross-modal image translation.

[0025] Specifically, the present application breaks through the bottleneck of insufficient modeling of cross-modal feature alignment in existing SAR-to-optical methods, and introduces the intermediate state in the forward diffusion process to guide the feature to gradually migrate from the SAR domain to the optical domain, significantly reducing the learning difficulty and improving the semantic consistency. At the same time, a conditional modulator is proposed, which can adaptively normalize and modulate the SAR image pixel by pixel, effectively enhancing the fine-grained semantic expression of the image and ensuring the consistency of the generated image with the original scene in the semantic level. In addition, the present application designs a wavelet-enhanced feature alignment mechanism, which uses wavelet transform to separate high and low frequency information, respectively optimizes the low frequency alignment and high frequency detail reconstruction between modalities, and further improves the structure restoration ability and visual realism of the generated image in complex scenes.

[0026] The application provides a feasible and efficient path for solving the problem that SAR images are difficult to understand intuitively, expands the application boundary of the diffusion model in the remote sensing cross-modal generation task, and provides technical support for further research and practical application of remote sensing image translation technology. BRIEF DESCRIPTION OF DRAWINGS

[0027] Figure 1 The SAR-to-optical image conversion method based on diffusion driving and wavelet enhancement is a general flowchart;

[0028] Figure 2 The structure diagram of the noise domain modal mapping network is shown in the figure;

[0029] Figure 3 The structure diagram of the reverse diffusion network is shown in the figure;

[0030] Figure 4 The structure diagram of the residual block is shown in the figure;

[0031] Figure 5 The structure diagram of the wavelet down-sampling block is shown in the figure;

[0032] Figure 6 The structure diagram of the wavelet up-sampling block is shown in the figure;

[0033] Figure 7 The structure diagram of the conditional modulator module is shown in the figure. DETAILED DESCRIPTION

[0034] Embodiment one: the specific process of the synthetic aperture radar-to-optical image conversion method based on diffusion driving and wavelet enhancement is as follows:

[0035] Step one: obtain SAR remote sensing and optical paired images of the same target area, and construct a training sample pair;

[0036] Step two: input the SAR remote sensing image in the training sample pair into the encoder of the pre-trained variational autoencoder (VAE), and the encoder outputs the feature vector x0 of the SAR remote sensing image after encoding; map the training sample pair to a low-dimensional feature space;

[0037] Input the optical image in the training sample pair into the encoder of the pre-trained variational autoencoder (VAE), and the encoder outputs the feature vector y0 of the optical image after encoding; map the training sample pair to a low-dimensional feature space;

[0038] Step three: add the same Gaussian noise to the feature vector x0 of the SAR remote sensing image after encoding output in step two and the feature vector y0 of the optical image after encoding, respectively, to generate the noisy SAR remote sensing image x t and the noisy optical image y t in the time step t noise domain.

[0039] Step four, constructing a modal mapping network;

[0040] The noisy SAR remote sensing image x t The noisy SAR remote sensing image x t The noisy SAR remote sensing image x G The modal mapping network is trained until the loss function L

[0041] SAR images are converted into optical images, which are easier for people to understand.

[0042] Step five, the feature vector x0 of the SAR remote sensing image output by step two, the noisy optical image y j , and the noisy optical image y j The Gaussian noise added in the noisy optical image y

[0043] j = 1, 2, …, t; t = 1, 2, …, T;

[0044] T represents the total number of time steps;

[0045] Step six, constructing a backward diffusion network;

[0046] The feature vector x0 of the SAR remote sensing image output by step two and the noisy optical image y j The feature vector x0 of the SAR remote sensing image output by step two and the noisy optical image y j The noisy optical image y p The backward diffusion network is trained until the loss function L

[0047] Step seven, inputting the SAR remote sensing image to be tested into the encoder of the pre-trained variational autoencoder (VAE), and the encoder outputs the feature vector of the SAR remote sensing image after encoding;

[0048] Adding Gaussian noise to the output feature vector of the SAR remote sensing image after encoding to generate a noisy SAR remote sensing image x t ;

[0049] Inputting the noisy SAR remote sensing image x t into the trained modal mapping network, and the modal mapping network outputs the noisy optical image

[0050] a noisy optical image and a real noise-free SAR remote sensing image x0input into the trained reverse diffusion network, and the trained reverse diffusion network outputs a noisy optical image noise added in the current time step;

[0051] Based on the noise, the real noise-free optical image y0is calculated step by step,

[0052] The real noise-free optical image y0is input into the decoder of the pre-trained variational autoencoder (VAE) to obtain the feature space of the decoded optical image.

[0053] The method divides the modal conversion task into two collaborative sub-tasks by utilizing the intermediate state in the forward diffusion process: noise domain modal mapping and reverse diffusion modeling, gradually aligns the feature distribution, reduces the learning complexity, and improves the conversion accuracy.

[0054] Specific implementation method two: the difference between this implementation method and the specific implementation method one is that in step three, the same Gaussian noise is added to the SAR remote sensing image encoded feature vector and the optical image encoded feature vector output by step two, respectively, to generate the noisy SAR remote sensing and noisy optical paired image in the time step t noise domain.

[0055] The specific process is as follows:

[0056] T: In the diffusion model, the diffusion process is usually discretized into T time steps, representing the total number of steps in the process of gradually adding Gaussian noise from the original data until it approaches the original Gaussian distribution. Each step corresponds to a stage of noise disturbance, and the more steps, the smoother the noise addition and removal process. In this paper, the total time step length T = 1000 is set for the process of gradually adding Gaussian noise from the original data until it approaches the original Gaussian distribution, which is widely used in existing research and has achieved a good balance between generation quality and computational efficiency.

[0057] The total time step length T = 1000 is set;

[0058] y T represents the optical image at time step T = 1000, which can be considered as a pure noise image conforming to the Gaussian distribution;

[0059] x T represents the SAR remote sensing image at time step T = 1000;

[0060] t represents a predefined time step for balancing the difficulty of the modal mapping network and the reconstruction quality of the backward diffusion network. The selection of the time step t plays a key role in balancing the computational complexity of the modal mapping network and the backward diffusion network. A larger t value means that the input SAR feature has been significantly added with noise, and the difference between different modalities is effectively suppressed, so that the modal mapping network is easier to map from the SAR feature domain with higher degradation to the corresponding optical feature domain. However, at this time, the noise has dominated in the feature space, limiting the effective information that the backward diffusion network can use, and thus affecting the details of the final reconstructed image. On the contrary, under a smaller t value, the intermediate state retains more original modal differences, resulting in a more challenging modal mapping task and increasing the risk of instability in network training;

[0061] The selection criteria of t are as follows: in order to achieve a reasonable balance between task difficulty and alleviate the problems of mode collapse and cumulative reasoning errors, a structural similarity index (SSIM) is introduced as a measurement tool to measure the cross-modal difference, and is used to guide the selection of the diffusion time step t. Specifically, we calculate the SSIM distance between (x t ,y t ) and (y0,y t ) under different t values, and take the distance between the two as close as possible as the selection of the balance point. This balance point indicates that the difficulty of the modal mapping task at the current time step is moderate, and the backward diffusion process can still retain enough semantic information for high-quality reconstruction.

[0062] The selection criteria of t are as follows:

[0063] Calculate the SSIM distance of (x t ,y t ), t = 1, 2, …, T;

[0064] Calculate the SSIM distance of (y0,y t );

[0065] Select t of (x t ,y t ) corresponding to the minimum value of the SSIM distance of (x t ,y t ) and (y0,y t ) as the selected t;

[0066] The process of obtaining (x t ,y t ) is as follows:

[0067]

[0068] where ∈ t represents Gaussian noise, represents a normal distribution, 0 represents a mean value, and I represents a variance;

[0069] α i is an intermediate variable, is an intermediate variable, and α t = 1-β t , β t is a constant, representing the noise intensity at the t time step.

[0070] y t predicted by the modal mapping network in the reasoning stage

[0071] y j : y j represents any intermediate time step between y0and y t during the diffusion process, where 0 < j < t. In Figure 1 , the purposes of y j and y j-1 are to illustrate that the reverse diffusion model gradually predicts the state of the previous time step to achieve a gradual restoration process from noise to clear images.

[0072] The other steps and parameters are the same as in the first embodiment.

[0073] In this embodiment, the modal mapping network is constructed in step four, which is different from the first or second embodiment. The specific process is as follows:

[0074] The modal mapping network includes a first residual block, a second residual block, a third residual block, a first 1x1 convolutional layer, a second 1x1 convolutional layer, a third 1x1 convolutional layer, a fourth residual block, a fifth residual block, and a sixth residual block.

[0075] The working process of the modal mapping network is as follows:

[0076] The noisy SAR remote sensing image x t in the noise domain generated in step three is input into the first residual block, and the first residual block outputs a feature A;

[0077] The feature A is input into the second residual block, and the second residual block outputs a feature B;

[0078] The feature B is input into the third residual block, and the third residual block outputs a feature C;

[0079] The feature A is input into the first 1x1 convolutional layer, and the first 1x1 convolutional layer outputs a feature D;

[0080] The feature B is input into the second 1x1 convolutional layer, and the second 1x1 convolutional layer outputs a feature E;

[0081] The feature C is input into a third 1x1 convolutional layer, and the third 1x1 convolutional layer outputs a feature F;

[0082] The feature F and random Gaussian noise are concatenated and fused, and then input into a fourth residual block, and the fourth residual block outputs a feature G;

[0083] The feature G and the feature E are concatenated and fused, and then input into a fifth residual block, and the fifth residual block outputs a feature H;

[0084] The feature H and the feature D are concatenated and fused, and then input into a sixth residual block, and the sixth residual block outputs a feature I;

[0085] The feature I output by the sixth residual block is taken as the output of the modal mapping network, and the output of the modal mapping network is x t The noisy optical image y t .

[0086] The other steps and parameters are the same as those in the first or second embodiment.

[0087] In the fourth embodiment, the first residual block, the second residual block, the third residual block, the fourth residual block, the fifth residual block, and the sixth residual block each include a fourth 3x3 convolutional layer, a first BN layer, a first ReLU activation function layer, a fifth 3x3 convolutional layer, a second BN layer, and a second ReLU activation function layer.

[0088] The working process of each residual block is as follows:

[0089] The feature 1 is sequentially input into the fourth 3x3 convolutional layer, the first BN layer, and the first ReLU activation function layer, and the first ReLU activation function layer outputs a feature 2;

[0090] The feature 2 is sequentially input into the fifth 3x3 convolutional layer and the second BN layer, and the second BN layer outputs a feature 3;

[0091] The feature 1 and the feature 3 are added element by element to obtain a feature 4;

[0092] The feature 4 is input into the second ReLU activation function layer, and the second ReLU activation function layer outputs a feature 5;

[0093] The feature 5 is the output feature of each residual block.

[0094] The other steps and parameters are the same as those in the first to third embodiments.

[0095] In the fifth embodiment, the inverse diffusion network includes:

[0096] the first conditional down-sampling block, the second conditional down-sampling block, the third conditional down-sampling block, the sixth 1x1 convolutional layer, the seventh residual block, the first conditional modulator, the eighth residual block, the second conditional modulator, the first wavelet down-sampling block, the ninth residual block, the third conditional modulator, the tenth residual block, the fourth conditional modulator, the second wavelet down-sampling block, the eleventh residual block, the fifth conditional modulator, the first wavelet up-sampling block, the twelfth residual block, the sixth conditional modulator, the thirteenth residual block, the seventh conditional modulator, the second wavelet up-sampling block, the fourteenth residual block, the eighth conditional modulator, the fifteenth residual block, the ninth conditional modulator, the seventh 1x1 convolutional layer;

[0097] The specific working process of the reverse diffusion network is as follows:

[0098] The real noise-free SAR image x0 is input into the first conditional down-sampling block, and the first conditional down-sampling block outputs a feature J;

[0099] The feature J is input into the second conditional down-sampling block, and the second conditional down-sampling block outputs a feature

[0100] The feature is input into the third conditional down-sampling block, and the third conditional down-sampling block outputs a feature L;

[0101] The noisy optical image y j is input into the sixth 1x1 convolutional layer, and the sixth 1x1 convolutional layer outputs a feature M;

[0102] The feature M is input into the seventh residual block, and the seventh residual block outputs a feature N;

[0103] The feature J and the feature N are input into the first conditional modulator, and the first conditional modulator outputs a feature O;

[0104] The feature O is input into the eighth residual block, and the eighth residual block outputs a feature P;

[0105] The feature J and the feature P are input into the second conditional modulator, and the second conditional modulator outputs a feature

[0106] The feature is input into the first wavelet down-sampling block, and the first wavelet down-sampling block outputs a feature R;

[0107] The feature R is input into the ninth residual block, and the ninth residual block outputs a feature S;

[0108] The feature S and the feature are input into the third conditional modulator, and the third conditional modulator outputs a feature T;

[0109] The feature T is input into the tenth residual block, and the tenth residual block outputs a feature U;

[0110] feature U and feature a fourth conditional modulator, the fourth conditional modulator outputting a feature

[0111] feature a second wavelet down-sampling block, the second wavelet down-sampling block outputting a feature W;

[0112] feature W inputting an eleventh residual block, the eleventh residual block outputting a feature X;

[0113] feature L and feature X inputting a fifth conditional modulator, the fifth conditional modulator outputting a feature Y;

[0114] feature Y and the second wavelet down-sampling block outputting a high-frequency component inputting a first wavelet up-sampling block, the first wavelet up-sampling block outputting a feature A';

[0115] feature A' inputting a twelfth residual block, the twelfth residual block outputting a feature B';

[0116] feature B' and the second conditional down-sampling block outputting a feature a sixth conditional modulator, the sixth conditional modulator outputting a feature C';

[0117] feature C' inputting a thirteenth residual block, the thirteenth residual block outputting a feature D';

[0118] feature D' and the second conditional down-sampling block outputting a feature a seventh conditional modulator, the seventh conditional modulator outputting a feature E';

[0119] feature E' and the first wavelet down-sampling block outputting a high-frequency component inputting a second wavelet up-sampling block, the second wavelet up-sampling block outputting a feature F';

[0120] feature F' inputting a fourteenth residual block, the fourteenth residual block outputting a feature G';

[0121] feature G' and the first conditional down-sampling block outputting a feature J inputting an eighth conditional modulator, the eighth conditional modulator outputting a feature H';

[0122] feature H' inputting a fifteenth residual block, the fifteenth residual block outputting a feature I';

[0123] feature I' and the first conditional down-sampling block outputting a feature J inputting a ninth conditional modulator, the ninth conditional modulator outputting a feature J';

[0124] feature J' inputting a seventh 1x1 convolutional layer, the seventh 1x1 convolutional layer outputting a feature K';

[0125] feature K' as a noisy optical image y in a noise domain at a time step j outputted by the inverse diffusion networkj noise added in the middle.

[0126] The other steps and parameters are the same as one of the first to fourth embodiments.

[0127] The seventh embodiment is different from one of the first to sixth embodiments in that each of the first conditional down-sampling block, the second conditional down-sampling block, and the third conditional down-sampling block is a convolutional layer with a step of 2.

[0128] The other steps and parameters are the same as one of the first to fifth embodiments.

[0129] The seventh embodiment is different from one of the first to sixth embodiments in that each of the first conditional down-sampling block, the second conditional down-sampling block, and the third conditional down-sampling block is a convolutional layer with a step of 2.

[0130] an eighth 1x1 convolutional layer, a first deep convolutional layer, a second deep convolutional layer, and a first group of normalization layers;

[0131] The working processes of the first conditional modulator, the second conditional modulator, the third conditional modulator, the fourth conditional modulator, the fifth conditional modulator, the sixth conditional modulator, the seventh conditional modulator, the eighth conditional modulator, and the ninth conditional modulator are the same;

[0132] The working process of the first conditional modulator is as follows:

[0133] The feature J is input into the eighth 1x1 convolutional layer, and the eighth 1x1 convolutional layer outputs a feature

[0134] The feature is input into the first deep convolutional layer, and the first deep convolutional layer outputs a feature

[0135] The feature is input into the second deep convolutional layer, and the first deep convolutional layer outputs a feature

[0136] The feature N output by the seventh residual block is input into the first group of normalization layers, and the first group of normalization layers outputs a feature

[0137] The feature a is multiplied element by element with the feature γ to obtain a feature

[0138] The feature δ is added element by element with the feature β to obtain a feature

[0139] The feature a is added element by element with the feature θ to obtain a feature

[0140] The feature θ' is the output of the conditional modulator.

[0141] Other steps and parameters are the same as one of embodiments one to six.

[0142] Embodiment eight: different from one of embodiments one to seven, the first wavelet downsampling block and each wavelet downsampling block in the second wavelet downsampling block comprise:

[0143] the second set of normalization layers, the first discrete wavelet transform, the third set of normalization layers, the third deep convolutional layer, the self-attention layer, the point-wise convolutional layer, the cross-attention layer, the ninth convolutional layer, the tenth convolutional layer, the second discrete wavelet transform;

[0144] The working process of the first wavelet downsampling block is as follows:

[0145] The feature is input into the second set of normalization layers, and the second set of normalization layers outputs the feature

[0146] The feature is input into the first discrete wavelet transform for processing, and the first discrete wavelet transform decomposes the input feature into a low-frequency component and a high-frequency component

[0147] The low-frequency component is input into the third set of normalization layers, and the third set of normalization layers outputs the feature

[0148] The high-frequency component is input into the third deep convolutional layer, and the third deep convolutional layer outputs the feature

[0149] The feature is input into the self-attention layer, and the self-attention layer outputs the feature Q;

[0150] The feature is input into the point-wise convolutional layer, and the point-wise convolutional layer outputs the features K and V;

[0151] The features Q, K and V are input into the cross-attention layer, and the cross-attention layer outputs the feature

[0152] The feature is spliced with the time encoding to obtain the feature

[0153] The feature is input into the ninth convolutional layer, and the ninth convolutional layer outputs the feature

[0154] The feature inputted into a second discrete wavelet transform, the second discrete wavelet transform decomposes the inputted feature into a low-frequency component and a high-frequency component

[0155] The low-frequency component is inputted into a tenth convolutional layer, the tenth convolutional layer outputs a feature

[0156] The feature and the feature are element-wise added to obtain a feature R;

[0157] The feature R is outputted as a feature of the first wavelet down-sampling block;

[0158] The working process of the second wavelet down-sampling block is as follows:

[0159] The feature is inputted into a second group of normalization layers, the second group of normalization layers outputs a feature

[0160] The feature is inputted into a first discrete wavelet transform for processing, the first discrete wavelet transform decomposes the inputted feature into a low-frequency component and a high-frequency component

[0161] The low-frequency component is inputted into a third group of normalization layers, the third group of normalization layers outputs a feature

[0162] The high-frequency component is inputted into a third deep convolutional layer, the third deep convolutional layer outputs a feature

[0163] The feature is inputted into a self-attention layer, the self-attention layer outputs a feature Q;

[0164] The feature is inputted into a point-wise convolutional layer, the point-wise convolutional layer outputs a feature K and a feature V;

[0165] The feature Q, the feature K and the feature V are inputted into a cross-attention layer, the cross-attention layer outputs a feature

[0166] The feature is spliced with a time encoding to obtain a feature

[0167] The feature is inputted into a ninth convolutional layer, the ninth convolutional layer outputs a feature

[0168] The feature is input into a second discrete wavelet transform for processing, and the second discrete wavelet transform decomposes the input feature into a low-frequency component and a high-frequency component

[0169] The low-frequency component is input into a tenth convolutional layer, and the tenth convolutional layer outputs a feature

[0170] The feature and the feature are element-wise added to obtain a feature W.

[0171] The feature W is taken as an output feature of the second wavelet down-sampling block.

[0172] The other steps and parameters are the same as one of the first to seventh embodiments.

[0173] The ninth embodiment is different from one of the first to eighth embodiments in that each wavelet up-sampling block in the first wavelet up-sampling block and the second wavelet up-sampling block comprises:

[0174] a fourth set of normalization layers, a wavelet inverse transformation layer, an eleventh convolutional layer, a twelfth convolutional layer, a pixel recombination layer, and a thirteenth convolutional layer.

[0175] The working process of the first wavelet up-sampling block is as follows:

[0176] The feature Y is input into the fourth set of normalization layers, and the fourth set of normalization layers outputs a feature Y1.

[0177] The feature Y1 and the high-frequency component input into the wavelet inverse transformation layer, and the wavelet inverse transformation layer outputs a feature Y2.

[0178] The feature Y2 is input into the eleventh convolutional layer, and the eleventh convolutional layer outputs a feature Y3.

[0179] The feature Y1 is input into the twelfth convolutional layer, and the twelfth convolutional layer outputs a feature Y4.

[0180] The feature Y4 is subjected to pixel recombination to obtain a feature Y5.

[0181] The feature Y3 and the feature Y5 are concatenated by Concat to obtain a feature Y6.

[0182] The feature Y6 and the time encoding are element-wise added to obtain a feature Y7.

[0183] The feature Y7 is input into the thirteenth convolutional layer, and the thirteenth convolutional layer outputs a feature A';

[0184] The feature A' is the output feature of the first wavelet up-sampling block;

[0185] The working process of the second wavelet up-sampling block is as follows:

[0186] The feature E' is input into the fourth group of normalization layers, and the fourth group of normalization layers outputs a feature E1';

[0187] The feature E1' is concatenated with the high-frequency component of the first discrete wavelet transform decomposition in the first wavelet down-sampling block to obtain a feature E2'; The feature E2' is input into the wavelet inverse transformation layer, and the wavelet inverse transformation layer outputs a feature E'2;

[0188] The feature E'2 is input into the eleventh convolutional layer, and the eleventh convolutional layer outputs a feature E'3;

[0189] The feature E1' is input into the twelfth convolutional layer, and the twelfth convolutional layer outputs a feature E'4;

[0190] The feature E'4 is subjected to pixel shuffle to obtain a feature E'5;

[0191] The feature E'3 is concatenated with the feature E'5 to obtain a feature E'6;

[0192] The feature E'6 is element-wise added with the time encoding to obtain a feature E'7;

[0193] The feature E'7 is input into the thirteenth convolutional layer, and the thirteenth convolutional layer outputs a feature F';

[0194] The feature F' is the output feature of the second wavelet up-sampling block.

[0195] The other steps and parameters are the same as one of the first to eighth embodiments.

[0196] The tenth embodiment is different from one of the first to ninth embodiments in that the training of the modal mapping network adopts a GAN form, and the training target is represented as:

[0197]

[0198] The first term is used to evaluate the authenticity of the generated image, and the second term is used to evaluate the similarity between the generated image and the real image;

[0199] The total loss function of the generator is represented as:

[0200] The input sample xt The expected value of the output, which represents the average of the outputs of all samples;

[0201] The output of the generator network G, parameterized by

[0202] The discriminative result of the generator output by the discriminator D, representing the probability that the generated sample is judged to be real;

[0203] The expected value of the output of the input sample pair (x t ,y t );

[0204] y t The true noisy optical image feature at the t-th time step, which is used as a reference image for supervised learning during training;

[0205] || ||1 represents the L1 norm, which is used to measure the pixel-level difference between two noisy images;

[0206] The training of the back diffusion network uses the L1 loss function, and the training target is represented as:

[0207]

[0208] Where p θ represents the trained model, and ∈ represents random Gaussian noise;

[0209] The joint expectation of the output of the input SAR image x0, the true optical image y0, the random Gaussian noise ∈, and the time step j, which represents the average of the outputs of the samples;

[0210] The output from time step 1 to time step j is represented as α j =1-β j ,β j is a constant, and β j represents the noise intensity of time step j; α i represents an intermediate variable;

[0211] || ||1 represents the L1 norm, which is used to measure the pixel-level difference between two noisy images.

[0212] The other steps and parameters are the same as one of the first to ninth embodiments.

[0213] Figure 1The core process of the synthetic aperture radar (SAR) to optical image conversion (DDW-SOT) framework based on diffusion driving and wavelet enhancement proposed in the application is illustrated. In order to improve the calculation efficiency, the method is completed in the latent feature space in the training stage and the inference stage. Specifically, in the training stage, first, the paired SAR images and optical images are respectively mapped to the low-dimensional latent space through the pre-trained variational autoencoder (VAE) encoder; then, the model training is carried out in the latent space, and the training process is divided into two cooperative sub-tasks: noise domain modal mapping and reverse diffusion modeling, which are respectively optimized by using independent objective functions.

[0214] The encoder and decoder modules in the pre-trained variational autoencoder (VAE) are respectively composed of 3 cascaded down-sampling convolutions and 3 cascaded up-sampling convolutions;

[0215] The pre-trained variational autoencoder is introduced to compress the features of the input image, aiming to map the high-dimensional image data to a continuous and controllable low-dimensional latent space, so as to reduce the modeling complexity, improve the stability and efficiency of cross-modal feature conversion, and enhance the structural consistency and expression accuracy of the generated image while keeping the integrity of the key information.

[0216] In the inference stage, the SAR image to be converted is first mapped to the latent space through the same VAE encoder; then, the forward diffusion process is used to convert it to the Gaussian noise domain, and the trained noise domain modal mapping network is used to convert its features to the corresponding noise domain optical features; then, the trained reverse diffusion network is used to denoise it step by step to restore the clear optical features; finally, the pre-trained VAE decoder is used to map the features back to the pixel space to generate the final optical image.

[0217] Figure 2 The noise domain modal mapping network structure used in the application is shown. The network is designed to be lightweight as a whole, mainly composed of a condition injection branch and a generator branch. Among them, the condition injection branch is composed of three cascaded residual blocks and down-sampling modules, which are used to extract multi-scale conditional features from the input SAR image; the generator branch is composed of three cascaded residual blocks and up-sampling modules, which are used to generate the corresponding noise domain optical features step by step. In the feature fusion process, the features of the same scale from the two branches are fused by splicing to realize the effective integration of cross-modal information. The training of the noise domain modal mapping network adopts the form of GAN, and the training target can be expressed as:

[0218]

[0219] The first term is used to evaluate the authenticity of the generated image, and the second term is used to evaluate the similarity between the generated image and the real image.

[0220] The wavelet scale diffusion UNet model (WSD-UNet) proposed in this paper is as follows Figure 3 As shown in Figure 2, this model innovatively expands upon the traditional UNet architecture, employing a lightweight encoder-decoder structure designed specifically for low-resolution latent spaces. It comprises two encoding layers, two decoding layers, and a bottleneck module based on residual blocks, optimizing computational efficiency while maintaining feature integrity. The core of this innovation lies in two components integrated within the multi-scale network: the wavelet upsampling block significantly improves the distribution consistency between SAR and optical features by eliminating low-frequency feature differences and optimizing high-frequency details, thereby achieving superior SAR-to-optical (S2O) image conversion accuracy and visual quality in complex scenarios; and the conditional modulator (CPM) enhances precise control of optical image reconstruction through adaptive feature modulation. These two core components are described in detail below.

[0221] like Figure 4 As shown, the wavelet downsampling block uses discrete wavelet transform (DWT) to decompose the input features into low-frequency features F low and high-frequency characteristic components F high The low-frequency features capture global structural information and use the self-attention mechanism (SA) to extract long-range dependencies, while the high-frequency features retain fine-grained details and use deep convolution (DWConv) and point-wise convolution (PWConv) to enhance feature representation. This process can be expressed as follows:

[0222] F low ,F high =DWT(GroupNorm(F in ))

[0223] Q=SA(GroupNorm(F low ))

[0224] K,V=PWConv(DWConv(F high ))

[0225] Subsequently, the Cross Attention (CA) mechanism is implemented by using F low As query (Q) and F high Enhance F as key and value (K, V) low Represents that the high-frequency details are integrated into the low-frequency context. The resulting features are then refined through temporal embedding and convolution, and combined with the original low-frequency components through residual connections to obtain the final output F out , the formula is:

[0226] F refine =MLP(sin(t))+CA(Q,K,V)

[0227] F out=Conv(F refine )+Conv(DWT(F in ))

[0228] Among them, MLP(sin(t)) represents the time embedding representation obtained by encoding the time parameter with a sine function and then inputting it into a multi-layer perceptron for processing.

[0229] In order to improve the feature recovery effect in the decoding stage, the wavelet upsampling module is as follows Figure 5 As shown in Figure 1, it can fuse low-frequency information and high-frequency details from the UNet encoder stage to ensure the integrity and structural fidelity of the generated image. Specifically, the input features are first processed by group normalization (GroupNorm), and then feature enhancement is performed in two parallel branches: the first branch upsamples the low-frequency features by pixel reorganization; the second branch uses the inverse wavelet transform (IWT) to introduce the high-frequency features F output by the encoder. high The output features of the two branches are then concatenated and further fused and optimized by combining temporal embedding information with convolution operations, thereby maintaining both structural integrity and detail clarity in the generated image.

[0230] The structure of the proposed conditional modulator is shown in Figure 6 As shown in the figure, the input features are first processed by a residual block, and then normalized features are obtained by applying group normalization on the channel dimension. The SAR image is then used as the modulation condition input, and the normalized features are dynamically modulated pixel by pixel. The specific steps are as follows: the SAR conditional features are first dimensionally adjusted through a 1×1 convolution, and then input into two depthwise convolutions to generate the scaling coefficient parameter γ and the offset coefficient parameter β, which are used to modulate the normalized features pixel by pixel, thus achieving accurate feature generation under the condition guidance.

[0231] For the training of the back diffusion network, the L1 loss function is used, and its training objective can be expressed as:

[0232]

[0233] where p θ represents the trained model, and ∈ represents the predicted noise.

[0234] The following examples are used to verify the beneficial effects of the present invention:

[0235] The beneficial effects of the present application are verified on the SEN12MS dataset, which contains 180,662 pairs of optical-SAR image pairs, a spatial resolution of 10 meters, covers the global human-inhabited continental area, and covers the observation data of four seasons. Among them, the optical image data comes from the Sentinel-2 satellite, containing 13 spectral bands; the SAR image data comes from the Sentinel-1 satellite, containing VH and VV two polarization modes. In this paper, three bands of RGB are selected from the optical image, and two polarization channels of VH and VV are selected from the SAR image. In order to consider the seasonal variation factor, we randomly sample 4,500 pairs of images in each season, and finally build a test set containing 18,000 pairs of images, and the remaining non-overlapping data samples are used for model training.

[0236] The comparison results of the proposed DDW-SOT model and the comparison method are shown in Table 1, and it can be seen that the proposed DDW-SOT model has a more obvious leading in each index compared with the comparison model;

[0237] Table 1 SAR-to-optical image reconstruction results of SEN12MS dataset

[0238]

[0239] In summary, in order to improve the interpretability and application value of remote sensing SAR images, the present embodiment provides a synthetic aperture radar image to optical image conversion method and system, which can effectively realize the visual reconstruction and semantic enhancement of SAR images, thereby providing more intuitive and easy-to-interpret image information, and significantly improving the utilization efficiency of data and the task execution effect.

[0240] The present application can also have other various embodiments, and those skilled in the art can make various corresponding changes and modifications according to the present application without departing from the spirit and essence of the present application. However, these corresponding changes and modifications should all belong to the protection scope of the claims attached to the present application.

Claims

1. A method for converting synthetic aperture radar to optical image based on diffusion drive and wavelet enhancement, characterized by: The specific process of the method is: Step 1: Obtain SAR remote sensing and optical paired images of the same target area to construct training sample pairs; Step 2: Input the SAR remote sensing image in the training sample into the encoder of the pre-trained variational autoencoder, and the encoder outputs the feature vector x0 after encoding the SAR remote sensing image; Input the optical image in the training sample pair into the encoder of the pre-trained variational autoencoder, and the encoder outputs the feature vector y0 after encoding the optical image; Step 3: Add the same Gaussian noise to the eigenvector x0 after encoding the SAR remote sensing image and the eigenvector y0 after encoding the optical image output in step 2, and generate the noisy SAR remote sensing image x in the noise domain at time step t. t and noisy optical image y t ; Step 4: Construct a modal mapping network; The noisy SAR remote sensing image x generated in step 3 under the noise domain at time step t t As the input of the modality mapping network, the noisy optical image y generated in the noise domain at time step t in step 3 is t As the output of the modality mapping network, the modality mapping network is trained until the loss function L G Converge and obtain a trained modality mapping network; Step 5: Encode the SAR remote sensing image output in step 2, and the noisy optical image y in the noise domain at time step j. j , and the noisy optical image y in the noise domain at time step j j The Gaussian noise added in is used as training data for the back diffusion network; j=1,2,…,t; t=1,2,…,T; T represents the total time step; Step 6: Construct a reverse diffusion network; The SAR remote sensing image encoded feature vector x0 output in step 2 and the noisy optical image y in the noise domain at time step j are j As the input of the back diffusion network, the noisy optical image y in the noise domain at time step j j The noise added in is used as the output of the back diffusion network, and the back diffusion network is trained until the loss function L p Converge and obtain a trained back diffusion network; Step 7: Input the SAR remote sensing image to be measured into the encoder of the pre-trained variational autoencoder, and the encoder outputs the feature vector of the encoded SAR remote sensing image; Add Gaussian noise to the eigenvector of the output SAR remote sensing image after encoding to generate a noisy SAR remote sensing image x in the noise domain t ; The noisy SAR remote sensing image x t Input the trained modal mapping network, and the modal mapping network outputs a noisy optical image Noisy optical images The trained back diffusion network is input with the real noise-free SAR remote sensing image x0, and the trained back diffusion network outputs a noisy optical image The noise added in the current time step; Based on the noise, the real noise-free optical image y0 is calculated step by step. The real noise-free optical image y0 is input into the decoder of the pre-trained variational autoencoder to obtain the feature space of the decoded optical image.

2. The method for converting synthetic aperture radar to optical image based on diffusion drive and wavelet enhancement according to claim 1, characterized in that: In the step 3, the same Gaussian noise is added to the encoded eigenvectors of the SAR remote sensing image and the encoded eigenvectors of the optical image output in the step 2, respectively, to generate noisy SAR remote sensing and noisy optical paired images in the noise domain at time step t. The specific process is as follows: Set the total time step T = 1000; y T represents the optical image when the time step T is 1000; x T represents the SAR remote sensing image when the time step T is 1000; t represents the predefined time step; The selection criteria for t are as follows: Calculate (x t ,y t )’s SSIM distance, t = 1, 2, …, T; Calculate (y0,y t )’s SSIM distance; Select (x t ,y t ) and the SSIM distance of (y0,y t ) corresponds to the minimum SSIM distance of (x t ,y t ) as the selected t; (x t ,y t ) is obtained as follows: Among them, ∈ t represents Gaussian noise, represents normal distribution, 0 represents mean, and I represents variance; α i is an intermediate variable, is the intermediate variable, α t =1-β t , β t Represents the noise intensity at time step t.

3. The method for converting synthetic aperture radar to optical image based on diffusion drive and wavelet enhancement according to claim 2, characterized in that: In step 4, a modal mapping network is constructed; the specific process is as follows: The modality mapping network includes a first residual block, a second residual block, a third residual block, a first 1×1 convolutional layer, a second 1×1 convolutional layer, a third 1×1 convolutional layer, a fourth residual block, a fifth residual block, and a sixth residual block; The working process of the modal mapping network is: The noisy SAR remote sensing image x generated in the noise domain in step 3 t Input the first residual block, and the first residual block outputs feature A; Input feature A into the second residual block, and the second residual block outputs feature B; Input feature B into the third residual block, and the third residual block outputs feature C; Input feature A into the first 1×1 convolutional layer, and the first 1×1 convolutional layer outputs feature D; Input feature B into the second 1×1 convolutional layer, and the second 1×1 convolutional layer outputs feature E; Input feature C into the third 1×1 convolutional layer, and the third 1×1 convolutional layer outputs feature F; The feature F and random Gaussian noise are concatenated and input into the fourth residual block, which outputs the feature G; Concat the features G and E and input them into the fifth residual block, which outputs the feature H. Concat the features H and D and input them into the sixth residual block, which outputs the feature I. The sixth residual block outputs feature I as the output of the modality mapping network, and the output of the modality mapping network is t The corresponding noisy optical image y in the noise domain t .

4. The method for converting synthetic aperture radar to optical image based on diffusion drive and wavelet enhancement according to claim 3, characterized in that: Each of the first residual block, the second residual block, the third residual block, the fourth residual block, the fifth residual block, and the sixth residual block includes a fourth 3×3 convolutional layer, a first BN layer, a first ReLU activation function layer, a fifth 3×3 convolutional layer, a second BN layer, and a second ReLU activation function layer; The working process of each residual block is as follows: Feature 1 is sequentially input into the fourth 3×3 convolutional layer, the first BN layer, and the first ReLU activation function layer. The first ReLU activation function layer outputs feature 2. Feature 2 is input into the fifth 3×3 convolutional layer and the second BN layer in sequence, and the second BN layer outputs feature 3; Add feature 1 and feature 3 element by element to get feature 4; Feature 4 is input into the second ReLU activation function layer, and the second ReLU activation function layer outputs feature 5; Feature 5 is the output feature of each residual block.

5. The method for converting synthetic aperture radar to optical image based on diffusion drive and wavelet enhancement according to claim 4, characterized in that: The reverse diffusion network comprises: first conditional downsampling block, second conditional downsampling block, third conditional downsampling block, sixth 1×1 convolutional layer, seventh residual block, first conditional modulator, eighth residual block, second conditional modulator, first wavelet downsampling block, ninth residual block, third conditional modulator, tenth residual block, fourth conditional modulator, second wavelet downsampling block, eleventh residual block, fifth conditional modulator, first wavelet upsampling block, twelfth residual block, sixth conditional modulator, thirteenth residual block, seventh conditional modulator, second wavelet upsampling block, fourteenth residual block, eighth conditional modulator, fifteenth residual block, ninth conditional modulator, seventh 1×1 convolutional layer; The specific working process of the reverse diffusion network is: The real noise-free SAR image x0 is input into the first conditional downsampling block, and the first conditional downsampling block outputs feature J; Feature J is input into the second conditional downsampling block, and the second conditional downsampling block outputs feature feature Input the third conditional downsampling block, and the third conditional downsampling block outputs feature L; The noisy optical image y in the noise domain at time step j is j Input the sixth 1×1 convolutional layer, and the sixth 1×1 convolutional layer outputs feature M; Feature M is input into the seventh residual block, and the seventh residual block outputs feature N; Feature J and feature N are input into the first conditional modulator, and the first conditional modulator outputs feature O; Feature O is input into the eighth residual block, and the eighth residual block outputs feature P; Feature J and feature P are input into the second conditional modulator, and the second conditional modulator outputs feature feature Input the first wavelet downsampling block, and the first wavelet downsampling block outputs feature R; Feature R is input into the ninth residual block, and the ninth residual block outputs feature S; Feature S and Features Input a third conditional modulator, and the third conditional modulator outputs a feature T; Feature T is input into the tenth residual block, and the tenth residual block outputs feature U; Feature U and Feature Input the fourth conditional modulator, the fourth conditional modulator output characteristics feature Input the second wavelet downsampling block, and the second wavelet downsampling block outputs feature W; The feature W is input into the eleventh residual block, and the eleventh residual block outputs the feature X; Feature L and feature X are input into the fifth conditional modulator, and the fifth conditional modulator outputs feature Y; Feature Y and the high-frequency component output by the second wavelet downsampling block are input into the first wavelet upsampling block, and the first wavelet upsampling block outputs feature A′; Feature A′ is input into the twelfth residual block, and the twelfth residual block outputs feature B′; Feature B′ and the output feature of the second conditional downsampling block Input the sixth conditional modulator, the sixth conditional modulator outputs a characteristic C′; Feature C′ is input into the thirteenth residual block, and the thirteenth residual block outputs feature D′; Feature D′ and the output feature of the second conditional downsampling block Input the seventh conditional modulator, and the seventh conditional modulator outputs the characteristic E′; Feature E′ and the high-frequency component output by the first wavelet downsampling block are input into the second wavelet upsampling block, and the second wavelet upsampling block outputs feature F′; Feature F′ is input into the fourteenth residual block, and the fourteenth residual block outputs feature G′; The feature G′ and the feature J output by the first conditional downsampling block are input into the eighth conditional modulator, and the eighth conditional modulator outputs the feature H′; Feature H′ is input into the fifteenth residual block, and the fifteenth residual block outputs feature I′; Feature I′ and feature J output by the first conditional downsampling block are input into the ninth conditional modulator, and the ninth conditional modulator outputs feature J′; Feature J′ is input into the seventh 1×1 convolutional layer, and the seventh 1×1 convolutional layer outputs feature K′; Feature K′ is the noisy optical image y in the noise domain at time step j as the output of the back diffusion network j Noise added in.

6. The method for converting synthetic aperture radar to optical image based on diffusion drive and wavelet enhancement according to claim 5, characterized in that: Each of the first conditional downsampling block, the second conditional downsampling block, and the third conditional downsampling block is a convolutional layer with a stride of 2.

7. The method for converting synthetic aperture radar to optical image based on diffusion drive and wavelet enhancement according to claim 6, characterized in that: Each of the first conditional modulator, the second conditional modulator, the third conditional modulator, the fourth conditional modulator, the fifth conditional modulator, the sixth conditional modulator, the seventh conditional modulator, the eighth conditional modulator, and the ninth conditional modulator includes: The eighth 1×1 convolutional layer, the first depthwise convolutional layer, the second depthwise convolutional layer, and the first set of normalization layers; The first conditional modulator, the second conditional modulator, the third conditional modulator, the fourth conditional modulator, the fifth conditional modulator, the sixth conditional modulator, the seventh conditional modulator, the eighth conditional modulator, and the ninth conditional modulator have the same working process; The working process of the first conditional modulator is: Feature J is input into the eighth 1×1 convolution layer, and the eighth 1×1 convolution layer outputs feature feature Input the first depth convolution layer, the first depth convolution layer outputs feature γ; feature Input the second depth convolution layer, and the first depth convolution layer outputs feature β; The output feature N of the seventh residual block is input into the first group of normalization, and the first group of normalization layer outputs the feature α; Feature α and feature γ are multiplied element by element to obtain feature δ; Feature δ and feature β are added element by element to obtain feature θ; Feature α and feature θ are added element by element to obtain feature θ′; The feature θ′ is the output of the conditional modulator.

8. The method for converting synthetic aperture radar to optical image based on diffusion drive and wavelet enhancement according to claim 7, characterized in that: Each of the first wavelet downsampling block and the second wavelet downsampling block includes: The second group of normalization layers, the first discrete wavelet transform, the third group of normalization layers, the third depthwise convolutional layer, the self-attention layer, the point-by-point convolutional layer, the cross-attention layer, the ninth convolutional layer, the tenth convolutional layer, and the second discrete wavelet transform; The working process of the first wavelet downsampling block is: The features Input the second set of normalization layers, the second set of normalization layers output features The features Input the first discrete wavelet transform for processing, the first discrete wavelet transform input features Decompose into low-frequency components With high frequency components The low-frequency components Input the third set of normalization layers, the third set of normalization layers output features The high frequency components Input the third depth convolution layer, the third depth convolution layer outputs features The features Input the self-attention layer, and the self-attention layer outputs feature Q; The features After the point-by-point convolution layer, the point-by-point convolution layer outputs features K and V; Input feature Q, feature K, and feature V into the cross attention layer, and the cross attention layer outputs feature The features Splice with time code to get features The features Input the ninth convolution layer, the ninth convolution layer outputs features The features Input the second discrete wavelet transform for processing, the second discrete wavelet transform input features Decompose into low-frequency components With high frequency components The low-frequency components Input the tenth convolutional layer, the tenth convolutional layer outputs features The features and features Perform element-by-element summation to obtain feature R; Take feature R as the output feature of the first wavelet downsampling block; The working process of the second wavelet downsampling block is: The features Input the second set of normalization layers, the second set of normalization layers output features The features Input the first discrete wavelet transform for processing, the first discrete wavelet transform input features Decompose into low-frequency components With high frequency components The low-frequency components Input the third set of normalization layers, the third set of normalization layers output features The high frequency components Input the third depth convolution layer, the third depth convolution layer outputs features The features Input the self-attention layer, and the self-attention layer outputs feature Q; The features After the point-by-point convolution layer, the point-by-point convolution layer outputs features K and V; Input feature Q, feature K, and feature V into the cross attention layer, and the cross attention layer outputs feature The features Splice with time code to get features The features Input the ninth convolution layer, the ninth convolution layer outputs features The features Input the second discrete wavelet transform for processing, the second discrete wavelet transform input features Decompose into low-frequency components With high frequency components The low-frequency components Input the tenth convolutional layer, the tenth convolutional layer outputs features The features and features Perform element-by-element summation to obtain feature W; The feature W is used as the output feature of the second wavelet downsampling block.

9. The method for converting synthetic aperture radar to optical image based on diffusion drive and wavelet enhancement according to claim 8, characterized in that: Each of the first wavelet upsampling block and the second wavelet upsampling block includes: The fourth group of normalization layers, wavelet inverse change layers, the eleventh convolution layer, the twelfth convolution layer, the pixel reconstruction layer, and the thirteenth convolution layer; The working process of the first wavelet upsampling block is: Input feature Y into the fourth group of normalization layers, and the fourth group of normalization layers outputs feature Y1; Combine feature Y1 with the high-frequency component of the first discrete wavelet transform decomposition in the second wavelet downsampling block Input wavelet inverse change layer, wavelet inverse change layer output feature Y2; Input feature Y2 into the eleventh convolutional layer, and the eleventh convolutional layer outputs feature Y3; Input feature Y1 into the twelfth convolutional layer, and the twelfth convolutional layer outputs feature Y4; Perform pixel reorganization on feature Y4 to obtain feature Y5; Concatenate feature Y3 and feature Y5 to obtain feature Y6; Add feature Y6 and time code element by element to obtain feature Y7; Input feature Y7 into the thirteenth convolutional layer, and the thirteenth convolutional layer outputs feature A′; Feature A′ is used as the output feature of the first wavelet upsampling block; The working process of the second wavelet upsampling block is: Input the feature E′ into the fourth group of normalization layers, and the fourth group of normalization layers outputs the feature E1′; Combine the feature E1′ with the high-frequency component of the first discrete wavelet transform decomposition in the first wavelet downsampling block Input wavelet inverse change layer, wavelet inverse change layer output feature E′2; Input feature E′2 into the eleventh convolutional layer, and the eleventh convolutional layer outputs feature E′3; Input feature E1′ into the twelfth convolutional layer, and the twelfth convolutional layer outputs feature E′4; Perform pixel reorganization on feature E′4 to obtain feature E′5; Concatenate feature E′3 and feature E′5 to obtain feature E′6; Add feature E′6 and time code element by element to obtain feature E′7; Input feature E′7 into the thirteenth convolutional layer, and the thirteenth convolutional layer outputs feature F′; The feature F′ is used as the output feature of the second wavelet upsampling block.

10. The method for converting synthetic aperture radar to optical image based on diffusion drive and wavelet enhancement according to claim 9, characterized in that: The training of the modality mapping network adopts the GAN form, and the training objective is expressed as: Among them, the first item is used to evaluate the authenticity of the generated image, and the second item is used to evaluate the similarity between the generated image and the real image; represents the total loss function of the generator; Represents the input sample x t The expected value of the output; Represents the output of the generator network G, and the parameters are It represents the judgment result of the discriminator D on the output of the generator, indicating the probability that the generated sample is judged to be true; Represents the input sample pair (x t ,y t )Expected value of the output; y t represents the true noisy optical image features at the t-th time step; || ||1 represents L1 norm; The training of the reverse diffusion network adopts the L1 loss function, and the training objective is expressed as: Among them, p θ represents the trained model, ∈ represents random Gaussian noise; Represents the joint expectation of the input SAR image x0, the real optical image y0, the random Gaussian noise ∈, and the output of time step j; Represents the output from time step 1 to time step j, expressed as α j =1-β j , β j represents the noise intensity at time step j; α i represents an intermediate variable; || ||1 represents the L1 norm.

Citation Information

Cited By

  • Synthetic aperture radar image conversion network construction method

    CN121458563A