SAR Image to Optical Image Translation Method and Device Based on Discrete Feature Mapping
By building a compression-reconstructor and a cross-modal feature mapper, the problem of poor quality of generated results in remote sensing image data translation is solved, and high-precision translation from SAR images to optical images is realized, which improves the processing and conversion effect of image data.
Patent Information
- Application Number
- CN202510581118.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The existing remote sensing image data translation methods have poor quality and inaccurate characteristics, and cannot fully utilize the complementary advantages of SAR and optical images.
Using a discrete feature map-based method, the translation of SAR images to optical images is realized by constructing a compression-reconstructor and a cross-modal feature mapper, including a spatial encoder, a feature discreteer, a frequency-attention fusion generator and a pixel discriminator.
The processing and conversion accuracy of image data is improved, and is suitable for multi-source fusion and application of remote sensing image data. The generated optical images show significant advantages in structure, color and detail recovery.
Smart Images

Figure CN120088619B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and specifically to a method and device for translating SAR images into optical images based on discrete feature mapping. Background Art
[0002] Synthetic Aperture Radar (SAR) data and Optical Image data are two of the most commonly used data types in the field of remote sensing. SAR is an active earth observation system whose operation is not restricted by weather and time, and has all-weather and all-time imaging capabilities. However, the complex scattering characteristics of SAR data lead to relatively complex processing and analysis, which requires certain professional knowledge and skills. On the other hand, optical images can provide high-resolution and color-rich images. With its intuitive visibility, rich texture features, and advantages of being easy to analyze and interpret, it occupies a dominant position in remote sensing tasks. However, it is affected by weather conditions and lighting, and cannot obtain effective data at night or in bad weather, resulting in insufficient continuity and reliability of data acquisition. To give full play to the complementary advantages of SAR and optical images, combine their advantages, and overcome their respective limitations, it is very important to conduct research on SAR-to-Optical (S2O) translation technology. However, remote sensing image data has a large amount of information and complex features, and the existing translation methods produce poor-quality results with inaccurate features. Summary of the Invention
[0003] Aiming at the above problems existing in the prior art, the present invention expands the traditional image generation model, and provides a method and device for translating SAR images into optical images based on discrete feature mapping.
[0004] The above object of the present invention is achieved by the following technical means:
[0005] A method for translating SAR images into optical images based on discrete feature mapping, comprising the following steps:
[0006] Step S1, obtaining the SAR image and the corresponding real optical image, and constructing a training set and a test set;
[0007] Step S2, constructing and training a compression-reconstructor, the compression-reconstructor includes a spatial encoder, a feature discretizer, and a spatial decoder, and there are three compression-reconstructors, namely an optical image compression-reconstructor, a high compression ratio SAR image compression-reconstructor, and a low compression ratio SAR image compression-reconstructor;
[0008] Step S3: Construct and train a cross-modal feature mapper. The cross-modal feature mapper includes a frequency-attention fusion generator. The high-compression-ratio SAR discrete feature vectors and low-compression-ratio SAR discrete feature vectors output by the feature discretizer of the high-compression-ratio SAR image compression-reconstructor and the low-compression-ratio SAR image compression-reconstructor are input into the frequency-attention fusion generator to obtain corresponding discrete feature vectors. The obtained discrete feature vectors are input into the spatial decoder of the optical image compression-reconstructor to obtain a generated optical image.
[0009] Step S4: Input the SAR image to be translated into the feature discretizers of the trained high-compression-ratio SAR image compression-reconstructor and the low-compression-ratio SAR image compression-reconstructor. The high-compression-ratio SAR discrete feature vectors and low-compression-ratio SAR discrete feature vectors obtained are input into the frequency-attention fusion generator to obtain corresponding discrete feature vectors. The obtained discrete feature vectors are input into the spatial decoder of the optical image compression-reconstructor to obtain the translated generated optical image.
[0010] As described above, the number of convolutional layers in the spatial encoder and spatial decoder of the high-compression-ratio SAR image compression-reconstructor is more than that in the spatial encoder and spatial decoder of the low-compression-ratio SAR image compression-reconstructor.
[0011] As described above, the optical image compression-reconstructor takes a real optical image as the input image and a reconstructed optical image as the output image; both the high-compression-ratio SAR image compression-reconstructor and the low-compression-ratio SAR image compression-reconstructor take a SAR image as the input image and a reconstructed SAR image as the output image.
[0012] The input image obtains a continuous feature vector through the spatial encoder. The continuous feature vector is input into the feature discretizer. The feature discretizer searches for the discrete feature vector closest to the continuous feature vector in the feature dictionary. The discrete feature vector is input into the spatial decoder, and the spatial decoder outputs a reconstructed image with the same size as the input image.
[0013] As described above, training the compression-reconstructor includes using a discriminator to discriminate the optical image compression-reconstructor, the high-compression-ratio SAR image compression-reconstructor, and the low-compression-ratio SAR image compression-reconstructor, and training the optical image compression-reconstructor, the high-compression-ratio SAR image compression-reconstructor, and the low-compression-ratio SAR image compression-reconstructor based on minimizing the reconstruction loss.
[0014] As described above, the frequency-attention fusion generator includes a frequency enhancement-attention fusion module and a discrete feature mapper.
[0015] In the frequency enhancement-attention fusion module, the high-compression-ratio SAR discrete feature vectors and the low-compression-ratio SAR discrete feature vectors are enhanced by frequency filters respectively to obtain the high-compression-ratio enhancement result and the low-compression-ratio enhancement result; the high-compression-ratio enhancement result and the low-compression-ratio enhancement result are stacked to obtain the stacked result; the stacked result is multiplied by the weight matrix respectively to obtain the query vector , the key vector , and the value vector . Through , the attention weight matrix is calculated, where T represents the transpose operation; the attention weight matrix is multiplied by to obtain the fusion result
[0016] . The discrete feature mapper performs multi-layer convolution processing on the fusion result to obtain the mapped discrete feature vectors.
[0017] As described above, the cross-modal feature mapper further includes a pixel discriminator, and the real optical image and the generated optical image are respectively input into the pixel discriminator for determination.
[0018] As described above, the loss function of the frequency-attention fusion generator is as follows:
[0019] ,
[0020] where is the adversarial loss, is the Smooth L1 loss, is the structural similarity loss, are the weights respectively,
[0021] ,
[0022] ,
[0023] where is the number of samples, is the generated optical image, is the pixel value at the -th position of the generated optical image, is the target matrix with all values equal to 1 and the same size as , is the target matrix , is the pixel value at the -th position of the target matrix, is the spatial decoder of the optical image compression-reconstructor, is the frequency-attention fusion generator. The feature discretizer for the low compression ratio SAR compression-reconstructor, The spatial encoder for the low compression ratio SAR compression-reconstructor, The SAR image, The feature discretizer for the high compression ratio SAR compression-reconstructor, The spatial encoder for the high compression ratio SAR compression-reconstructor,
[0024] ,
[0025] Among them, is the pixel value at the th position of the real optical image, is the pixel value at the th position of the generated optical image, is to take the absolute value,
[0026] ,
[0027] ,
[0028] ,
[0029] Among them, is the pixel mean of the real optical image, is the pixel mean of the generated optical image, is the variance of the real optical image, is the variance of the generated optical image, is the covariance between the real optical image and the generated optical image, is the dynamic range of the real optical image. and are both intermediate parameters.
[0030] As described above, the loss function of the pixel discriminator is:
[0031] ,
[0032] Among them, is the number of samples, is the Sigmod function, is the real optical image, is the pixel value at the th position of the real optical image, is the generated optical image, is the pixel value at the th position of the generated optical image, is A target matrix with all values of the same size being 1, is the target matrix at the th position of the pixel value, is a target matrix with all values of the same size being 0, is the target matrix at the th position of the pixel value.
[0033] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above translation method are implemented.
[0034] A computer program product includes a computer program. When the computer program is executed by a processor, the steps of the above translation method are implemented.
[0035] The present invention has the following beneficial effects compared with the prior art:
[0036] Based on the feature mapping method of pixel-feature conversion and frequency-attention fusion, the present invention realizes cross-modal translation from SAR images to optical images. The aim is to improve the processing and conversion accuracy of image data and is applicable to multi-source fusion and application of remote sensing image data. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 is a schematic diagram of the implementation process provided by an embodiment of the present invention;
[0038] Figure 2 is a training process diagram of the compression-reconstructor used in an embodiment of the present invention;
[0039] Figure 3 is a training process diagram of the cross-modal feature mapper used in an embodiment of the present invention;
[0040] Figure 4 is a result comparison diagram of the SAR-optical image translation of the method of the present invention and other methods in an embodiment of the present invention,
[0041] where (a1) is the first SAR image, and (b1), (c1), (d1), (e1), and (f1) are the optical images obtained by processing the first SAR image by the Pix2pix method, BicycleGAN method, CycleGAN method, CUT method, and the method of the present invention respectively, and (g1) is the corresponding real optical image;
[0042] (a2) is the second SAR image, and (b2), (c2), (d2), (e2), and (f2) are the optical images obtained by processing the second SAR image with the Pix2pix method, BicycleGAN method, CycleGAN method, CUT method, and the method of the present invention, respectively. (g2) is the corresponding real optical image;
[0043] (a3) is the third SAR image, and (b3), (c3), (d3), (e3), and (f3) are the optical images obtained by processing the third SAR image with the Pix2pix method, BicycleGAN method, CycleGAN method, CUT method, and the method of the present invention, respectively. (g3) is the corresponding real optical image;
[0044] (a4) is the fourth SAR image, and (b4), (c4), (d4), (e4), and (f4) are the optical images obtained by processing the fourth SAR image with the Pix2pix method, BicycleGAN method, CycleGAN method, CUT method, and the method of the present invention, respectively. (g4) is the corresponding real optical image;
[0045] (a5) is the fifth SAR image, and (b5), (c5), (d5), (e5), and (f5) are the optical images obtained by processing the fifth SAR image with the Pix2pix method, BicycleGAN method, CycleGAN method, CUT method, and the method of the present invention, respectively. (g5) is the corresponding real optical image. Detailed implementation manners
[0046] To facilitate the understanding and implementation of the present invention by those of ordinary skill in the art, the present invention will be further described in detail below in conjunction with implementation examples. It should be understood that the implementation examples described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.
[0047] Example 1
[0048] Figure 1 It is a schematic flowchart of the SAR image to optical image translation method based on discrete feature mapping provided by the present invention, including the following steps:
[0049] Step S1: Obtain high-resolution SAR images and real optical images of the same area, and construct a training set and a test set, specifically including:
[0050] Step S11: Obtain high-resolution SAR images in the same area as samples, and use the corresponding real optical images of the SAR images as the labels for the SAR images. In this embodiment, the MSAW (Multi-Sensor All Weather Mapping) dataset is obtained. This dataset combines real optical images and corresponding SAR images, and is dedicated to promoting high-precision land use classification research. The dataset covers 120 square kilometers and is collected in the Port of Rotterdam, the Netherlands. The real optical images in the dataset are from the Maxar WorldView-2 satellite, including real optical images of the panchromatic band with a resolution of 0.5 meters, multi-spectral band images (RGBN) with a resolution of 2 meters, and sharpened multi-spectral band images (RGBN) with a resolution of 0.5 meters; the SAR images are provided by Capella Space, with a resolution of 0.5 meters and four polarization mode bands of VV, VH, HV, and HH.
[0051] Step S12: Process the MSAW dataset. Obtain paired SAR images and real optical images from the above MSAW dataset. Among them, three bands of VV band, HH band, and VV / HH band are selected for the SAR images, and the RGB bands of the sharpened multi-spectral images with a resolution of 0.5 meters are selected for the real optical images; perform Lee filtering on the SAR images to reduce speckle noise and retain edges and details; crop the SAR images and the corresponding real optical images to a size of 256×256 pixels, and divide the SAR images into a training set and a validation set according to 8:2.
[0052] Step S2: Construct a compression-reconstructor (pixel-discrete feature compression-reconstructor) to achieve the conversion of images in different data spaces. In this embodiment, there are 3 compression-reconstructors, namely the optical image compression-reconstructor, the high-compression-ratio SAR image compression-reconstructor, and the low-compression-ratio SAR image compression-reconstructor. The compression-reconstructor includes a spatial encoder, a feature discretizer, and a spatial decoder. The present invention needs to train the optical image compression-reconstructor, the high-compression-ratio SAR image compression-reconstructor, and the low-compression-ratio SAR image compression-reconstructor. Among them, the compression ratios of the high- and low-compression-ratio SAR image compression-reconstructors are controlled by the number of convolutional layers in the spatial encoder and the spatial decoder. The number of convolutional layers in the spatial encoder and the spatial decoder inside the high-compression-ratio SAR image compression-reconstructor is more than that in the spatial encoder and the spatial decoder of the low-compression-ratio SAR image compression-reconstructor; the optical image compression-reconstructor has the same structure as the low-compression-ratio SAR image compression-reconstructor. The optical image compression-reconstructor takes the real optical image as the input image and the reconstructed optical image as the output image. Both the high- and low-compression-ratio SAR image compression-reconstructors take the SAR image as the input image and the reconstructed SAR image as the output image. Specifically:
[0053] Step S21: The input image gradually reduces the spatial dimension and increases the number of channels through the spatial encoder of the compression-reconstructor, obtains the continuous feature vectors of the input image in the feature space, and inputs them into the feature discretizer of the compression-reconstructor.
[0054] The real optical image obtains the continuous feature vectors of the optical image in the feature space through the spatial encoder of the optical image compression-reconstructor. The SAR image obtains the continuous feature vectors with a high compression ratio of the SAR image in the feature space through the spatial encoder of the high-compression-ratio SAR image compression-reconstructor. The SAR image obtains the continuous feature vectors with a low compression ratio of the SAR image in the feature space through the spatial encoder of the low-compression-ratio SAR image compression-reconstructor.
[0055] Step S22: Using the feature discretizer, each continuous feature vector output by the spatial encoder finds the closest discrete feature vector in the feature dictionary, and the discrete feature vectors are input into the spatial decoder.
[0056] The continuous feature vectors of the optical image are replaced by optical discrete feature vectors through the feature discretizer of the optical image compression-reconstructor. The continuous feature vectors with a high compression ratio of the SAR are replaced by discrete feature vectors with a high compression ratio of the SAR through the feature discretizer of the high-compression-ratio SAR image compression-reconstructor. The continuous feature vectors with a low compression ratio of the SAR are replaced by discrete feature vectors with a low compression ratio of the SAR through the feature discretizer of the low-compression-ratio SAR image compression-reconstructor.
[0057] Step S23: The spatial decoder uses the discrete feature vectors as the input, gradually increases the spatial dimension and reduces the number of channels through a series of transposed convolution layers or upsampling layers, and finally outputs a reconstructed image with the same size as the input image. The optical discrete feature vectors obtain the optical reconstructed image through the spatial decoder of the optical image compression-reconstructor. The discrete feature vectors with a high compression ratio of the SAR obtain the corresponding high-compression-ratio SAR reconstructed image through the spatial decoder of the high-compression-ratio SAR image compression-reconstructor. The discrete feature vectors with a low compression ratio of the SAR obtain the corresponding low-compression-ratio SAR reconstructed image through the spatial decoder of the low-compression-ratio SAR image compression-reconstructor.
[0058] Step S24: Use the discriminator to discriminate each compression-reconstructor, and train the optical image compression-reconstructor, the high-compression-ratio SAR image compression-reconstructor, and the low-compression-ratio SAR image compression-reconstructor based on minimizing the reconstruction loss, that is: calculate the reconstruction loss based on the SAR image and the high-compression-ratio SAR reconstructed image, and optimize the high-compression-ratio SAR image compression-reconstructor based on minimizing the reconstruction loss; calculate the reconstruction loss based on the SAR image and the low-compression-ratio SAR reconstructed image, and train the low-compression-ratio SAR image compression-reconstructor based on minimizing the reconstruction loss; calculate the reconstruction loss based on the real optical image and the optical reconstructed image, and optimize the optical image compression-reconstructor based on minimizing the reconstruction loss.
[0059] Step S3: Train the cross-modal feature mapper based on the already trained compression-reconstructor. The cross-modal feature mapper includes a frequency-attention fusion generator and a pixel discriminator. Specifically:
[0060] Step S31: Use the spatial encoder and the feature discretizer of the trained high-compression-ratio SAR compression-reconstructor to compress the SAR image into a high-compression-ratio SAR discrete feature vector, and use the spatial encoder and the feature discretizer of the trained low-compression-ratio SAR compression-reconstructor to compress the SAR image into a low-compression-ratio SAR discrete feature vector;
[0061] Step S32: Construct a frequency-attention fusion generator, which includes a frequency enhancement-attention fusion module and a discrete feature mapper;
[0062] Step S321: Construct a frequency enhancement-attention fusion module: In the frequency enhancement-attention fusion module, the high-compression-ratio SAR discrete feature vector and the low-compression-ratio SAR discrete feature vector are respectively enhanced by frequency filters to obtain a high-compression-ratio enhancement result and a low-compression-ratio enhancement result; stack the high-compression-ratio enhancement result and the low-compression-ratio enhancement result to obtain a stacked result; the stacked result is respectively multiplied by three weight matrices to obtain a query vector , a key vector , and a value vector , where the three weight matrices are generated by random initialization and are optimized in step 35; calculate the similarity between the query vector and the key vector through (T is the transpose operation) to obtain the attention weight matrix ; multiply the attention weight matrix by Multiply them to obtain the final output, i.e., the fusion result. The discrete feature vectors with a high compression ratio preserve more semantic information but lose spatial information, while the discrete feature vectors with a low compression ratio preserve more spatial information but less semantic information. By fusing the effective information in the two types of data, the performance of the model can be improved.
[0063] Step S322: Construct a discrete feature mapper: Process the fusion result of step S321 to obtain the mapped discrete feature vectors. In this embodiment, the fusion result input to the discrete feature mapper passes through 10 convolutional blocks composed of a convolutional layer + a BatchNorm layer + a ReLU layer to obtain the mapped discrete feature vectors.
[0064] Step S33: Input the mapped discrete feature vectors into the spatial decoder of the optical image compression-reconstructor to reconstruct and obtain the generated optical image;
[0065] Step S34: Construct a pixel discriminator to discriminate between the generated optical image and the real optical image. The real optical image and the generated optical image are respectively input into the pixel discriminator.
[0066] In this embodiment, the images input to the pixel discriminator first pass through 1 convolutional layer and 1 ReLU layer, then pass through 3 convolutional blocks composed of a convolutional layer + a BatchNorm layer + a ReLU layer, and finally pass through 1 convolutional layer to obtain the discrimination result.
[0067] Step S35: Calculate the loss between the generated optical image and the real optical image, and optimize the frequency-attention fusion generator and the pixel discriminator.
[0068] The frequency-attention fusion generator is optimized according to the loss function of the frequency-attention fusion generator as follows: Specifically:
[0069] ,
[0070] where is the adversarial loss, is the Smooth L1 loss, is the structural similarity loss, are respectively the weights of.
[0071] The adversarial loss :
[0072] ,
[0073] ,
[0074] where is the number of samples, is to generate an optical image, is the th pixel value of the position for generating the optical image, is a target matrix with all values of 1 and the same size as is the target matrix the th pixel value of the position of is the Sigmod function, is the spatial decoder of the optical image compression-reconstructor, is the frequency-attention fusion generator, is the feature discretizer of the low compression ratio SAR compression-reconstructor, is the spatial encoder of the low compression ratio SAR compression-reconstructor, is the SAR image, is the feature discretizer of the high compression ratio SAR compression-reconstructor, is the spatial encoder of the high compression ratio SAR compression-reconstructor,
[0075] Smooth L1 loss :
[0076] The Smooth L1 loss calculates the loss for each pixel point at each position of the real optical image and the generated optical image separately, and finally calculates an average loss value.
[0077] ,
[0078] where, is the th pixel value of the real optical image, is the th pixel value of the generated optical image, is to take the absolute value.
[0079] Structural similarity loss :
[0080] ,
[0081] ,
[0082] ,
[0083] where, is the pixel mean of the real optical image, is the pixel mean of the generated optical image, is the variance of the real optical image, is the variance of the generated optical image, is the covariance between the real optical image and the generated optical image, is the dynamic range of the real optical image. and are both intermediate parameters.
[0084] The pixel discriminator is optimized according to the loss function of the pixel discriminator as follows: Specifically:
[0085] ,
[0086] wherein, is the number of samples, is the Sigmod function, is the real optical image, is the th pixel value at the position of the real optical image, is the generated optical image, is the th pixel value at the position of the generated optical image, is a target matrix with all values equal to 1 and the same size as , is the pixel value at the th position of the target matrix , is a target matrix with all values equal to 0 and the same size as th position of the target matrix
[0087] After the optical image compression-reconstructor, the high compression ratio SAR image compression-reconstructor, the low compression ratio SAR image compression-reconstructor, and the cross-modal feature mapper are trained, the SAR image to be translated is subjected to steps S31 - S33 to obtain the translated generated optical image.
[0088] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods.
[0089] Figure 4This is a comparison chart of the results of the method of the present invention with the Pix2pix method, BicycleGAN method, CycleGAN method, and CUT method. The generated results of the Pix2pix method look relatively blurred, with serious loss of details, especially in the edge and building areas; the generated results of the BicycleGAN method are poor, with problems of blurred edges, and inaccurate color and structure reconstruction; the CycleGAN method generates clearer structures to a certain extent, but there are still large deviations in color accuracy and overall texture distribution; the CUT method has improvements in color and edge structure, looking more natural, but there are still problems of incorrect reconstruction of image objects; the method of the present invention shows significant advantages in terms of structure, color, and detail restoration of the generated optical images. Especially in the reconstruction of storage tanks, houses, and ground textures, this method can better restore the characteristics of real optical images, and the overall generation quality is better than other methods.
[0090] Table 1 shows the quantitative evaluation results of the method of the present invention and other methods. The quantitative evaluation indicators selected are the structural similarity index and the peak signal-to-noise ratio. The Structural Similarity Index Measure (SSIM) is an indicator used to measure the structural similarity between two images. It evaluates the similarity between images by considering luminance, contrast, and structural information. SSIM is based on the perceptual characteristics of the human visual system, especially its performance in image detail and overall structure perception. It is mainly used for image quality evaluation, especially in the fields of image compression, image enhancement, and image restoration. The Peak Signal-to-Noise Ratio (PSNR) is an indicator used for image quality evaluation, mainly measuring the difference between the reconstructed image and the original image. PSNR gives a quantitative measure of image quality by calculating the ratio between the maximum possible value of the image and the noise. It is usually used in the fields of image compression, image restoration, and image processing, etc. The generated results of this method on the MSAW dataset are more prominent than other methods, achieving the best results in both SSIM and PSNR, the two image quality evaluation indicators. Specifically, the SSIM of this method reaches 0.3605, slightly improved compared to the sub-optimal BicycleGAN method (0.3542), indicating that it has better performance in maintaining the structural similarity of the generated images. In terms of PSNR, this method significantly leads other methods, reaching 17.6309, nearly 2 points higher than the sub-optimal BicycleGAN method (15.7278), indicating that the noise level of the generated image is lower and the quality is higher.
[0091] Table 1
[0092]
[0093] Example 2
[0094] This embodiment provides a computer device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0095] Example 3
[0096] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0097] Example 4
[0098] This embodiment provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0099] It should be noted that the embodiments described in the present invention are only illustrative of the spirit of the present invention. Those skilled in the art to which the present invention pertains can make various modifications or supplements to the described embodiments or use similar ways to replace them, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
Claims
1. A method for translating SAR images to optical images based on discrete feature mapping, characterized in that, It includes the following steps: Step S1: Obtain SAR images and corresponding real optical images, and construct a training set and a test set; Step S2: Construct and train a compression-reconstructor, which includes a spatial encoder, a feature discretizer, and a spatial decoder. There are three compression-reconstructors, namely an optical image compression-reconstructor, a high compression ratio SAR image compression-reconstructor, and a low compression ratio SAR image compression-reconstructor; Step S3: Construct and train a cross-modal feature mapper, which includes a frequency-attention fusion generator. The high compression ratio SAR discrete feature vector and the low compression ratio SAR discrete feature vector output by the feature discretizers of the high compression ratio SAR image compression-reconstructor and the low compression ratio SAR image compression-reconstructor are input into the frequency-attention fusion generator to obtain corresponding discrete feature vectors, and the obtained discrete feature vectors are input into the spatial decoder of the optical image compression-reconstructor to obtain a generated optical image; Step S4: Input the SAR image to be translated into the feature discretizers of the trained high compression ratio SAR image compression-reconstructor and the low compression ratio SAR image compression-reconstructor. The high compression ratio SAR discrete feature vector and the low compression ratio SAR discrete feature vector obtained are input into the frequency-attention fusion generator to obtain corresponding discrete feature vectors, and the obtained discrete feature vectors are input into the spatial decoder of the optical image compression-reconstructor to obtain the translated generated optical image.
2. The SAR image to optical image translation method based on discrete feature mapping according to claim 1, wherein The number of convolutional layers in the spatial encoder and spatial decoder of the high compression ratio SAR image compression-reconstructor is more than that in the spatial encoder and spatial decoder of the low compression ratio SAR image compression-reconstructor.
3. The SAR image to optical image translation method based on discrete feature mapping according to claim 1, wherein The optical image compression-reconstructor takes the real optical image as the input image and the reconstructed optical image as the output image; both the high compression ratio SAR image compression-reconstructor and the low compression ratio SAR image compression-reconstructor take the SAR image as the input image and the reconstructed SAR image as the output image; The input image obtains a continuous feature vector through the spatial encoder, the continuous feature vector is input into the feature discretizer, the feature discretizer finds the discrete feature vector closest to the continuous feature vector in the feature dictionary, and the discrete feature vector is input into the spatial decoder, and the spatial decoder outputs a reconstructed image with the same size as the input image.
4. The SAR image to optical image translation method based on discrete feature mapping according to claim 3, wherein, The training of the compression-reconstructor includes using a discriminator to discriminate the optical image compression-reconstructor, the high compression ratio SAR image compression-reconstructor, and the low compression ratio SAR image compression-reconstructor, and training the optical image compression-reconstructor, the high compression ratio SAR image compression-reconstructor, and the low compression ratio SAR image compression-reconstructor based on minimizing the reconstruction loss.
5. The SAR image to optical image translation method based on discrete feature mapping according to claim 3, wherein The frequency-attention fusion generator includes a frequency enhancement-attention fusion module and a discrete feature mapper, In the frequency enhancement-attention fusion module, the high-compression-ratio SAR discrete feature vectors and the low-compression-ratio SAR discrete feature vectors are respectively enhanced through frequency filters to obtain high-compression-ratio enhancement results and low-compression-ratio enhancement results; the high-compression-ratio enhancement results and the low-compression-ratio enhancement results are stacked to obtain a stacked result; The stacking results are respectively multiplied by the weight matrix to obtain the query vector , the key vector , and the value vector . Through , the attention weight matrix is calculated, where T represents the transpose operation; the attention weight matrix is multiplied by to obtain the fusion result The discrete feature mapper performs multi-layer convolution processing on the fusion result to obtain the mapped discrete feature vectors.
6. The method for translating SAR images to optical images based on discrete feature mapping according to claim 5, wherein The cross-modal feature mapper further includes a pixel discriminator, and the real optical image and the generated optical image are respectively input into the pixel discriminator for determination.
7. The method for translating SAR images to optical images based on discrete feature mapping according to claim 6, wherein, The loss function of the frequency-attention fusion generator is as follows: , wherein is the adversarial loss, is the Smooth L1 loss, is the structural similarity loss, are the weights respectively, , , Among them, is the number of samples, is to generate an optical image, is the th pixel value at the position of generating the optical image, is a target matrix with all values equal to 1 and the same size as , is the th pixel value at the position of the target matrix, is the Sigmod function, is the spatial decoder of the optical image compression-reconstructor, is the frequency-attention fusion generator, is the feature discretizer of the low compression ratio SAR compression-reconstructor, is the spatial encoder of the low compression ratio SAR compression-reconstructor, is the SAR image, is the feature discretizer of the high compression ratio SAR compression-reconstructor, is the spatial encoder of the high compression ratio SAR compression-reconstructor, , Among them, is the pixel value at the th position of the real optical image, is the pixel value at the th position of the generated optical image, is to take the absolute value, , , , Among them, is the pixel mean of the real optical image, is the pixel mean of the generated optical image, is the variance of the real optical image, is the variance of the generated optical image, is the covariance between the real optical image and the generated optical image, is the dynamic range of the real optical image, and are both intermediate parameters.
8. The SAR image to optical image translation method based on discrete feature mapping according to claim 6, wherein The loss function of the pixel discriminator is as follows: , Among them, is the number of samples, is the Sigmod function, is the real optical image, is the pixel value at the -th position of the real optical image, is the generated optical image, is the pixel value at the -th position of the generated optical image, is a target matrix with all values equal to 1 and the same size as , is the pixel value at the -th position of the target matrix , is a target matrix with all values equal to 0 and the same size as , is the pixel value at the -th position of the target matrix .
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the translation method according to any one of claims 1 to 8 are implemented.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the translation method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Method for translating SAR image into optical image
CN105809194A
Optical image translation method based on ViT-Pix2Pix
CN115272787A