Image conversion method and system based on self-attention mechanism and generative adversarial network
Through the self-attention mechanism and the two-step generation method of generating adversarial networks, the problem of insufficient generalization ability in SAR image conversion is solved, high-quality optical image generation is achieved, the interpretability and controllability of the network are enhanced, and the clarity of the image and spectral information performance are improved.
Patent Information
- Application Number
- CN202510567410.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art has limited generalization capabilities in SAR to optical image conversion, especially poor performance when dealing with complex textures and long-distance dependencies, resulting in a lack of texture details that conform to optical imaging rules for the conversion results.
Using a two-step generation method based on self-attention mechanism and generative adversarial network, the SAR image is converted into grayscale images through the joint processing of global and local generators, and then converted into high-quality optical images. A global feature extractor is introduced for information supplementation and training with a multi-scale discriminator.
It significantly improves the quality and detailed performance of the generated images, enhances the interpretability and controllability of the network, and can accurately generate the details of optical images, especially in edge clarity and spectral information.
Smart Images

Figure CN120339442A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and deep learning, and relates to an image conversion method and system based on a self-attention mechanism and a generative adversarial network, which is applicable to the SAR-to-optical image conversion scenario. Background Art
[0002] With the continuous improvement of the global remote sensing satellite observation system (such as the Luojia series and Sentinel series satellite constellations), multi-source remote sensing data has become the core support for surface monitoring and ground object interpretation. As two core remote sensing data sources, SAR (Synthetic Aperture Radar) images and optical images play an irreplaceable role in the fields of natural resource monitoring, disaster emergency response, military reconnaissance, etc.
[0003] In terms of remote sensing imaging, SAR images and optical images are typical active and passive remote sensing data respectively. Synthetic Aperture Radar (SAR) images can obtain surface structure information under complex conditions such as cloud cover and night by virtue of microwave active imaging characteristics. However, its gray-scale imaging mechanism leads to the lack of spectral features and low visual interpretability; although optical images can intuitively represent the spectral attributes of ground objects, they are easily restricted by environmental factors such as light and weather. This cross-modal data representation gap severely limits the application efficiency of remote sensing technology in geological surveys, crop monitoring, and other scenarios such as surface vegetation coverage. How to efficiently utilize the internal correlation of multi-source remote sensing data has become a key challenge for improving the ground object interpretation ability.
[0004] To break through the above bottleneck, the SAR-to-optical image conversion technology has opened up a new path for the application of multi-source remote sensing images and has become a research hotspot in the field of remote sensing images in recent years. This technology establishes a non-linear mapping relationship between SAR and optical modalities through a deep neural network, and converts the microwave scattering characteristics of SAR images into optical-style images with pseudo-true color characteristics.
[0005] Thanks to the rapid development of remote sensing image registration technology, it provides solid data support for the SAR-to-optical image conversion technology, making the SAR-to-optical image conversion technology show great potential. However, existing methods still have significant problems: traditional image conversion networks are usually constructed based on the Generative Adversarial Network (GAN), and their generalization ability is limited in scenarios with extremely large cross-modal differences such as SAR and optics. Specifically, GAN does not make full use of key physical attributes such as polarization characteristics and incident angle effects in SAR images, and may perform poorly when dealing with complex textures and long-range dependencies, resulting in the conversion results lacking texture details that conform to the optical imaging law.
[0006] The deficiencies of the GAN network limit the practicality of the converted images in subsequent application scenarios. For example, low-quality pseudo-optical images may introduce misleading spectral features and cannot provide effective inputs for subsequent image fusion methods. To solve this problem, a deep network dedicated to SAR-optical cross-modal conversion can be constructed to break through the inherent limitations of traditional image conversion models. Summary of the Invention
[0007] Aiming at the problem that the generalization ability of traditional image conversion networks is limited in scenarios with extremely large cross-modal differences such as SAR and optics, the present invention proposes a two-step generation network, namely an image conversion model based on the self-attention mechanism and the generative adversarial network. The network mainly consists of two generation stages: grayscale image generation and optical image generation. In addition, a global feature extractor is introduced to supplement information in the first generation stage. Correspondingly, there are also two discriminators in the discrimination stage, both of which adopt a multi-scale discrimination strategy.
[0008] The technical solution adopted by the present invention is: an image conversion model based on the self-attention mechanism and the generative adversarial network. At the same time, in order to reduce the task complexity, the network is disassembled into a two-step generation method. In the first step, through the joint processing of the global and local generators, the input SAR image is converted into a grayscale image with clear edge and structural information. In the second step, the grayscale image is converted into a high-quality optical image. This two-step generation design not only optimizes the model structure but also enhances the interpretability and controllability of the network.
[0009] The generation in both stages adopts a generator design from rough to fine. The grayscale image generator includes a global generator and a local generator, which cooperate through the residual network architecture. The global generator processes the downsampled SAR image to capture global features and generate a rough conversion image; the local generator further extracts detailed information to generate a clear grayscale image. In addition, the global feature extractor is based on the Swin Transformer and realizes the organic fusion of local and global information through the window multi-head self-attention (W-MSA) and shifted window multi-head self-attention (SW-MSA) mechanisms, and improves the expression ability through feature normalization and residual connections.
[0010] Then, the fused image is input into the optical image generator that also adopts a generation mode from rough to fine. The finally generated optical image not only reflects the consistency of global features but also retains the richness of local details, making the generated image have high clarity and accurate color information. Correspondingly, the discriminator also adopts a dual discrimination strategy to discriminate the images output by the two-step generator respectively.
[0011] The method includes the following steps: Step 1, construct a grayscale image generator. Through the joint processing of the global generator and the local generator, convert the input SAR image into a grayscale image with clear edges and structural information; Step 2, construct a global feature extractor. Use multiple global feature extraction modules based on the Swin Transformer encoder to extract deep global features through skip connections to obtain global feature information; Step 3, construct an optical image generator. Take the fused grayscale image and global feature information as the input. Through the joint processing of the global generator and the local generator, map the fused features to a color optical image to obtain the target optical image; Step 4, use a dual discriminator structure to discriminate the grayscale image generated by the grayscale image generator and the optical image generated by the optical image generator, and perform generative adversarial training through multiple loss functions; Step 5, after the generative adversarial training is completed, use Steps 1 - 3 to convert the grayscale image to be converted into an optical image.
[0012] Furthermore, the network structure of the global generator includes three 3×3 convolutional layers with a stride of 2, three IN layers, three ReLU activation layers, and a series of residual blocks, as well as three 3×3 transposed convolutional layers with a stride of 2, two IN layers, two ReLU activation layers, and one Tanh activation layer. The input is the downsampled SAR image, and its output is the corresponding rough conversion image, which is further input into the local generator.
[0013] Furthermore, the network structure of the local generator includes two 3×3 convolutional layers with a stride of 2, three IN layers, three ReLU activation layers, and a series of residual blocks, as well as three 3×3 transposed convolutional layers with a stride of 2, two IN layers, two ReLU activation layers, and one Tanh activation layer. The input is the SAR image, and the rough conversion image is also input into a series of residual blocks for processing. The final output is a clear grayscale image.
[0014] Furthermore, the global feature extraction module adopts the window multi-head self-attention (W-MSA) and shifted window multi-head self-attention (SW-MSA) mechanisms. Among them, local features are extracted by the W-MSA layer, and cross-window global features are captured by the SW-MSA layer; After three stacked global feature extraction modules, the output Y3 is obtained. Through upsampling and activation layers, finally, the dimensions are aligned and its size is adjusted to be consistent with the input features to obtain the global feature information Y.
[0015] Furthermore, the specific processing process of the global feature extraction module in Step 2 is as follows: Step 2.1, SAR image First, it passes through two 3×3 initial convolutional layers and ReLU activation layers, and then undergoes downsampling to obtain , and enters the first global feature extraction module; Step 2.2, within the local window of Module 1, a W-MSA layer is introduced. For the input feature matrix , first perform layer normalization, and then make a skip connection between the output of the self-attention module and the original input to obtain , as shown in the following formula:
[0016] Then it passes through the MLP module. The MLP introduces non-linear transformation through the activation function to make the feature representation more abundant, and adopts skip connection to obtain the output of Module 1:
[0017] Step 2.3, enter Module 2. In order to capture the global dependency information across windows, an SW-MSA layer is introduced. First, perform a shifting operation on the input feature , so that there is an overlap between originally non-adjacent windows, and then perform the same calculation as W-MSA to calculate the self-attention to obtain , and perform a skip connection output to obtain :
[0018] Restore the original arrangement order before the MLP module to perform an inverse shifting operation , and adopt skip connection to obtain the output of Module 2, that is, the output of the first global feature extraction module :
[0019] Repeat Step 2.2 and Step 2.3, and output after passing through three stacked global feature extraction modules .
[0020] Furthermore, each discriminator in the dual discriminator structure contains three stacked PatchGAN networks, which respectively discriminate the original scale, twice downsampled, and four times downsampled images of the grayscale image and the target optical image; Each PatchGAN network is composed of three 4×4 convolutional layers with a stride of 2, two 4×4 convolutional layers with a stride of 1, three IN normalization layers, and four ReLU activation layers.
[0021] Furthermore, the multiple loss functions include the generator loss and the discriminator loss; the generator loss is composed of three parts: the grayscale image generation network loss, the optical image generation network loss, and the structural similarity loss, and its total loss function is:
[0022] ① Grayscale image generation network loss : This loss term is used to measure the difference between the features of the generated high-resolution image and the real image. At the same time, a feature matching loss is added. The expression is as follows:
[0023] Among them, is the reconstruction loss, which measures the pixel difference between the generated image and the real grayscale image; is the result of the th downsampling of the generated grayscale image; is the corresponding real grayscale image; is the number of layers of the discriminator and also the scale number of downsampling; is the weight coefficient of the feature matching loss, is the feature matching loss, which is used to maintain the consistency between multi-scale features. The specific calculation formula is as follows:
[0024] Among them, is the output of the grayscale image discriminator, refer to the generated grayscale image and the real grayscale image respectively; ② Optical image generation network loss : It is used to optimize the similarity between the generated optical image and the real color image, and also includes a feature matching loss term; the calculation formula is:
[0025] Among them, is the result of the th downsampling of the generated color optical image, represents the real color image, is the feature matching loss for optical image generation, and the calculation formula is the same as ; ③ Structural similarity loss : To enhance the structural consistency of the generated image, SSIM is used to calculate the structural similarity between the generated image and the target image. The calculation formula is:
[0026] Among them, is the generated grayscale image; is the real grayscale image; is the generated color optical image; is the real optical image.
[0027] Further, the discriminator loss consists of two parts: the grayscale image discrimination loss and the optical image discrimination loss, and its total loss function is the formula:
[0028] Grayscale image discrimination loss : Used to measure the discrimination ability of the grayscale image discriminator for real images and generated images, and the calculation formula is:
[0029] Among them, is the result of the th downsampling of the generated grayscale image, is determined as a real grayscale image; is determined as a generated grayscale image; Optical image discrimination loss : Used to measure the effect of the optical image discriminator, and the calculation formula is:
[0030] Among them, is the result of the th downsampling of the generated optical image, is determined as a real optical image; is determined as a generated optical image.
[0031] Further, step 5 also includes evaluating the generated optical image using the peak signal-to-noise ratio PSNR, cross-correlation CC, structural similarity SSIM, and visual information fidelity VIF.
[0032] The present invention also provides an image conversion system based on the self-attention mechanism and the generative adversarial network, including: A processor and a memory. The memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the image conversion method based on the self-attention mechanism and the generative adversarial network as described in the above technical solution.
[0033] Advantages and beneficial effects of the present invention compared with the prior art: One of the innovations of the present invention is to propose an image conversion method based on self-attention mechanism and generative adversarial, aiming to improve the conversion quality from SAR images to optical images. That is, by introducing the self-attention mechanism into the generator, the deficiencies of traditional generative adversarial networks (GANs) in dealing with complex textures and long-range dependencies are successfully overcome. Another innovation is the adoption of a two-step generation strategy, which divides the image conversion process into two independent steps, thus effectively improving the quality and detail performance of the generated images. Specifically, in the first step, through the joint processing of the global and local generators, the input SAR image is converted into a grayscale image with clear edges and structural information. In addition, a global feature extractor is designed to supplement the information of the grayscale image. In the second step, the global feature information of the grayscale image is fused and converted into a high-quality optical image. This two-step generation design not only optimizes the model structure, but also enhances the interpretability and controllability of the network, providing convenience for the independent optimization of each step. Experimental results show that the present invention is significantly superior to other comparison models in multiple evaluation indicators. In terms of the reproduction of ground object information, the present invention can accurately generate the details of optical images, especially outstanding in the performance of edge sharpness and spectral information. In addition, the residual images also show the advantages of the present method in spectral information and texture generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is the overall structure diagram of the SAR-to-optical image network.
[0035] Figure 2 It is the structure diagram of the grayscale image generator.
[0036] Figure 3 It is the structure diagram of the global feature extractor.
[0037] Figure 4 It is the structure diagram of the optical image generator.
[0038] Figure 5 It is the structure diagram of the multi-scale discriminator.
[0039] Figure 6 It is the subjective visual comparison result of the results of different methods on the dataset SEN1-2.
[0040] Figure 7 It is the residual comparison result of the results of different methods on the dataset SEN1-2. DETAILED DESCRIPTION OF THE INVENTION
[0041] For the convenience of those of ordinary skill in the art to understand and implement the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0042] In view of the problem that the generalization ability of traditional image conversion networks is limited in scenarios with extremely large cross-modal differences such as SAR and optical, a two-step generation network is proposed, namely an image conversion model based on the self-attention mechanism and the generative adversarial network. The network mainly consists of two generation stages: grayscale image generation and optical image generation. In addition, a global feature extractor is introduced to supplement information for the first generation stage. Correspondingly, there are also two discriminators in the discrimination stage, both of which adopt a multi-scale discrimination strategy. The overall structure diagram of the SAR-to-optical image network is as Figure 1 shown. The method specifically includes the following steps: The method includes the following steps: Step 1, grayscale image generation. In order to improve the clarity of the generated image, the generator is designed in a rough-to-fine process, consisting of a global generator and a local generator, both based on the residual network architecture. The model structure is as Figure 2 shown. First, the global generator needs to be pre-trained during the training process. The input is the downsampled SAR image, and its output is the corresponding rough conversion image. By downsampling the input image, the global generator can capture image information in a larger range, thus focusing on the global consistency of the generated image at a larger scale. After completing the preliminary image generation, the model removes the last two layers of the global generator, and the output features are input into the local generator. The input of the local generator includes not only the feature map output by the global generator, but also performs a two-fold downsampling operation on the original input image, integrating information at multiple scales into the network. This design ensures that the model can process image details at a finer level, making the finally generated image show excellent clarity and visual effects at all levels.
[0043] Step 2, global feature supplementation. The global feature extractor uses three global feature extraction modules based on the Swin Transformer encoder to extract deep global features through skip connections, obtaining a model structure such as Figure 3 shown. The global feature extraction module can effectively capture local details and global dependencies in the image. Each global feature extraction module adopts the W-MSA (window multi-head self-attention) and SW-MSA (shifted window multi-head self-attention) mechanisms. Among them, local features are extracted by the W-MSA layer, and cross-window global features are captured by the SW-MSA layer. Repeat the above process, and after three stacked global feature extraction modules, the output Y3 is obtained. Through upsampling and activation layers, finally, the dimensions are aligned and adjusted to a size of 256×256×1 to keep it consistent with the input features, obtaining the global feature information Y.
[0044] Step 3, optical image generation. As Figure 4As shown in the figure, through the generator and global feature extractor in the first step, a grayscale image with clear edges and structural information and global feature information are obtained. The two are fused to obtain As input into the network in the second step. In this way, the generator can not only retain the global information of the original input and richer local details, improve the quality of the generated image, but also enhance the interpretability of the entire neural network.
[0045] Step 4: Use a dual discriminator structure to discriminate the grayscale image generated by the grayscale image generator and the optical image generated by the optical image generator, and perform generative adversarial training through multiple loss functions; Corresponding to the two-step generation, the discriminator also adopts a dual discrimination strategy to discriminate the images output by the two-step generator respectively. The discriminator model combines a multi-scale discrimination mechanism from coarse to fine, which can independently evaluate the global consistency and local details of the image at different resolutions, thereby effectively improving the quality and authenticity of the generated image.
[0046] Specifically, the discriminator structure contains three stacked PatchGAN networks, which respectively discriminate the original scale, twice downsampled, and four times downsampled images of the grayscale image and the target optical image. Its model structure is as Figure 5 shown.
[0047] Each PatchGAN network consists of three 4×4 convolutional layers with a stride of 2, two 4×4 convolutional layers with a stride of 1, three IN normalization layers, and four ReLU activation layers. This multi-scale discrimination design enables the model to maintain a high discrimination ability at different levels of image details, ensuring that the generator has consistency and fineness in both global structure and local details, thereby motivating the generator to generate more delicate and realistic images.
[0048] Step 5: After the generative adversarial training is completed, use Steps 1 - 3 to convert the grayscale image to be converted into an optical image.
[0049] Furthermore, the specific implementation of Step 1 includes the following sub-steps, Step 1.1: Global generator: The network structure includes three 3×3 convolutional layers with a stride of 2, three IN layers, three ReLU activation layers, and a series of residual blocks, as well as three 3×3 transposed convolutional layers with a stride of 2, two IN layers, two ReLU activation layers, and finally a Tanh activation layer. These components can effectively extract the global features in the image. In the global generator, after multiple convolutional and normalization processes, the important features in the image are gradually extracted, and the features are retained through the residual blocks to ensure the stability of information transmission.
[0050] Step 1.2, Local Generator: After the initial image generation is completed, the model removes the last two layers of the global generator, and the output features are input into the local generator. In the local generator, the network still consists of convolutional layers, ReLU activation layers, residual blocks, and deconvolutional layers, and finally uses a Tanh activation layer. These network layers work together to further extract local detail information.
[0051] The mathematical expressions of the two sub-steps are shown as follows:
[0052]
[0053] where, is the downsampled version of the input image, and represent the function mappings of the global generator and the local generator in the grayscale image generator respectively, and the finally output feature is the clear grayscale image.
[0054] Furthermore, the specific implementation of Step 2 includes the following sub-steps. Step 2.1, SAR Image First, it passes through two 3×3 initial convolutional layers and ReLU activation layers, and then is downsampled to obtain , and enters the first global feature extraction module.
[0055] Each global feature extraction module uses the W-MSA (Window Multi-Head Self-Attention) and SW-MSA (Shifted Window Multi-Head Self-Attention) mechanisms. Among them, local features are extracted by the W-MSA layer, and cross-window global features are captured by the SW-MSA layer.
[0056] Step 2.2, Within the local window of Module 1, for the input feature matrix , first perform layer normalization, and then make a skip connection between the output of the self-attention module and the original input to obtain , as shown in the following formula:
[0057] Then, it passes through the MLP module. The MLP introduces non-linear transformation through the activation function to make the feature representation more abundant, and uses a skip connection to obtain the output of Module 1:
[0058] Step 2.3, Enter Module 2. In order to capture cross-window global dependency information, SW-MSA is introduced here. First, a shift operation is performed on the input feature , causing overlap between originally non - adjacent windows, and then performing the same W - MSA calculation to obtain self - attention as , performing skip connection output to obtain :
[0059] Restore the original arrangement order before the MLP module to perform an anti - shift operation , using skip connection to obtain the output of module 2, that is, the output of the first global feature extraction module :
[0060] Repeat the above process, and output through three stacked global feature extraction modules , through upsampling and activation layers, and finally align the dimensions and adjust its size to 256×256×1, keeping it consistent with the input features to obtain global feature information .
[0061] Generally speaking, first, module 1 uses W - MSA to capture the features of the local area. Next, through skip connection, the features are further processed by normalization (LN layer) and multi - layer perceptron (MLP) modules. Then, the processed features are input into module 2, and the adopted SW - MSA performs a shift operation between windows, breaking the independence between windows and allowing global information interaction across windows. In this way, while maintaining the local feature extraction ability, it can also capture global dependencies and enhance the feature expression ability.
[0062] Furthermore, the specific implementation of step 3 includes, Through the generator and global feature extractor in the first step, a grayscale image with clear edges and structural information and global feature information are obtained. The two are fused to obtain as input into the network in the second step:
[0063] where is the fusion coefficient, set to 0.9. In this way, the generator can not only retain the global information of the original input and more abundant local details, improving the quality of the generated image, but also enhancing the interpretability of the entire neural network. The structure of the optical image generator in the second step is the same as that of the grayscale generator, and it is also composed of convolutional layers, downsampling, activation layers, residual blocks, upsampling, and transposed convolutional layers, etc. It can map the fused features to a color optical image, and the mathematical expressions are respectively as follows:
[0064]
[0065] Among them, is the downsampled version of the input image , and respectively represent the function mappings of the global generator and the local generator in the optical image generator, and the finally output feature is the target optical image.
[0066] In step 4, the following loss function is used as the guidance for network optimization: The loss functions of the generator and the discriminator are the key parts of the SAR-optical image conversion model, and through the combination of multiple loss terms, the quality and consistency of the generated images are ensured.
[0067] (1) Generator loss: The generator loss consists of three parts: the grayscale image generation network loss, the optical image generation network loss, and the structural similarity (SSIM) loss. Its total loss function is:
[0068] ① Grayscale image generation network loss : It is used to measure the difference between the features of the generated grayscale image and the real image, and at the same time, a feature matching loss is added to improve the stability of the generated image. The calculation formula is:
[0069] Among them, is the reconstruction loss, which measures the pixel difference between the generated image and the real grayscale image; is the result of the -th downsampling of the generated grayscale image. In this embodiment, takes 0, 1, 2. When i takes 0, it represents the original image. When i takes 1, it represents 2-fold downsampling. When i takes 2, it represents 4-fold downsampling; is the real grayscale image corresponding to the scale; is the number of layers of the discriminator and also the number of downsampling scales. In this embodiment, it takes 3; is the feature matching loss, which is used to maintain the consistency between multi-scale features; is the weight coefficient of the feature matching loss. Its specific calculation formula is as follows:
[0070] Among them, is the output of the grayscale image discriminator, respectively refer to the generated grayscale image and the real grayscale image.
[0071] ② Optical Image Generation Network Loss : It is used to optimize the similarity between the generated optical image and the real color image, and also contains a feature matching loss term. The calculation formula is:
[0072] Among them, is the result of the -th downsampling of the generated color optical image, is the feature matching loss generated for the optical image, and the calculation formula is the same as above and will not be elaborated here.
[0073] ③ Structural Similarity Loss : To enhance the structural consistency of the generated image, SSIM is used to calculate the structural similarity between the generated image and the target image. The calculation formula is:
[0074] Among them, is the generated grayscale image; is the real grayscale image; is the generated color optical image; is the real optical image.
[0075] (2) Discriminator Loss: It consists of two parts: grayscale image discrimination loss and optical image discrimination loss. Its total loss function is the formula:
[0076] Grayscale Image Discrimination Loss : It is used to measure the discrimination ability of the grayscale image discriminator for real images and generated images. The calculation formula is:
[0077] Among them, is the result of the -th downsampling of the generated grayscale image, is determined to be a real grayscale image; is determined to be a generated grayscale image. Among them, when the discriminator discriminates a real image, it takes 1, and when it is a generated image, it takes 0.
[0078] Optical Image Discrimination Loss : It is used to measure the effect of the optical image discriminator. The calculation formula is:
[0079] Among them, is the result of the -th downsampling of the generated optical image, is determined to be a real optical image; It is determined to generate an optical image.
[0080] This multi-level loss design enables the model to perform excellently in grayscale and optical image generation tasks, ensuring that the generated images have high fidelity and consistency in both global structure and local details.
[0081] Based on the above steps, the SAR-to-optical results are obtained. To compare with other methods, we use four classic image conversion methods, Pix2Pix, Pix2PixHD, CycleGAN, and NICE-GAN, to conduct a comparison on the publicly available remote sensing dataset SEN1-2. The results are shown in the appendix Figure 6 - Appendix Figure 7 as shown.
[0082] To quantitatively evaluate the experimental results, we use peak signal-to-noise ratio (PSNR), cross-correlation (CC), structural similarity (SSIM), and visual information fidelity (VIF) as evaluation metrics for optical image generation. The quantitative comparison results on the publicly available SEN1-2 dataset are as follows: Table 1 Quantitative comparison results on the SEN1-2 dataset
[0083] Among them, bold indicates the best, and italic indicates the second best; the quantitative index results show that the optical images obtained by the method proposed in the present invention are superior to the existing methods and can accurately generate the spatial and spectral details of the optical images.
[0084] On the other hand, the embodiment of the present invention also provides an image conversion system based on self-attention mechanism and generative adversarial network, including: A processor and a memory. The memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the image conversion method based on self-attention mechanism and generative adversarial network as described in the above technical solution.
[0085] It should be understood that the parts not elaborated in detail in this specification all belong to the prior art.
[0086] It should be understood that the above description of the embodiments is relatively detailed, and it should not be considered as a limitation to the protection scope of the present invention patent. Those of ordinary skill in the art, under the inspiration of the present invention, without departing from the protection scope defined by the claims of the present invention, can also make substitutions or modifications, which all fall within the protection scope of the present invention. The scope of protection requested by the present invention shall be subject to the appended claims.
Claims
1. An image conversion method based on self-attention mechanism and generative adversarial network, characterized in that It includes the following steps: Step 1, construct a grayscale image generator. Through the joint processing of the global generator and the local generator, convert the input SAR image into a grayscale image with clear edges and structural information; Step 2, construct a global feature extractor. Use multiple global feature extraction modules based on the Swin Transformer encoder to extract deep global features through skip connections to obtain global feature information; Step 3, construct an optical image generator. Take the fused grayscale image and global feature information as the input. Through the joint processing of the global generator and the local generator, map the fused features to a color optical image to obtain the target optical image; Step 4, use a dual discriminator structure to discriminate the grayscale image generated by the grayscale image generator and the optical image generated by the optical image generator, and perform generative adversarial training through multiple loss functions; Step 5, after the generative adversarial training is completed, use Steps 1 - 3 to convert the grayscale image to be converted into an optical image.
2. The image conversion method based on the self-attention mechanism and the generative adversarial network according to claim 1, characterized in that: The network structure of the global generator includes three 3×3 convolutional layers with a stride of 2, three IN layers, three ReLU activation layers and a series of residual blocks, and three 3×3 transposed convolutional layers with a stride of 2, two IN layers, two ReLU activation layers, and one Tanh activation layer. The input is the downsampled SAR image, and its output is the corresponding rough conversion image, which is further input into the local generator.
3. The image conversion method based on the self-attention mechanism and the generative adversarial network according to claim 2, characterized in that: The network structure of the local generator includes two 3×3 convolutional layers with a stride of 2, three IN layers, three ReLU activation layers and a series of residual blocks, and three 3×3 transposed convolutional layers with a stride of 2, two IN layers, two ReLU activation layers, and one Tanh activation layer. The input is the SAR image, and the rough conversion image is also input into a series of residual blocks for processing. The final output is a clear grayscale image.
4. The image conversion method based on the self-attention mechanism and the generative adversarial network according to claim 1, wherein: The global feature extraction module adopts the window multi-head self-attention W-MSA and shifted window multi-head self-attention SW-MSA mechanisms. Among them, local features are extracted by the W-MSA layer, and cross-window global features are captured by the SW-MSA layer; after three stacked global feature extraction modules, the output Y3 is obtained. Through upsampling and activation layers, finally, the dimensions are aligned and its size is adjusted to be consistent with the input features to obtain the global feature information Y.
5. The image conversion method based on the self-attention mechanism and the generative adversarial network according to claim 1 or 4, characterized in that: The specific processing process of the global feature extraction module in Step 2 is as follows: Step 2.1, SAR image First, it passes through two 3×3 initial convolutional and ReLU activation layers, and then is downsampled to obtain , and enters the first global feature extraction module; Step 2.2, within the local window of Module 1, introduce the W-MSA layer and input the feature matrix , first perform layer normalization, and then make a skip connection between the output of the self-attention module and the original input to obtain , as shown in the following formula: Then pass through the MLP module. The MLP introduces non-linear transformation through the activation function to make the feature representation more abundant, and uses skip connections to obtain the output of Module 1: Step 2.3, Enter Module 2 To capture global dependency information across windows, an SW-MSA layer is introduced. First, a shifting operation is performed on the input features to make the originally non-adjacent windows overlap, and then the same W-MSA calculation is carried out to obtain self-attention , and a skip connection is made to output : Restore the original permutation order before the MLP module for the reverse shift operation , and use skip connections to obtain the output of Module 2, which is the output of the first global feature extraction module : Repeat steps 2.2 and 2.3, and output after passing through three stacked global feature extraction modules .
6. The image conversion method based on the self-attention mechanism and the generative adversarial network according to claim 1, characterized in that: Each discriminator in the dual discriminator structure contains three stacked PatchGAN networks, which respectively discriminate the original scale, twice downsampled, and four times downsampled images of the grayscale image and the target optical image; Each PatchGAN network is composed of three 4×4 convolutional layers with a stride of 2, two 4×4 convolutional layers with a stride of 1, three IN normalization layers, and four ReLU activation layers.
7. The image conversion method based on the self-attention mechanism and the generative adversarial network according to claim 1, characterized in that: The multiple loss functions include the generator loss and the discriminator loss; the generator loss consists of three parts: the grayscale image generation network loss, the optical image generation network loss, and the structural similarity loss, and its total loss function is: ① Grayscale image generation network loss : This loss term is used to measure the difference between the features of the generated high-resolution image and the features of the real image. At the same time, a feature matching loss is added, and the expression is as follows: Among them, is the reconstruction loss, which measures the pixel difference between the generated image and the real grayscale image; is the result of the -th downsampling of the generated grayscale image. i starts from 0. When i = 0, it represents the original image; is the corresponding real grayscale image; is the number of layers of the discriminator and also the scale number of downsampling; is the weight coefficient of the feature matching loss, is the feature matching loss, which is used to maintain the consistency between multi-scale features. The specific calculation formula is as follows: Among them, is the output of the grayscale image discriminator, which respectively refer to the generated grayscale image and the real grayscale image; ② Optical image generation network loss : Used to optimize the similarity between the generated optical image and the real color image, and also includes a feature matching loss term; the calculation formula is: Among them, is the th downsampling result of the generated color optical image, represents the real color image, is the feature matching loss generated for the optical image, and the calculation formula is the same as ; ③ Structural similarity loss : To enhance the structural consistency of the generated images, the SSIM is used to calculate the structural similarity between the generated images and the target images. The calculation formula is as follows: Among them, is the generated grayscale image; is the real grayscale image; is the generated color optical image; is the real optical image.
8. The image conversion method based on the self-attention mechanism and the generative adversarial network according to claim 7, wherein: The discriminator loss consists of two parts: the grayscale image discrimination loss and the optical image discrimination loss, and its total loss function is the formula: Gray image discrimination loss : Used to measure the discrimination ability of the gray image discriminator for real images and generated images. The calculation formula is as follows: Among them, is the result of the th downsampling of the generated grayscale image, is determined to be a real grayscale image; is determined to be a generated grayscale image. Optical image discrimination loss : Used to measure the effect of the optical image discriminator, and the calculation formula is: Among them, is the result of the th downsampling of the generated optical image, is determined to be a real optical image; is determined to be a generated optical image.
9. The image conversion method based on the self-attention mechanism and the generative adversarial network according to claim 1, characterized in that: Step 5 further includes evaluating the generated optical image using the peak signal-to-noise ratio PSNR, the cross-correlation CC, the structural similarity SSIM, and the visual information fidelity VIF.
10. An image conversion system based on self-attention mechanism and generative adversarial network, characterized in that, Including: A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the image conversion method based on the self-attention mechanism and the generative adversarial network according to any one of claims 1-9.