Method and system for converting SAR (Synthetic Aperture Radar) into optical image based on segmentation guidance
By using a segmentation-guided high-resolution conditional GAN network, combined with Pix2PixHD and CycleGAN networks, the image mismatch problem when converting SAR images to optical images is solved, achieving accurate generation of high-resolution optical images. This addresses the issues of unclear texture details and loss of small-area targets in the generated images, thereby improving the effectiveness of image analysis.
Patent Information
- Application Number
- CN202511153818.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-11-28
AI Technical Summary
In existing technologies, when SAR images are converted into optical images, the resulting images have insufficient texture details and edge information, abnormal colors, and small targets are easily lost. This leads to a mismatch between the generated optical images and the real scene, affecting information interpretation and analysis.
A high-resolution conditional GAN network based on segmentation guidance is adopted, which combines Pix2PixHD and CycleGAN networks. A pre-trained segmentation network is introduced, and the segmentation loss is calculated by extracting feature maps through the generator. The generative adversarial network is optimized by combining a multi-scale discriminator and cycle consistency loss to generate high-resolution optical images.
It effectively preserves small target areas, improves the accuracy and clarity of generated images, ensures that generated images match real-world scenes, and enhances the effectiveness of image analysis.
Smart Images

Figure CN121032832A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image conversion technology, and in particular to a segmentation-guided SAR-to-optical image conversion method and system. Background Technology
[0002] With the development of remote sensing technology, Synthetic Aperture Radar (SAR) and optical remote sensing images are increasingly used in mapping, navigation, target detection, and earth science. SAR is a microwave imaging radar that can generate high-resolution images using electromagnetic waves. Due to the penetrating properties of electromagnetic waves, SAR images can compensate for the limitations of visible light electromagnetic waves in penetrating fog, clouds, and obstructions. This all-weather, all-day operation allows SAR images to contain more unique information than optical images.
[0003] While SAR offers the advantage of capturing images around the clock and in all weather conditions, several challenges remain. First, the color information of land cover types presented in SAR images differs significantly from that in optical images, hindering the visual interpretation of the observed landscape by the human eye, whereas optical images align perfectly with human observation. Second, SAR images suffer from geometric distortion and speckle noise, causing a mismatch between SAR images and the physical structures of the real environment. Furthermore, the noise interference makes it difficult to extract valuable information from SAR images, making them more challenging to analyze than optical images.
[0004] With the rapid development of deep learning, techniques for converting SAR images into optical images to aid SAR image interpretation have gradually improved. However, existing methods for converting SAR images to optical images suffer from several drawbacks. The generated images often lack clear and complete texture details and edge information. Furthermore, due to the lack of color information in SAR images, color anomalies may occur during the conversion process. Improved methods primarily focus on pixel-level and scene-level similarity, which can lead to the loss of small-area targets in the generated images. All these issues result in the ineffective representation of SAR image information and a mismatch between the generated optical images and the real-world scene, significantly impacting researchers' interpretation and analysis of the information contained within SAR images. Summary of the Invention
[0005] The purpose of this invention is to provide a segmentation-guided SAR to optical image conversion method and system to solve the problem of mismatch between the generated optical image and the real scene in the prior art.
[0006] To achieve the above objectives, the present invention provides a segmentation-guided SAR-to-optical image conversion method, comprising:
[0007] S1. Acquire SAR images, input the SAR images into the generator of the improved high-resolution conditional GAN network, extract SAR image features and generate predicted optical images, and input the multi-scale feature maps extracted from the SAR images by the cascaded residual module in the generator into the pre-trained segmentation loss network to obtain the segmentation loss.
[0008] S2. Input the predicted optical image and the paired optical image and its downsampled image into the multi-scale discriminator to calculate the adversarial loss;
[0009] S3. Introduce a recurrent consistency structure using the same network structure as the generative adversarial network described above. Use the real optical image as the generator input, and the predicted SAR image and paired SAR images along with their downsampled images as the multi-scale discriminator input, to obtain the recurrent consistency loss.
[0010] S4. Optimize the generative adversarial network based on segmentation loss, adversarial loss, and feature matching loss, and use the generative adversarial network to generate high-resolution optical images obtained from SAR image conversion.
[0011] In some embodiments of this application, in S1, the improved high-resolution conditional GAN network is a combined network of the improved conditional GAN network and the CycleGAN network.
[0012] In some embodiments of this application, in S1, the DeepLabv3 network is pre-trained to obtain the target accuracy network parameters after freezing the pre-trained target accuracy network parameters, and the segmentation loss network is obtained based on the target accuracy network parameters.
[0013] In some embodiments of this application, in S4, the generative adversarial network is obtained based on the Pix2PixHD model.
[0014] In some embodiments of this application, in S4, the generator of the generative adversarial network includes a global generation module, a local enhancement module, and a residual module, and a dilated convolution module is added after the residual module to increase the receptive field of the upsampling part.
[0015] The discriminator in the generative adversarial network employs a multi-scale discriminator to generate high-resolution images.
[0016] In some embodiments of this application, in S4, the feature matching loss is the result of the L1 norm between the feature map generated by each layer of multiple multi-scale discriminators and the feature map generated by the discriminator acting on the actual image, which is used to supervise the generated image to contain feature details of the original image.
[0017] In some embodiments of this application, it also includes:
[0018] The optical image generated by S4 is input into a preset optical-SAR image generation network to obtain a pseudo-SAR image for verification. The L1 loss between the pseudo-SAR image and the SAR image obtained by S1 is calculated, and the generative adversarial network is optimized.
[0019] In some embodiments of this application, a segmentation-guided SAR-to-optical image conversion system is also disclosed, comprising:
[0020] The acquisition module is used to acquire SAR images. It inputs the SAR images into the generator of the improved high-resolution conditional GAN network, extracts SAR image features and generates a predicted optical image. The multi-scale feature map extracted from the SAR image by the cascaded residual module in the generator is input into a pre-trained segmentation loss network to obtain the segmentation loss.
[0021] The computation module is used to input the predicted optical image and the paired optical image and its downsampled image into the multi-scale discriminator to calculate the adversarial loss;
[0022] The recurrent module is used to predict SAR images as input from optical images and calculate the recurrent consistency loss together with the above network.
[0023] The generation module is used to optimize the generative adversarial network based on segmentation loss, adversarial loss, and feature matching loss, and to generate high-resolution optical images obtained from SAR image conversion using the generative adversarial network.
[0024] The advantages and beneficial effects of this invention compared to the prior art are:
[0025] 1. The SAR-to-optical image conversion method based on a segmentation-guided high-resolution conditional GAN network provided by this invention introduces a pre-trained segmentation network, calculates the segmentation loss using the original image feature map extracted by the generator, and incorporates the segmentation loss into the total loss function. Compared to the problem of potential loss of small-area targets in previous SAR-to-optical image conversion methods, this method can alleviate the problem of small-area target loss by constraining the effective preservation of segmented targets during the generation process.
[0026] 2. This invention combines the ideas of Pix2PixHD and CycleGAN networks. It utilizes an improved Pix2PixHD network to generate high-resolution images. The number of convolutional layers and cascaded residual modules in the Pix2PixHD network can be manually adjusted according to the resolution of the input image to obtain more detailed features. Adding dilated convolutions captures more detailed information from the real image, improving the accuracy of the generated image. Simultaneously, the use of a multi-scale discriminator allows for a larger receptive field, enabling better differentiation between high-resolution real images and synthetic images.
[0027] 3. This invention utilizes a method of first training a network on low-resolution images, and then using a combination of the pre-trained network and the original network to train a network on high-resolution images, which can improve the quality of the generated high-resolution images. Furthermore, the cycle consistency loss introduced by the CycleGAN model can effectively constrain the consistency of individual image types between the original and generated images.
[0028] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0029] Figure 1 This is a schematic diagram illustrating the steps of a segmentation-guided SAR-to-optical image conversion method in an embodiment of the present invention;
[0030] Figure 2 This is a structural diagram of a segmentation-guided SAR-to-optical image conversion system according to an embodiment of the present invention;
[0031] Figure 3 This is a schematic diagram of the overall network structure according to an embodiment of the present invention;
[0032] Figure 4 This is a detailed view of the improved residual portion according to an embodiment of the present invention;
[0033] Figure 5 This is a detailed view of the dilated convolution part in an embodiment of the present invention. Detailed Implementation
[0034] In the description of this invention, it should be noted that the terms "upper," "lower," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product is in use. They are used only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," and "connect" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0035] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0036] like Figure 1As shown, this invention provides a segmentation-guided SAR-to-optical image conversion method, comprising:
[0037] S1. Acquire SAR images, input the SAR images into the generator of the improved high-resolution conditional GAN network, extract SAR image features and generate predicted optical images, and input the multi-scale feature maps extracted from the SAR images by the cascaded residual module in the generator into the pre-trained segmentation loss network to obtain the segmentation loss.
[0038] S2. Input the predicted optical image and the paired optical image and its downsampled image into the multi-scale discriminator to calculate the adversarial loss.
[0039] S3. Introduce a recurrent consistency structure using the same network structure as the generative adversarial network described above. Use the real optical image as the generator input, and the predicted SAR image and paired SAR images along with their downsampled images as the multi-scale discriminator input, to obtain the recurrent consistency loss.
[0040] S4. Optimize the generative adversarial network based on segmentation loss, adversarial loss, feature matching loss and cycle consistency loss, and use the generative adversarial network to generate high-resolution optical images obtained from SAR image conversion.
[0041] In some embodiments of this application, in S1, the improved high-resolution conditional GAN network is a combined network of the improved conditional GAN network and the CycleGAN network.
[0042] In some embodiments of this application, in S1, the DeepLabv3 network is pre-trained to obtain the target accuracy network parameters after freezing the pre-trained target accuracy network parameters, and the segmentation loss network is obtained based on the target accuracy network parameters.
[0043] In some embodiments of this application, in S4, the generative adversarial network is obtained based on the Pix2PixHD model.
[0044] In some embodiments of this application, in S4, the generator of the generative adversarial network includes a global generation module, a local enhancement module, and a residual module, and a dilated convolution module is added after the residual module to increase the receptive field of the upsampling part.
[0045] The discriminator in the generative adversarial network employs a multi-scale discriminator to generate high-resolution images.
[0046] In some embodiments of this application, in S4, the feature matching loss is the result of the L1 norm between the feature map generated by each layer of multiple multi-scale discriminators and the feature map generated by the discriminator acting on the actual image, which is used to supervise the generated image to contain feature details of the original image.
[0047] In some embodiments of this application, it also includes:
[0048] The optical image generated by S4 is input into a preset optical-SAR image generation network to obtain a pseudo-SAR image for verification. The L1 loss between the pseudo-SAR image and the SAR image obtained by S1 is calculated, and the generative adversarial network is optimized.
[0049] In some embodiments of this application, a segmentation-guided SAR-to-optical image conversion system is also disclosed, comprising:
[0050] The acquisition module is used to acquire SAR images. The SAR images are input into the generator of the improved high-resolution conditional GAN network to extract SAR image features and generate predicted optical images. The multi-scale feature maps extracted from the SAR images by the cascaded residual module in the generator are input into a pre-trained segmentation loss network to obtain the segmentation loss.
[0051] The computation module is used to input the predicted optical image and the paired optical image and its downsampled image into the multi-scale discriminator to calculate the adversarial loss.
[0052] The recurrent module is used to take the optical image as input, predict the SAR image, and calculate the recurrent consistency loss together with the above network.
[0053] The generation module is used to optimize the generative adversarial network based on segmentation loss, adversarial loss, and feature matching loss, and to generate high-resolution optical images obtained from SAR image conversion using the generative adversarial network.
[0054] The advantages and beneficial effects of this invention compared to the prior art are:
[0055] 1. The SAR-to-optical image conversion method based on a segmentation-guided high-resolution conditional GAN network provided by this invention introduces a pre-trained segmentation network, calculates the segmentation loss using the original image feature map extracted by the generator, and incorporates the segmentation loss into the total loss function. Compared to the problem of potential loss of small-area targets in previous SAR-to-optical image conversion methods, this method can alleviate the problem of small-area target loss by constraining the effective preservation of segmented targets during the generation process.
[0056] 2. This invention combines the ideas of Pix2PixHD and CycleGAN networks. It utilizes an improved Pix2PixHD network to generate high-resolution images. The number of convolutional layers and cascaded residual modules in the Pix2PixHD network can be manually adjusted according to the resolution of the input image to obtain more detailed features. Adding dilated convolutions captures more detailed information from the real image, improving the accuracy of the generated image. Simultaneously, the use of a multi-scale discriminator allows for a larger receptive field, enabling better differentiation between high-resolution real images and synthetic images.
[0057] 3. This invention utilizes a method of first training a network on low-resolution images, and then using a combination of the pre-trained network and the original network to train a network on high-resolution images, which can improve the quality of the generated high-resolution images. Furthermore, the cycle consistency loss introduced by the CycleGAN model can effectively constrain the consistency of individual image types between the original and generated images.
[0058] The system implementation of the present invention will be described in detail below with reference to specific embodiments.
[0059] This embodiment provides a SAR-to-optical image conversion method based on a segmentation-guided high-resolution conditional GAN network. The overall network structure is as follows: Figure 3 As shown.
[0060] This scheme employs a near-symmetrical network structure to further constrain the accuracy of the generated images. The network is first trained on downsampled versions of real SAR images. The downsampled SAR images are then input into a low-resolution SAR-OPT image generator, and the synthesized low-resolution optical image and the downsampled version of the real optical image are simultaneously input into a discriminator for training. The symmetrical part inputs the downsampled optical image and uses both the downsampled real SAR image and the generated low-resolution SAR image to train the generative adversarial network. This process yields two pre-trained low-resolution image conversion networks, referred to as the low-resolution generation network.
[0061] During the overall training process, the input is a real SAR image. The generator of the network consists of two parts: a low-resolution generation network and a high-resolution enhancement network. The high-resolution enhancement network comprises four parts: a convolutional part, an improved segmentation-constrained residual network, a dilated convolutional part, and an upsampling part. The output of the low-resolution generation network is superimposed with the output of the convolutional part and then input into the improved residual part. The output of the generator is a synthesized high-resolution optical image, which, along with the real optical image, is input into the multi-scale discriminator. Simultaneously, the synthesized optical image is input into the OPT-SAR generative adversarial network to generate a synthesized SAR image for comparison and verification with the real SAR image, thereby verifying the effectiveness of the synthesized optical image.
[0062] Similarly, the input for the symmetrical part is a real high-resolution optical image. After passing through the OPT-SAR generator, a synthetic SAR image is output and input together with the real high-resolution SAR image into the discriminator for differentiation and training. At the same time, the synthetic SAR image is input into the SAR-OPT generative adversarial network to generate an optical image for verification.
[0063] First, the strategy of training a low-resolution generative network first and then adding a global generative adversarial network to fine-tune all networks together can effectively combine global and local detail information of the input image, ensuring the quality of the generated high-resolution image. Second, the symmetric structure network can ensure that the transformation is reversible through the constraint of cycle consistency loss, so that the target domain accurately preserves the source domain information and improves the accuracy of generating high-resolution images.
[0064] The details of each part of the overall network and the objective function of this scheme are as follows:
[0065] The input SAR image is a 3D image created by copying two dimensions from its 1D image. It first enters the convolutional part of the generator. This part contains two downsampling convolutional layers to extract features from the input image. After the convolutional part, the original image dimensions change from H×W×C to H / 2×W / 2×128.
[0066] 1. Improved residual component:
[0067] The extracted feature map is superimposed with the feature map of the same scale as the low-resolution generative network and then input into the improved residual. Detailed images of the improved residual are shown below. Figure 4 As shown.
[0068] As can be seen, the input first passes through a reflection filling layer to preserve the image's edge information during subsequent convolutional operations. Then, convolutional layers are used for feature extraction. Repeating this process twice can be considered a residual module. The residual part contains three downsampling residual modules and six ordinary residual modules. When the input and output sizes are different, this residual module is called a downsampling residual module; otherwise, it is an ordinary residual module. The feature maps output from the three downsampling residual modules are input into a pre-trained segmentation network with frozen parameters to obtain the segmentation loss. This operation helps reduce the probability of targets in the original image being lost in the generated image.
[0069] 2. Dentular convolution part:
[0070] Detailed image of the dilated convolution part is shown below. Figure 5 As shown.
[0071] Using dilated convolutions after the residual network can increase the receptive field of the discriminator, improve the ability to understand global information, and reduce the impact of speckle noise on SAR images. This part processes the input feature map using four different dilation rates, and the results are concatenated together through 1×1 convolutional layers as the output.
[0072] 3. Multi-scale discriminator:
[0073] The discriminator part adopts a multi-scale discriminator similar to the Pix2PixHD model. It first constructs a three-layer image pyramid structure based on real images, synthetic images, and their average pooled versions. A discriminator is trained for each layer to perform discrimination. These three discriminators have the same structure but different input sizes. The multi-scale discriminator can improve the generation of images at both the overall and detailed levels.
[0074] 4. Loss function:
[0075] This invention involves three types of loss functions: generative adversarial network loss function, segmentation loss function, cycle consistency loss function, and feature matching loss function. The generative adversarial network loss function for generating optical images from SAR images is expressed as follows:
[0076]
[0077] Ultimately, the problem to be solved is the following formula:
[0078]
[0079] Where G is the SAR-to-optical image generator, and DK Y A discriminator for different SAR to optical images.
[0080] The principle of the loss function F in generative adversarial networks (GANs) for generating SAR images from optical images is similar to that of G. The loss function F in GANs for generating SAR images from optical images addresses the following problem:
[0081]
[0082] Where F is the SAR-to-optical image generator, and DK Y The discriminator section for different optical images to SAR images.
[0083] The feature matching loss function calculated in the multi-scale discriminator is based on the comparison of features extracted from the multi-layer structure of the multi-scale discriminator, and its expression is:
[0084]
[0085] Where T is the total number of layers in the discriminator, and N i denoted as the number of elements in each layer, and x as the input SAR or optical image.
[0086] The segmentation loss is the result directly output by the network. The segmentation loss calculated in the SAR to optical image generator is the sum of the segmentation losses calculated by the three segmentation networks, denoted as:
[0087] L seg (G)=L seg1 +L seg2 +L seg3 ;
[0088] The cycle consistency loss is:
[0089]
[0090] In summary, the overall loss function of the segmentation-guided high-resolution conditional GAN network is:
[0091]
[0092] In this application, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. In case of any inconsistency, the meaning set forth in this specification or derived from the content described herein shall prevail. Furthermore, the terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit the scope of this application.
[0093] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A segmentation-guided SAR-to-optical image conversion method, characterized in that, include: S1. Acquire SAR images, input the SAR images into the generator of the improved high-resolution conditional GAN network, extract SAR image features and generate predicted optical images, and input the multi-scale feature maps extracted from the SAR images by the cascaded residual module in the generator into the pre-trained segmentation loss network to obtain the segmentation loss. S2. Input the predicted optical image and the paired optical image and its downsampled image into the multi-scale discriminator to calculate the adversarial loss; S3. Introduce a recurrent consistency structure using the same network structure as the generative adversarial network described above. Use the real optical image as the generator input, and the predicted SAR image and paired SAR images along with their downsampled images as the multi-scale discriminator input, to obtain the recurrent consistency loss. S4. Optimize the generative adversarial network based on segmentation loss, adversarial loss, and feature matching loss, and use the generative adversarial network to generate high-resolution optical images obtained from SAR image conversion.
2. The SAR-to-optical image conversion method based on segmentation guidance according to claim 1, characterized in that, In S1, the improved high-resolution conditional GAN network is the SAR-to-optical image conversion part of a combined network of an improved conditional GAN network and a CycleGAN network.
3. The SAR-to-optical image conversion method based on segmentation guidance according to claim 2, characterized in that, In step S1, the DeepLabv3 network is pre-trained using known paired high-resolution SAR images, corresponding optical image datasets, and segmentation labels to obtain frozen target accuracy network parameters. Based on the target accuracy network parameters, a segmentation loss network is obtained.
4. The SAR-to-optical image conversion method based on segmentation guidance according to claim 3, characterized in that, In S4, the generative adversarial network is obtained based on the Pix2PixHD model.
5. The SAR-to-optical image conversion method based on segmentation guidance according to claim 4, characterized in that, In S4, the generator of the generative adversarial network includes a global generation module, a local enhancement module, and a residual module, and a dilated convolution module is added after the residual module to increase the receptive field of the upsampling part. The discriminator in the generative adversarial network employs a multi-scale discriminator to generate high-resolution images.
6. The SAR-to-optical image conversion method based on segmentation guidance according to claim 5, characterized in that, In S4, the feature matching loss is the result of the L1 norm between the feature map generated by each layer of multiple multi-scale discriminators and the feature map generated by the discriminator acting on the actual image, which is used to supervise whether the generated image contains feature details of the original image.
7. The SAR-to-optical image conversion method based on segmentation guidance according to claim 6, characterized in that, Also includes: The optical image generated by S4 is input into a preset optical-SAR image generation network to obtain a pseudo-SAR image for verification. The cycle consistency loss between the pseudo-SAR image and the SAR image obtained by S1 is calculated, and the generative adversarial network is optimized based on the cycle consistency loss.
8. A segmentation-guided SAR-to-optical image conversion system, characterized in that, include: The acquisition module is used to acquire SAR images. It inputs the SAR images into the generator of the improved high-resolution conditional GAN network, extracts SAR image features and generates a predicted optical image. The multi-scale feature map extracted from the SAR image by the cascaded residual module in the generator is input into a pre-trained segmentation loss network to obtain the segmentation loss. The computation module is used to input the predicted optical image and the paired optical image and its downsampled image into the multi-scale discriminator to calculate the adversarial loss; The recurrent module is used to predict SAR images as input from optical images and calculate the recurrent consistency loss together with the above network. The generation module is used to optimize the generative adversarial network based on segmentation loss, adversarial loss, feature matching loss and cycle consistency loss, and use the generative adversarial network to generate high-resolution optical images obtained from SAR image conversion.