An optical-sar image translation method based on a scattering feature enhancement strategy
By constructing an image translation framework SFEG with a scattering feature enhancement strategy, and combining it with SCG and GFE modules, the problems of missing scattering characteristics and insufficient local features in optical-SAR image translation are solved, and high-performance SAR image generation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HARBIN INST OF TECH
- Filing Date
- 2025-06-04
- Publication Date
- 2026-05-26
Smart Images

Figure CN120634889B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image processing and relates to a synthetic aperture radar (SAR) image translation method, specifically an optical-SAR image translation method based on a scattering feature enhancement strategy. Background Technology
[0002] SAR, as an active microwave remote sensing imaging technology, boasts the advantage of all-weather, all-day operation, stably acquiring target and background information under complex weather conditions, making it invaluable for remote sensing applications. However, the currently available publicly available SAR image datasets are limited in scale, particularly lacking image samples containing high-value targets such as aircraft, hindering further development in both military and civilian fields. In contrast, optical remote sensing image databases, with their mature imaging technology and widespread application needs, have accumulated a large amount of sample data containing aircraft targets. These samples possess advantages such as clear edges, intuitive textures, and rich colors. Therefore, SAR image transfer generation methods based on optical data have become an exploratory approach to overcome the bottleneck in SAR data generation. Currently, researchers have proposed a series of optical-SAR image translation methods based on image style transfer technology, alleviating the problem of insufficient SAR image data to some extent. Image translation technology refers to a method of mapping a source image to a target image domain, so that the converted image has the visual features of the target domain while preserving the structure and semantic content of the source image as much as possible.
[0003] However, current optical-SAR image translation technology still has the following key defects: (1) Existing methods can only ensure that the SAR images generated are close to the real SAR images in terms of global statistical characteristics, and cannot effectively restore the scattering characteristics of key targets in the input image. (2) Existing methods are difficult to enhance the local feature representation of targets in the generated image while suppressing background interference, and the actual effect of the generated SAR images cannot meet the needs of related research. Summary of the Invention
[0004] To address the shortcomings of existing optical-SAR image translation methods, such as the lack of key target scattering characteristics and insufficient representation of local image features, this invention provides an optical-SAR image translation method based on a scattering feature enhancement strategy. This method effectively provides prior guidance on the physical characteristics of aircraft targets, preserving the true physical scattering characteristics of targets in SAR images while ensuring pixel-level matching between the generated and input images, thus enabling high-performance SAR image generation.
[0005] The objective of this invention is achieved through the following technical solution:
[0006] An optical-SAR image translation method based on a scattering feature enhancement strategy includes the following steps:
[0007] Step 1: Construct an image translation framework SFEG (Scattering Feature Enhancement GAN) based on a scattering feature enhancement strategy. Use GAN loss, cycle consistency loss and identity loss to constrain the training process of SFEG, so that SFEG can receive unpaired optical and SAR image datasets for training and generate corresponding SAR images based on input optical images.
[0008] Step 2: Construct the Scattering Characteristic Guidance (SCG) module. Considering the similarity of the geometric structure of the same target in optical and SAR images, the edge is first extracted from the input optical image. Then, the edge mask focused on the aircraft target area is generated by combining the target annotation information. Next, the scattering field is modeled by combining the prior knowledge of the scattering characteristics of the aircraft target in the SAR image. Finally, the generated scattering field and the corresponding optical image are input into the generator of the image translation model to perform the subsequent training process.
[0009] Step 3: Construct a Guided Feature Enhancement (GFE) module to enhance the nonlinear and refined output capabilities of SFEG, thereby improving the actual quality of the generated SAR images. This module includes two attention branches. The channel attention branch compresses the channel dimensions to varying degrees while maintaining the spatial resolution of the input depth features, and reorganizes the feature information using a nonlinear function to obtain channel-enhanced features. The spatial attention branch compresses the spatial dimensions of the output features and reorganizes the feature information using a nonlinear function to obtain spatially enhanced features.
[0010] Compared with the prior art, the present invention has the following advantages:
[0011] (1) An SCG module capable of simulating the scattering characteristics of a target in an input optical image is proposed. This module first extracts the edges of the input image, generates an edge mask by combining the target annotation information, and models the scattering field by combining prior knowledge of the SAR target scattering characteristics. Finally, the generated scattering field and the corresponding optical image are input into the generator. This module can effectively restore the scattering characteristics of the target in the input image.
[0012] (2) A GFE module that integrates and enhances the target scattering features of the input image is proposed. This module includes a channel attention branch and a spatial attention branch, which can enhance the information of the input features from the channel dimension and the spatial dimension, respectively. It also combines a nonlinear function to realize feature recombination, further enhancing the nonlinear and refined output capability of SFEG. This module can enhance the target feature representation in the generated image while suppressing background interference, thus improving the actual effect of the generated SAR image. Attached Figure Description
[0013] Figure 1 This is a schematic diagram of the optical-SAR image translation model structure based on a scattering feature enhancement strategy;
[0014] Figure 2 A comparison of SAR image results generated by different unpaired image translation methods. Detailed Implementation
[0015] The technical solution of the present invention will be further described below with reference to the accompanying drawings, but it is not limited thereto. Any modifications or equivalent substitutions to the technical solution of the present invention that do not depart from the spirit and scope of the technical solution of the present invention should be covered within the protection scope of the present invention.
[0016] This invention provides an optical-SAR image translation method based on a scattering feature enhancement strategy. The method designs an optical-SAR image translation model SFEG based on the scattering feature enhancement strategy, mainly including a scattering characteristic-guided SCG module and a guided feature enhancement GFE module. This model can establish a mapping relationship between the optical domain and the SAR domain, generating a SAR image that matches the input optical image at the pixel level and retains the true physical scattering characteristics of the target. Figure 1 As shown, the specific steps include the following:
[0017] Step 1: Construct the image translation framework SFEG based on a scattering feature enhancement strategy. The training process of SFEG is constrained by GAN loss, cycle consistency loss, and identity loss, ultimately enabling SFEG to be trained on unpaired optical and SAR image datasets and generate corresponding SAR images from input optical images. The specific steps are as follows:
[0018] Step 11: Construct a basic unpaired image translation framework. The SFEG model consists of two generators and two discriminators, where the generators... The discriminant generates a corresponding SAR image based on the input optical image. Another set of generators is used to evaluate whether the input image is a real SAR image or a generated SAR image. and discriminator Then it will perform the exact opposite function.
[0019] Steps one and two: During training, the two sets of generators and discriminators perform an adversarial process, which can be achieved using the GAN loss function. It is expressed as follows:
[0020]
[0021] In the formula, and These represent image samples in the optical domain and the SAR domain, respectively. and These represent the data distributions in the optical and SAR domains, respectively. The construction of the GAN loss enables the generator to generate a target domain image based on the source domain image.
[0022] Step 13: Propose the cycle consistency loss To further constrain the generated target domain image, this loss ensures that the details of the generated image, such as the orientation and shape of the target, maintain a pixel-level correspondence with the input image. The loss is defined as follows:
[0023]
[0024] The purpose of cycle consistency loss is to ensure and This ensures that the generated image corresponds to the input image at the pixel level in terms of local details.
[0025] Step 14: Propose Identity Loss This is to ensure the continuity and consistency of the generated target domain image. The loss is defined as follows:
[0026]
[0027] Construct the overall loss function by combining the above three types of loss functions. This ensures that the generated image closely approximates the real SAR image in terms of overall feature distribution. The overall loss function of the SFEG model is defined as follows:
[0028]
[0029] In the formula, , and It is a weight hyperparameter.
[0030] Step 2: Construct the Scattering Characteristics Guidance Module (SCG). Considering the similarity in geometry between the same target in optical and SAR images, the input optical image is first edge-extracted. Then, an edge mask focusing on the aircraft target region is generated by combining the target annotation information. Next, the scattering field is modeled using prior knowledge of the aircraft target scattering characteristics in the SAR image. Finally, the generated scattering field and the corresponding optical image are input into the SFEG generator to perform the subsequent training process. The specific steps are as follows:
[0031] Step 21: To simulate the sensitivity of SAR imaging results of aircraft targets to the incident azimuth angle of radar waves, the gradient vector field generated by the Sobel operator during edge extraction is utilized. To construct orientation-sensitive features Based on this, the direction-sensitive weights are obtained. :
[0032]
[0033] in, This represents a 1×1 convolutional layer. This indicates a fully connected operation. Represents the hyperbolic tangent function. and These are learnable parameters.
[0034] Step 22: To simulate the effect of target scattering intensity attenuating with increasing distance during SAR imaging, a negative exponential spatial attenuation distribution weight is constructed. Specifically, based on edge masks... Construct an initial distance field and extract local spatial features through multi-level dilated convolution:
[0035]
[0036] In the formula, It is a local average pooling operation that preserves spatial resolution. Indicates the expansion rate The convolution operation is then performed. Subsequently, the obtained multi-level spatial local features are aggregated to generate negative exponential spatial decay distribution weights. The details are as follows:
[0037]
[0038] In the formula, These are learnable parameters.
[0039] Steps 2 and 3: Adjust the overall scattering intensity to obtain the final output of the aircraft target scattering field. :
[0040]
[0041] In the formula, Indicates global average pooling. This represents the Hadamard multiplication operator. These are learnable parameters.
[0042] Step 24: Input the target scattering field and the corresponding optical image into the SFEG generator to perform the model training process.
[0043] Step 3: Construct the Guided Feature Enhancement (GFE) module to enhance the nonlinear and refined output capabilities of SFEG, thereby improving the actual effect of generated SAR images. This module contains two attention branches. The channel attention branch compresses the channel dimensions to varying degrees while maintaining the spatial resolution of the input depth features, and reorganizes the feature information using a nonlinear function to obtain channel-enhanced features. The spatial attention branch compresses the spatial dimensions of the output features and reorganizes the feature information using a nonlinear function to obtain spatially enhanced features. The specific steps are as follows:
[0044] Step 31: First, apply channel attention branches to the given deep feature map. The operation is performed, where R represents the set of real numbers, C represents the number of channels, H represents the image height, and W represents the image width. This feature is then passed to two 1×1 convolutional layers. While maintaining the feature space resolution, the channel dimensions are compressed to varying degrees to obtain the output features. and and use the softmax function to The information in the data is reorganized; then, the reorganized features are compared with... Matrix multiplication is performed to enhance features; then, the calculation results are sequentially fed into 1×1 convolutional layers. The layerNorm(LN) layer and sigmoid function are used to obtain the channel weight vector. The formula is as follows:
[0045]
[0046] In the formula, This represents the sigmoid function. This represents the LayerNorm (LN) layer. This represents the softmax function. This represents matrix multiplication.
[0047] Finally, and Channel-by-channel multiplication yields deep features enhanced with channel information. The equation is as follows:
[0048]
[0049] In the formula This represents the channel-level multiplication operator.
[0050] Step 32: First, apply the spatial attention branch to the input deep feature map. The process involves feeding the input into two 1×1 convolutional layers and performing global average pooling (GAP) on the output features to compress the spatial dimensions and provide new output features. and After obtaining and Then, use the softmax function to... The information in the data is reorganized; then, the reorganized features are compared with... Perform matrix multiplication; then, input the calculation result into the sigmoid function to obtain the spatial weight distribution vector. The formula is as follows:
[0051]
[0052] In the formula, This indicates a feature dimension transformation operation.
[0053] Finally, for and Spatial multiplication yields depth features enhanced by spatial information. The equation is expressed as follows:
[0054]
[0055] In the formula, This represents a multiplication operator in spatial direction.
[0056] Through steps 31 and 32, a SAR output image corresponding to the input optical image at the pixel level can be obtained.
[0057] Step 4: The performance of the proposed SFEG is verified using the CORS-ADD dataset (an optical aircraft target dataset applied to remote sensing scenarios) and the SAR-AIRcraft-1.0 dataset (based on high-resolution SAR remote sensing images). The specific steps are as follows:
[0058] Validation was performed using the CORS-ADD and SAR-AIRcraft-1.0 datasets. The CORS-ADD dataset has image resolutions ranging from 0.31 to 1.02 meters, with a fixed image size of 640×640 pixels. It contains 32,285 labeled aircraft target instances. This dataset was divided into training and testing sets in a 7:3 ratio, with the training set containing 3,764 images and the testing set containing 1,722 images. SAR-AIRcraft-1.0 is an aircraft dataset based on high-resolution SAR remote sensing images. It uses a single polarization mode, beam focusing imaging mode, and a spatial resolution of 1 meter. The dataset contains 4,368 images and 16,463 aircraft target instances. The image sizes cover four specifications: 800×800, 1000×1000, 1200×1200, and 1500×1500 pixels.
[0059] The model training process was implemented on an Ubuntu 23.04 server based on the PyTorch 1.9.1 framework. The CPU was an Intel(R) Xeon(R) Platinum 8375C @ 2.90GHz, and the GPU was an Nvidia RTX 4090 24G. All images from the source and target domains were input at their original sizes, with a batch size of 1, and the weights were updated using the Adam optimizer. The training process lasted for 50 epochs. The first 25 epochs were trained with an initial learning rate of 0.0002, and the learning rate was linearly decayed to 0 for the next 25 epochs. Experimental results are as follows: Figure 2 As shown, it can be found that the optical-SAR image translation method based on scattering feature enhancement strategy proposed in this invention has excellent performance, can provide prior guidance on the physical characteristics of aircraft targets, and can achieve pixel-level matching between the generated image and the input image while maintaining the true physical scattering characteristics of the target in the SAR image.
Claims
1. An optical-SAR image translation method based on a scattering feature enhancement strategy, characterized in that... The method includes the following steps: Step 1: Construct the image translation framework SFEG based on the scattering feature enhancement strategy. Use GAN loss, cycle consistency loss and identity loss to constrain the training process of SFEG, so that SFEG can receive unpaired optical and SAR image datasets for training and generate corresponding SAR images based on the input optical images. Step 2: Construct a scattering characteristic-guided SCG module. Considering the similarity of the geometric structure of the same target in optical and SAR images, firstly, the edge is extracted from the input optical image. Then, the edge mask focused on the aircraft target area is generated by combining the target annotation information. Next, the scattering field is modeled by combining the prior knowledge of the scattering characteristics of the aircraft target in the SAR image. Finally, the generated scattering field and the corresponding optical image are input into the generator of the image translation model to perform the subsequent training process. Step 3: Construct the Guided Feature Enhancement (GFE) module to enhance the nonlinear and refined output capabilities of SFEG and improve the actual effect of generating SAR images. This module contains two attention branches. The channel attention branch compresses the channel dimension to different degrees while keeping the spatial resolution of the input depth features unchanged, and combines nonlinear functions to reorganize the feature information to obtain channel-enhanced features. The spatial attention branch compresses the spatial dimension of the output features and combines nonlinear functions to reorganize the feature information to obtain spatially enhanced features.
2. The optical-SAR image translation method based on scattering feature enhancement strategy according to claim 1, characterized in that... The specific steps of step one are as follows: Step 11: Constructing a basic unpaired image translation framework: The SFEG model consists of two generators and two discriminators, where the generators... The discriminant generates a corresponding SAR image based on the input optical image. Another set of generators is used to evaluate whether the input image is a real SAR image or a generated SAR image. and discriminator Then it will perform the exact opposite function; Steps one and two: During training, the two sets of generators and discriminators perform an adversarial learning process using the GAN loss function. It is expressed as follows: (1) (2) In the formula, and These represent image samples in the optical domain and the SAR domain, respectively. and The data distributions in the optical and SAR domains are represented respectively. The construction of the GAN loss enables the generator to generate a target domain image based on the source domain image. Step 13: Propose the cycle consistency loss The loss is further defined to restrict the generated target domain image as follows: (3) The purpose of cycle consistency loss is to ensure and This ensures that the generated image corresponds to the input image at the pixel level in terms of local details; Step 14: Propose Identity Loss To ensure the continuity and consistency of the generated target domain image, the loss is defined as follows: (4) Construct the overall loss function by combining the above three types of loss functions. To ensure that the generated image closely approximates the real SAR image in terms of overall feature distribution, the overall loss function is defined as follows: (5) In the formula, , and It is a weight hyperparameter.
3. The optical-SAR image translation method based on scattering feature enhancement strategy according to claim 1, characterized in that... The specific steps of step two are as follows: Step 21: Utilize the gradient vector field generated during edge extraction using the Sobel operator. Constructing orientation-sensitive features Based on this, the direction-sensitive weights are obtained. : (6) (7) in, This represents a 1×1 convolutional layer. This indicates a fully connected operation. Represents the hyperbolic tangent function. and These are learnable parameters; Step 22: Construct the negative exponential spatial decay distribution weights : Steps 2 and 3: Adjust the overall scattering intensity to obtain the final output of the aircraft target scattering field. : (10) In the formula, Indicates global average pooling. This represents the Hadamard multiplication operator. These are learnable parameters. Indicates the edge mask; Step 24: Input the target scattering field and the corresponding optical image into the SFEG generator to perform the model training process.
4. The optical-SAR image translation method based on scattering feature enhancement strategy according to claim 3, characterized in that... The specific steps of step two are as follows: First, based on edge mask Construct an initial distance field and extract local spatial features through multi-level dilated convolution: (8) In the formula, It is a local average pooling operation that preserves spatial resolution. Indicates the expansion rate Convolution operations; Subsequently, the obtained multi-level spatial local features are aggregated to generate negative exponential spatial decay distribution weights. The details are as follows: (9) In the formula, These are learnable parameters.
5. The optical-SAR image translation method based on scattering feature enhancement strategy according to claim 1, characterized in that... The specific steps of step three are as follows: Step 31: First, apply channel attention branches to the given deep features. The operation is performed, where R represents the set of real numbers, C represents the number of channels, H represents the image height, and W represents the image width. This feature is passed to two 1×1 convolutional layers. While maintaining the feature space resolution, the channel dimensions are compressed to different degrees to obtain the output feature. and and use the softmax function to The information in the data is reorganized; then, the reorganized features are compared with... Matrix multiplication is performed to enhance features; then, the calculation results are sequentially fed into 1×1 convolutional layers. The LN layer and sigmoid function are used to obtain the channel weight vector. The formula is as follows: (11) In the formula, This represents the sigmoid function. Indicates LN layer, This represents the softmax function. This represents matrix multiplication. Finally, and Channel-by-channel multiplication yields deep features enhanced with channel information. The equation is as follows: (12) In the formula Represents a channel-level multiplication operator; Step 32: First, apply spatial attention branch to the deep features of the input. The process involves feeding the input into two 1×1 convolutional layers, performing global average pooling on the output features to compress the spatial dimensions, and providing new output features. and After obtaining and Then, use the softmax function to... The information in the data is reorganized; then, the reorganized features are compared with... Perform matrix multiplication; then, input the calculation result into the sigmoid function to obtain the spatial weight distribution vector. The formula is as follows: (13) In the formula, This indicates a feature dimension transformation operation; Finally, for and Spatial multiplication yields depth features enhanced by spatial information. The equation is expressed as follows: (14) In the formula, Represents a multiplication operator in spatial direction; Through steps 31 and 32, a SAR output image corresponding to the input optical image at the pixel level is obtained.