Lightweight fusion method of SAR and visible light images based on double-branch GAN network
Patent Information
- Application Number
- CN202311310352.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-11
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-10-11
AI Technical Summary
(详见:H.Yin and J.Xiao,"Laplacian Pyramid Generative AdversarialNetwork for Infrared and Visible Image Fusion,"in IEEE Signal ProcessingLetters,vol.29,pp.1988-1992,2022,doi:10.1109/LSP.2022.3207621.)但是目前融合图像并不能很好的包含SAR图像的细节信息,而且对比度较低,不利于人类视觉感知
[0020]本发明提供的基于双分支GAN网络的SAR与可见光图像轻量化融合方法,对不同对比度的原图像,DB-GAN网络均能获得良好的融合效果,克服了DD-GAN网络对不同对比度图像适应性较低的缺点;模型参数较少,实现了模型的轻量化,可节省硬件开销,有利于实际应用。本发明使融合图像在保留可见光图像光谱信息、SAR与可见光图像细节信息的同时,更加符合人类视觉感知,且算法模型实现轻量化。
Smart Images

Figure CN117495690B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multi-source image fusion, and more particularly to a method for fusing SAR and visible light images. Background Technology
[0002] The fusion of SAR (Synthetic Aperture Radar) and visible light images has wide applications in target detection, disaster prediction, and land resource statistics, and has significant research value. It can support subsequent image analysis and information extraction, and lay the foundation for better and more accurate target identification and detection. However, SAR images suffer from low contrast, severe speckle noise, and unclear details, resulting in fused images with unclear details and low contrast, which does not conform to human visual perception. Therefore, how to make the fused image retain the spectral information of the visible light image while containing more detailed information from both SAR and visible light images, and having good contrast that conforms to human visual perception, is currently a research hotspot in the fusion of SAR and visible light images.
[0003] Deep learning is extremely popular in the field of computer vision, and it has shown outstanding performance in object detection and recognition. Because of its excellent learning ability, researchers have applied deep learning to the field of image fusion, resulting in many excellent methods. H.Li et al. proposed the DenseFuse network to fuse infrared and visible light images. The backbone network is an encoder and decoder. This network incorporates dense blocks, in which the output of each layer is connected to every other layer. This structure can obtain more features from the source image and uses two fusion strategies to fuse the obtained features, achieving a better fusion effect. (See: Li H, Wu X J. DenseFuse: A fusion approach to infrared and visible images[J].IEEE Transactions on Image Processing,2018,28(5):2614-2623.) Ji et al. proposed a method to fuse infrared and visible light images using three-scale decomposition and feature transfer. The three-scale decomposition method is used to refine the source image into layers through two decompositions. Then, the ResNet feature transfer method is used for detail layer fusion to extract deeper contour structures and other detailed information. (See: Ji J, Zhang Y, Hu Y, et al. Fusion of Infrared and Visible Images Based on Three-Scale Decomposition and ResNet Feature Transfer[J]. Entropy, 2022, 24(10): 1356.) Wang et al. proposed a multi-level fusion network (MSFNet) for infrared and visible light image fusion. It uses an encoder-decoder architecture with downsampling operation to learn contextual features and introduces a cross-level fusion module (CSFM) to propagate multi-scale contextual features from the early stage to the later stage, thereby improving the utilization of features. (See: Wang C, Wu J, Zhu Z, et al. MSFNet: MultiStage Fusion Network for infrared and visible image fusion[J]. Neurocomputing, 2022, 507: 26-39.) However, currently, image fusion using convolutional neural networks still requires the design of fusion strategies. The quality of the fusion strategy design directly affects the fusion result, and manually designed fusion strategies often fail to achieve the best fusion effect. Therefore, some researchers have also tried to use generative adversarial networks for image fusion.
[0004] Deng et al. proposed a spatial frequency consistent generative adversarial network (SFGAN) model framework for fusion of SAR and visible light images, resulting in fused images with more realistic texture details and clearer contour features of roads and rivers. (See: Deng B, Lv H. Research on Image Fusion Method of SAR and Visible Image Based on CNN[C] / / 2022 IEEE 4th International Conference on Civil Aviation Safety and Information Technology (ICCASIT). IEEE, 2022: 1400-1403.) In 2022, Haitao Yin et al. addressed the problems of incomplete feature extraction and unstable network training in existing algorithms by proposing the Laplacian Pyramid GAN, constructing a generator consisting of a shallow feature extraction module, a Laplacian pyramid module, and a reconstruction module. Furthermore, an attention module is included in the decoder to effectively decode salient features. Then, two discriminators are used to distinguish between the fused image and two different modalities. (See: H. Yin and J. Xiao, "Laplacian Pyramid Generative Adversarial Network for Infrared and Visible Image Fusion," in IEEE Signal Processing Letters, vol. 29, pp. 1988-1992, 2022, doi:10.1109 / LSP.2022.3207621.) However, current fused images cannot effectively capture the detailed information of SAR images, and their low contrast is detrimental to human visual perception. Furthermore, existing algorithm models have large parameters, which increases hardware costs and hinders practical applications. Summary of the Invention
[0005] The purpose of this invention is to propose a lightweight fusion method for SAR and visible light images based on a dual-branch GAN network. For the fusion of SAR and visible light images, this invention proposes for the first time a lightweight fusion algorithm based on a dual-branch GAN network. This method constructs a dual-branch GAN network and utilizes the generative adversarial mechanism of GAN networks. It also introduces a pre-fused image as guidance for the generator to avoid manually designing complex fusion rules. Compared with other deep learning-based fusion algorithms, the DB-GAN network can obtain higher quality fused images in the field of SAR and visible light image fusion, effectively preserving the spectral information of the visible light image as well as the detailed information of both SAR and visible light images. The model has fewer parameters, achieving lightweight modeling, saving hardware costs, and is beneficial for practical applications.
[0006] This invention is achieved through the following technical solutions.
[0007] The present invention discloses a lightweight fusion method for SAR and visible light images based on a dual-branch GAN network, comprising the following steps:
[0008] Step 1: Constructing the Dataset: First, visible light images and SAR images are acquired from Landsat 8 and Sentinel-1 satellites, respectively. Then, the acquired images are registered using ENVI 5.2 software. After registration, the images are cropped using ENVI 5.2 software to obtain SAR and visible light image pairs. Finally, the BM3D algorithm is used to denoise the cropped SAR images. After denoising, the original dataset is obtained. Next, the original dataset is divided into training and test sets: First, a portion of image pairs are selected from the original dataset and further cropped to form part of the training set; finally, the remaining original dataset is used as the test set.
[0009] Step 2: Generate pre-fused images using the GFCE algorithm: The original GFCE algorithm only enhances the contrast of visible light images during fusion. Considering the low contrast and unclear details of SAR images, the contrast of SAR images is also enhanced during the generation of pre-fused images, resulting in pre-fused images with good spectral information and clear details. The same pre-fused image pairs as the training set are selected and cropped, and combined with the part of the training set obtained in Step 1 to form a complete training set. At this point, both the training set and the test set are completed.
[0010] Step 3: Construct a dual-branch GAN network (DB-GAN): DB-GAN consists of two generators and two discriminators. To reduce model parameters and improve the quality of generated images, weights are shared between the two generators and between the two discriminators. Simultaneously, the traditional convolutional networks in the generators and discriminators are replaced with the lightweight module Ghost to further reduce model parameters. The pre-fused image obtained in Step 2 is then used to guide the generator's generation.
[0011] Step 4: Train the DB-GAN network using the training set from Step 1.
[0012] Step 5: Fusion of SAR and visible light images: Input the test set from Step 1 into the DB-GAN network trained in Step 4 to obtain a fused image that retains the spectral information of the visible light image and the detailed information of the SAR and visible light image, and evaluate the quality of the fused image using subjective and objective evaluation metrics.
[0013] Furthermore, the DB-GAN network described in step 3 includes two generators and two discriminators:
[0014] The two generators have identical network structures, each containing three Ghost modules (Ghost1, Ghost2, Ghost3), one Resblock_CBAM module, and one Denseblock module. The input is a channel-connected SAR and visible light image. After passing through Ghost1, the output feature map is input to the Denseblock module for full feature extraction. The output of the Denseblock then passes through the Resblock_CBAM module to obtain a weighted feature map, which is then input to Ghost2 for further feature extraction. Finally, it is input to Ghost3 to output the fused image. The two GAN networks share weights among Ghost1, Denseblock, and Resblock_CBAM to reduce the number of parameters and help the generator learn the joint distribution between the SAR and visible light images.
[0015] Traditional convolutions produce feature maps containing a large amount of redundant information. While this helps improve network performance, it also increases the computational cost and thus the number of model parameters. The Ghost module is not proposed to avoid redundant information, but rather to generate it in a low-cost way, reducing computational cost. DB-GAN uses three Ghost modules: Ghost1, Ghost2, and Ghost3. Each of Ghost1, Ghost2, and Ghost3 consists of two parts: a primary convolution and a cheap operation. The primary convolution obtains the intrinsic feature map of the input image, while the cheap operation generates the redundant information in the intrinsic feature map. Finally, the results of the primary convolution and the cheap operation are concatenated along the channel dimension to obtain the final output. Ghost1 and Ghost2 have identical network structures. The primary convolution consists of a convolutional layer and an activation function layer. Because convolution operations reduce the size of the feature map, padding is applied to the input feature map to ensure that the feature map size remains unchanged after the primary convolution. The kernel size of the convolutional layer is set to 3×3, and the activation function is LeakyReLU. The cheap convolution consists of convolutional layers, batch normalization (BN) layers, and the LeakyReLU activation function. The kernel size of the convolutional layers remains 3×3. Unlike the previous two Ghost modules, the initial convolution in the Ghost3 module does not include an activation function; the activation function chosen for the cheap operation is Tanh.
[0016] The Denseblock module comprises three Bottlenecks. The output of each Bottleneck is concatenated with its input in the channel dimension to fully utilize the features. Each Bottleneck contains two Ghost modules. The network structure of the Ghost modules is the same as that of the Ghost1 modules. The first Ghost module acts as an extension layer to increase the number of feature map channels, while the second Ghost module reduces the number of feature map channels to match the shortcut.
[0017] The Resblock_CBAM described is an improved residual block consisting of three parts: Bottleneck, CBAM, and shortcut. Bottleneck comprises three convolutional layers. The first convolutional layer uses a 1×1 kernel to compress the number of channels, reducing the dimensionality of the input feature map and thus decreasing computation. This invention reduces the number of channels to half of the original number. This layer is followed by a batch normalization (BN) layer and a ReLU activation function. The second convolutional layer uses a 3×3 kernel for feature extraction, capturing local feature information from the input feature map. This layer is also followed by a batch normalization (BN) layer and a ReLU activation function. The third convolutional layer uses a 1×1 kernel to restore the number of channels. The Bottleneck output feature map F of size (120, 120, 64) is input into the CBAM, where (120, 120) represents the height and width of the feature map, and 64 is the number of channels. First, it enters the channel attention module, where the height and width of the feature map are max-pooled and average-pooled to obtain two feature maps of size (1, 1, 4). These are then fed into the multilayer perceptron to obtain channel weights. Finally, the obtained weights are summed and fed into the sigmoid activation function, outputting a channel attention feature map M of size (1, 1, 64). C Then M C Multiplying the input feature map F by the input feature map yields a feature map F' of size (120, 120, 64), which serves as the input to the spatial attention module. After inputting into the spatial attention module, max pooling and average pooling are first performed on each channel of the input feature map, resulting in two feature maps of size (120, 120, 1). These two feature maps are then concatenated along the channel dimension to obtain a feature map of size (120, 120, 2). A convolutional layer is then used to reduce the dimensionality of the feature map, resulting in a feature map of size (120, 120, 1). Finally, the obtained feature map is fed into the sigmoid activation function to obtain the spatial attention feature map M. S Then M S Multiplying by F' yields the final output of CBAM, which consists of channel-oriented and spatial attention feature maps. Finally, the input of Bottleneck and the output of CBAM are shortened to obtain the final feature map.
[0018] After the generator generates the fused image, it is fed into the discriminator for discrimination. The two discriminators have identical structures and share weights, allowing the generator to generate a fused image that incorporates information from another source image. This also reduces the number of model parameters and accelerates model convergence. Ghost modules are used instead of traditional convolutional layers to achieve network lightweighting. Both discriminators have identical structures, consisting of four Ghost modules and one linear layer. The four Ghost modules have the same network structure as Ghost1 in the generator, consisting of convolutional layers, batch normalization layers, and activation function layers. The convolutional kernel size is set to 3×3, the stride to 2, the batch normalization layer uses BN, and the activation function is LeakyReLU. Finally, a linear layer transforms the flattened feature map into output, obtaining the relative distance between the generated image and the real image. To reduce the number of parameters, the weights of the third and fourth Ghost modules and the linear layer are shared.
[0019] Further, the process of training the DB-GAN network in step 4 includes: inputting the SAR and visible light image training set samples obtained in step 2 into the DB-GAN network for training; setting the learning rate of the generator to 0.0001 and the learning rate of the discriminator to 0.0001 to balance the learning rates of the generator and discriminator, and selecting Adam as the optimizer. The training ratio of the discriminator to the generator is 1:1, the batch size is 16, and a total of 10 epochs are trained. By continuously iterating and updating the network hyperparameters, the network training process is completed when the set number of iterations is reached.
[0020] This invention provides a lightweight fusion method for SAR and visible light images based on a dual-branch GAN network. The DB-GAN network achieves good fusion results for original images with varying contrasts, overcoming the limitation of the DD-GAN network's low adaptability to images with different contrasts. The method uses fewer model parameters, achieving lightweight design and saving hardware costs, which is beneficial for practical applications. This invention enables the fused image to retain spectral information of the visible light image and detailed information of both SAR and visible light images, while also better conforming to human visual perception, and the algorithm model is lightweight. Attached Figure Description
[0021] Figure 1 This is a diagram of the overall framework of the proposed dual-branch GAN network (DB-GAN). gSAR I is a fused image of the biased SAR images generated by the generator. gVIS I is a fused image of the biased visible light image generated by the generator. gV (I gSAR -I VIS) is the first discriminator's value for determining whether the fused image of the SAR image is a visible light image. gS (I gVIS -I SAR The second discriminator determines whether the fused image, biased towards visible light images, is a SAR image; I g This is the final merged image.
[0022] Figure 2 It is the network structure of the generator.
[0023] Figure 3 It is the network structure of the Ghost1 and Ghost2 modules in the generator.
[0024] Figure 4 It is the network structure of the Ghost3 module in the generator.
[0025] Figure 5 It is the network structure of the Denseblock module in the generator.
[0026] Figure 6 This refers to the network structure of the Bottleneck module within the Denseblock module.
[0027] Figure 7 It is the network structure of CBAM in the Resblock_CBAM module.
[0028] Figure 8 It is the network structure of the channel attention module in the CBAM module.
[0029] Figure 9 It is the network structure of the spatial attention module in the CBAM module.
[0030] Figure 10 It is the network structure of the discriminator.
[0031] In the diagram, ⊕ represents the sum of the obtained weights. This is the Sigmoid activation function. Detailed Implementation
[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.
[0033] As a feasible embodiment of the present invention, images from three regions were selected as experimental data. These regions include features such as water bodies, soil, vegetation, mountains, and buildings, which allows for better verification of the fusion effect and makes the experimental results more reliable. The image data from the three regions are: images of the Bohai Sea area in Northeast China, images of the Tianjin Binhai New Area, and images of the Yujiang River area in Nanning, Guangxi. The visible light images were obtained from the Geospatial Data Cloud website (…). https: / / www.gscloud.cn / search Data was obtained from the Landsat 8 satellite, and a true-color image was created by sequentially selecting the red, green, and blue bands, with a resolution of 5m. SAR images were obtained from the website operated by the European Space Agency. https: / / scihub.copernicus.eu / dhus / # / home To acquire data, select Sentinel-1B GRD level data, choose VV polarization mode, and set the resolution to 30m. Then, use the registration tool in ENVI 5.2 software to register and resample the SAR and visible light images. The main steps are as follows:
[0034] 1. First, click File->open…, select the folder containing the images to be registered, and load the SAR and visible light images to be registered respectively.
[0035] 2. Then, in the toolbox on the right, select Geometric Correction->Registration->ImageRegistration Workflow, use the SAR image as the reference image and the visible light image as the image to be registered, and click Next.
[0036] 3. Begin registration. First, click "Main," select the matching algorithm. SAR and visible light images are different image types, so mutual information is selected as the matching algorithm. Set the tie point precision value to discard registration points that do not meet the preset precision; leave other options at their default values. Next, click "Seed Tie Points" to manually select several reference points in the center and around the edges of the image. Finally, click "Advanced," set the registration bands of the visible light image to match those of the SAR image, leave other options at their default values, and click Next to generate registration points. Check the scores of the registration points, delete points with low scores, then click "Warping," select the resampling method (cubic convolution interpolation is selected in this invention), and click Next to obtain the registered image.
[0037] After image registration was completed, the three sets of images were cropped using ENVI 5.2 software to create SAR and visible light image pairs with a width of 408 and a height of 555. To expand the dataset, each selected cropped region was rotated by 45 degrees each time. °Eight image pairs were obtained, for a total of 215 image pairs. After cropping, the BM3D algorithm was used to denoise the SAR images. After denoising, 36 image pairs were selected and divided into four groups. The first group was left untouched; the second group underwent contrast enhancement; the third group only enhanced the visible light image; and the fourth group only enhanced the SAR image. The four groups were merged and cropped to 120×120 pixel images, totaling 22,488 pairs. The selected 36 image pairs were then fused using the GFCE algorithm to obtain pre-fused images, which were also cropped to 120×120 pixel images, totaling 22,488 pairs. The 22,488 SAR images, visible light images, and pre-fused images were cut and pasted into three folders: SAR, VIS, and PF, respectively, to form the training set. The remaining 179 image pairs with a pixel size of 555×408 were used as the test set.
[0038] Step 1: Constructing the dataset: First, visible light images and SAR images were acquired from Landsat 8 and Sentinel-1 satellites, respectively. Then, the acquired images were registered using ENVI 5.2 software, and the images were cropped into 215 pairs with a pixel size of 555×408. Finally, the BM3D algorithm was used to denoise the SAR images. After denoising, 36 pairs of images were selected and cropped into 22,488 pairs of SAR and visible light image pairs with a pixel size of 120×120.
[0039] Step 2: Generate pre-fused images using the GFCE algorithm: The original GFCE algorithm only enhances the contrast of the visible light image during fusion. Considering the low contrast and unclear details of the SAR image, the contrast of the SAR image is also enhanced during the generation of the pre-fused image, resulting in a pre-fused image with good spectral information and clear details. The same 36 pairs of pre-fused images as in Step 1 are selected and cropped to obtain 22,488 pairs of SAR and visible light image pairs with a pixel size of 120×120. These 22,488 pairs of images obtained in Step 1 form the training set, and the remaining 179 pairs of images with a pixel size of 555×408 are used as the test set.
[0040] Step 3: Construct a dual-branch GAN network (DB-GAN): DB-GAN consists of two generators and two discriminators. To reduce model parameters and improve image generation quality, weights are shared between the two generators and between the two discriminators. Simultaneously, the traditional convolutional networks in the generators and discriminators are replaced with the lightweight module Ghost to further reduce model parameters. The pre-fused image obtained in Step 2 is then used to guide the generator's generation.
[0041] Step 4: Train the DB-GAN network using the training set from Step 1.
[0042] Step 5: Fusion of SAR and visible light images: Input the test set from Step 1 into the DB-GAN network trained in Step 4 to obtain a fused image that retains the spectral information of the visible light image and the detailed information of the SAR and visible light image, and evaluate the quality of the fused image using subjective and objective evaluation metrics.
[0043] The DB-GAN network described in step 3 includes two generators and two discriminators:
[0044] Both generators have identical network structures, each containing three Ghost modules (Ghost1, Ghost2, Ghost3), one Resblock_CBAM module, and one Denseblock module. The input is a channel-connected SAR and visible light image. After passing through Ghost1, the output feature map is input to the Denseblock module for full feature extraction. The output of the Denseblock then passes through the Resblock_CBAM module to obtain a weighted feature map, which is then input to Ghost2 for further feature extraction. Finally, it is input to Ghost3 to output the fused image. The two GAN networks share weights among Ghost1, Denseblock, and Resblock_CBAM to reduce the number of parameters and help the generator learn the joint distribution between the SAR and visible light images.
[0045] Traditional convolutions produce feature maps containing a large amount of redundant information. While this helps improve network performance, it also increases the computational cost and thus the number of model parameters. The Ghost module is not proposed to avoid redundant information, but rather to generate it in a low-cost way, reducing computational cost. DB-GAN uses three Ghost modules: Ghost1, Ghost2, and Ghost3. Each of Ghost1, Ghost2, and Ghost3 consists of two parts: a primary convolution and a cheap operation. The primary convolution obtains the intrinsic feature map of the input image, while the cheap operation generates the redundant information in the intrinsic feature map. Finally, the results of the primary convolution and the cheap operation are concatenated along the channel dimension to obtain the final output. Ghost1 and Ghost2 have identical network structures. The primary convolution consists of a convolutional layer and an activation function layer. Because convolution operations reduce the size of the feature map, padding is applied to the input feature map to ensure that the feature map size remains unchanged after the primary convolution. The kernel size of the convolutional layer is set to 3×3, and the activation function is LeakyReLU. The cheap convolution consists of convolutional layers, batch normalization (BN) layers, and the LeakyReLU activation function. The kernel size of the convolutional layers remains 3×3. Unlike the previous two Ghost modules, the initial convolution in the Ghost3 module does not include an activation function; the activation function chosen for the cheap operation is Tanh.
[0046] The Denseblock module contains three Bottlenecks. The output of each Bottleneck is concatenated with its input in the channel dimension to fully utilize the features. Each Bottleneck contains two Ghost modules. The network structure of the Ghost modules is the same as that of the Ghost1 modules. The first Ghost module acts as an extension layer to increase the number of feature map channels, while the second Ghost module is used to reduce the number of feature map channels to match the shortcut.
[0047] Resblock_CBAM is an improved residual block consisting of three parts: Bottleneck, CBAM, and shortcut. Bottleneck contains three convolutional layers. The first convolutional layer uses a 1×1 kernel to compress the number of channels, reducing the dimensionality of the input feature map and thus reducing the computational cost, shrinking the number of channels to half of the original number. This layer is followed by a batch normalization (BN) layer and a ReLU activation function. The second convolutional layer uses a 3×3 kernel to extract features, capturing local feature information from the input feature map. This layer is also followed by a batch normalization (BN) layer and a ReLU activation function. The third convolutional layer uses a 1×1 kernel to restore the number of channels. The Bottleneck output feature map F of size (120, 120, 64) is input into the CBAM, where (120, 120) represents the height and width of the feature map, and 64 is the number of channels. First, it enters the channel attention module, where the height and width of the feature map are max-pooled and average-pooled to obtain two feature maps of size (1, 1, 4). These are then fed into the multilayer perceptron to obtain channel weights. Finally, the obtained weights are summed and fed into the sigmoid activation function, outputting a channel attention feature map M of size (1, 1, 64). C Then M C Multiplying the input feature map F by the input feature map yields a feature map F' of size (120, 120, 64), which serves as the input to the spatial attention module. After inputting into the spatial attention module, max pooling and average pooling are first performed on each channel of the input feature map, resulting in two feature maps of size (120, 120, 1). These two feature maps are then concatenated along the channel dimension to obtain a feature map of size (120, 120, 2). A convolutional layer is then used to reduce the dimensionality of the feature map, resulting in a feature map of size (120, 120, 1). Finally, the obtained feature map is fed into the sigmoid activation function to obtain the spatial attention feature map M. S Then M S Multiplying by F' yields the final output of CBAM, which consists of channel-oriented and spatial attention feature maps. Finally, the input of Bottleneck and the output of CBAM are shortened to obtain the final feature map.
[0048] After the generator generates the fused image, it is fed into the discriminator for discrimination. The two discriminators have identical structures and share weights, allowing the generator to generate a fused image that incorporates information from another source image. This also reduces the number of model parameters and accelerates model convergence. Ghost modules are used instead of traditional convolutional layers to achieve network lightweighting. Both discriminators have identical structures, consisting of four Ghost modules and one linear layer. The four Ghost modules have the same network structure as Ghost1 in the generator, consisting of convolutional layers, batch normalization layers, and activation function layers. The convolutional kernel size is set to 3×3, the stride to 2, the batch normalization layer uses BN, and the activation function is LeakyReLU. Finally, a linear layer transforms the flattened feature map into output, obtaining the relative distance between the generated image and the real image. To reduce the number of parameters, the weights of the third and fourth Ghost modules and the linear layer are shared.
[0049] Step 4, the process of training the DB-GAN network, includes: inputting the SAR and visible light image training set samples obtained in Step 2 into the DB-GAN network for training; setting the learning rate of the generator to 0.0001 and the learning rate of the discriminator to 0.0001 to balance the learning rates of the generator and discriminator; selecting Adam as the optimizer. The training ratio of the discriminator to the generator is 1:1, the batch size is 16, and the training lasts for 10 epochs. By continuously iterating and updating the network hyperparameters, the network training process is completed when the set number of iterations is reached.
[0050] The final results of the fused image quality and the time required to generate one fused image obtained in this embodiment are as follows. To further demonstrate the advantages of this invention, the method of this invention was compared with other deep learning-based fusion algorithms and traditional image fusion algorithms. The fusion quality of the fused images obtained by each algorithm under the same experimental conditions and the average time required to generate one image were compared. Tables 1 and 2 show that the lightweight fusion algorithm for SAR and visible light images based on a dual-branch GAN network obtains high-quality fused images while taking less time to generate one image. Furthermore, the network model size is 863KB, while the network model size of DD-GAN is 9.63MB, approximately 11 times the size of the DB-GAN model. This demonstrates that a lightweight model is achieved while maintaining high-quality fused images.
[0051] Table 1 Evaluation results of fused image quality under different algorithms
[0052]
[0053]
[0054] Table 2 Training and testing times for different algorithms
[0055]
[0056] The above description merely illustrates preferred embodiments of the present invention, and while the description is relatively specific and detailed, it should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications, improvements, and substitutions without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the present invention should be determined by the appended claims.
Claims
1. A lightweight fusion method for SAR and visible light images based on a dual-branch GAN network, characterized by: Includes the following steps: Step 1: Constructing the dataset: First, acquire visible light images and SAR images from Landsat 8 and Sentinel-1 satellites, respectively; then, register the acquired images using ENVI 5.2 software. After registration, continue to crop the images using ENVI 5.2 software to obtain SAR and visible light image pairs; finally, use the BM3D algorithm to denoise the cropped SAR images. After denoising, the original dataset is obtained; Next, the original dataset is divided into a training set and a test set: First, a portion of image pairs are selected from the original dataset and further cropped to form part of the training set; then, the remaining original dataset is used as the test set. Step 2: Generate pre-fused images using the GFCE algorithm: During the generation of pre-fused images, the contrast of the SAR images is simultaneously enhanced to obtain pre-fused images with good spectral information and clear details; and select the same pre-fused image pairs as the training set for cropping, and combine them with the part of the training set obtained in Step 1 to form a complete training set. At this point, both the training set and the test set have been completed. Step 3: Construct a dual-branch GAN network DB-GAN: DB-GAN consists of two generators and two discriminators, with weights shared between the two generators and between the two discriminators; at the same time, the traditional convolutional networks in the generators and discriminators are replaced with the lightweight module Ghost to further reduce the model parameters; and the pre-fused image obtained in Step 2 is used to guide the generation of the generator. Step 4: Train the DB-GAN network using the training set from Step 1; Step 5: Fusion of SAR and visible light images: Input the test set from Step 1 into the DB-GAN network trained in Step 4 to obtain the final fused image; wherein, the two generators of the DB-GAN network generate a fused image biased towards SAR image and a fused image biased towards visible light image respectively. The final fused image is obtained by combining the fused image biased towards SAR image and the fused image biased towards visible light image. The final fused image retains the spectral information of the visible light image and the fused image of the SAR and visible light image details. The quality of the fused image is evaluated using subjective and objective evaluation metrics.
2. The lightweight fusion method for SAR and visible light images based on a dual-branch GAN network as described in claim 1. Its characteristic is that the DB-GAN network described in step 3 includes two generators and two discriminators: The two generators have identical network structures, each containing three Ghost modules (Ghost1, Ghost2, and Ghost3), one Resblock_CBAM module, and one Denseblock module. The input is a channel-connected SAR and visible light image. After passing through Ghost1, the output feature map is input to the Denseblock module for full feature extraction. The output of Denseblock then passes through the Resblock_CBAM module to obtain a weighted feature map, which is then input to Ghost2 for further feature extraction. Finally, it is input to Ghost3 to output the fused image. The two GAN networks share weights among Ghost1, Denseblock, and Resblock_CBAM to reduce the number of parameters and help the generator learn the joint distribution between the SAR and visible light images. Ghost1, Ghost2, and Ghost3 all consist of two parts: a primary convolution and a cheap operation. The primary convolution is used to obtain the intrinsic feature map of the input image, while the cheap operation is used to generate redundant information in the intrinsic feature map. Finally, the results of the primary convolution and the cheap operation are concatenated in the channel dimension to obtain the final output. The network structures of Ghost1 and Ghost2 are exactly the same. The primary convolution consists of a convolutional layer and an activation function layer. To ensure that the feature map size remains unchanged after the primary convolution, a padding operation is first performed on the input feature map. The kernel size of the convolutional layer is set to 3×3, and the activation function is LeakyReLU. The cheap convolution consists of a convolutional layer, a batch normalization (BN) layer, and the LeakyReLU activation function. The kernel size of the convolutional layer is also set to 3×3. Unlike the first two Ghost modules, the initial convolution of the Ghost3 module does not contain an activation function; the activation function chosen for cheap computation is Tanh. The Denseblock module contains three Bottlenecks. The output of each Bottleneck is concatenated with its input in the channel dimension to fully utilize the features. Each Bottleneck contains two Ghost modules. The network structure of the Ghost module is the same as that of the Ghost1 module. The first Ghost module is used as an extension layer to increase the number of feature map channels, and the second Ghost module is used to reduce the number of feature map channels to match the shortcut. The Resblock_CBAM described is an improved residual block, comprising three parts: Bottleneck, CBAM, and shortcut. Bottleneck contains three convolutional layers. The first convolutional layer uses a 1×1 kernel for channel compression, reducing the dimensionality of the input feature map and thus decreasing computation. This layer is followed by a batch normalization (BN) layer and a ReLU activation function. The second convolutional layer uses a 3×3 kernel for feature extraction, capturing local feature information from the input feature map. This layer is also followed by a batch normalization (BN) layer and a ReLU activation function. The third convolutional layer uses a 1×1 convolutional kernel to recover the number of channels. A feature map F of size (120, 120, 64) output from the Bottleneck is input into the CBAM, where (120, 120) represents the height and width of the feature map, and 64 represents the number of channels. First, it enters the channel attention module, where the height and width of the feature map are max-pooled and average-pooled to obtain two feature maps of size (1, 1, 4). These are then fed into the multilayer perceptron to obtain channel weights. Finally, the obtained weights are summed and fed into the sigmoid activation function, outputting a channel attention feature map M of size (1, 1, 64). C Then M C Multiplying the input feature map F by the input feature map yields a feature map F' of size (120, 120, 64), which serves as the input to the spatial attention module. After inputting into the spatial attention module, max pooling and average pooling are first performed on each channel of the input feature map, resulting in two feature maps of size (120, 120, 1). These two feature maps are then concatenated along the channel dimension to obtain a feature map of size (120, 120, 2). A convolutional layer is then used to reduce the dimensionality of the feature map, resulting in a feature map of size (120, 120, 1). Finally, the obtained feature map is fed into the sigmoid activation function to obtain the spatial attention feature map M. S Then M S Multiplying with F' yields the final output of CBAM, namely the attention feature maps in the channel and spatial directions; finally, the input of Bottleneck and the output of CBAM are shortened to obtain the final feature map. After the generator generates the fused image, it is fed into the discriminator for discrimination. The two discriminators have identical structures and share weights, allowing the generator to generate a fused image that includes information from another source image. This also reduces the number of model parameters and accelerates model convergence. Ghost modules are used to replace traditional convolutional layers, achieving network lightweighting. The two discriminators have identical structures, each consisting of four Ghost modules and one linear layer. The four Ghost modules have the same network structure as Ghost1 in the generator, consisting of convolutional layers, batch normalization layers, and activation function layers. The convolutional kernel size is set to 3×3, the stride is set to 2, the batch normalization layer uses BN, and the activation function is LeakyReLU. Finally, a linear layer transforms the flattened feature map into output, obtaining the relative distance between the generated image and the real image. To reduce the number of parameters, the weights of the third and fourth Ghost modules and the linear layer are shared.
3. The lightweight fusion method for SAR and visible light images based on a dual-branch GAN network as described in claim 1. The characteristic of step 4 is that the process of training the DB-GAN network includes: inputting the SAR and visible light image training set samples obtained in step 2 into the DB-GAN network for training; setting the learning rate of the generator to 0.0001 and the learning rate of the discriminator to 0.0001 to balance the learning rates of the generator and the discriminator, and selecting Adam as the optimizer; the ratio of the number of training iterations of the discriminator to the generator is 1:1, the batch size is 16, and a total of 10 epochs are trained; and the network hyperparameters are continuously updated iteratively, and the training process of the network is completed when the number of iterations reaches the set number of iterations.