Method for efficiently converting visible light image into infrared light image based on GAN
Through the improved GAN architecture and multi-scale feature fusion mechanism, the problems of poor confidence in infrared image conversion and lack of feature details in the prior art are solved, and high-quality and realistic infrared image generation is achieved.
Patent Information
- Application Number
- CN202411966214.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-30
AI Technical Summary
The existing GAN-based visible light image-to-infrared image conversion methods have problems of poor confidence and lack of semantic information feature details.
Using an improved GAN architecture, an efficient generator and discriminator architecture is designed, combining physical consistency loss function and multi-scale feature fusion mechanism to optimize the quality of generated images through adversarial training.
It realizes the generation of high-quality, realistic and physically consistent infrared images, improving the conversion effect.
Smart Images

Figure CN119941887A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and digital image processing, and in particular to a method for efficiently converting a visible light image into an infrared light image based on GAN. Background Art
[0002] In the field of computer vision, infrared imaging is becoming increasingly important in the fields of object detection, medical diagnosis, satellite remote sensing, etc. However, the scarcity of infrared datasets and the high cost of obtaining them limit the application of infrared images. Deep learning technology is used to convert visible light images into infrared images to meet the needs of infrared scene prediction. Existing GAN-based visible-to-infrared image conversion methods still have the following shortcomings: (1) The confidence of the infrared images converted by the cross-modal model is poor; (2) The infrared features of different semantic information lack details and have color differences. Summary of the invention
[0003] The present invention provides an improved visible light image to infrared image conversion method based on the generative adversarial network (GAN) architecture, which realizes high-quality infrared image generation by designing an efficient generator and discriminator architecture, a physically consistent loss function, and a multi-scale feature fusion mechanism. The method optimizes the quality of the generated image through adversarial training of the generator and the discriminator, so that the generated infrared image has a high sense of reality and physical consistency.
[0004] The present invention adopts the following technical solution: A method for efficiently converting a visible light image into an infrared light image based on GAN comprises the following steps:
[0005] Input a visible light image and convert it into an infrared image through an adversarial network;
[0006] For the generator, the InceptioNeXt V2 network is used as the encoder to extract multi-scale features from the input visible light image;
[0007] In the decoder, the collaborative extraction module SEM and the selective state space model Mamba module are used to extract global features, and the LFEM module is used to supplement local information and generate infrared images;
[0008] Use GFS-guided fusion strategy or direct concatenation of features by channel to fuse multi-scale features from the encoder and infrared features generated by the decoder;
[0009] For the discriminator, the authenticity of the generated infrared image is evaluated to identify the authenticity of the overall image and each pixel. The discriminator adopts the U-Net structure, uses the spectral normalization residual block SRB as the encoder and decoder, and optimizes the generator through adversarial training.
[0010] Use the GFS-guided fusion strategy to fuse the features from the encoder and the infrared features generated by the decoder;
[0011] A composite loss function is calculated and minimized, wherein the composite loss function includes structural consistency loss, mean absolute value loss, texture loss, color loss and adversarial loss to promote minimization of the difference between the generated infrared image and the real infrared image.
[0012] The visible light image is a preprocessed image, and the preprocessing step includes normalizing the image so that the image pixel values are within the range of [-1, 1].
[0013] The decoder of the generator performs the following steps:
[0014] The input features are successively extracted by the collaborative extraction module SEM, and processed by the selective state space model Mamba module for spatial and channel features, and then fused with the input features;
[0015] After the fusion features pass through the LFEM module, they are concatenated with the fused features as the output of the decoder.
[0016] The synergistic extraction module SEM performs the following steps:
[0017] The input features are sequentially passed through a depth-wise separable convolution, a normalization layer, and a linear layer, connected and activated by the SiLU function, and then output through another linear layer.
[0018] The Selective State Space Model Mamba module performs the following steps:
[0019] After the input features pass through the normalization layer, they are divided into two branches:
[0020] In the first branch, the input passes through the linear layer and activation function in sequence;
[0021] In the second branch, the input is processed by a linear layer, a depth-wise separable convolution, and an activation function, and then fed into a 2D selective scanning module for further feature extraction; subsequently, another normalization layer is used to normalize the features.
[0022] Perform element-wise production using the output of the first branch to merge the two paths of the first and second branches; finally, use a linear layer to mix the features and combine this result with a residual connection to form the output of the VSS block;
[0023] The 2D selective scanning module consists of three parts: scanning extension, S6 block and scanning merging;
[0024] The scan operation expands the input into sequences along four different directions;
[0025] Then, the S6 block extracts features from these sequences to ensure that the scanning process captures different features from each direction;
[0026] Scan merging fuses the features in four directions together by summing them, restoring the output to the same shape as the input.
[0027] The LFEM module performs the following steps:
[0028] The input passes through a linear layer to expand the channel to four times the original, and then is divided into two parts;
[0029] In the first part, multiple transformations are performed at different scales to capture coarse-scale features at different scales; these outputs are concatenated and activated using the SiLU function and then refined using the scSE block;
[0030] The second part is activated by the SiLU function and multiplies the output of the first part by concatenating the output elements to achieve interaction between multi-scale features;
[0031] Finally, the combined features are processed through a linear layer, returning the output to the original input dimension.
[0032] The discriminator adopts a U-Net structure, uses an SRB module to encode and extract mixed features, and introduces a GCM module after the spectral normalization residual block SRB in the middle layer to enhance the global contextual relationship.
[0033] The GFS guided fusion strategy includes the following steps:
[0034] First, in the generator or discriminator, the decoder feature F from the previous layer is e and encoder features F d Splicing is performed according to the channel dimension, and the spatial positions are averaged using global average pooling to generate global statistics that represent the overall characteristics of the channel;
[0035] Then, two convolutional layers are used to generate the weight W for each channel. c :
[0036] W c =Sigmoid(Conv(ReLU(Conv((AvgPool([F e ,F d ]))))))
[0037] Among them, Sigmoid represents Sigmoid function, Conv represents convolution, ReLU represents Relu function, and AvgPool represents average pooling;
[0038] Get the weight coefficient Wc After that, broadcast and multiply with the original concatenated features to re-weight the channels;
[0039]
[0040] Among them, F c is the weighted feature;
[0041] The decoder features F e and encoder features F d Add and fuse, and get the spatial attention coefficient W through the Sigmoid function s :
[0042]
[0043] Finally, the spatial attention coefficient W s And the feature map F after channel attention c Broadcast multiplication is performed to correct the features again.
[0044] The texture loss is calculated based on the texture difference between the converted infrared image and the real infrared image to maintain the texture consistency of the generated image, and is expressed as follows:
[0045]
[0046] Among them, L texture represents texture loss, N i Represents the normalization factor of the i-th layer, the product of the number of channels, height and width of the feature map of this layer, F i represents the feature map obtained after the image passes through the i-th layer of VGG-19, X represents the visible light image, Y represents the real infrared image, G represents the generator, and G(X) represents the converted infrared image;
[0047] The color loss is calculated based on the color difference between the color distribution of the converted infrared image and the real infrared image to maintain the color consistency of the generated image, and is expressed as follows:
[0048]
[0049] Among them, L color represents color loss, and B represents the Gaussian blur convolution operation.
[0050] A GAN-based system for efficiently converting visible light images into infrared images, using an adversarial network, including:
[0051] The generator uses the InceptioNeXt V2 network as an encoder to extract multi-scale features from the input visible light image;
[0052] In the decoder, it is used to use the collaborative extraction module SEM and the selective state space model Mamba module to extract global features, the LFEM module to supplement local information, and generate infrared images;
[0053] Use GFS-guided fusion strategy or direct concatenation of features by channel to fuse multi-scale features from the encoder and infrared features generated by the decoder;
[0054] The discriminator is used to evaluate the authenticity of the generated infrared image and identify the authenticity of the overall image and each pixel. The discriminator adopts the U-Net structure, uses the spectral normalization residual block SRB as the encoder and decoder, adds the global context module GCM, and optimizes the generator through adversarial training.
[0055] Use the GFS-guided fusion strategy to fuse the features from the encoder and the infrared features generated by the decoder;
[0056] A composite loss function, wherein the composite loss function includes structural consistency loss, mean absolute value loss, texture loss, color loss and adversarial loss to promote minimization of the difference between the generated infrared image and the real infrared image.
[0057] The present invention produces the following beneficial effects and advantages:
[0058] 1. The present invention adopts the InceptioNeXt V2 network as the encoder, combined with the SEM, Mamba module and LFEM module as the decoder, to effectively extract and restore the multi-scale features in the image and enhance the image details.
[0059] 2. The present invention adopts an improved U-Net structure as a discriminator, adds a spectrally normalized residual block SRB, evaluates the authenticity of the generated image through the discriminator, and promotes the generator to generate more realistic infrared images.
[0060] 3. The present invention introduces the GFS guided fusion strategy to effectively fuse multi-level feature information.
[0061] 4. The present invention designs a composite loss function including texture loss, color loss and adversarial loss to improve the texture consistency and physical consistency of the generated image. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 Diagram of the loss structure of the visible light to infrared image conversion system.
[0063] Figure 2 The generated network structure diagram of IMGAN, as well as the composition and connection diagram of IN module, SEM module, Mamba module and LFEM module.
[0064] Figure 3 The discriminant network structure diagram and schematic diagrams of the GFS module, SRB module, and GCM module are shown. DETAILED DESCRIPTION
[0065] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.
[0066] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0067] The present invention designs a method based on GAN to efficiently convert visible light images into infrared images, called IMGAN. In the generative network, the InceptioNeXt V2 network, which is lighter and has improved performance than ConvNeXt V2, is proposed as the encoder. The decoder introduces a selective state space model represented by Mamba, and combines the designed LFEM module for extracting local information to improve the fidelity of the image. A guided fusion strategy GFS is designed to ensure that the generated infrared image not only retains the structural information of the visible light image, but also reflects the physical characteristics of infrared imaging. Texture loss and color loss are designed as additional auxiliary functions to constrain bilateral mapping to solve the problems of texture roughness and color difference. The method is evaluated on four different datasets, and the results are compared with existing architectures (including Pix2Pix, ThermalGAN and InfraGAN). Compared with the InfraGAN architecture, the SSIM and PSNR of the method on the VEDAI dataset are improved by 2.38% and 7.28%, respectively.
[0068] Input a pair of visible light images with size (B×C×H×W), where B is the batch size, C is the number of channels, and H and W are the height and width of the image respectively;
[0069] The generator uses the InceptioNeXt V2 network as an encoder to extract multi-scale deep features from the input visible light image;
[0070] In the decoder of the generator, the collaborative extraction module SEM and the selective state space model Mamba module are used to extract global features, and the LFEM module is used to supplement local information, gradually restore the details of the input image, and generate an infrared image;
[0071] Use GFS-guided fusion strategy or direct concatenation of features by channel to fuse features from the encoder and infrared features generated by the decoder to improve network performance;
[0072] The discriminator evaluates the authenticity of the generated image and identifies the authenticity of the overall image and each pixel. The discriminator adopts the U-Net structure, uses the spectral normalization residual block SRB as the encoder and decoder, adds the global context module GCM to improve the discriminator performance, and optimizes the generator through adversarial training to make the generated image more realistic.
[0073] A composite loss function is calculated and minimized, wherein the composite loss function includes structural consistency loss, mean absolute value loss, texture loss, color loss and adversarial loss, so as to minimize the difference between the generated infrared image and the real infrared image.
[0074] like Figure 1 As shown, the specific technical solution is as follows:
[0075] The visible light to infrared image conversion method based on a generative adversarial network of the present invention comprises the following steps:
[0076] 1. Input data: Input a pair of visible light images with the size of , where B is the batch size, C is the number of channels (3 for RGB), and H and W are the height and width of the image (for example, 256×256 pixels). All input images are normalized so that the pixel values of the image are in the range of [-1,1].
[0077] 2. Generator architecture:
[0078] Encoder: InceptioNeXt V2 network is used as encoder. The network uses multi-scale convolution kernels to extract global features and local details of the image through convolution operations of different sizes. The output of the encoder is a multi-dimensional deep feature that contains high-level information of the input image.
[0079] Decoder: The decoder consists of the SEM module, the Mamba module, and the LFEM module. The Mamba module gradually recovers the low-resolution features of the image and maps them to higher-resolution image details. The LFEM module is used to enhance the local details of the image and improve the spatial resolution of the generated image.
[0080] Guided Fusion Strategy: This module is used to fuse the global features from the encoder and the local features from the decoder. Through an adaptive weighting mechanism, the GFS module fuses features at multiple scales to optimize the quality of the generated image.
[0081] 3. Discriminator architecture:
[0082] The discriminator adopts an improved U-Net structure, which designs an SRB module, i.e. a residual block with spectral normalization technology, to improve the stability of the network and the accuracy of discrimination. Spectral normalization can effectively suppress mode collapse during the training process of the discriminator and improve the authenticity assessment of the generated images.
[0083] The task of the discriminator is to determine whether the input image is a real image or a generated image. By comparing the infrared image generated by the generator with the real infrared image, the discriminator provides feedback signals to the generator, prompting the generator to generate more realistic infrared images.
[0084] 4. Loss function:
[0085] Adversarial loss: Through adversarial training of generative adversarial networks, the generator "cheats" the discriminator by generating images, and the discriminator guides the optimization of the generator by distinguishing between real and fake images. Adversarial loss can drive the generator to generate more realistic infrared images.
[0086] Texture loss: Calculate the texture difference between the generated image and the real image to ensure that the texture structure of the generated image is similar to that of the real image.
[0087] Color loss: Calculate the difference between the color distribution of the generated image and the real infrared image to ensure that the color characteristics of the generated image conform to the physical characteristics of the infrared image.
[0088] 5. Training process:
[0089] The generator and discriminator are optimized by alternating training. First, the generator generates an infrared image based on the input visible light image, and then the discriminator evaluates the generated image and provides error feedback to the generator. The generator uses the feedback signal to adjust the weights and gradually improve the image generation quality.
[0090] During the training process, the generator and discriminator are optimized simultaneously through a composite loss function to ensure that the generated images reach a high level in texture, color, and physical consistency.
[0091] 6. Output:
[0092] After training, the generator can generate corresponding infrared images based on the input visible light images. The generated infrared images are highly consistent with the real infrared images in details, textures and colors, and are suitable for applications such as target detection and image enhancement in low visibility environments.
[0093] The present invention comprises the following steps:
[0094] (1) Generator network design, such as Figure 2 As shown:
[0095] Encoder: ConvNeXt is a modern convolutional neural network (CNN) architecture inspired by the design principles of visual transformers (ViTs). It aims to bridge the gap between traditional CNNs and VITs by combining several key design elements of transformers while retaining the efficiency and simplicity of convolution operations. ConvNeXt V2 is an evolution of the ConvNeXt architecture. V2 discards Layer Scale and uses GRN. On multiple benchmarks, ConvNeXt V2 surpasses ConvNeXt, proving that it is more adaptable and has better performance in different tasks, and effectively alleviates the problem of feature collapse.
[0096] In the ConvNeXt V2 block, features are first processed through 7×7 deep convolution to achieve spatial information interaction. The network introduces GRN (Global Response Normalization) to improve the traditional batch normalization (BatchNormalization), normalize the response of each channel, improve the stability during training, and enhance the convergence of the model. Compared with Batch Normalization, GRN can better adapt to large-scale image datasets, especially for deep network structures, with significant effects. Later, multi-layer perceptron MLP is used to help the network to integrate features efficiently and further improve accuracy. At the same time, a mechanism called Shifted Window Attention is introduced to perform attention calculation on the local window of the image. This method can effectively capture spatial relationships while avoiding the computational overhead brought by global attention. Finally, DropPath is used for regularization, which is similar to Dropout. It prevents overfitting by randomly discarding some paths during training, thereby improving the generalization ability of the model. Through these designs, ConvNeXt V2 further integrates modern design ideas while retaining the classic convolutional neural network architecture, improving performance and training efficiency.
[0097] InceptionNeXt introduces an innovative approach to decompose the large kernel deep convolution of ConvNext into four parallel branches: a small square kernel, two orthogonal band kernels, and a unit map. This decomposition not only reduces the memory access cost, but also significantly improves the throughput. Inspired by this design, a new variant of the ConvNeXtV2 architecture, called InceptionioNext V2, is proposed. The separable convolution in ConvNeXt v2 is replaced with the Inception block. Similarly, the deep convolution in ConvNeXt V2 is decomposed by dividing the input features into four channel partitions. The first three partitions each consist of one-eighth of the input feature channels, while the fourth partition accounts for five-eighths. Deep convolutions with different kernel sizes are applied to the first three features, using convolutions of different sizes of 3×3, 1×11, and 11×1 as kernels, while the fourth feature remains unchanged to retain the original information and reduce memory overhead, which is defined as IDConv. InceptionNeXt V2 achieves excellent performance while reducing computational overhead.
[0098] The encoder part of the generator includes the InceptioNeXt V2 network, which uses multiple convolution kernels of different sizes to extract multi-scale features of the image. The network input is a three-channel RGB image I∈R W×H×3 Takes as input and outputs a global feature map Where d is the downsampling coefficient, W and H are the width and height of the image, C is the number of channels, and R is a real number.
[0099] Decoder: A global feature extractor based on SSM, using Mamba as a key component. After LayerNormalization, the input is divided into two branches. In the first branch, the input passes through a linear layer followed by an activation function. In the second branch, the input is processed through a linear layer, a depthwise separable convolution, and an activation function, and then fed into the 2D Selective Sweep (SS2D) module for further feature extraction. Subsequently, the features are normalized using LayerNormalization, and then an element-wise production is performed using the output of the first branch to merge the two paths. Finally, the features are mixed using a linear layer, and this result is combined with a residual connection to form the output of the VSS block. The SS2D module consists of three main parts: Scan Expansion, S6 Block, and Scan Merge. The Scan operation expands the input into sequences along four different directions (top left to bottom right, bottom right to top left, top right to bottom left, and bottom left to top right). The S6 block then extracts features from these sequences, ensuring that the scanning process captures different features from each direction. The Scan Merge operation fuses the features from the four directions together by summing, restoring the output to the same shape as the input.
[0100] Local Feature Extractor LFEM: This model is designed for efficient and effective local feature extraction, leveraging multi-scale processing by combining depthwise separable convolution, pooling, and attention mechanisms. Initially, the input is dilated to four times the original channel through a linear layer and then split into two parts. The first part undergoes multiple transformations at different scales: “dwconv5×5” and “conv1×1” capture the mid-scale features, “dwconv3×3” and “conv1×1” capture the fine-scale features, and “max-pooling” and “conv1×1” capture the coarse-scale features. These outputs are concatenated and activated using the SiLU function and then refined using the scSE block, which emphasizes important features through spatial and channel recalibration. The second part of the split input is activated by SiLU and multiplied element-wise through the refined concatenated output, thereby enabling interactions between multi-scale features. Finally, the combined features are processed through a linear layer, and the output is returned to the original input dimension. This structure effectively captures multi-scale features, reduces the computational cost through depthwise separable convolution, and enhances the model's focus on important features using the attention mechanism, thereby achieving efficient and robust feature extraction.
[0101] The convolution operation in the decoder SEM enhances feature extraction by capturing the spatial hierarchy before passing the data to Mamba. Mamba is good at modeling long-term dependencies and sequential relationships. Compared with the linear complexity of Transformer, this synergy enables the model to effectively process spatial and channel features, which can improve the performance of image generation tasks. The LFEM module is used to enhance the details of local areas of the image and improve the performance of the generated image in local areas. The decoder is used to gradually restore the low-resolution features of the image and generate higher-resolution infrared image details.
[0102] See Figure 2 (a) Fuse in the generator, using GFS-guided fusion strategy or direct concatenation of features by channel.
[0103] Guided Fusion Strategy GFS:
[0104] This paper proposes a guided fusion strategy (GFS) as a core feature fusion module, which is seamlessly integrated into the generator and discriminator. The main function of GFS is to use the rich features extracted by the encoder as guidance and effectively fuse them with the upsampled features, while adaptively adjusting the weights to highlight the key features of the infrared image. In the generator, GFS ensures that the generated infrared image retains the structural details of the visible light image while capturing the physical properties of infrared imaging. In the discriminator, GFS promotes guided feature fusion and enhances the discriminator's ability to evaluate the authenticity and physical consistency of infrared images. This dual-role strategy strengthens the constraints on the generated images and improves the synergy between the generation and recognition processes.
[0105] The GFS feature fusion module adopts an adaptive weighting strategy to effectively fuse the features extracted by the encoder and the infrared features generated by the decoder. Experimental results show that in the generator, the GFS-guided fusion strategy or direct channel concatenation of features for different data sets is used to fuse the features from the encoder and the infrared features generated by the decoder; while in the discriminator, GFS fusion is better.
[0106] First, the decoder feature F from the previous layer e and encoder features F d The channels are spliced together according to the channel dimension, and the spatial positions are averaged using global average pooling to generate global statistics, which represent the overall characteristics of the channel and reflect the importance of the channel in the entire feature map. Then, the weight W of each channel is generated through two convolutional layers. c .
[0107] W c =Sigmoid(Conv(ReLU(Conv((AvgPool([F e ,F d ]))))))
[0108] Get the weight coefficient W c Afterwards, these weights reflect the importance of the channels and are broadcast and multiplied with the original concatenated features to re-weight the channels, highlight important features, suppress unimportant features, and thus enhance the ability of feature representation. After a layer of convolution operation, the number of channels is reduced to one-half to accommodate the input from spatial attention. The weight of each channel is dynamically adjusted according to the global context information, so that the model can adaptively select the most relevant features under different inputs.
[0109]
[0110] The two features are added and fused, and only the middle feature is convolved to simplify the calculation, so as to compress the number of channels to 1 and obtain the spatial attention map, focus on the important spatial position, and obtain the spatial attention coefficient W after the Sigmoid function. s .
[0111]
[0112] Finally, the spatial attention coefficient is combined with the feature map F after channel attention. c Broadcast multiplication is performed to correct the features again and make them focus on spatial features. In this way, the network can focus on more important areas in channels and space at the same time, and complete adaptive adjustment of weights to help the model better integrate features.
[0113]
[0114] (2) Discriminator network design, such as Figure 3 As shown:
[0115] The discriminator adopts the U-Net structure, in which SRB includes residual blocks and spectral normalization to suppress the mode collapse phenomenon during training and enhance the ability to discriminate generated images.
[0116] The discriminator is based on the U-Net discriminator and has been modified according to the training requirements of . The discriminator takes a 512×512×4 tensor as input. Downsampling is performed using average pooling, while the spectral residual block (SRB) encodes and extracts mixed features. In the intermediate layer with 256 channels, a global context module (GCM) is introduced to enhance the global context relationship. In a certain SRB in the middle, the SRB is performed before the GCM. On the encoder side, the model outputs a 4×4×2048 tensor, which is simplified to a scalar statistic by element-wise summation. This statistic passes through a fully connected layer and then a sigmoid activation function to produce a 1×1 scalar output to evaluate the overall authenticity of the image. The decoder is symmetric to the encoder. Upsampling is performed using linear interpolation, which generates a 512×512×1 matrix for evaluating the authenticity of each pixel. In order to enhance the ability of the discriminator, GFS is used between layers to effectively guide feature integration.
[0117] The detailed process is as follows: Inspired by the denoising diffusion probability model, a self-attention mechanism is added to the intermediate layers of both the encoder and the decoder. The attention mechanism may enhance the discriminator's ability to capture long-term dependencies in the input data, thereby improving the discrimination ability by considering the global context rather than relying solely on local features. In addition, self-attention enables the discriminator to dynamically adjust the importance of various features according to the relevance of the recognition task, thereby more effectively processing complex and structured inputs. This adaptability enhances the robustness of the model and the generalization ability across different datasets. And a guided fusion strategy is added between layers. The discriminator produces two outputs: one evaluates the authenticity of the overall image based on low-frequency information, and the other evaluates the authenticity of each pixel, focusing on high-frequency details. This approach has been shown to be very effective in image generation tasks. In the architecture, SRB consists of a convolutional layer and an activation function, and batch normalization is omitted. This design enables the network to learn a residual function (difference) instead of a direct mapping, which facilitates the training of deeper networks. Spectral normalization is applied to the residual block within the discriminator. Spectral normalization is a technique that controls the Lipschitz constant of a network by normalizing the spectral norm (largest singular value) of the weight matrix of each layer. By applying spectral normalization to the discriminator, it is prevented from becoming disproportionately stronger than the generator. Maintaining this balance is critical to the training dynamics of GANs as it ensures that neither network dominates. Failure to maintain this balance can lead to problems such as mode collapse or vanishing gradients.
[0118] (3) Loss function design:
[0119] The composite loss function consists of the following five parts:
[0120] Structural consistency loss, using the SSIM loss function to keep the structure of visible light images and infrared images consistent. Mean absolute value loss L1, mainly acts on the low-frequency information of the entire image, constraining the infrared image at the pixel level.
[0121] Texture loss, which is calculated based on the texture difference between the converted infrared image and the real infrared image to maintain the texture consistency of the generated image;
[0122] Color loss, which is calculated based on the color difference between the color distribution of the converted infrared image and the real infrared image to maintain the color consistency of the generated image;
[0123] Adversarial loss, through adversarial training of generative adversarial networks, pushes the generator to generate higher quality infrared images.
[0124] 1. Texture loss:
[0125] In the image generation task, L1 loss focuses on the precise matching of pixel values, and calculates the absolute difference in pixel values between the generated image and the real image. It is an accurate representation of low-frequency components, but has a weaker ability to capture high-frequency details. SSIM loss focuses on the similarity of local structures. It measures the similarity between the generated image and the real image in brightness, contrast, and structure. It focuses on the perceptual quality and local structure of the image, and is more sensitive to mid-frequency and high-frequency details. Although SSIM has better capture capabilities in mid- and high-frequency than L1 loss, it is still not enough to capture complex texture details. Therefore, the introduction of texture loss can further enhance the detail quality of the generated image and improve the realism of the generated image, especially when dealing with images with complex textures and rich details.
[0126] The VGG-19 model has a wide range of applications in the field of computer vision, such as image classification, object detection, and image style transfer. Compared with other convolutional neural network models, VGG-19 has a deeper network structure, which enables it to better learn the complex features of the image and enhance the model's expressiveness. The deep features of VGG-19 can capture higher-level semantic information, and the shallower features can capture lower-level detail information. However, choosing too deep a layer may cause the gradient vanishing problem of the loss function, and choosing too shallow a layer may cause overfitting or loss of some important semantic information. Therefore, when constructing texture loss, selecting the middle layer to extract features as loss can capture the required texture semantic information and will not cause the gradient to vanish. According to experience, the conv4_2 layer of VGG-19, that is, the first 22 layers, is selected, the first layer of the VGG-19 network is modified to adapt to single-channel input, and then the generated image and the real label image are input into the VGG-19 network to obtain their feature representations at each layer until this layer. As shown in the figure, it is a visualization image of the extracted features of VGG-19. The squared error loss between the feature representations of the generated image and the target image is calculated at each layer and accumulated so that the final total loss can reflect the differences at multiple levels. In this way, the generative model will simultaneously consider the feature differences at multiple levels during the optimization process to ensure that the generated image is similar to the target image at different scales and levels of detail. After the above process, the texture loss is obtained, and the formula is shown below.
[0127]
[0128] Where N i Represents the normalization factor of the i-th layer, the product of the number of channels, height and width of the feature map of this layer, F i Represents the feature map obtained after the image passes through the i-th layer of VGG-19.
[0129] 2. Color loss:
[0130] In order to ensure that the color style between the generated image and the real image is similar, color loss is used to constrain the generated image to have the same color distribution as the real image. In order to reduce the color difference between the generated image and the real image, Gaussian blur is applied and the Euclidean distance between the obtained feature representations is calculated. The implementation is as follows: First, a Gaussian blur kernel is constructed, and then the image is convolved using the Gaussian blur kernel as the convolution kernel to obtain the blurred image; then the mean square error function between the input image and the target image is calculated as the loss function. As shown in the formula.
[0131]
[0132] Among them, B represents the Gaussian blur convolution operation.
[0133] 3. Other losses: Structural consistency loss, using the SSIM loss function to keep the structure of the visible light image and the infrared image consistent. SSIM is an indicator for evaluating the local contrast, brightness, and structural similarity between two images. The mean absolute value loss L1 mainly acts on the low-frequency information of the entire image and constrains the infrared image at the pixel level.
Claims
1. A method for efficiently converting visible light images into infrared light images based on GAN, characterized in that: The following steps are involved: Input a visible light image and convert it into an infrared image through an adversarial network; For the generator, the InceptioNeXt V2 network is used as the encoder to extract multi-scale features from the input visible light image; In the decoder, the collaborative extraction module SEM and the selective state space model Mamba module are used to extract global features, and the LFEM module is used to supplement local information and generate infrared images; Use GFS-guided fusion strategy or direct concatenation of features by channel to fuse multi-scale features from the encoder and infrared features generated by the decoder; For the discriminator, the authenticity of the generated infrared image is evaluated to identify the authenticity of the overall image and each pixel; The discriminator adopts the U-Net structure, uses the spectral normalized residual block SRB as the encoder and decoder, and optimizes the generator through adversarial training; Use the GFS-guided fusion strategy to fuse the features from the encoder and the infrared features generated by the decoder; A composite loss function is calculated and minimized, wherein the composite loss function includes structural consistency loss, mean absolute value loss, texture loss, color loss and adversarial loss to promote minimization of the difference between the generated infrared image and the real infrared image.
2. The method for efficiently converting a visible light image into an infrared light image based on GAN as claimed in claim 1, characterized in that: The visible light image is a preprocessed image, and the preprocessing step includes normalizing the image so that the image pixel values are within the range of [-1, 1].
3. The method for efficiently converting a visible light image into an infrared light image based on GAN as claimed in claim 1, characterized in that: The decoder of the generator performs the following steps: The input features are successively extracted by the collaborative extraction module SEM, and processed by the selective state space model Mamba module for spatial and channel features, and then fused with the input features; After the fusion features pass through the LFEM module, they are concatenated with the fused features as the output of the decoder.
4. The method for efficiently converting a visible light image into an infrared light image based on GAN as claimed in claim 3, characterized in that: The synergistic extraction module SEM performs the following steps: The input features are sequentially passed through a depth-wise separable convolution, a normalization layer, and a linear layer, connected and activated by the SiLU function, and then output through another linear layer.
5. The method for efficiently converting a visible light image into an infrared light image based on GAN as claimed in claim 3, characterized in that: The Selective State Space Model Mamba module performs the following steps: After the input features pass through the normalization layer, they are divided into two branches: In the first branch, the input passes through the linear layer and activation function in sequence; In the second branch, the input is processed by a linear layer, a depth-wise separable convolution, and an activation function, and then fed into a 2D selective scanning module for further feature extraction; subsequently, another normalization layer is used to normalize the features. Perform element-wise production using the output of the first branch to merge the two paths of the first and second branches; finally, use a linear layer to mix the features and combine this result with a residual connection to form the output of the VSS block; The 2D selective scanning module consists of three parts: scanning extension, S6 block and scanning merging; The scan operation expands the input into sequences along four different directions; Then, the S6 block extracts features from these sequences to ensure that the scanning process captures different features from each direction; Scan merging fuses the features in four directions together by summing them, restoring the output to the same shape as the input.
6. The method for efficiently converting a visible light image into an infrared light image based on GAN as claimed in claim 3, characterized in that: The LFEM module performs the following steps: The input passes through a linear layer to expand the channel to four times the original, and then is divided into two parts; In the first part, multiple transformations are performed at different scales to capture coarse-scale features at different scales; These outputs are concatenated and activated using the SiLU function and then refined using the scSE block; The second part is activated by the SiLU function and multiplies the output of the first part by concatenating the output elements to achieve interaction between multi-scale features; Finally, the combined features are processed through a linear layer, returning the output to the original input dimension.
7. The method for efficiently converting a visible light image into an infrared light image based on GAN as claimed in claim 1, characterized in that: The discriminator adopts a U-Net structure, uses a spectral normalized residual block (SRB) module to encode and extract mixed features, and introduces a GCM module after the middle spectral normalized residual block (SRB) module to enhance the global contextual relationship.
8. The method for efficiently converting a visible light image into an infrared light image based on GAN as claimed in claim 1, characterized in that: The GFS guided fusion strategy includes the following steps: First, in the generator or discriminator, the decoder feature F from the previous layer is e and encoder features F d Splicing is performed according to the channel dimension, and the spatial positions are averaged using global average pooling to generate global statistics that represent the overall characteristics of the channel; Then, two convolutional layers are used to generate the weight W for each channel. c : W c =Sigmoid(Conv(ReLU(Conv((AvgPool([F e ,F d ])))))) Among them, Sigmoid represents Sigmoid function, Conv represents convolution, ReLU represents Relu function, and AvgPool represents average pooling; Get the weight coefficient W c After that, broadcast and multiply with the original concatenated features to re-weight the channels; Among them, F c is the weighted feature; The decoder features F e and encoder features F d Add and fuse, and get the spatial attention coefficient W through the Sigmoid function s : Finally, the spatial attention coefficient W s And the feature map F after channel attention c Broadcast multiplication is performed to correct the features again.
9. The method for efficiently converting a visible light image into an infrared light image based on GAN as claimed in claim 1, characterized in that: The texture loss is calculated based on the texture difference between the converted infrared image and the real infrared image to maintain the texture consistency of the generated image, and is expressed as follows: Among them, L texture represents texture loss, N i Represents the normalization factor of the i-th layer, the product of the number of channels, height and width of the feature map of this layer, F i represents the feature map obtained after the image passes through the i-th layer of VGG-19, X represents the visible light image, Y represents the real infrared image, G represents the generator, and G(X) represents the converted infrared image; The color loss is calculated based on the color difference between the color distribution of the converted infrared image and the real infrared image to maintain the color consistency of the generated image, and is expressed as follows: Among them, L color represents color loss, and B represents the Gaussian blur convolution operation.
10. A system based on GAN for efficiently converting visible light images into infrared light images, using an adversarial network, characterized in that: include: The generator uses the InceptioNeXt V2 network as an encoder to extract multi-scale features from the input visible light image; In the decoder, it is used to use the collaborative extraction module SEM and the selective state space model Mamba module to extract global features, the LFEM module to supplement local information, and generate infrared images; Use GFS-guided fusion strategy or direct concatenation of features by channel to fuse multi-scale features from the encoder and infrared features generated by the decoder; The discriminator is used to evaluate the authenticity of the generated infrared image and identify the authenticity of the overall image and each pixel; The discriminator adopts the U-Net structure, uses the spectral normalized residual block SRB as the encoder and decoder, and optimizes the generator through adversarial training; Use the GFS-guided fusion strategy to fuse the features from the encoder and the infrared features generated by the decoder; A composite loss function, wherein the composite loss function includes structural consistency loss, mean absolute value loss, texture loss, color loss and adversarial loss to promote minimization of the difference between the generated infrared image and the real infrared image.
Citation Information
Patent Citations
Enhancement method for crop disease data
CN112488963A
Dual-light image fusion method and system
CN118196583A
Infrared and visible light image fusion method based on progressive multi-branch and improved UNet3 + deep supervision
CN118229548A
Cited By
Processing method and processing device for low-visible-light image
CN120235774A
High-precision infrared image temperature expression method based on U-Net architecture
CN120599432A
Infrared image generation method and device, equipment and storage medium
CN120612387A
Infrared image enhancement method and system based on Mamba2
CN120823117A
An infrared image enhancement method and system based on Mamba2
CN120823117B