Infrared image generation method and device based on multi-mode self-attention mechanism generative adversarial network
By introducing a multimodal self-attention mechanism and adaptive correlation mask into the generative adversarial network, the background and target features in the visible light image are decoupled, and the problem of incorrect infrared features when generating infrared images is solved, achieving high-quality and correct infrared image generation.
Patent Information
- Application Number
- CN202510100455.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-22
AI Technical Summary
When generating infrared images using a generative adversarial network, it is difficult for the prior art to correctly decouple infrared features of the target and background, resulting in the background and target in the generated infrared image having the same infrared characteristics, contrary to the characteristics of the real infrared image.
A generative adversarial network based on a multimodal self-attention mechanism is adopted to segment and linearly compress the input visible light image, local and global self-attention calculations are performed, and an adaptive correlation mask is introduced, local and global correlation expressions are fused, and the features of the background and target are decoupled to generate the correct infrared image.
It realizes the generation of infrared images with both visual effects and correct infrared features, improves the quality and authenticity of infrared images, and enhances the robustness of the model on different data sets.
Smart Images

Figure CN120070213A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of infrared image generation, and more specifically, relates to an infrared image generation method and device based on a multi-modal self-attention mechanism generative adversarial network. Background Art
[0002] Infrared imaging utilizes the energy difference of thermal radiation of an object and its environment to obtain its image information. In an infrared imaging system, an infrared detector receives thermal radiation, converts it into an electrical signal, and through signal processing and imaging techniques, finally presents the image of the object being photographed. Thanks to this imaging mechanism, infrared images can be used at any time and under any weather conditions, because the energy of thermal radiation of an object and its environment is not restricted by natural conditions such as day and night, sunny and rainy days. Therefore, using infrared images as training data sets for various deep learning tasks is of great significance.
[0003] Traditional infrared simulation systems need to consider various influencing factors such as the target geometric model, target physical property parameters, target temperature calculation, target self-infrared radiation, infrared characteristics of the target surface material, reflection of the target on the background, and environmental radiation when performing infrared simulation, which brings great inconvenience to the simulation. And after considering factors such as the above, under ideal conditions, the simulation results are often too ideal and do not conform to the actual situation, especially the infrared background is quite different from the actual background.
[0004] Compared with traditional infrared simulation systems, deep learning methods such as generative adversarial networks do not need to consider any factors in the simulation process, nor do they need to perform complex theoretical calculations. They only need to input a visible light image to output the corresponding infrared image; in terms of the results, the work done by deep learning is to transfer the target image from the visible light domain to the infrared domain, and the quality of the generated infrared image is higher. However, due to the different imaging mechanisms of visible light and infrared images, directly performing the conversion from visible light to infrared image without decoupling the target and the background will result in the background and the target in the generated infrared image having the same infrared characteristics, which does not conform to the characteristics of real infrared images. Summary of the Invention
[0005] Aiming at the defects of the related technologies, the purpose of the present invention is to provide an infrared image generation method and device based on a multi-modal self-attention mechanism generative adversarial network, aiming to solve the problem that when generating an infrared image from a complex visible light image containing a target through a generative adversarial network, the infrared features are incorrect.
[0006] To achieve the above purpose, the present invention provides an infrared image generation method based on a multi-modal self-attention mechanism generative adversarial network, including:
[0007] S1. Segment the input visible light image to obtain image patches of the same size, and perform linear compression;
[0008] S2. Perform multi-layer window-based local self-attention calculation within the compressed image patches and introduce an adaptive correlation mask to obtain the local correlation expression of the image;
[0009] S3. Perform multi-layer fast global self-attention calculation among the compressed image patches to obtain the global correlation expression of the image;
[0010] S4. Fuse the local correlation expression and the global correlation expression of the image to obtain a feature map containing decoupled background features and target features;
[0011] S5. Decode and upsample the feature map to output an infrared image with the same size as the input image.
[0012] Optionally, in step S2, use W-MSA and SW-MSA in Swin Transformer to perform self-attention calculation inside the window, and introduce an adaptive correlation mask in SW-MSA at the same time.
[0013] Optionally, step S3 includes:
[0014] Adopt the fast global self-attention mechanism FG-MSA among the compressed image patches, compress each window into a vector available for dot product calculation of the self-attention mechanism, and perform multi-head self-attention calculation among the windows to obtain the global correlation expression of the image.
[0015] Optionally, step S4 includes:
[0016] S41. Concatenate and fuse the output of W-MSA and the output of the first FG-MSA to obtain a first result, and perform residual calculation on the first result and its output after passing through a layer normalization layer and a multi-layer perceptron;
[0017] S42. Use the residual result as the input of SW-MSA and the second FG-MSA, concatenate and fuse the output of SW-MSA and the output of the second FG-MSA to obtain a second result, and perform residual calculation on the second result and its output after passing through a layer normalization layer and a multi-layer perceptron;
[0018] S43. Perform linear embedding on the residual result calculated in S42, use it as the input of W-MSA and the first FG-MSA in S41, and perform 6 iterations in sequence to obtain a feature map containing decoupled background features and target features.
[0019] Optionally, step S5 includes:
[0020] The transposed convolution is used to decode the feature map to obtain a high-resolution feature map, and then upsampling is performed to obtain an infrared image with the same size as the input image.
[0021] In a second aspect, the present invention further provides an infrared image generation device based on a multi-modal self-attention mechanism generative adversarial network, including:
[0022] An image segmentation module, configured to segment the input visible light image to obtain image blocks of the same size, and perform linear compression;
[0023] A local processing module, configured to perform multi-layer window-based local self-attention calculation inside the compressed image blocks and introduce an adaptive correlation mask to obtain a local correlation expression of the image;
[0024] A global processing module, configured to perform multi-layer fast global self-attention calculation between the compressed image blocks to obtain a global correlation expression of the image;
[0025] A fusion module, configured to fuse the local correlation expression and the global correlation expression of the image to obtain a feature map containing decoupled background features and target features;
[0026] A decoupling module, configured to decode and upsample the feature map, and output an infrared image with the same size as the input image.
[0027] In a third aspect, the present invention further provides an electronic device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any one of the first aspects are implemented.
[0028] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in any one of the first aspects are implemented.
[0029] Through the above technical solutions conceived by the present invention, compared with the prior art, the following beneficial effects can be achieved:
[0030] 1. An infrared image generation method based on a multi-modal self-attention mechanism generative adversarial network provided by the present invention constructs a generative adversarial network model based on the multi-modal self-attention mechanism (Multi Modal Self-Attention GAN, MMSA-GAN), respectively constructs local and global relevance expressions of an image, and fuses them. During the feature extraction stage, the target and background in the image block are decoupled to generate an infrared image with both visual effects and correct infrared features. Compared with traditional infrared image generation algorithms, the generated infrared image has more correct infrared features and better visual effects.
[0031] 2. An infrared image generation method based on a multi-modal self-attention mechanism generative adversarial network provided by the present invention fuses the local and global relevance expressions of an image to obtain an expression in the latent feature space of the background and target, introduces a feature extraction method including decoupling, and the output infrared image result contains better infrared image evaluation indicators; constructs a fast global self-attention mechanism, which enables the MMSA-GAN model to obtain a global receptive field at the beginning while reducing the computational amount; in local self-attention calculation, by introducing an adaptive correlation mask, the robustness of the MMSA-GAN model on different data sets is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a schematic flowchart of an infrared image generation method based on a multi-modal self-attention mechanism generative adversarial network provided by an embodiment of the present invention;
[0033] Figure 2 is a model framework diagram of an infrared image generation method based on a multi-modal self-attention mechanism generative adversarial network provided by an embodiment of the present invention, where (a) is an image preprocessing module, (b) is a multi-modal self-attention module, (c) is a decoding module, and (d) is an adaptive correlation mask module;
[0034] Figure 3 is a model diagram of the multi-modal self-attention module in an infrared image generation method based on a multi-modal self-attention mechanism generative adversarial network provided by an embodiment of the present invention; where (a) is the overall architecture of the multi-modal self-attention module, and (b) is the internal details of the multi-modal self-attention Block;
[0035] Figure 4 is a model diagram of the fast self-attention module in an infrared image generation method based on a multi-modal self-attention mechanism generative adversarial network provided by an embodiment of the present invention;
[0036] Figure 5This is the training process of an infrared image generation method based on a multi-modal self-attention mechanism generative adversarial network provided by an embodiment of the present invention.
[0037] Figure 6 This is a comparison chart of the results of an infrared image generation method based on a multi-modal self-attention mechanism generative adversarial network provided by an embodiment of the present invention and other methods. Detailed implementation manners
[0038] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0039] The following describes the content involved in the above embodiments in conjunction with a preferred embodiment.
[0040] Embodiment 1
[0041] An infrared image generation method based on a multi-modal self-attention mechanism generative adversarial network, comprising:
[0042] S1. Segment the input visible light image to obtain image blocks of the same size, and perform linear compression;
[0043] S2. Perform multi-layer window-based local self-attention calculation inside the compressed image blocks and introduce an adaptive correlation mask to obtain a local correlation expression of the image;
[0044] S3. Perform multi-layer fast global self-attention calculation between the compressed image blocks to obtain a global correlation expression of the image;
[0045] S4. Fuse the local correlation expression and the global correlation expression of the image to obtain a feature map containing decoupled background features and target features;
[0046] S5. Decode and upsample the feature map to output an infrared image with the same size as the input image.
[0047] For the calculation process of the self-attention mechanism, the input feature is represented as a matrix X, and then through three learnable linear transformation matrices W Q 、W K and W VQuery, key, and value matrices Q, K, and V are generated respectively. The similarity between Q and K is calculated through dot product and the result is scaled by dividing it by the square root of the vector dimension to obtain the attention score matrix A. Subsequently, A is normalized by Softmax to generate the attention weight matrix α, which represents the correlation of each input element with other elements. Then, the value matrix V is weighted and summed using the attention weights α to generate the final output, where each output vector captures the global dependencies in the input sequence. The multi-head self-attention mechanism calculates multiple different attention heads in parallel and finally concatenates the outputs of all heads and fuses them through a linear layer.
[0048] As Figure 2 shown in (a) of H × W , for the visible light image of H×W×3, the Multi Modal Self-Attention GAN (MMSA-GAN) provided by this solution first divides the input visible light image into visible light image patches of size n H × W . Then, the image patches are linearly compressed, and the n H × W three-channel pixel points are compressed into a tensor of 1×1×N. Subsequently, only the compressed tensor needs to be processed, greatly reducing the training time required.
[0049] Optionally, in step S2, W-MSA and SW-MSA in Swin Transformer are used for self-attention calculation inside the window, and an adaptive correlation mask is introduced in SW-MSA.
[0050] W-MSA is a multi-head self-attention mechanism based on local windows. The input features are divided into multiple non-overlapping windows of a fixed size, and then self-attention is calculated independently within each window. This method reduces the computational complexity from the usual O(N 2 ) to O(M×w 2 ) by restricting the attention calculation range, where M is the number of windows and w is the window size. This localized calculation is very efficient, but due to the non-overlap between windows, it limits the information interaction between windows and is difficult to capture the global context.
[0051] To solve the problem of lack of interaction between windows in W-MSA, Swin Transformer introduces SW-MSA. The core idea is that when dividing the input feature map into windows, the window positions are shifted by a certain amount relative to the previous layer. SW-MSA shifts the overall window division one-half window size to the right and down. This shift causes an overlapping area between windows, enabling information between windows to interact through the overlapping area.
[0052] When calculating self-attention within the window in W-MSA and SW-MSA of Swin Transformer, only SW-MSA adopts a mask, and the parameters of the mask are fixed according to the position. The role of this mask is only to provide an efficient way for the calculation of SW-MSA, without considering the possible correlations between non-adjacent tokens within the patched window and the possible different strengths of the correlations between adjacent tokens in ordinary windows. Relying solely on the self-attention matrix to determine the correlations between different tokens is not efficient during optimization. To enable the mask to provide certain guidance during the self-attention calculation in W-MSA and SW-MSA, we propose an adaptive correlation mask. The size of the mask is consistent with the size of the self-attention matrix, and the mask parameters between different stages are independent. After calculating the self-attention matrix, the mask is fused with the self-attention matrix, and then through softmax, the strong and weak correlations between different tokens are determined. The adaptive correlation mask will be specifically optimized for different datasets to efficiently complete the correlation guidance during self-attention calculation.
[0053] As Figure 3 shown in (a) in Figure 3 it, MMSA-GAN processes the input in four stages, each stage containing a pair of multi-modal self-attention modules and a linear embedding calculation. As Figure 2The output in (a) undergoes consecutive self-attention calculations. The multi-modal self-attention module uses W-MSA and SW-MSA in Swin Transformer for self-attention calculations within the window; the calculation method of W-MSA is consistent with that of Swin Transformer. Through the multi-head self-attention mechanism of local windows, the input features are divided into multiple non-overlapping windows of a fixed size, and self-attention calculations are independently performed within each window. By restricting the calculation range of attention, the computational efficiency of the MMSA-GAN model is improved. However, due to the non-overlap between windows, the information interaction between windows is restricted, making it difficult to capture global context. After calculating W-MSA, its output will be fused with the output of FG-MSA and used as the input for the second part of the multi-modal self-attention Block. When SW-MSA divides the input feature map into windows, the window positions are offset by a certain amount relative to W-MSA. SW-MSA will offset the overall window division one-half window size to the right and down. This offset causes a window of SW-MSA to contain parts of multiple windows divided by W-MSA, thus realizing the information interaction between different windows of W-MSA. As the depth of the MMSA-GAN model increases, the receptive field obtained by the local self-attention calculation based on windows will continuously increase. During the calculation of SW-MSA, the adaptive correlation mask is fine-tuned according to the currently trained dataset, as shown in Figure 2 as shown in (d). This enables the self-attention calculation of SW-MSA to have mask values more suitable for the current dataset, resulting in higher robustness during subsequent inference for generating infrared images.
[0054] Optionally, step S3 includes:
[0055] Adopt the fast global self-attention mechanism FG-MSA between the compressed image patches, compress each window into a vector available for the dot product calculation of the self-attention mechanism, and perform multi-head self-attention calculations between the windows to obtain the global relevance expression of the image.
[0056] To enable the MMSA-GAN model to have a global receptive field from the beginning, the MMSA-GAN model is made to obtain a global receptive field from the start. Considering the resource consumption of global self-attention calculation, FG-MSA is proposed. As shown in Figure 4 as shown. FG-MSA divides the MMSA-GAN model into multiple windows and obtains a tensor representing the window through the encoder. By calculating the multi-head self-attention of each feature vector with all tensors representing the windows, fast global self-attention calculation is achieved.
[0057] This solution proposes a Fast Global Self-Attention Mechanism (FG-MSA). FG-MSA compresses each window into a vector that can be used for the dot product calculation of the self-attention mechanism and performs multi-head self-attention calculation among the windows. Assuming there are h×w windows, each window has a size of M×M, and the depth of the feature vector inside the window is C, then the time complexities of the global self-attention calculation and FG-MSA are respectively:
[0058] O(G-MSA) = 4hwC 2 + 2(hw) 2 C
[0059]
[0060] At a fixed size, for an input image of any size, O(FG-MSA) will be a linear complexity of hw.
[0061] Optionally, step S4 includes:
[0062] S41. Concatenate and fuse the output of W-MSA and the output of the first FG-MSA to obtain a first result, and perform a residual calculation on the first result and its output after passing through a layer normalization layer and a multi-layer perceptron;
[0063] S42. Use the residual result as the input of SW-MSA and the second FG-MSA, concatenate and fuse the output of SW-MSA and the output of the second FG-MSA to obtain a second result, and perform a residual calculation on the second result and its output after passing through a layer normalization layer and a multi-layer perceptron;
[0064] S43. Perform a linear embedding on the residual result calculated in S42, use it as the input of W-MSA and the first FG-MSA in S41, and perform 6 iterations in sequence to obtain a feature map containing decoupled background features and target features.
[0065] Referring to Figure 3 in (b), the multi-modal self-attention block takes the output of the previous stage as the input. First, concatenate the output of W-MSA and the output of the first FG-MSA through the contact module to obtain a first result, and perform a residual calculation on the first result and its output after passing through a layer normalization layer and a multi-layer perceptron to obtain the self-attention calculation fusion result of W-MSA and the first FG-MSA. Then, use the self-attention calculation fusion result of W-MSA and the first FG-MSA as the input of SW-MSA and the second FG-MSA to perform a new round of fusion self-attention calculation, realizing the deep fusion of local self-attention and global self-attention.
[0066] The fusion result of the self-attention calculations of W-MSA and the second FG-MSA is used as the input to SW-MSA and the first FG-MSA in the next-level module. Through continuous deepening of the depth and 6 iterations in sequence, the MMSA-GAN model can decouple the target and background in the visible light image and output the latent space representations of the background and the target.
[0067] Optionally, the step S5 includes:
[0068] The feature map is decoded using transposed convolution to obtain a high-resolution feature map, and then upsampled to obtain an infrared image with the same size as the input image.
[0069] Using the transposed convolution technique, the latent space representations of the background and the target can be decoded and upsampled to an infrared image with the same size as the input. Transposed convolution expands the spatial dimension of the feature map by applying a convolutional kernel to the feature map. The latent space representations of the background and the target are initially regarded as the output generated by conventional convolution, and then upsampled to obtain an infrared image with the same size as the input image. The MMSA-GAN model adopts the same training strategy as CycleGAN and performs unsupervised training under the condition of an unpaired dataset by introducing a cycle consistency loss.
[0070] Different from standard convolution, the purpose of transposed convolution is to expand the spatial dimension of the feature map, usually achieved by applying a convolutional kernel to the feature map. The latent space representations of the background and the target are first regarded as the output generated by conventional convolution, and then the original high-resolution feature map is "restored" through an inverse operation. Specifically, transposed convolution uses the same convolutional kernel, but moves and pads in a different way during the calculation process. Usually, zeros are padded around the input feature map, and new data is inserted between adjacent positions, which can effectively capture and reconstruct the details of the image.
[0071] The above ensures that the generated infrared image has the correct infrared characteristics. To make the results generated by the MMSA-GAN model under unsupervised conditions consistent with the content of the input visible light vehicle image. The entire training process is as Figure 5 shown. The training mainly consists of two parts. The first part generates an infrared image from a visible light image, and the second part generates a visible light image from an infrared image. The training processes of these two parts are exactly the same, and each part includes two generator models and one discriminator model. The parameters of the generators in each part are shared. In the first part, the generator G AB is a network that generates an infrared image from a visible light image output; the second generator G BA is a network that generates a visible light image from an infrared image. The discriminator model D BThe input is the infrared images generated by the generative model and the real infrared images, and the output is a scalar used to evaluate whether the input is a real infrared image. The second part is the process of training the generation of visible light from infrared. The discriminative model D A The input is the visible light images generated by the generative model and the real visible light images, and the output is a scalar used to evaluate whether the input is a real visible light image. After each part generates an image, the generated image is used as the input, and through another generator, an image with the same band as the original input image is generated, and the cyclic consistency loss L1 between the two is calculated to ensure the consistency of the content of the input and output images. Using the cyclic consistency loss as one item of the loss function realizes unsupervised training under unpaired data sets.
[0072] To verify the effectiveness of the infrared image generation method based on the multi-modal self-attention mechanism generative adversarial network proposed in the present invention, in a specific embodiment, model training and performance verification are to be combined with the actual scenario.
[0073] The specific steps are as follows:
[0074] (1) Construction of the training data set
[0075] To comprehensively evaluate the advantages of this solution in generating infrared images containing targets, the experiment will be carried out on the VisDrone-DroneVehicle data set. VisDrone-DroneVehicle contains 28,439 pairs of RGB-infrared paired images, covering urban roads, residential areas, parking lots and other scenarios, from day to night. Each scenario contains a large number of vehicle targets. Considering the training and testing costs, this solution uses 8,000 pairs of visible-infrared images for training and 1,000 visible light images for testing.
[0076] (2) Evaluation metrics and training platform
[0077] For the evaluation metrics, this solution selects SSIM, PSNR and LPIPS. These metrics have their own characteristics and comprehensively reflect different aspects of image quality. SSIM focuses on the structural information in the image and effectively evaluates the human perception of image details. PSNR provides a quantitative evaluation of the overall brightness and contrast and is often used to measure the distortion degree in compressed images. LPIPS calculates the perceptual difference between images using a deep learning model and provides an approximation closer to the experience of the human visual system. All experiments are carried out on NVIDIA GeForce RTX 3090 using the PyTorch framework.
[0078] (3) Display of experimental results
[0079] This solution compares MWSA-GAN with current visible-to-infrared image conversion methods.
[0080] Reference Figure 6 , intuitively demonstrating the detailed differences in the generated features and results among various methods. CycleGAN can retain the original shape information of the input image as much as possible to achieve a rough visible-to-infrared image conversion. However, due to the lack of decoupling between the target and the background in the feature extraction stage, both show the same infrared features in the generated image. Although pix2pix can partially retain the infrared features of the target overall through supervised training and emphasizing in the loss function, it ignores the shape information of the input image during the generation process, resulting in serious distortion of the shapes of the target and the background in the generated infrared image. The results generated by MUNIT also face the problems of shape distortion and loss of infrared feature information. The infrared image generation method based on the multi-modal self-attention mechanism generative adversarial network proposed in this solution not only retains the original content information of the image but also ensures that both the generated infrared background and target contain accurate infrared features.
[0081] Table 1 shows the evaluation of the quality of the generated images using four metrics: SSIM, PSNR, and LPIPS. SSIM is mainly used to measure the similarity in brightness, contrast, and structure between two images. Its calculation method takes into account the local features of the image, comparing brightness, contrast, and structural information. This method enables SSIM to more accurately reflect human perception of image quality and thus becomes an important metric for evaluating the fidelity of generated images. PSNR is an objective image quality evaluation method, and its basic concept is to evaluate image quality by comparing the mean square error between the real image and the generated image. LPIPS is a deep learning-based metric that focuses on the perceptual features of images and aims to more accurately measure the perceptual similarity between them. It uses a pre-trained convolutional neural network to extract deep features from images and then calculates the distance between these features.
[0082] It can be clearly seen from Table 1 that when performing visible-to-infrared conversion in complex scenarios, MMSA-GAN outperforms all other methods in terms of the SSIM, PSNR, and LPIPS metrics. This indicates that MMSA-GAN effectively maintains image quality and perceptual similarity while enhancing data augmentation capabilities.
[0083] Table 1
[0084] SSIM PSNR LPIPS CycleGAN 0.5722 27.86 0.2444 pix2pix 0.6583 28.31 0.1841 munit 0.5539 28.25 0.2971 MMSA-GAN 0.6907 28.85 0.1835
[0085] A method for generating infrared images based on a multi-modal self-attention mechanism generative adversarial network, namely MMSA-GAN, is proposed in this solution. Relying on multi-head self-attention, including window-based local self-attention and fast global self-attention calculations, the MMSA-GAN model has a global receptive field from the beginning. In the feature extraction stage, by fusing the local and global correlation expressions of the image, the decoupling of the background and the target is achieved, enabling the generated infrared images to have more accurate infrared features. In the calculation of local self-attention, an adaptive correlation mask is proposed, which makes the mask vary according to different training datasets, enhancing the stability of the MMSA-GAN model on different datasets. It solves the technical problem that when generating infrared images from complex visible light images containing targets through a generative adversarial network, the infrared features are incorrect, realizes high-quality infrared images with consistent output content, and at the same time ensures that under unsupervised training conditions, the content of the generated infrared vehicle images is consistent with the input, without the need to use paired training data.
[0086] Embodiment 2
[0087] The present invention also provides an infrared image generation device based on a multi-modal self-attention mechanism generative adversarial network, including:
[0088] An image segmentation module for segmenting the input visible light image into image blocks of the same size and performing linear compression;
[0089] A local processing module for performing multi-layer window-based local self-attention calculations within the compressed image blocks and introducing an adaptive correlation mask to obtain the local correlation expression of the image;
[0090] A global processing module for performing multi-layer fast global self-attention calculations between the compressed image blocks to obtain the global correlation expression of the image;
[0091] A fusion module for fusing the local and global correlation expressions of the image to obtain a feature map containing decoupled background and target features;
[0092] A decoupling module for decoding and upsampling the feature map to output an infrared image with the same size as the input image.
[0093] The infrared image generation device based on a multi-modal self-attention mechanism generative adversarial network provided by the embodiments of the present invention is used to execute the infrared image generation method based on a multi-modal self-attention mechanism generative adversarial network provided by any embodiment of the present invention, and has corresponding beneficial effects.
[0094] Embodiment 3
[0095] The present invention also provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any one of the first embodiments are implemented.
[0096] Embodiment Four
[0097] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the first embodiments are implemented.
[0098] It is easy for those skilled in the art to understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for generating infrared images based on a multimodal self-attention mechanism generative adversarial network, characterized in that: include: S1, segment the input visible light image to obtain image blocks of the same size and perform linear compression; S2, perform multi-layer window-based local self-attention calculations within the compressed image block, and introduce an adaptive correlation mask to obtain the local correlation expression of the image; S3, performing multiple layers of fast global self-attention calculations between compressed image blocks to obtain a global correlation expression of the image; S4, fusing the local correlation expression and the global correlation expression of the image to obtain a feature map including decoupled background features and target features; S5. Decode and upsample the feature map to output an infrared image with the same size as the input image.
2. The method according to claim 1, characterized in that In step S2, W-MSA and SW-MSA in Swin Transformer are used to perform self-attention calculation inside the window, and an adaptive correlation mask is introduced in SW-MSA.
3. The method according to claim 2, characterized in that The step S3 comprises: A fast global self-attention mechanism FG-MSA is used between compressed image blocks to compress each window into a vector that can be used for dot product calculation of the self-attention mechanism, and multi-head self-attention calculation is performed between windows to obtain the global correlation expression of the image.
4. The method according to claim 3, characterized in that The step S4 comprises: S41, splicing and fusing the output of the W-MSA and the output of the first FG-MSA to obtain a first result, and performing residual calculation on the first result and the output after passing through a layer of normalization layer and a multi-layer perceptron; S42, using the residual result as the input of the SW-MSA and the second FG-MSA, concatenating and fusing the output of the SW-MSA and the output of the second FG-MSA to obtain a second result, and performing residual calculation on the second result and the output after passing through a layer of normalization layer and a multi-layer perceptron; S43, linearly embed the residual result calculated in S42, use it as the input of W-MSA and the first FG-MSA in S41, and perform 6 iterations in sequence to obtain a feature map containing decoupled background features and target features.
5. The method according to claim 1, characterized in that The step S5 comprises: The feature map is decoded by using transposed convolution to obtain a high-resolution feature map, and is upsampled to obtain an infrared image with the same size as the input image.
6. An infrared image generation device based on a multimodal self-attention mechanism to generate an adversarial network, characterized in that: include: An image segmentation module is used to segment the input visible light image to obtain image blocks of the same size and perform linear compression; The local processing module is used to perform multi-layer window-based local self-attention calculations within the compressed image block and introduce adaptive correlation masks to obtain the local correlation expression of the image; The global processing module is used to perform multi-layer fast global self-attention calculations between compressed image blocks to obtain the global correlation expression of the image; A fusion module is used to fuse the local correlation expression and the global correlation expression of the image to obtain a feature map containing decoupled background features and target features; The decoupling module is used for decoding and up-sampling the feature map, and outputting an infrared image with the same size as the input image.
7. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Multi-modal image fusion method, device and equipment
CN116363037A
Three-dimensional seismic data fault identification method and device
CN118011484A
Multi-modal end-to-end automatic driving method and system based on unified aerial view representation
CN119049000A
Method for generating pseudo-infrared thermal imaging images in batches by visible light images
CN119205968A