Thermal infrared image optimization method based on multi-channel fusion and semantic information

By adopting the MCFGAN algorithm based on multi-channel fusion and semantic information in infrared image processing, the problems of poor detail recovery and high noise in image processing in infrared image field are solved, and the image quality and resolution are improved.

CN120163710AActive Publication Date: 2025-06-17CHINA THREE GORGES UNIV

Patent Information

Application Number
CN202510226690.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-17
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

In the prior art, image processing in the infrared image field has problems such as poor detail recovery, high noise and lack of semantic information guidance.

Method used

Using a thermal infrared image super-resolution algorithm (MCFGAN) based on multi-channel fusion and semantic information, the image is optimized using semantic information by building a model including a generator and a semantic discriminator of a multi-channel fusion attention module.

Benefits of technology

It effectively eliminates noise in thermal infrared images, improves the efficiency of high-frequency details, improves the information utilization of information mining depth and attention, and thus improves image quality and resolution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163710A_ABST
    Figure CN120163710A_ABST
Patent Text Reader

Abstract

The invention provides a thermal infrared image optimization method based on multichannel fusion and semantic information, and relates to the technical field of image processing. Comprising the steps that a VIS and a corresponding TIR data set are obtained, an HR image obtained after preprocessing is input into an MCFGAN model, the MCFGAN model comprises a semantic discriminator and a generator, and the generator comprises three DISR modules skipping connection; the method comprises the following steps: inputting an HR image data set into an infrared degradation model to obtain an LR image data set, training a generator by using the LR and HR image data sets, obtaining an SR image data set, training the generator and a discriminator to play a game with each other in combination with a semantic discriminator through a generative adversarial network mode, and obtaining an optimized MCFGAN model, thereby effectively eliminating noise in a thermal infrared image. According to the method, the extraction efficiency of high-frequency details of the thermal infrared image is improved, the information mining depth and the utilization rate of information in attention are improved, so that the information fusion efficiency is improved, and the image quality and resolution are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a method for optimizing thermal infrared images based on multi-channel fusion and semantic information. Background Art

[0002] Currently, infrared thermal imaging technology can perform non-contact and high-resolution temperature imaging, providing rich information about the measurement target. Therefore, it has been widely applied in industries such as power systems, civil engineering, automobiles, metallurgy, petrochemicals, and medicine. Limited by the hardware performance of the thermal imager, compared with visible light (VIS, Visible Image) images, thermal infrared (TIR, Thermal Infrared) images are characterized by low resolution, high noise, blurred edges, random bright spots, and irregular stripes. To overcome the limitations of hardware facilities, it is more cost-effective to improve the resolution of TIR images by using super-resolution (SR, Super Resolution) reconstruction technology.

[0003] Among them, single-image super-resolution is a classic computer vision problem, aiming to restore a given low-resolution (LR, Low Resolution) image to a high-resolution (HR, High Resolution) image. In the field of VIS images, with the rapid development of deep neural networks, SR has made quite a breakthrough. Currently, the mainstream SR methods include: based on convolutional neural networks, based on generative adversarial networks (GANs), based on Transformers, blind super-resolution, and real-world super-resolution.

[0004] Currently, in the research of infrared images, a single-image super-resolution method for infrared images that combines compressed sensing theory and deep learning is adopted. By using the sparsity in the fact that a low-resolution image can be regarded as the compressed sampling result of a high-resolution image in compressed sensing, a higher-resolution image is reconstructed. Combining a Transformer with spatial and channel dual attention mechanisms can solve the problem that infrared SR requires more global edge structure information.

[0005] Existing research has problems such as poor detail recovery and high noise. The present invention proposes a super-resolution algorithm for thermal infrared images based on multi-channel fusion and semantic information (MCFGAN, Multi-Channel Feature Generative Adversarial Network), and uses a training strategy of guiding network training with a semantic discriminator and grayscale visible light images, showing significant advantages in thermal infrared image super-resolution compared with other super-resolution networks. Summary of the Invention

[0006] The main purpose of the present invention is to provide a thermal infrared image optimization method based on multi-channel fusion and semantic information, so as to solve the technical problems of poor detail restoration, large noise and lack of semantic information guidance in image processing in the field of infrared images in the prior art.

[0007] In order to solve the above technical problems, the technical solution adopted by the present invention is: a thermal infrared image optimization method based on multi-channel fusion and semantic information, comprising the following steps: Obtain VIS and corresponding TIR datasets, perform data preprocessing, and obtain HR image datasets; Construct an MCFGAN model; wherein the MCFGAN model includes a generator module with a multi-channel fusion attention module and a semantic discriminator: the generator includes a multi-channel fusion attention module with three skip connections, namely, a DISR module; the semantic discriminator includes a semantic module and a semantic perception fusion module; wherein the semantic module is used to generate semantic information, and the semantic perception fusion module is used to integrate the semantic information into the discriminator; Input the HR image dataset into the infrared degradation model to obtain the LR image dataset; The generator is trained using the HR image dataset and the LR image dataset to generate the SR image dataset, and then the generated SR image dataset is sent to the semantic discriminator to judge whether it is true or false; In the semantic discriminator, the semantic module generates semantic information based on the HR image dataset, and integrates the semantic information into the discriminator through the semantic perception fusion module, so as to perform adversarial training on the MCFGAN model and obtain the optimized MCFGAN model; The TIR image to be tested is input into the optimized MCFGAN model to obtain the optimized thermal infrared image.

[0008] In the preferred embodiment, the data preprocessing includes graying the VIS and TIR data sets to generate HR image data sets, and then performing CLAHE processing; CLAHE processing: The core of CLAHE contrast limited adaptive histogram equalization processing lies in local histogram equalization and contrast limitation. The formula of local histogram equalization is: (1.1); In the formula, The gray value in the local area is The number of pixels, To indicate that the gray value is less than or equal to The cumulative sum of all pixel numbers, and is the pixel value of the image; The local histogram is equalized through linear transformation to make the pixel value distribution of the image as uniform as possible. The formula is: (1.2); Wherein, is the pixel value of the original image at position ; is the pixel value of the new image after local histogram equalization processing at position ; is the cumulative distribution function value of the minimum pixel value in the local histogram, is the cumulative distribution function value of the maximum pixel value in the local histogram, is the cumulative distribution function value of the pixel value in the local histogram, is the maximum value of the new pixel value range; The contrast limit formula is: (1.3); Wherein, is the frequency value of the local histogram after contrast limit processing, is the histogram of the local area, and ClipLimit is the contrast limit parameter.

[0009] In the preferred solution, in the semantic discriminator, the semantic module Semantic Model is used to extract accurate and rich semantic features of the infrared image; then the semantic perception fusion module combines the extracted semantic features with the discriminator; Among them, the semantic module adopts Modified Resnet-50 and extracts the semantic feature S of the third layer h Input the semantic perception fusion block, and the discriminator conducts semantic guidance, embedding the extracted image semantic information into the feature space of the discriminator, providing key guidance for its discrimination of texture authenticity, thereby forcing the discriminator to focus on the distribution of semantic perception textures, specifically: A1: After the semantics is transmitted to the self-attention module, the query is obtained: (7.1); Wherein, LN, GN, SA, and RA are layer normalization, group normalization, self-attention module, and rearrangement respectively; A2: Transmit the SR image generated by the generator and the HR image in the HR dataset into the convolutional layer to obtain the original enhanced features. After the original enhanced features are rearranged, the keys and of the SR image and the HR image, as well as the values and , and then input the query Q, the keys of the SR image and the HR image and as well as the values and into the Cross Attention module; A3: Concatenate the semantic-aware image features output by the Cross Attention with the original enhanced features to obtain the features of the final SR image and the HR image and , and the formula is: (7.2); (7.3); (7.4); (7.5); (7.6); (7.7); (7.8); (7.9); In the formula, is the output after the multi-head attention mechanism of the Cross Attention module, and are the outputs after the BasicTransformerBlock, is the scaling factor, is the feed-forward network, is the original enhanced feature.

[0010] In the preferred solution, the generator module includes a shallow feature extraction module, a multi-channel fusion attention module for depth feature extraction, and a reconstruction module, and the formula is: The shallow feature extraction module uses a convolutional layer with a 3×3 convolutional kernel, and the formula is: , is the extracted feature map; The multi-channel fusion attention module, that is, the MCF module, is used to extract depth features, and the formula is: (6.1); In the formula, is the MCF block, and the MCF block is composed of three DISR blocks with skip connections, is the output of the nth DISR block; The DISR block contains a densely connected module branch, a context attention module branch, and a statistical spatial channel attention module branch. The dynamic weights of these three branches are determined by the input features. Then, features are extracted from using a 3×3 convolutional kernel: , completing the deep feature extraction part; The equation for image reconstruction is: (6.2); In the formula, is interpolation upsampling, is the final SR output.

[0011] In the preferred solution, each DISR module generates weights for three independent channels by using the same input features. The formula for the specific operation is: (5.1); (5.2); (5.3); (5.4); (5.5); (5.6); In the formula, is the output of the densely connected module, is the output of the context attention module, is the output of the statistical spatial channel attention module; is the weights of the three channels generated by the dynamic weight module according to the input features. In the preferred solution, each DISR module controls the balance through weighted summation, including the dynamic weighted contributions of the densely connected module, the context attention module, and the statistical spatial channel attention module; The densely connected module Dense is adopted in the DISR module, and the formula is: (2.1); (2.2); (2.2); (2.4); (2.5); (2.6); (2.7); (2.8); (2.9); (2.10); Wherein, is the convolutional layer, is the activation function, is the feature obtained by processing layer-by-layer convolution and concatenating with the previous layer, Y is the final output, is the weighting coefficient.

[0012] In a preferred embodiment, a context attention module is adopted in the DISR module, and the formula is as follows: (3.1); (3.2); (3.3); (3.4); (3.5); (3.6); (3.7); Wherein, is the global feature, is the compressed feature, is the context information in the horizontal direction, is the context information in the vertical direction, is the length of the horizontal convolution kernel, is the height of the vertical convolution kernel, is the feature remapped to the original channel dimension, is the attention weight, is the final output, is the element-wise multiplication.

[0013] In a preferred embodiment, in the DISR module, a statistical spatial channel attention module, the specific formula is as follows: (4.1); (4.2); (4.3); (4.4); (4.5); (4.6); Wherein, is the mean within the channel, is the squared deviation from the mean, is the stable normalization factor after adding the regularization parameter to the normalization factor, is the normalization factor, is the regularization parameter to prevent the denominator from being too small when calculating the normalization factor, is the attention weight.

[0014] In the preferred embodiment, in the DISR module, the input features are first compressed using global average pooling, then passed through a connection layer containing two fully connected layers and a ReLU activation layer, and finally the weights of each branch are calculated using the Softmax function.

[0015] In the preferred embodiment, the generator is trained using the HR image dataset and the LR image dataset. During the adversarial training process of the MCFGAN model, the discriminator is used to distinguish whether the image is real or fake, and the discriminator is optimized using the loss function. The formula is: (9.1); In the formula, and are the distributions of the high-quality image and the generated image respectively, is the expectation, and D is the discriminator; The generator is optimized through the combination of three losses. The loss function includes the L1 loss, the perceptual loss, and the adversarial loss. The formula is: (9.2); In the formula, and are the weight coefficients of the perceptual loss and the adversarial loss respectively, is the adversarial loss, is the perceptual loss, is the pixel-level supervision loss; The goal of the adversarial loss is to make the generator deceive the discriminator, and the optimization goal is opposite to that of : (9.3); In the formula, is the expectation, and are the distributions of the high-quality image and the generated image respectively.

[0016] The present invention provides a method for optimizing thermal infrared images based on multi-channel fusion and semantic information. By obtaining VIS and corresponding TIR data sets, and then inputting the HR image data set obtained after preprocessing into the MCFGAN model. The MCFGAN model includes a semantic discriminator and a generator. The generator includes three DISR modules with skip connections. Input the HR image data set into the infrared degradation model to obtain the LR image data set. Then, use the LR and HR image data sets to train the generator to obtain the SR image data set. Combine with the semantic discriminator, and through the way of generative adversarial network, train the generator and the discriminator to play against each other to obtain the optimized MCFGAN model, effectively eliminating the noise in the thermal infrared image, improving the extraction efficiency of high-frequency details in the thermal infrared image, increasing the depth of information mining and the utilization rate of information in attention, thereby improving the information fusion efficiency, image quality and resolution. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The present invention will be further described below with reference to the drawings and embodiments: Figure 1 is the flow chart of the thermal infrared image optimization method of the present invention; Figure 2 is the schematic diagram of the generator model structure of the present invention; Figure 3 is the Semantic Model diagram of the semantic module structure of the present invention; Figure 4 is the MST structure diagram of the semantic perception fusion block of the present invention; Figure 5 is the structure diagram of the semantic discriminator of the present invention; Figure 6 is the overall network architecture diagram of the generator of the present invention; Figure 7 is the comparison effect diagram of MCFGAN and the existing super-resolution network when the scaling factor is 4 of the present invention; Figure 8 is the comparison effect diagram of MCFGAN and the existing super-resolution network when the scaling factor is 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] Embodiment 1 As Figure 1-8 shown, a method for optimizing thermal infrared images based on multi-channel fusion and semantic information includes the following steps: S1: Obtain VIS and corresponding TIR data sets, and after data preprocessing, obtain the HR image data set.

[0019] S2: Construct the MCFGAN model. The MCFGAN model includes a generator (MCFNet) module with a multi-channel fusion attention (DISR) module and a semantic discriminator. The generator (MCFNet) includes three DISR modules with skip connections. The semantic discriminator includes a semantic module and a semantic perception fusion module (MST). The semantic module is used to generate semantic information, and the MST module is used to integrate the semantic information into the discriminator.

[0020] S3: Input the HR image dataset into the infrared degradation model to obtain the LR image dataset.

[0021] S4: Use the HR image dataset and the LR image dataset to train the generator to generate the SR image dataset, and then send the generated SR image dataset into the semantic discriminator to judge true or false.

[0022] S5: In the semantic discriminator, the semantic module generates semantic information based on the HR image dataset, and integrates the semantic information into the discriminator through the semantic perception fusion module, so as to perform adversarial training on the MCFGAN model to obtain an optimized MCFGAN model.

[0023] S6: Input the VIS to be measured and the corresponding TIR images into the optimized MCFGAN model to obtain the optimized thermal infrared image.

[0024] In this embodiment, MCFNet is a Multi-Channel Feature Network, and DISR is Deep Image Super-Resolution with Multi-Channel Fusion.

[0025] Furthermore, Figure 1 The infrared degradation model used in this embodiment is the high-order thermal infrared degenerator used in the patent application with the application number CN 117372254A and the name "A Thermal Infrared Image Super-Resolution Algorithm Based on a Generative Adversarial Network with Multi-Structure Fusion", and other modules that can obtain LR images can be adaptively used instead.

[0026] In this embodiment, by obtaining the VIS and corresponding TIR data sets, and then inputting the HR image data set obtained after preprocessing into the MCFGAN model. The MCFGAN model includes a semantic discriminator and a generator. The generator includes three DISR modules with skip connections. The HR image data set is input into the generator to obtain the LR image data set. Then, the generator (MCFNet) is trained using the LR and HR image data sets to obtain the SR image data set. Combined with the semantic discriminator, through the generative adversarial network (GAN) method, the generator and the discriminator play against each other. The semantic discriminator provides semantic information about the image content, and the optimized MCFGAN model is obtained for testing, effectively eliminating the noise in the thermal infrared image. The multi-channel fusion structure improves the depth of information mining and the utilization rate of information in the attention, thereby improving the efficiency of information fusion, the image quality, and the resolution.

[0027] The semantic discriminator is composed of a semantic module, a semantic perception fusion block, and an ordinary discriminator.

[0028] In the preferred solution, in step S1, the data preprocessing includes grayscaling the VIS and TIR data sets to generate the HR image data set, and then performing CLAHE (Contrast Limited Adaptive Histogram Equalization) processing.

[0029] CLAHE processing: A technique for enhancing the local contrast of an image through local histogram equalization and contrast limitation, aiming to avoid the over-amplification of noise while enhancing the detailed information of the image. The core of CLAHE contrast-limited adaptive histogram equalization processing lies in local histogram equalization and contrast limitation. The formula for local histogram equalization is: (1.1); In the formula, is the number of pixels with a gray value of in the local area, represents the cumulative sum of all pixels with a gray value less than or equal to , and are the pixel values of the image.

[0030] Then, the local histogram is equalized through a linear transformation to make the pixel value distribution of the image as uniform as possible. The formula is: (1.2); In the formula, is the pixel value of the original image at position , is the pixel value of the new image after local histogram equalization processing at position The pixel value at is the cumulative distribution function value of the minimum pixel value in the local histogram, is the cumulative distribution function value of the maximum pixel value in the local histogram, is the pixel value is the cumulative distribution function value in the local histogram, is the maximum value of the new pixel value range; The contrast limit formula is: (1.3); In the formula, is the frequency value of the local histogram after contrast limit processing, is the histogram of the local area, the frequency corresponding to the gray value of each pixel in this area, ClipLimit is the contrast limit parameter, used to control the contrast enhancement of the local area.

[0031] In this embodiment, a grayscale visible light image is used to guide network training: first, an LR image is generated by a thermal infrared degradation generator constructed by combining a grayscale visible light image with a simply sharpened thermal infrared image, and MCFNet is trained with the LR and HR images, and then on this basis, a semantic module is combined to guide the training of MCFGAN.

[0032] In this embodiment, the generator (MCFNet) module includes a shallow feature extraction module, a DISR module for deep feature extraction, and a reconstruction module.

[0033] The generator includes three DISR modules with skip connections.

[0034] In an infrared image, there may be large differences in subtle temperature changes or object shapes. Dense connections, through the close cooperation between layers, help capture these details, thereby improving the ability to perceive minute changes.

[0035] In a preferred solution, in the multi-channel fusion attention module, each DISR module controls the balance through weighted summation, including the dynamic weighted contributions of a dense connection module (Dense), a context attention module (DCA), and a statistical spatial channel attention module (SSCA); The Dense module is adopted in the DISR module, and the formula is: (2.1); (2.2); (2.2); (2.4); (2.5); (2.6); (2.7); (2.8); (2.9); (2.10); In the formula, represents the convolutional layer, is the activation function, is the feature obtained by processing layer-by-layer convolution and concatenating with the previous layers, Y is the final output, is the weighting coefficient for enhancing residual learning.

[0036] In this embodiment, the Dense module processes features through layer-by-layer convolution and concatenates with the previous layers, thereby enhancing the utilization of features and the flow of information, promoting the flow of information and the transmission of gradients, and reducing the problem of gradient disappearance.

[0037] The DCA module combines multiple convolutional operations to extract context information, including horizontal convolution and vertical convolution, and generates adaptive weights related to the input feature map through these convolutions.

[0038] In the preferred solution, the DCA module is adopted in the DISR module, and the formula is as follows: (3.1); (3.2); (3.3); (3.4); (3.5); (3.6); (3.7); In the formula, is the global feature, is the compressed feature, is the context information in the horizontal direction, is the context information in the vertical direction, is the length of the horizontal convolution kernel, is the height of the vertical convolution kernel, is the feature remapped to the original channel dimension, is the attention weight, is the final output, is the element-wise multiplication.

[0039] In this embodiment, as Figure 6 shown, the main purpose of the DCA module is to extract spatial context information through convolution operations in the horizontal and vertical directions, and fuse this information into the original features by means of an attention mechanism. Finally, the context information and the original features are combined through element-wise multiplication, thereby enhancing the representation of key regions or channels and improving the network's ability to process complex features.

[0040] The SSCA module generates adaptive weights by calculating the mean difference kernel variance information of the image and applies them to the input to enhance the feature response of key regions.

[0041] In the preferred solution, in the DISR module, the specific formula of the SSCA module is as follows: (4.1); (4.2); (4.3); (4.4); (4.5); (4.6); In the formula, is the mean within the channel, is the squared deviation of the mean, is the stable normalization factor after adding a regularization parameter to the normalization factor, is the normalization factor, is the regularization parameter to prevent the denominator from being too small when calculating the normalization factor, is the attention weight.

[0042] In this embodiment, the SSCA module calculates the mean and variance of each channel, performs stable normalization to prevent instability caused by too small values during calculation; calculates the attention weight, calculates the attention weight of each channel through the Sigmoid activation function, so that the features of the channel can be weighted according to the weight during output, the features of important channels will be amplified, and unimportant channels will be suppressed. Finally, the original features and the calculated attention weights are weighted and fused through element-wise multiplication to obtain the output features. The SSCA module dynamically focuses on channels with large statistical variations and suppresses channels with relatively stable information, thereby improving the effectiveness of feature representation and network performance.

[0043] In this embodiment, as Figure 6As shown, in the DISR module, the dense connection module enhances the details of the thermal infrared image by fusing features from different layers, the context attention module enhances the feature responses at different scales, optimizes the performance in low-contrast scenarios, and the statistical spatial channel attention module captures the distribution of global and local features, enhances the attention to key regions, and adaptively identifies the degraded features of the thermal infrared image.

[0044] In the generator, the DISR module is adopted. Each DISR module controls and balances the dynamic weighted contributions of the dense connection module, the context attention module, and the statistical spatial channel attention module through weighted summation. The DISR module generates weights by using the same input features of its blocks as three independent channels.

[0045] In the preferred solution, each DISR module generates weights by using the same input features as three independent channels, and the formula for the specific operation is: (5.1); (5.2); (5.3); (5.4); (5.5); (5.6); In the formula, is the output of the dense connection, is the output of the context attention module, is the output of the statistical spatial channel attention module. is trained and generated by the dynamic weight module according to the input features instead of setting fixed values to calculate the weights of the three channels.

[0046] To simplify the learning process of the DISR module, in the embodiment, is set, and it can be adaptively modified according to the actual situation.

[0047] In this embodiment, a generator with a multi-channel fusion attention module is adopted, which is called MCFNet. MCFNet is as Figure 6 shown.

[0048] In the preferred solution, the generator (MCFNet) module includes a shallow feature extraction module, a multi-channel fusion attention module for deep feature extraction, and a reconstruction module, and the formula is: The shallow feature extraction module uses a convolutional layer with a 3×3 convolutional kernel, and the formula is , is the extracted feature map; The multi-channel fusion attention module, namely the MCF module, is used to extract deep features, and the formula is: (6.1); In the formula, is the MCF block. The MCF block consists of three skipped-connected DISR blocks. The DISR block contains a densely connected module branch, a context attention module branch, and a statistical spatial channel attention module branch. The dynamic weights of these three branches are determined by the input features; then, a 3×3 convolutional kernel is used to extract features from : , completing the deep feature extraction part.

[0049] The equation for image reconstruction is: (6.2); In the formula, represents interpolation upsampling, is the final SR output.

[0050] In the preferred solution, as Figure 4 and Figure 5 shown, in the semantic discriminator, the semantic module is used to extract accurate and rich semantic features of the infrared image; then, the MST module is used to combine the extracted semantic features with the discriminator.

[0051] In the semantic discriminator, the semantic module adopts Modified Resnet-50 and extracts the semantic feature S of the third layer h input to MST.

[0052] Obtain the semantic , adopt MST to conduct semantic guidance on the discriminator, embed the extracted image semantic information into the feature space of the discriminator, and provide key guidance for its identification of texture authenticity, specifically: A1: After the semantic is passed to the self-attention module, the query is obtained: (7.1); In the formula, LN, GN, SA, and RA are layer normalization, group normalization, self-attention module, and rearrangement respectively.

[0053] A2: Pass the SR image generated by the generator and the HR image in the HR dataset into the convolutional layer to obtain the original enhanced features. After rearrangement, the keys and of the SR image and the HR image and the values and , and then input the query Q, the keys of the SR image and the HR image and as well as the values and into the Cross Attention module.

[0054] A3: Concatenate the semantic-aware image features output by the Cross Attention with the original enhanced features to obtain the features of the final SR image and the HR image and , and the formula is: (7.2); (7.3); (7.4); (7.5); (7.6); (7.7); (7.8); (7.9); In the formula, and are the outputs after the multi-head attention mechanism of the Cross Attention module, and are the outputs after the BasicTransformerBlock, is the scaling factor, is the feed-forward network, is the original enhanced feature.

[0055] In step A1, the relationship between each pixel and other pixels is calculated through the self-attention mechanism, and the semantic features pass through the cross-attention module, which cross-attends the semantic features with other relevant features of the image, focusing on the connection between texture and semantics, to obtain the super-resolution / high-quality image. In step A2, the convolutional layer processes the features and generates key-value pairs. The semantic-aware image features output in step A3 will be concatenated with the original enhanced features. Through feature concatenation, the final feature map will contain information from semantic perception and the enhanced image, thus obtaining high-quality image features and improving the discrimination accuracy of the discriminator.

[0056] Such as Figure 2As shown, it is the structural diagram of the Semantic Model. In this embodiment, the semantic module Semantic Model is based on Modified Resnet-50. Resnet-50 originally consisted of four layers. However, as the number of layers increased, the resolution of the features was downsampled and the semantics became more abstract.

[0057] The semantic module Semantic Model extracts accurate and rich semantic features of the infrared image. Then, MST is used to combine the extracted semantic features with the discriminator. The structure of the semantic discriminator is as Figure 5 shown.

[0058] In the preferred solution, in the DISR module, the input features are first compressed using global average pooling, then passed through a connection layer containing two fully connected layers and a ReLU activation layer, and finally the weights of each branch are calculated through the Softmax function.

[0059] In this embodiment, the semantic module is denoted as ϕ, aiming to achieve more fine-grained semantic-aware texture generation. Its goal is: (8.1).

[0060] In the adversarial training of this embodiment, the loss function consists of pixel-level supervision loss , perceptual loss, and adversarial loss. Among them, the pixel-level supervision loss is used to constrain the pixel-level consistency and ensure the realism of the pixel values; the perceptual loss utilizes the features of VGG to provide a coarse-grained perceptual constraint between the generated image and the real image; the adversarial loss is introduced through the adversarial training process. During the adversarial training process, a discriminator D is used to distinguish whether the image is real or fake and is optimized through the following loss function.

[0061] In the preferred solution, the HR image dataset and the LR image dataset are used to train the generator. During the adversarial training process of the MCFGAN model, the discriminator is used to distinguish whether the image is real or fake. The discriminator is optimized using the loss function, and the formula is: (9.1); In the formula, and are the distributions of the high-quality image and the generated image respectively, is the expectation, and D is the discriminator; The generator is optimized through the combination of three losses. The loss function includes L1 loss, perceptual loss, and adversarial loss, and the formula is: (9.2); In the formula, and are the weight coefficients of the perceptual loss and the adversarial loss respectively, is the adversarial loss, is the perceptual loss, is the pixel-level supervision loss; The goal of the adversarial loss is to make the generator deceive the discriminator, and the optimization goal is opposite to that of : (9.3); In the formula, is the expectation, and are the distributions of the high-quality image and the generated image respectively.

[0062] In this embodiment, the generator generates high-quality images by optimizing the combination of the three losses, and minimizes these losses to optimize the image quality, making it closer to the real image in terms of pixels, perceptual characteristics and adversariality, and being able to "deceive" the discriminator. The goal of the discriminator is to maximize this loss, so as to correctly identify the difference between the generated image and the real image. Finally, the generator and the discriminator continuously play games in the adversarial training, promoting the improvement of the quality of the generated image.

[0063] After the above steps, an optimized MCFGAN model is obtained and used for testing.

[0064] Verify the above steps with actual images: Select 2 images from the thermal infrared dataset and use the optimized MCFGAN model for prediction. Then, compare with the existing super-resolution networks, which are: Bicubic (Bicubic Interpolation), BSRGAN (a generative adversarial network for image super-resolution), ESRGAN (Enhanced Super-Resolution Generative Adversarial Network), SwinIR (Swin Transformer for Image Restoration), PSRGAN (progressive super-resolution generative adversarial network), Real-ESARGAN (Real-World Enhanced Super-Resolution Generative Adversarial Network), SRDGRL (A blind super-resolution framework with degradation reconstruction loss), DARSR (Unsupervised Degradation Representation Learning for Blind Super-Resolution); the datasets are CVC-14 (Visible-FIR Day-Night Pedestrian Sequence Dataset), CVC-09, OSU TPD (OSU Thermal Pedestrian Database), OSU CTD (OSU Color-Thermal Database), and TWIRD (Terravic Motion IR Database).

[0065] The scale factor is 4 asFigure 7 , and the scaling factor is 2 as Figure 8 shown: 1) The optimized MCFGAN model can effectively remove noise while better restoring the detailed information of thermal infrared images, achieving the best effect. 2) Two metrics, NIQE (Natural Image Quality Evaluator) and BRISQUE (Blind / Referenceless Image Spatial Quality Evaluator), are used for objective analysis. As shown in Table 1 and Table 2, the technical solution of this embodiment obtains the lowest NIQE and BRISQUE values.

[0066] Table 1 NIQE / BRISQUE values of different super-resolution algorithms under five datasets with a scaling factor of 4

[0067] Table 2 NIQE / BRISQUE values of different super-resolution algorithms under five datasets with a scaling factor of 2

[0068] In use, in the process of super-resolution of this embodiment, the noise in thermal infrared images is effectively eliminated; the multi-channel fusion structure improves the depth of information mining and the utilization rate of information in attention, thereby improving the efficiency of information fusion; and by comparing with other super-resolution networks, it can be concluded that this embodiment can generate more realistic image textures and achieve better visual effects.

[0069] The above embodiments are only the preferred technical solutions of the present invention and should not be regarded as limitations on the present invention. The protection scope of the present invention should be the technical solutions recorded in the claims, including the equivalent replacement solutions of the technical features in the technical solutions recorded in the claims. That is, the equivalent replacement improvements within this scope are also within the protection scope of the present invention.

Claims

1. A thermal infrared image optimization method based on multi-channel fusion and semantic information, characterized in that: The following steps are involved: Obtain VIS and corresponding TIR datasets, perform data preprocessing, and obtain HR image datasets; Construct an MCFGAN model; wherein the MCFGAN model includes a generator module with a multi-channel fusion attention module and a semantic discriminator: the generator includes a multi-channel fusion attention module with three skip connections, namely, a DISR module; the semantic discriminator includes a semantic module and a semantic perception fusion module; wherein the semantic module is used to generate semantic information, and the semantic perception fusion module is used to integrate the semantic information into the discriminator; Input the HR image dataset into the infrared degradation model to obtain the LR image dataset; The generator is trained using the HR image dataset and the LR image dataset to generate the SR image dataset, and then the generated SR image dataset is sent to the semantic discriminator to judge whether it is true or false; In the semantic discriminator, the semantic module generates semantic information based on the HR image dataset, and integrates the semantic information into the discriminator through the semantic perception fusion module, so as to perform adversarial training on the MCFGAN model and obtain the optimized MCFGAN model; The TIR image to be tested is input into the optimized MCFGAN model to obtain the optimized thermal infrared image.

2. The thermal infrared image optimization method based on multi-channel fusion and semantic information according to claim 1, characterized in that: The data preprocessing includes graying the VIS and TIR data sets to generate HR image data sets, and then performing CLAHE processing; CLAHE processing: The core of CLAHE contrast limited adaptive histogram equalization processing lies in local histogram equalization and contrast limitation. The formula of local histogram equalization is: (1.1); In the formula, The gray value in the local area is The number of pixels, To indicate that the gray value is less than or equal to The cumulative sum of all pixel numbers, and is the pixel value of the image; The local histogram is equalized through linear transformation to make the pixel value distribution of the image as uniform as possible. The formula is: (1.2); In the formula, is the original image at position The pixel value at is the new image after local histogram equalization at position The pixel value at is the cumulative distribution function value of the minimum pixel value in the local histogram, is the cumulative distribution function value of the maximum pixel value in the local histogram, is the pixel value The cumulative distribution function value in the local histogram, is the maximum value of the new pixel value range; The contrast limit formula is: (1.3); In the formula, is the frequency value of the local histogram after contrast limitation processing, is the histogram of the local area, and ClipLimit is the contrast limitation parameter.

3. The thermal infrared image optimization method based on multi-channel fusion and semantic information according to claim 1, characterized in that: In the semantic discriminator, the semantic module Semantic Model is used to extract accurate and rich semantic features of the infrared image; then the semantic perception fusion module is used to combine the extracted semantic features with the discriminator; The semantic module uses Modified Resnet-50 and extracts the semantic features S of the third layer. h The input semantic perception fusion block performs semantic guidance on the discriminator, embeds the extracted image semantic information into the feature space of the discriminator, and provides key guidance for it to identify the authenticity of textures, thereby forcing the discriminator to pay attention to the distribution of semantic perception textures. Specifically: A1: Semantics After passing to the self-attention module, the query : (7.1); Where LN, GN, SA and RA are layer normalization, group normalization, self-attention module and rearrangement respectively; A2: Generate SR images from the generator and HR images in the HR dataset Passed to the convolutional layer, the original enhanced features are obtained. The original enhanced features are rearranged to obtain the keys of the SR image and the HR image respectively. and and value and , then the query Q, the SR image and the HR image key and and value and Input to the Cross Attention module; A3: Connect the semantically-aware image features output by Cross Attention with the original enhanced features to obtain the features of the final SR image and HR image and , the formula is: (7.2); (7.3); (7.4); (7.5); (7.6); (7.7); (7.8); (7.9); In the formula, is the output after the multi-head attention mechanism of the Cross Attention module. and is the output after BasicTransformerBlock, is the scale factor, is a feed-forward network, It is the original enhancement feature.

4. The thermal infrared image optimization method based on multi-channel fusion and semantic information according to claim 1, characterized in that: The generator module includes a shallow feature extraction module, a multi-channel fusion attention module for deep feature extraction, and a reconstruction module, and the formula is: The shallow feature extraction module uses a convolution layer with a 3×3 convolution kernel, and the formula is: , is the extracted feature map; The multi-channel fusion attention module, namely the MCF module, is used to extract deep features. The formula is: (6.1); In the formula, The MCF block consists of three DISR blocks with skip connections. is the output of the nth DISR block; The DISR block contains a dense connection module branch, a context attention module branch, and a statistical spatial channel attention module branch. The dynamic weights of these three branches are determined by the input features. Then, a 3×3 convolution kernel is used to Extract features from: , complete the deep feature extraction part; The equation for image reconstruction is: (6.2); In the formula, For interpolation upsampling, is the final SR output.

5. The thermal infrared image optimization method based on multi-channel fusion and semantic information according to claim 1, characterized in that: Each DISR module generates weights for three independent channels by using the same input features. The specific operation formula is: (5.1); (5.2); (5.3); (5.4); (5.5); (5.6); In the formula, is the output of dense connection, is the output of the contextual attention module, is the output of the statistical spatial channel attention module; The weights of the three channels are generated by the dynamic weight module training according to the input features.

6. The thermal infrared image optimization method based on multi-channel fusion and semantic information according to claim 5 is characterized in that: Each DISR module controls the balance through weighted summation, which includes the dynamic weighted contributions of the dense connection module, the contextual attention module, and the statistical spatial channel attention module; The dense connection module Dense is used in the DISR module, and the formula is: (2.1); (2.2); (2.2); (2.4); (2.5); (2.6); (2.7); (2.8); (2.9); (2.10); In the formula, is the convolutional layer, is the activation function, It is the feature obtained by convolution layer by layer and concatenating it with the previous layer. Y is the final output. is the weighting coefficient.

7. The thermal infrared image optimization method based on multi-channel fusion and semantic information according to claim 5, characterized in that: The context attention module is adopted in the DISR module, and the formula is as follows: (3.1); (3.2); (3.3); (3.4); (3.5); (3.6); (3.7); In the formula, is a global feature, is the compression feature, is the context information in the horizontal direction, is the context information in the vertical direction, is the length of the horizontal convolution kernel, is the height of the vertical convolution kernel, To remap the features to the original channel dimension, is the attention weight, For the final output, is element-wise multiplication.

8. The thermal infrared image optimization method based on multi-channel fusion and semantic information according to claim 5, characterized in that: In the DISR module, the statistical spatial channel attention module has the following specific formula: (4.1); (4.2); (4.3); (4.4); (4.5); (4.6); In the formula, is the mean value within the channel, is the square of the deviation from the mean, The stable normalization factor after adding the regularization parameter to the normalization factor, is the normalization factor, To prevent the regularization parameter from having a too small denominator when calculating the normalization factor, is the attention weight.

9. The thermal infrared image optimization method based on multi-channel fusion and semantic information according to claim 5, characterized in that: In the DISR module, the input features are first compressed using global average pooling, then pass through a connection layer consisting of two fully connected and one ReLU activation layer, and finally the weight of each branch is calculated by the Softmax function.

10. The thermal infrared image optimization method based on multi-channel fusion and semantic information according to claim 1, characterized in that: The HR image dataset and the LR image dataset are used to train the generator. In the adversarial training process of the MCFGAN model, a discriminator is used to distinguish whether the image is real or fake. The discriminator is optimized using a loss function, and the formula is: (9.1); In the formula, and High-quality images and generate images The distribution of is the expectation, D is the discriminator; The generator is optimized through a combination of three losses. The loss function includes L1 loss, perceptual loss and adversarial loss. The formula is: (9.2); In the formula, and are the weight coefficients of perceptual loss and adversarial loss, respectively. To combat losses, is the perceived loss, is the pixel-level supervision loss; Fighting Losses The goal is to make the generator deceive the discriminator and optimize the objective The opposite purpose: (9.3); In the formula, For expectations, and High-quality images and generate images distribution.

Citation Information

Patent Citations

  • Facial expression recognition method based on multi-channel fusion and lightweight neural network

    CN113989890A

  • Image reconstruction method and device, equipment and storage medium

    CN115526773A

  • Multi-clue-guided unmanned aerial vehicle thermal infrared image super-resolution reconstruction method and system

    CN116523753A

  • Remote sensing image super-resolution reconstruction method based on MT-SRGAN

    CN117196950A

  • Thermal infrared image super-resolution algorithm of generative adversarial network based on multi-structure fusion

    CN117372254A

Cited By

  • Image super-resolution reconstruction method based on semantic joint perception

    CN120318078A

  • A semantic-aware image super-resolution reconstruction method

    CN120318078B