Infrared image super-resolution reconstruction method based on visible light correlation feature fusion

Through the visible light correlation feature fusion method, the problem of insufficient calculation and fusion of inter-span modal feature correlation in infrared image super-resolution reconstruction is solved, and efficient detail recovery and clarity improvement of infrared images is achieved, meeting the application needs of fine analysis and recognition.

CN120339079AActive Publication Date: 2025-07-18JIANGXI NORMAL UNIV

Patent Information

Application Number
CN202510829165.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-18
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

The existing infrared image super-resolution reconstruction methods have shortcomings in cross-modal feature correlation calculation and fusion, resulting in blurred details of infrared images and poor texture information, which makes it difficult to meet the application needs of fine analysis and recognition.

Method used

Using a method based on visible light correlation feature fusion, the visible light correlation feature modulation, cross-modal feature fusion and Transformer image reconstruction modules are used to accurately calculate and robustly migrate visible light texture information, and integrate multi-scale channel attention blocks in the Transformer image reconstruction module to achieve efficient depth fusion and detail recovery.

Benefits of technology

It significantly improves the detail richness and clarity of infrared images, improves the reconstruction quality, and enhances the realism and visual effects of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339079A_ABST
    Figure CN120339079A_ABST
Patent Text Reader

Abstract

The invention discloses an infrared image super-resolution reconstruction method based on visible light correlation feature fusion, and belongs to the technical field of image processing. The method comprises the following steps: inputting an up-sampling low-resolution infrared image into a learnable texture extraction module, and extracting query features from the learnable texture extraction module; the visible light image and the high-resolution visible light image which are subjected to down-sampling and up-sampling processing are input into a texture feature coding module, and key features and value features are extracted from the visible light image and the high-resolution visible light image; generating a correlation guide graph and a corresponding feature weighted graph according to the query features and the key features; obtaining a migration feature graph according to the corresponding feature weighted graph and the value feature; and inputting the shallow layer features, the correlation guide map and the migration feature map into a cross-modal feature fusion module to obtain fusion features, and inputting the fusion features into a Transform image reconstruction module to generate a super-resolution infrared image. According to the method, the guiding effect of visible light on infrared detail reconstruction is remarkably improved, so that the reconstructed image is clearer and sharper in texture details.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly relates to a method for infrared image super-resolution reconstruction based on visible light correlation feature fusion. Background Art

[0002] Infrared images have unique application advantages in night, bad weather or occlusion environments because they can obtain the thermal radiation information of objects, and are widely used in fields such as security monitoring, intelligent driving, industrial inspection, and medical diagnosis. However, limited by the physical characteristics and costs of infrared detectors, it is usually challenging to obtain high-resolution infrared images. Existing low-resolution infrared images generally have problems such as blurred details and poor texture information, which seriously restricts their performance in application scenarios that require fine analysis and recognition. Therefore, researching efficient infrared image super-resolution (SR) reconstruction technology has important research significance and practical application value.

[0003] In recent years, significant progress has been made in single-image super-resolution technology based on deep learning, showing powerful detail recovery capabilities in visible light image super-resolution reconstruction tasks. However, when applying these methods directly to infrared images, the performance is often unsatisfactory, mainly because the inherently low contrast, weak texture and high noise characteristics of infrared images make it difficult to recover rich high-frequency details only from their own information.

[0004] To overcome the limitation of insufficient information in infrared images themselves, using other modality images, especially visible light images with higher resolution, as guiding information has become a promising research direction for infrared image super-resolution. Visible light images usually contain rich texture, edges, and structural details, which have potential guiding value for the super-resolution reconstruction of infrared images. However, existing visible light-guided infrared image super-resolution methods still face many challenges: First, the challenges of calculating the correlation of visible light features and effective texture migration: How to accurately calculate the cross-modal correlation between low-resolution infrared features and high-resolution visible light features, and based on this, accurately and robustly transfer the beneficial texture in visible light to the infrared domain is a key problem. Existing methods may ignore the feature amplitude or overly rely on local single source points when calculating the correlation, resulting in inaccurate and inflexible transferred texture and limited guiding quality. Second, the refinement and enhancement of the fused features need to be optimized: After fusing the visible light-guided features and the infrared's own features, how to design an efficient network module to refine and enhance this hybrid feature, especially using effective activation and attention mechanisms to highlight and restore the key details in the infrared image, is the key to improving the final reconstruction quality. Traditional feature processing modules may not be fully adapted to the complex feature characteristics after this cross-modal fusion. Third, the lack of efficient deep fusion of multi-modal features: How to effectively deeply fuse the features from high-resolution visible light guidance and the basic features extracted from the original low-resolution infrared image to maximize the use of their complementary information while avoiding redundancy and conflict is a key factor affecting the reconstruction performance. Simple fusion strategies are difficult to handle complex cross-modal feature interactions, limiting the detail richness and clarity of the final reconstructed image. Summary of the Invention

[0005] Aiming at the problem of blurred details caused by the low resolution of infrared images and the deficiencies of existing visible light-guided methods in accurate matching and robust transfer of cross-modal information (for example, traditional correlation ignores feature amplitude, and hard attention relies on a single source point) and effective fusion, the present invention proposes a super-resolution reconstruction method for infrared images based on visible light correlation feature fusion.

[0006] The present invention is realized by the following technical solutions: A super-resolution reconstruction method for infrared images based on visible light correlation feature fusion, the steps are as follows: S1: Input the low-resolution infrared image into the shallow feature extraction module to extract its shallow feature F; input the upsampled low-resolution infrared image into the learnable texture extraction module (LTE) to extract the query feature Q; input the visible light image that has been downsampled and then upsampled and the high-resolution visible light image , it is input into the Texture Feature Encoding module (TFE) to extract the key feature K and the value feature V therefrom; the query feature Q and the key feature K are input into the correlation calculation module, and the correlation calculation module calculates the similarity between the two and generates a correlation guidance map S and a corresponding feature weighting map C for weighting; the corresponding feature weighting map C and the value feature V are input into the Corresponding Feature Weighting Module (CFWM), and the corresponding feature weighting map C and the value feature V are weighted to obtain a migration feature map M; S2: Input the shallow feature F, the correlation guidance map S, and the migration feature map M into the cross-modal feature fusion module to fuse and obtain a fused feature; S3: Input the fused feature into the Transformer image reconstruction module, and use the Transformer image reconstruction module to further process, upsample, and restore details of the fused feature, and finally generate a super-resolution infrared image .

[0007] Further preferably, the specific process of calculating the corresponding feature weighting map C is as follows: ; where, represents the calculated weight and is also an element in the corresponding feature weighting map C, k is the spatial position index in the high-resolution visible light image, is the similarity between the i-th position in the low-resolution infrared image and the j-th position in the high-resolution visible light image, is the similarity between the feature at the i-th position in the low-resolution infrared image and the k-th position in the high-resolution visible light image.

[0008] Further preferably, the correlation guidance map S is calculated according to the following formula: ; where, represents the value at the i-th position in the correlation guidance map S.

[0009] More preferably, the cross-modal feature fusion module (CFFM) first performs channel concatenation on the shallow feature F and the transfer feature map M; then, the element-wise multiplication operation is performed on the concatenated features using the correlation guidance map S to achieve intelligent modulation and guidance of the fusion process. The modulated features then pass through a layer normalization block; the layer-normalized features are fed into two parallel branches: one branch contains a 3×3 convolution and a ReLU activation function for capturing local spatial features; the other branch contains a 1×1 convolution and a ReLU activation function for channel dimension interaction and feature projection; the outputs of the two branches are then concatenated again in channels and finally integrated through a 1×1 convolution; finally, the integrated features are added element-wise to the shallow feature F of the original input to form a local residual connection to obtain the final fusion features.

[0010] More preferably, the Transformer image reconstruction module consists of n cascaded residual multi-scale attention groups, 3 3×3 convolutions, a sub-pixel convolution layer, and a global residual connection.

[0011] More preferably, the residual multi-scale attention group consists of m cascaded residual multi-scale attention blocks, an overlapping cross-fusion attention block (OCFAB), a 3×3 convolution, and a local residual in sequence.

[0012] More preferably, each residual multi-scale attention block is a functional unit. The residual multi-scale attention block first performs the first layer normalization on the original input features, and then feeds the first-normalized features into the multi-scale channel attention block and the shifted window multi-head self-attention block in parallel; after summing the outputs of the multi-scale channel attention block and the shifted window multi-head self-attention block, a first residual connection is made with the original input features to obtain intermediate features; then, the intermediate features pass through the second layer normalization and a multi-layer perceptron (MLP) in sequence; finally, a second residual connection is made between the output of the multi-layer perceptron and the intermediate features to obtain the output features of the residual multi-scale attention block.

[0013] More preferably, the processing flow of the multi-scale channel attention block is as follows: the input features first pass through a 3×3 convolution for feature extraction and then are fed into the multi-scale large kernel attention module for processing. The output of the multi-scale large kernel attention module is divided into two paths. The first path passes through the GELU activation function and the multi-layer perceptron for processing, and the second path passes through the GELU activation function and the max pooling operation. The results of the two paths are added element-wise, then passed through the Sigmoid activation function, and then the output of the Sigmoid activation function is multiplied element-wise with the output features of the multi-scale large kernel attention module. The multiplication result then passes through a 3×3 convolution and is then fed into the channel attention block for processing.

[0014] The present invention also provides an electronic device, including a memory and a processor. Computer-readable instructions are stored in the memory. When the instructions are executed by the processor, the processor is caused to implement the above infrared image super-resolution reconstruction method when executed.

[0015] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above infrared image super-resolution reconstruction method is implemented.

[0016] The beneficial effects of the present invention are as follows: (1) The visible light correlation feature modulation module (VCFM) designed by the present invention can efficiently transfer high-resolution visible light textures, effectively cope with cross-modal parallax, and significantly improve the detail richness and realism of the reconstructed infrared images. Through the improved scaled dot product correlation calculation and the weighted average mechanism based on Softmax, more accurate and robust transfer of high-resolution visible light texture information is achieved, effectively overcoming the problems of ignoring amplitude information and relying on a single source point in traditional methods.

[0017] (2) The cross-modal feature fusion module (CFFM) designed by the present invention cleverly uses the correlation guidance map S generated by the correlation calculation module in the visible light correlation feature modulation module to intelligently modulate and fuse the infrared self-features and the visible light transferred textures, maximizing the utilization of complementary information; and integrating multi-scale channel attention blocks (MS-CAB) in the Transformer reconstruction module to enhance the feature expression after fusion. The deep and effective integration of visible light-guided features and infrared self-features is realized, greatly enhancing the representation ability and robustness of the fused features, and laying a solid foundation for high-quality reconstruction.

[0018] (3) The multi-scale channel attention block (MS-CAB) introduced by the present invention enhances the model's perception and utilization ability of key feature channels at different scales, and significantly improves the clarity and sharpness of the reconstructed infrared images in terms of edges and textures. Description of the Drawings

[0019] Figure 1 It is a structural diagram of the visible light-guided super-resolution infrared image reconstruction network of the present invention.

[0020] Figure 2 It is a structural diagram of the cross-modal feature fusion module of the present invention.

[0021] Figure 3 It is a structural diagram of the Transformer image reconstruction module of the present invention.

[0022] Figure 4 It is a structural diagram of the residual multi-scale attention group of the present invention.

[0023] Figure 5 This is the structural diagram of the residual multi-scale attention block of the present invention.

[0024] Figure 6 This is the structural diagram of the multi-scale channel attention block of the present invention. Specific Embodiments

[0025] The present invention will be further elaborated in detail below in conjunction with embodiments and drawings.

[0026] The first embodiment of the present invention introduces an infrared image super-resolution reconstruction method based on visible light correlation feature fusion. This method relies on a visible light-guided super-resolution infrared image reconstruction network as shown in Figure 1 . The core of this network includes: a shallow feature extraction module, a visible light correlation feature modulation module, a cross-modal feature fusion module, and a Transformer image reconstruction module. Among them, the shallow feature extraction module is used to extract shallow feature F from the low-resolution infrared image . The visible light correlation feature modulation module receives multiple inputs and uses the corresponding feature weighting module and correlation guidance module to generate cross-modal guidance features for guiding reconstruction. The cross-modal feature fusion module is used to deeply fuse the shallow feature F and the cross-modal guidance features to obtain fused features. The Transformer image reconstruction module further processes the fused features and completes upsampling, and finally generates a super-resolution infrared image .

[0027] The specific steps are as follows: S1: First, input the low-resolution infrared image into the shallow feature extraction module to extract its shallow feature F. At the same time, execute the visible light correlation feature modulation module (VCFM), input the upsampled low-resolution infrared image into the learnable texture extraction module (LTE) to extract the query feature Q from it; input the visible light image that has been downsampled and then upsampled, as well as the high-resolution visible light image , into the texture feature encoding module (TFE) to extract the key feature K and the value feature V from it. Subsequently, input the query feature Q and the key feature K into the correlation calculation module. The correlation calculation module calculates the similarity between the two and generates a correlation guidance map S and a corresponding feature weighting map C for weighting. Then, input the corresponding feature weighting map C and the value feature V into the corresponding feature weighting module (CFWM), and use the corresponding feature weighting map C and the value feature V for weighting to obtain a transfer feature map M.

[0028] S2: Input the shallow features F extracted from the low-resolution infrared image, the obtained correlation guidance map S, and the migration feature map M into the cross-modal feature fusion module CFFM for deep fusion to obtain the fusion features.

[0029] S3: Input the fusion features into the Transformer image reconstruction module, and use the Transformer image reconstruction module to further process, upsample, and restore details of the fusion features, and finally generate the super-resolution infrared image. 。

[0030] The following will further elaborate on each step or module in detail: Refer to Figure 1 First, perform shallow feature extraction on the infrared low-resolution image The specific process is as follows: (1); Among them, represents the shallow feature extraction module, represents the low-resolution infrared image, represents the shallow features.

[0031] At the same time, perform visible light feature extraction operations. The visible light correlation feature modulation module proposed in this application receives the low-resolution infrared image ( ), its upsampled version ( ), the high-resolution visible light image ( ), and the visible light image ( ) after domain consistency processing) as inputs to achieve accurate extraction and fusion of visible light feature information.

[0032] The processing flow of the visible light correlation feature modulation module is as follows: First, use the learnable texture extraction module (LTE) to process the input , and use the texture feature encoding module (TFE) to process the input , . The extracted texture features (query feature Q, key feature K, value feature V) are used as the three basic elements for cross-modal guidance. This process is described as: (2); (3); (4); Among them, represents the learnable texture extraction operation, Represents the operation of encoding and extracting texture features. From the extracted texture features, three elements Q, K, and V required for correlation calculation and feature weighting are obtained.

[0033] Subsequently, correlation calculation is performed. By calculating the similarity between the query feature Q and the key feature K, the similarity between the low-resolution infrared image and the high-resolution visible light image is calculated, effectively transferring the texture information related to the low-resolution infrared image in the high-resolution visible light image, while suppressing the transmission of those irrelevant textures. This process can be described as: (5); Among them, represents the similarity between the i-th position in the low-resolution infrared image and the j-th position in the high-resolution visible light image, represents the operation of vector inner product, and are the vectors after expanding the query feature Q and the key feature K respectively, is the feature vector 's dimension (or 's dimension), and the generated will be used for the subsequent calculation of the correlation guidance map S and the migration feature map M.

[0034] Then it enters the corresponding feature weighting module for processing. Before that, a corresponding feature weighting map C needs to be calculated. The specific process is as follows: (6); Among them, represents the calculated weight, which is also an element in the corresponding feature weighting map C. These weights can be regarded as correlation metrics, which represent the degree of texture contribution of all positions in the high-resolution visible light image to the i-th position in the low-resolution infrared image. k is the index of all possible spatial positions in the high-resolution visible light image, is the similarity between the feature at the i-th position in the low-resolution infrared image and the k-th position in the high-resolution visible light image.

[0035] Send the corresponding feature weighting map C and the value feature V to the corresponding feature weighting module, and this module transfers the value feature V. In the weighted texture migration module, for each position in the low-resolution image, according to its correlation with all positions in the high-resolution visible light image, the value feature V is weighted and summed to obtain the migration feature map M.

[0036] Finally, according to these weights, the weighted sum of the high-resolution features V of the visible light image is performed: (7); Among them, represents the texture feature finally migrated to the $i$-th position in the low-resolution infrared image, that is, the value at the $i$-th position in the migration feature map $M$. is the feature vector corresponding to the $j$-th position of the high-resolution visible light image in the value feature $V$.

[0037] At the same time, a correlation guidance map $S$ needs to be calculated. The specific process is as follows: (8); Among them, represents the value at the $i$-th position in the correlation guidance map.

[0038] Then, the shallow feature $F$, the correlation guidance map $S$, and the migration feature map $M$ are simultaneously input into the cross-modal feature fusion module (CFFM). Referring to Figure 2 , this module first performs channel concatenation on the shallow feature $F$ and the migration feature map $M$. Then, it uses the correlation guidance map $S$ to perform an element-wise multiplication operation on the concatenated features to achieve intelligent modulation and guidance of the fusion process. The modulated features are then passed through a layer normalization block. The layer-normalized features are fed into two parallel branches: one branch contains a $3\times3$ convolution and a ReLU activation function for capturing local spatial features; the other branch contains a $1\times1$ convolution and a ReLU activation function for channel dimension interaction and feature projection. The outputs of the two branches are then concatenated again in the channel dimension and passed through a $1\times1$ convolution for final feature integration. Finally, the integrated features are added element-wise to the original input shallow feature $F$ to form a local residual connection, obtaining the final fusion feature. Through the above series of operations, the cross-modal feature fusion module (CFFM) can deeply and efficiently fuse the infrared shallow features and the visible light migration texture features, and use the correlation guidance map $S$ to perform intelligent modulation on the fusion process, thereby generating fusion features containing rich cross-modal complementary information and texture details, providing a more powerful representation for subsequent image reconstruction.

[0039] This process is described as: (9); (10); Among them, represents the intermediate feature after fusion, represents the fusion feature output by the cross-modal feature fusion module, represents the channel concatenation operation, and the operator represents the element-wise multiplication between features, represents the $3\times3$ convolution, represents the $1\times1$ convolution, denotes the ReLU activation function, represents the layer normalization operation, represents the shallow features.

[0040] Finally, it enters the image reconstruction stage, and the fused features are input into the subsequent Transformer image reconstruction module. Refer to Figure 3 , the Transformer image reconstruction module is composed of n cascaded residual multi-scale attention groups, 3 3×3 convolutions, a sub-pixel convolution layer, and a global residual connection. The residual multi-scale attention group can perform multi-scale extraction on the fused features, make greater use of the features for image reconstruction, and thus generate clearer images. In this embodiment, n = 10 is taken, and the process is specifically described as: (11); (12); (13); Among them, represents the sub-pixel convolution layer, represents the upsampled features, represents the reconstructed super-resolution infrared image, represents the i-th residual multi-scale attention group of the residual multi-scale attention group, and n represents the number of residual multi-scale attention groups ( ), represents the features output by the (i - 1)-th residual multi-scale attention group, represents the features output by the i-th residual multi-scale attention group, represents the features output by the n-th residual multi-scale attention group.

[0041] As Figure 4 shown, the residual multi-scale attention group is successively composed of m cascaded residual multi-scale attention blocks, an overlapping cross-fusion attention block (OCFAB), a 3×3 convolution, and a local residual. In this embodiment, m = 20 is taken, and the processing process of the residual multi-scale attention group is described as: (14); (15); Among them, represents the j-th residual multi-scale attention block of the i-th residual multi-scale attention group, represents the features output by the (j - 1)-th residual multi-scale attention block of the i-th residual multi-scale attention group, represents the features output by the j-th residual multi-scale attention block of the i-th residual multi-scale attention group, Denote the features output by the m-th residual multi-scale attention block in the i-th residual multi-scale attention group, Denote the input features of the i-th residual multi-scale attention group, Denote the output features of the i-th residual multi-scale attention group, Denote the multi-scale extraction operation of the overlapping cross-fusion attention block.

[0042] As Figure 5 shown, each residual multi-scale attention block is a functional unit that first performs the first layer normalization on the original input features, and then parallelly sends the features after the first normalization into the multi-scale channel attention block and the shifted window multi-head self-attention block ; after summing the outputs of the multi-scale channel attention block and the shifted window multi-head self-attention block, perform the first residual connection with the original input features to obtain intermediate features; then, the intermediate features sequentially pass through the second layer normalization and the multi-layer perceptron (MLP); finally, perform the second residual connection between the output of the multi-layer perceptron and the intermediate features to obtain the output features of the residual multi-scale attention block, realizing the flexible activation of features, the adjustment of channel importance, and the capture of spatial dependence relationships, and enhancing the feature expression ability and training stability through two-level residual connections.

[0043] Generally speaking, for the given original input features , the entire calculation process of the residual multi-scale attention block is as follows: (16); (17); (18); Among them, denote the features after the first layer normalization, denote the intermediate features, denote the output of the residual multi-scale attention block, denote the layer normalization operation, denote the multi-layer perceptron. Denote the shifted window multi-head self-attention block, Denote the multi-scale channel attention block.

[0044] Referring to Figure 6 , the multi-scale channel attention block The processing flow is as follows: The input features first undergo feature extraction through a 3×3 convolution, and then are sent to the multi-scale large kernel attention module for processing. The output of the multi-scale large kernel attention module is divided into two paths. The first path is processed through the GELU activation function and a multi-layer perceptron, and the second path is processed through the GELU activation function and a max pooling operation. The results of the two paths are added element-wise, then passed through the Sigmoid activation function, and the output of the Sigmoid activation function is multiplied element-wise with the output features of the multi-scale large kernel attention module. The multiplication result then undergoes a 3×3 convolution and is sent to the channel attention block for processing. In the channel attention block, the features first compress the spatial dimension through a global pooling layer to obtain global context information; then, they sequentially pass through a 3×3 convolution, the ReLU activation function, and a 3×3 convolution; and then through the Sigmoid activation function to generate the weights for each channel. Finally, these channel weights are multiplied element-wise with the input features of the channel attention block (i.e., the convolution result sent to the channel attention block) to obtain the channel-weighted features, which are used as the output features of the multi-scale channel attention block. Its processing process can be described as: (19); (20); (21); Among them, represents the multi-scale large kernel attention module, represents the channel attention operation, represents the GELU activation function, represents the Sigmoid activation function, represents the 3×3 convolution, represents the element-wise multiplication operation between features, represents the multi-layer perceptron, represents the max pooling operation, represents the input features of the multi-scale channel attention block, represents the output features after passing through the multi-scale large kernel attention module, represents the multi-scale large kernel attention features after passing through the Sigmoid, represents the output features of the multi-scale channel attention block.

[0045] To comprehensively verify the effectiveness of the infrared image super-resolution reconstruction method based on visible light correlation feature fusion proposed in the present invention, a series of comparative experiments were carried out on the publicly available FLIR-aligned infrared dataset. This dataset contains rich urban scene infrared images. 500 images were selected as the test set, and all experiments were carried out under the bicubic interpolation degradation model.

[0046] In the experiment, the method of the present invention was compared with a variety of mainstream and advanced image super-resolution algorithms, including: EDSR (Enhanced Deep Residual Network), RDN (Residual Dense Network), RCAN (Residual Channel Attention Network), SwinIR (Image Restoration Network Based on Swin Transformer), HAT (Hybrid Attention Transformer), ATD (Advanced Super-Resolution Transformer with Adaptive Token Dictionary), UGSR (Unaligned Guided Thermal Infrared Super-Resolution), and PAGSR (Multi-Scale Edge Attention Thermal Infrared Super-Resolution). Among them, VCFM is the visible light correlation feature modulation module designed by the present invention, CFFM is the cross-modal feature fusion module, and MS-CAB is the multi-scale channel attention block introduced by the present invention.

[0047] The peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) were selected as the core evaluation indicators to evaluate the quality of the super-resolution results. PSNR mainly measures the difference between the reconstructed image and the original high-resolution image at the pixel level, and the higher the value, the smaller the distortion; SSIM evaluates the similarity between images from three aspects: brightness, contrast, and structure, and the closer the value is to 1, the higher the similarity and the better the reconstruction quality. These two indicators were calculated on the Y channel of the YCbCr color space and the average value on the entire test set was reported. At the same time, the number of model parameters (Params) and running time (Runtime) were considered to evaluate the complexity and running efficiency of the model. The number of parameters (in millions) reflects the size of the model, and the running time refers to the average time required for the model to process one picture (in seconds). By comprehensively analyzing these indicators, the performance, efficiency, and practicality of a super-resolution algorithm can be more comprehensively measured. The specific experiment is shown in Table 1. The present invention achieved the best results in terms of PSNR and SSIM index values. Although the number of parameters and running time increased, it was exchanged for better improvement of index values and visual effects.

[0048] Table 1 Comparison Experiment Results

[0049] The second embodiment of the present invention provides an electronic device, including a memory and a processor. The memory stores computer-readable instructions, and when the instructions are executed by the processor, the processor is caused to implement the above infrared image super-resolution reconstruction method.

[0050] The third embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above infrared image super-resolution reconstruction method is implemented.

[0051] The above only expresses the preferred embodiments of the present invention and does not limit the present invention in other forms. Any person skilled in the relevant art may use the disclosed content above to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. An infrared image super-resolution reconstruction method based on visible light correlation feature fusion, characterized in that the steps are as follows As follows: S1: Input the low-resolution infrared image into the shallow feature extraction module to extract its shallow feature F; Input the upsampled low-resolution infrared image into the learnable texture extraction module to extract the query feature Q; Input the visible light image that has been downsampled and then upsampled, as well as the high-resolution visible light image into the texture feature encoding module to extract the key feature K and the value feature V; Input the query feature Q and the key feature K into the correlation calculation module. The correlation calculation module calculates the similarity between the two and generates a correlation guidance map S and a corresponding feature weighting map C for weighting; Input the corresponding feature weighting map C and the value feature V into the corresponding feature weighting module, and perform weighting using the corresponding feature weighting map C and the value feature V to obtain a migration feature map M; S2: Input the shallow feature F, the correlation guidance map S, and the migration feature map M into the cross-modal feature fusion module CFFM to fuse and obtain a fused feature; S3: Input the fused features into the Transformer image reconstruction module, and use the Transformer image reconstruction module to further process, upsample, and restore details of the fused features, and finally generate a super-resolution infrared image .

2. The infrared image super-resolution reconstruction method according to claim 1, wherein The specific process of calculating the corresponding feature weighting map C is as follows: ; Among them, represents the calculated weight, which is also an element in the corresponding feature weighted graph C. k is the spatial position index in the high-resolution visible light image. is the similarity between the i-th position in the low-resolution infrared image and the j-th position in the high-resolution visible light image. is the similarity between the feature of the i-th position in the low-resolution infrared image and the k-th position in the high-resolution visible light image.

3. The infrared image super-resolution reconstruction method according to claim 2, wherein The correlation guidance map S is calculated according to the following formula: ; wherein, represents the value at the i-th position in the correlation guidance diagram S.

4. The infrared image super-resolution reconstruction method according to claim 1, characterized in that, The cross-modal feature fusion module first performs channel splicing on the shallow feature F and the migration feature map M; Then, use the correlation guidance map S to perform an element-wise multiplication operation on the spliced features to achieve intelligent modulation and guidance of the fusion process. The modulated features are then passed through a layer normalization block; The layer-normalized features are fed into two parallel branches: one branch contains a 3×3 convolution and a ReLU activation function for capturing local spatial features; The other branch contains a 1×1 convolution and a ReLU activation function for channel dimension interaction and feature projection; The outputs of the two branches are then spliced again in the channel dimension and finally integrated through a 1×1 convolution; Finally, add the integrated features to the original input shallow feature F element-wise to form a local residual connection to obtain the final fused feature.

5. The infrared image super-resolution reconstruction method according to claim 1, characterized in that The Transformer image reconstruction module is composed of n cascaded residual multi-scale attention groups, 3 3×3 convolutions, a sub-pixel convolution layer, and a global residual connection.

6. The infrared image super-resolution reconstruction method according to claim 5, characterized in that, The residual multi-scale attention group is sequentially composed of m cascaded residual multi-scale attention blocks, an overlapping cross-fusion attention block, a 3×3 convolution, and a local residual.

7. The infrared image super-resolution reconstruction method according to claim 6, characterized in that, Each residual multi-scale attention block is a functional unit. The residual multi-scale attention block first performs the first layer normalization on the original input features, and then parallelly feeds the features after the first normalization into the multi-scale channel attention block and the shifted window multi-head self-attention block; After summing the outputs of the multi-scale channel attention block and the shifted window multi-head self-attention block, perform the first residual connection with the original input features to obtain intermediate features; Then, the intermediate features pass through the second layer normalization and a multi-layer perceptron in sequence; Finally, perform the second residual connection between the output of the multi-layer perceptron and the intermediate features to obtain the output features of the residual multi-scale attention block.

8. The infrared image super-resolution reconstruction method according to claim 7, wherein The processing flow of the multi-scale channel attention block is as follows: The input features first undergo a 3×3 convolution for feature extraction, and then are fed into the multi-scale large kernel attention module for processing; The output of the multi-scale large kernel attention module is split into two paths. The first path is processed by the GELU activation function and a multi-layer perceptron, and the second path is processed by the GELU activation function and a max pooling operation. The results of the two paths are added element-wise, then passed through the Sigmoid activation function. Next, the output of the Sigmoid activation function is multiplied element-wise with the output features of the multi-scale large kernel attention module. The multiplication result is then passed through a 3×3 convolution and then fed into a channel attention block for processing.

9. An electronic device, comprising a memory and a processor, wherein computer-readable instructions are stored in the memory, characterized in that, When the instruction is executed by the processor, it causes the processor to implement the infrared image super-resolution reconstruction method according to any one of claims 1-8 when executed.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the infrared image super-resolution reconstruction method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Fire behavior detection method and device based on deep learning, equipment and medium

    CN115761409A

  • Infrared image super-resolution reconstruction method based on complementary reference

    CN118096534A

  • CT image segmentation method and system based on improved Swinin-Unet

    CN119151963A

Cited By

  • Infrared image super-resolution method and device based on state space model and medium

    CN120746839A

  • Infrared and visible light image fusion method based on second-order attention mixed features

    CN120976034A

  • Infrared and visible image fusion method based on second-order attention mixed features

    CN120976034B

  • Remote sensing image super-division method based on texture migration guidance double diffusion model

    CN121458542A