Feature integration method for interactive convolution and dynamic focusing of infrared image
By building a two-layer feature extraction module and a two-way information interaction mechanism, the problem of insufficient fusion of multi-scale features in infrared image super-resolution technology is solved, efficient detail reconstruction and resolution improvement of infrared images are achieved, and target recognition and imaging effects are improved.
Patent Information
- Application Number
- CN202510816805.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-18
AI Technical Summary
There are problems in existing infrared image super-resolution technology that lack multi-scale feature fusion, poor dynamic adaptability and weak detail recovery capabilities, making it difficult to effectively improve image resolution and detail quality in complex backgrounds.
The feature integration method of interactive convolution and dynamic focus is adopted. By building a two-layer feature extraction module, combining the two-way information interaction mechanism and multi-task collaborative loss optimization, efficient fusion and detailed reconstruction of local and global features are achieved, including shallow feature extraction, deep feature fusion blocks and PixelShuffle operations.
It significantly improves the resolution and detail quality of infrared images, enhances the target recognition and imaging capabilities in complex environments, and provides a more efficient image reconstruction solution.
Smart Images

Figure CN120339779A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and more particularly to a feature integration method for interactive convolution and dynamic focusing of infrared images. Background Art
[0002] Infrared imaging technology has wide application values in the fields of military reconnaissance, night vision security, medical diagnosis, industrial inspection, etc. However, due to the inherent resolution limitations of infrared sensors, signal noise interference, and uneven thermal radiation in complex environments, the acquired infrared images often face problems such as low resolution, blurred details, blurred edges, and insufficient contrast, seriously affecting subsequent target recognition and information extraction. Traditional infrared image enhancement methods mainly rely on interpolation algorithms (such as bilinear interpolation, bicubic interpolation) or filtering processes (such as non-uniformity correction), but such methods can only smooth image textures and are difficult to restore the missing high-frequency detail information, resulting in limited super-resolution reconstruction effects.
[0003] In recent years, significant progress has been made in super-resolution technology based on deep learning. The super-resolution convolutional neural network (SRCNN) extracts features through multiple layers of convolution and reconstructs high-resolution images, but the receptive field of the fixed convolution kernel is difficult to adapt to the diverse feature distributions of infrared images. In addition, the introduction of the generative adversarial network (GAN) has improved the image perception quality, but its training instability and mode collapse phenomena limit the robustness. The introduction of dynamic filtering modules and attention mechanisms can enhance the extraction of key features, but most methods adopt a single manifold optimization strategy and lack the interactive fusion of multi-scale information, resulting in insufficient fine-grained texture reconstruction ability.
[0004] In the prior art, the dynamic focusing strategy dynamically adapts the feature weights through adjustable attention parameters, but most designs only focus on local feature interactions and ignore the collaborative modeling of cross-scale features. In addition, although the architecture based on residual connections can alleviate the problem of gradient disappearance, there is still room for improvement in its feature transfer efficiency in the complex background of infrared images. Therefore, there is an urgent need for a method that integrates dynamic feature interaction and focused attention mechanism guided by multi-directional information to improve the detail reconstruction ability while suppressing noise and provide a more efficient solution for infrared images. Summary of the Invention
[0005] The object of the present invention is to provide a feature integration method for interactive convolution and dynamic focusing of infrared images, which can effectively solve the problems of insufficient multi-scale feature fusion, poor dynamic adaptability, and weak detail recovery ability existing in traditional infrared image super-resolution technology.
[0006] To achieve the above object, the present invention provides a feature integration method for interactive convolution and dynamic focusing of infrared images, including the following steps: S1. Obtain infrared image pairs with different resolutions in various real scenarios through an infrared camera; S2. Degrade and preprocess a part of the original high-resolution images to obtain low-quality high-resolution images, forming a mixed low-resolution dataset, and divide the processed dataset into a training set and a test set; S3. Construct a two-layer feature extraction module for feature modeling; S4. Use the processed training set to train the network and optimize the loss function; S5. Input the low-resolution infrared image into the trained network and output the high-resolution reconstruction result.
[0007] Furthermore, in step S1, infrared image data of different scenarios are obtained through the adjustment of optical zoom and focal length using an infrared camera. Specifically: S1.1. For optical zoom, use a long-wave (8 - 14μm) and mid-wave (3 - 5μm) infrared camera that supports electric zoom. Complete the switch from wide-angle to telephoto by adjusting the lens focal length, and directly generate original image pairs with different spatial resolutions; S1.2. To obtain infrared image pairs with different resolutions, the long-wave camera outputs a resolution of 640×480 at a focal length of 50mm and 1280×1024 at a focal length of 200mm; the mid-wave camera generates images with different resolutions by replacing the lens with a focal length of 30mm / 90mm, directly using the hardware characteristics to generate real physical resolution differences, which conforms to the actual working mode of the infrared imaging system; Furthermore, the steps of degrading and preprocessing a part of the original high-resolution images in S2 to obtain low-quality high-resolution images, forming a mixed low-resolution dataset, and dividing the processed dataset into a training set and a test set are as follows: S2.1 - 1. Simulate the blurring effect of the optical system through a Gaussian function, and combine downsampling operations to achieve resolution degradation. Artificially degrade a part of the original high-resolution images. Among them, Gaussian blurring is achieved by convolving the image with a two-dimensional Gaussian kernel, and its kernel is defined as: ; In the formula, is the standard deviation of the Gaussian kernel, controlling the degree of blurring; S2.1 - 2. Simulate the optical system aberration through Gaussian blurring, and then generate low-resolution images through downsampling: ; In the formula, is the Gaussian kernel G for the high-resolution image Convolution, Downsample is a downsampling operation that takes the mean of adjacent pixels. The synthetic LR-HR pairs generated through controllable degradation preprocessing can provide accurate pixel-level supervision signals, helping the model learn the basic degradation patterns. Incorporating the synthetic pairs into the collected real LR-HR pairs ensures that the training data covers a variety of degradation scenarios and improves the model's robustness. S2.2. Further, the mixed low-resolution dataset needs to be divided into a training set and a test set for subsequent network training and validation. The data is divided into a training set (70%), a validation set (15%), and a test set (15%) according to a ratio to ensure that each subset covers different scenarios, resolutions, and noise levels and to verify the generalization ability. Further, in S3, a two-layer feature extraction module is constructed for feature modeling, including two stages: shallow feature extraction and deep feature extraction. First, a smaller convolution kernel is used to extract shallow features to capture local details and high-frequency information (such as edges and textures) in the image. Then, the deep extraction module is used to focus on the global context and semantic information to understand the overall structure and semantic relationships of the image. The specific steps are as follows: S3.1-1. The shallow feature extraction module uses a 3×3 convolutional layer to perform preliminary feature extraction on the input low-resolution infrared image and outputs a shallow feature map. The input low-resolution image is set as: ; In the formula, is the spatial resolution of the image (height × width), is the number of input channels. Infrared images are usually single-channel with 1. When the convolution kernel is set to , the pixel value of each pixel in the output feature map F is calculated by the following formula: ; In the formula, is the size space of the convolution kernel, are the pixel coordinates of the output feature map, c is the output channel index, where , and is the number of output channels; is the bias term, which is initialized to 0 here; is the boundary processing after padding; S3.1-2. The input size is , the convolution kernel size is k x k, the stride s = 1, and there is no padding, i.e., p = 0. Then the output size is as follows: ; ; In the formula, and are the output length and width without padding.
[0008] Further, the deep feature extraction module (DFB) in S3 is stacked by multiple feature fusion blocks (FB). The module uses a two-way information interaction mechanism to enhance the feature expression ability, including a source information branch based on depth convolution and an information branch based on the attention mechanism. The specific steps are as follows: S3.2-1. The deep feature fusion block (DFB) generates a channel attention map for enhancing the feature representation of the self-attention branch by passing the features of the convolution branch through global average pooling, a 1×1 convolutional layer, and a Sigmoid activation function. Among them, global average pooling can capture global context information, avoid local bias, calculate the global mean for each channel, and obtain a channel descriptor : ; Among them, is the c-th channel value of the input feature X at the position (i, j), are all channels; S3.2-2. Enhance the feature representation through the channel attention mechanism. The core idea is: use the convolution branch to extract local features, the self-attention branch models global dependencies, and the channel attention map can dynamically adjust the feature weights, integrating the advantages of both. Among them, the number of channels is compressed by a 1×1 convolution (usually reduced to C / r, where r is the compression ratio): ; In the formula, , s is the compressed feature vector, and ; In addition, the ReLU activation function is used in the channel attention map, and the formula is as follows: ; ReLU in the formula is a non-linear transformation, which can alleviate the problem of gradient disappearance; Another information branch of the two-way information interaction mechanism comes from the spatial interaction module. The features of the self-attention branch are passed through a 1×1 convolutional layer and a Sigmoid activation function to generate a spatial attention map for enhancing the feature representation of the convolution branch. When generating the spatial attention weights, mainly through a fully connected layer (FC) and a Sigmoid activation function, generate the spatial attention weights: ; ; In the formula, is the weight of the fully connected layer, belongs to the real number range and is the bias term; is the sigmoid function, ensuring that the weight value .
[0009] Furthermore, the deep feature fusion block DFB adopts a dual-branch parallel structure, including a convolutional branch and a self-attention branch, and realizes feature enhancement in the channel and spatial dimensions through a two-way information interaction mechanism. The specific steps are as follows: S3.3-1. The convolutional branch adopts a dual-path depth convolution structure. The first path uses a 3×3 depth convolution to extract local detail features, and the second path uses a 5×5 depth convolution to expand the receptive field and enhance texture information. The outputs of the two paths are concatenated after dimensionality reduction through a 1×1 convolution to form the final convolutional branch feature. The feature outputs of the two paths are calculated as follows: ; ; where X belongs to is the input feature, H×W is the spatial resolution, C is the number of channels, K represents the convolutional kernel, b is an optional bias term for this convolutional kernel size. Then, the output feature maps of the two branches are directly concatenated in the channel dimension to form the final fusion feature. The concatenation operation is an element-wise merge, without parameter calculation, only increasing the number of channels and having a small computational overhead. The formula for fusing the dual-path features is as follows: ; where and are the features output by the two paths, is a 1×1 dimensionality reduction convolution operation, whose convolutional kernel size is 1×1 (i.e., only covering a single pixel point of the input feature map), used to adjust the number of channels, fuse features or achieve cross-channel information interaction. The concatenation (Concat) fusion combines the dual-branch features through the channel dimension, simply and efficiently combines local and global information, and is one of the core operations to improve the model performance; S3.3-2. The self-attention branch adopts a window partitioning and multi-head self-attention mechanism to model global dependencies. Specifically, the input feature is divided into multiple non-overlapping windows, self-attention is calculated within each window, the multi-head self-attention mechanism is adopted to divide the features into multiple subspaces for parallel calculation, and a window offset strategy is adopted to enhance information interaction between different windows. In this step, the main task is to generate attention weights by calculating the correlation between features. The specific formula is as follows: ; where Z is the feature representation after self-attention weighting, containing global context information, D is the dimension of the feature, Softmax is the normalization exponential function, and Q, K, and V are the input feature X mapped to query, key, and value respectively.
[0010] Furthermore, in S4, the processed training set will be used to train the network and optimize the loss function to improve the quality and details of the reconstructed image. It combines pixel-level loss, cross-entropy loss, and minimizing the reconstruction loss. The specific steps are as follows: S4.1. To enhance the discriminative ability of the feature representation, the model can achieve this by minimizing the loss function. The expression of the pixel-level loss function is: ; In the formula, L is the pixel difference between the high-resolution image and the reconstructed image . refers to the pixel value of the high-resolution image at coordinates (i, h, w); S4.2. Based on the attention mechanism, the model is fine-tuned through a supervised segmentation task. The segmentation loss function is calculated, and the network weights are optimized through backpropagation, which can help the model learn richer spatial structure information from the data. This task uses the cross-entropy loss function to calculate the segmentation error: ; In the formula, given the input x and the corresponding label , where is the true label of the i-th pixel, is the probability that the model predicts the i-th pixel belongs to a certain category; S4.3. After data processing and feature extraction, the model needs to generate new samples by learning the reconstruction process. The training objective is to minimize the reconstruction error, and the following loss function is used: ; Among them, is the result of the network reconstructing the image, and the loss function is the error between the prediction result and the true original high-resolution data .
[0011] Furthermore, in S5, the low-resolution infrared image is input into the trained network. The shallow features and deep features are fused through long skip connections, and PixelShuffle operation is used for upsampling. Finally, a 3×3 convolutional layer is used to restore the target resolution and output the high-resolution reconstruction result. The specific steps are as follows: S5.1. Long skip connections can directly transfer shallow features and enhance the gradient backpropagation. The main task in this step is channel dimension concatenation, and the formula is as follows: ; In the formula, is the feature map of the l-th layer, is the feature map of the k-th layer, and concat is the fusion method; S5.2. Next, perform upsampling using the PixelShuffle operation. The core idea is to expand the spatial dimensions (H, W) of the low-resolution feature map while reducing the number of channels (C), thereby achieving efficient upsampling. The calculation formula is as follows: ; In the formula, r is the upsampling factor, , , are the number of channels, height, and width after recombination; Y and X are the output and input features respectively. Compared with traditional interpolation (such as bilinear interpolation), it can better retain high-frequency details. Thus, after the image reconstruction module uses pixel recombination (PixelShuffle) and a 3×3 convolutional layer, the feature map is restored to a high-resolution image.
[0012] Beneficial effects: The present invention first collects three-dimensional point cloud data of the underground space from multiple perspectives and directions by arranging multiple sensors including lidar and visual sensors at different positions; registers and aligns data from different perspectives or different time periods; extracts meaningful local features and removes the influence of dynamic objects (such as vehicles or pedestrians) on map construction to complete feature extraction and map construction; designs a sparse 3D U-Net module for three-dimensional feature extraction; on the basis of feature extraction, introduces a self-supervised learning module to enhance feature representation using unlabeled point cloud data; fills in missing point cloud data through an iterative generation process and supplements missing spatial details; uses a feature projection model to map features of different modalities into a shared latent space. Finally, it breaks through key technologies such as three-dimensional semantic segmentation and three-dimensional reconstruction under information missing conditions, multi-agent data stitching, etc., establishes a method for detecting and characterizing underground enclosed spaces, provides technical support for the development of new underground space autonomous detection and environmental characterization equipment, and improves the intelligence acquisition ability for underground enclosed space targets. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a schematic flow chart of the present invention; Figure 2 is a schematic diagram of two-way information interaction adopted by the present invention; Figure 3 is a schematic diagram of the principle of the feature fusion block FB proposed by the present invention; Figure 4 is a fusion structure diagram of two layers of features of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0014] The present invention will be further described below with reference to the accompanying drawings.
[0015] As shown in Figure 1As shown, a feature integration method for interactive convolution and dynamic focusing of infrared images includes the following steps: S1. Obtain infrared image pairs of different resolutions in various real scenarios through an infrared camera; S2. Degrade and preprocess a part of the original high-resolution images to obtain low-quality high-resolution images, forming a mixed low-resolution dataset, and divide the processed dataset into a training set and a test set; S3. Construct a two-layer feature extraction module for feature modeling; S4. Use the processed training set to train the network and optimize the loss function; S5. Input the low-resolution infrared image into the trained network and output the high-resolution reconstruction result.
[0016] As a preferred implementation manner, in step S1, infrared image data of different scenes are obtained through an infrared camera by adjusting the optical zoom and focal length. The steps are as follows: S1.1. Use a long-wave (8 - 14μm) and mid-wave (3 - 5μm) infrared camera that supports electric zoom for optical zoom. Complete the switch from wide-angle to telephoto by adjusting the lens focal length, and directly generate original image pairs of different spatial resolutions; S1.2. To obtain infrared image pairs of different resolutions, the long-wave camera outputs a resolution of 640×480 at a focal length of 50mm and 1280×1024 at a focal length of 200mm; the mid-wave camera generates images of different resolutions by replacing the lens with focal lengths of 30mm / 90mm, and directly utilizes the hardware characteristics to generate real physical resolution differences, which conforms to the actual working mode of the infrared imaging system.
[0017] As a preferred implementation manner, the steps of degrading and preprocessing a part of the original high-resolution images in S2 to obtain low-quality high-resolution images, forming a mixed low-resolution dataset, and dividing the processed dataset into a training set and a test set are as follows: S2.1 - 1. Simulate the blurring effect of the optical system through a Gaussian function, and combine the downsampling operation to achieve resolution degradation. Artificially degrade a part of the original high-resolution images. Among them, Gaussian blurring is achieved by convolving a two-dimensional Gaussian kernel with the image, and its kernel is defined as: ; In the formula, is the standard deviation of the Gaussian kernel, which controls the degree of blurring; S2.1 - 2. Simulate the optical system aberration through Gaussian blurring, and then generate low-resolution images through downsampling: ; In the formula, is the Gaussian kernel G convolving with the high-resolution image Convolution, Downsample is the downsampling operation, taking the mean of adjacent pixels. The synthetic LR-HR pairs generated through controllable degradation preprocessing can provide accurate pixel-level supervision signals, helping the model learn the basic degradation patterns. Incorporating the synthetic pairs into the collected real LR-HR pairs can ensure that the training data covers multiple degradation scenarios and improve the model's robustness; S2.2. It is necessary to divide the mixed low-resolution dataset into a training set and a test set for subsequent network training and verification. The data is divided into a training set (70%), a validation set (15%), and a test set (15%) according to a certain proportion to ensure that each subset covers different scenarios, resolutions, and noise levels and to verify the generalization ability.
[0018] As a preferred implementation, a two-layer feature extraction module is constructed in S3 for feature modeling, including two stages: shallow feature extraction and deep feature extraction. First, a smaller convolutional kernel is used to extract shallow features to capture local details and high-frequency information (such as edges and textures) in the image, and then the deep extraction module is used to focus on global context and semantic information to understand the overall structure and semantic relationship of the image. The specific steps are as follows: S3.1-1. The shallow feature extraction module uses a 3×3 convolutional layer to perform preliminary feature extraction on the input low-resolution infrared image and outputs a shallow feature map. The input low-resolution image is set as: ; In the formula, is the spatial resolution of the image (height × width), is the number of input channels. Infrared images are usually single-channel with 1. When the convolutional kernel is set as in the case of, each pixel value of the output feature map F is calculated by the following formula: ; In the formula, is the size space of the convolutional kernel, is the pixel coordinate of the output feature map, c is the output channel index, where , and is the number of output channels; is the bias term, which is initialized to 0 here; is the boundary processing after padding; S3.1-2. The input size is , the convolutional kernel size is k x k, the stride s = 1, and there is no padding, i.e., p = 0. Then the output size is as follows: ; ; In the formula, and Output the length and width without padding.
[0019] As a preferred embodiment, the deep feature extraction module (DFB) in S3 is stacked by multiple feature fusion blocks (FB). The module uses a bidirectional information interaction mechanism to enhance the feature expression ability, including a source information branch based on depth convolution and an information branch based on the attention mechanism, such as Figure 3 , the feature fusion block (FFB) realizes the efficient fusion of local and global features through a double-branch parallel structure. Its convolution branch uses 3×3 and 5×5 dual-path depth convolution to extract multi-scale local features, which are spliced after being compressed by 1×1 convolution; the attention branch models the long-range dependence relationship through window division and multi-head self-attention mechanism, and introduces window offset to enhance cross-window interaction. The two branches are dynamically complementary through a bidirectional interaction mechanism: the features of the convolution branch generate channel attention weights through global average pooling and a fully connected layer to enhance the value matrix of the attention branch; the features of the attention branch generate a spatial attention mask through 1×1 convolution to refine the output features of the convolution branch. Finally, feature fusion is achieved by element-wise addition, and the residual connection is retained to ensure the flow of gradients. This design optimizes the representation of both the channel and spatial dimensions on the feature map of H×W×C, where C is the number of channels, M is the attention window size, and h is the number of attention heads. The specific steps are as follows: S3.2-1. The deep feature fusion block (DFB) generates a channel attention map for the features of the convolution branch through global average pooling, a 1×1 convolutional layer, and a Sigmoid activation function to enhance the feature representation of the self-attention branch. Among them, global average pooling can capture global context information, avoid local bias, calculate the global mean for each channel, and obtain the channel descriptor : ; Among them, is the value of the c-th channel of the input feature X at the position (i, j), are all channels; S3.2-2. Enhance the feature representation through the channel attention mechanism. Its core idea is: use the convolution branch to extract local features, the self-attention branch models the global dependence relationship, and the channel attention map can dynamically adjust the feature weights, integrating the advantages of both. Among them, the number of channels is compressed through 1×1 convolution (usually reduced to C / r, where r is the compression ratio): ; In the formula, , s is the compressed feature vector, and ; in addition, the ReLU activation function is used in the channel attention map, and the formula is as follows: ; In the formula, ReLU is a non-linear transformation that can alleviate the vanishing gradient problem; S3.2-3. Another information branch of the two-way information interaction mechanism comes from the spatial interaction module. The features of the self-attention branch are used to generate a spatial attention map through a 1×1 convolutional layer and a Sigmoid activation function, which is used to enhance the feature representation of the convolutional branch. When generating the spatial attention weights, a fully connected layer (FC) and a Sigmoid activation function are mainly used to generate the spatial attention weights: ; ; In the formula, is the weight of the fully connected layer, belongs to the real number range and is the bias term; is the sigmoid function, ensuring that the weight value .
[0020] As a preferred implementation, the deep feature fusion block DFB adopts a double-branch parallel structure, including a convolutional branch and a self-attention branch, and realizes feature enhancement in the channel and spatial dimensions through a two-way information interaction mechanism. As Figure 2 shown, the convolutional branch generates channel attention weights (σ activation) through global average pooling (GAP) and a fully connected layer (FC), which are used to modulate the Value matrix of the attention branch; the attention branch generates a spatial attention mask through 1×1 convolution and σ activation, which is used to enhance the key spatial regions of the convolutional branch. Finally, the features of the two branches are fused by element-wise addition, and the residual connection is retained to ensure training stability. This design synchronously optimizes the channel and spatial dimension representations on the H×W×C feature map, where M is the attention window size and C is the number of feature channels. The specific steps are as follows: S3.3-1. The convolutional branch adopts a two-path deep convolutional structure. The first path uses a 3×3 deep convolution to extract local detail features, and the second path uses a 5×5 deep convolution to expand the receptive field and enhance the texture information. The outputs of the two paths are concatenated after dimensionality reduction through a 1×1 convolution to form the final convolutional branch feature. The feature outputs of the two paths are calculated as follows: ; ; In the formula, X belongs to is the input feature, H×W is the spatial resolution, C is the number of channels, K represents the convolutional kernel, and b is the optional bias term under the size of this convolutional kernel. Then, the output feature maps of the two branches are directly concatenated in the channel dimension to form the final fusion feature. The concatenation operation is an element merge, without parameter calculation, only increasing the number of channels, and the computational cost is small. The formula for fusing the two-path features is as follows: ; In the formula, and are the features of two path outputs. is a 1×1 dimensionality reduction convolution operation with a convolution kernel size of 1×1 (i.e., only covering a single pixel of the input feature map), which is used to adjust the number of channels, fuse features, or achieve cross-channel information interaction. The splicing (Concat) fusion combines the dual-branch features through the channel dimension, simply and efficiently combines local and global information, and is one of the core operations to improve the model performance. S3.3-2. The self-attention branch adopts window partitioning and multi-head self-attention mechanism to model global dependencies. Specifically, the input features are divided into multiple non-overlapping windows, self-attention is calculated within each window, the multi-head self-attention mechanism is adopted, the features are divided into multiple subspaces for parallel calculation, and the window offset strategy is adopted to enhance the information interaction between different windows. In this step, the main task is to generate attention weights by calculating the correlation between features, and the specific formula is as follows: ; In the formula, Z is the feature representation after self-attention weighting, which contains global context information, D is the dimension of the feature, Softmax is the normalized exponential function, and Q, K, and V are the input feature X mapped to query, key, and value respectively.
[0021] As a preferred implementation, in S4, the processed training set will be used to train the network and optimize the loss function to improve the quality and details of the reconstructed image. It combines pixel-level loss, cross-entropy loss, and minimization of reconstruction loss. The specific steps are as follows: S4.1. To enhance the discrimination ability of the feature representation, the model can be achieved by minimizing the loss function. The expression of the pixel-level loss function is: ; In the formula, L is the pixel difference between the high-resolution image and the reconstructed image . refers to the pixel value of the high-resolution image with coordinates (i, h, w); S4.2. Based on the attention mechanism, by fine-tuning the model through a supervised segmentation task, calculating the segmentation loss function, and optimizing the network weights through backpropagation, it can help the model learn richer spatial structure information from the data. This task uses the cross-entropy loss function to calculate the segmentation error: ; In the formula, given the input x and the corresponding label , here is the true label of the i-th pixel, is the probability that the model predicts the i-th pixel belongs to a certain category; S4.3. After data processing and feature extraction, the model needs to generate new samples through the learning reconstruction process. The training objective is to minimize the reconstruction error, and the following loss function is used: ; where, is the result of the network reconstructing the image, and the loss function is the error between the prediction result and the true original high-resolution data ; As a preferred implementation, such as Figure 4 , S5 inputs the low-resolution infrared image into the trained network, fuses the shallow features and the deep features through long skip connections, performs upsampling using the PixelShuffle operation, and finally uses a 3×3 convolutional layer to restore the target resolution to output a high-resolution reconstruction result. The specific steps are as follows: S5.1. The long skip connection can directly transfer the shallow features and enhance the gradient backpropagation. The main task in this step is channel dimension concatenation, and the formula is as follows: ; In the formula, is the feature map of the l-th layer, is the feature map of the k-th layer, and concat is the fusion method; S5.2. Next, the PixelShuffle operation is used for upsampling. The core idea is to expand the spatial dimensions (H, W) of the low-resolution feature map while reducing the number of channels (C), so as to achieve efficient upsampling. The calculation formula is as follows: ; In the formula, r is the upsampling factor, , , are the recombined number of channels, height, and width; Y and X are the output and input features respectively. Compared with traditional interpolation (such as bilinear interpolation), it can better retain high-frequency details. Therefore, after the image reconstruction module uses pixel recombination (PixelShuffle) and a 3×3 convolutional layer, the feature map is restored to a high-resolution image.
[0022] This infrared image super-resolution interactive convolution and dynamic focus feature integration method is based on multi-directional information feature interaction and fusion technology, adaptive attention dynamic focus technology, and multi-task collaborative loss optimization strategy. It can effectively extract and enhance key information in low-resolution infrared images, and then significantly improve the resolution of infrared images and the quality of details. It can provide core algorithm support and technical guarantee for target recognition, precise monitoring, and high-definition imaging in complex infrared environments.
Claims
1. An interactive convolution and dynamic focusing feature integration method for infrared images, characterized in that, It includes the following steps: S1. Obtain infrared image pairs with different resolutions in various real scenarios through an infrared camera; S2. Degrade and preprocess a part of the original high-resolution images to obtain a mixed low-resolution dataset composed of low-quality high-resolution images, and divide the processed dataset into a training set and a test set; S3. Construct a two-layer feature extraction module for feature modeling; S4. Use the processed training set to train the network and optimize the loss function; S5. Input the low-resolution infrared image into the trained network and output the high-resolution reconstruction result.
2. The feature integration method for interactive convolution and dynamic focusing of infrared images according to claim 1, wherein Step S1 obtains infrared image data of different scenarios through an infrared camera by adjusting optical zoom and focal length, specifically: S1.
1. Use a long-wave and mid-wave infrared camera supporting electric zoom for optical zoom, and complete the switch from wide-angle to telephoto by adjusting the lens focal length to directly generate original image pairs with different spatial resolutions; S1.
2. In order to obtain infrared image pairs with different resolutions, the long-wave camera outputs a resolution of 640×480 at a focal length of 50mm and 1280×1024 at a focal length of 200mm; the mid-wave camera generates images with different resolutions by replacing the lens with a focal length of 30mm / 90mm.
3. The feature integration method for interactive convolution and dynamic focusing of infrared images according to claim 1, wherein Step S2 is specifically: S2.1-1. Simulate the blurring effect of the optical system through the Gaussian function, and combine the downsampling operation to achieve resolution degradation. Artificially degrade a part of the original high-resolution images, where the Gaussian blur is realized by convolving the two-dimensional Gaussian kernel with the image, and its kernel is defined as: ; wherein, is the standard deviation of the Gaussian kernel, controlling the degree of blurring; S2.1-2. Simulate optical system aberration through Gaussian blur and then generate low-resolution images through downsampling; ; In the formula, is the convolution of the Gaussian kernel G with the high-resolution image , and Downsample is the downsampling operation where the mean value of adjacent pixels is taken; S2.
2. Divide the mixed low-resolution dataset into a training set and a test set for subsequent network training and verification, and divide the data into a training set, a verification set, and a test set according to a ratio to ensure that each subset covers different scenarios, resolutions, and noise levels to verify the generalization ability.
4. The feature integration method for interactive convolution and dynamic focusing of infrared images according to claim 1, characterized in that In step S3, a two-layer feature extraction module is constructed for feature modeling, including two stages of shallow feature extraction and deep feature extraction. First, a 3×3 convolution kernel is used to extract shallow features to capture local details and high-frequency information in the image, and then a deep extraction module is used to focus on global context and semantic information to understand the overall structure and semantic relationship of the image, specifically: S3.1-1. The shallow feature extraction module uses a 3×3 convolutional layer to perform preliminary feature extraction on the input low-resolution infrared image and outputs a shallow feature map. Set the input low-resolution image as: ; Wherein, is the spatial resolution of the image, i.e., height × width, is the number of input channels. For an infrared image, it is usually 1 for a single channel. When the convolutional kernel is set to each pixel value of the output feature map F is calculated by the following formula: ; In the formula, is the size space of the convolution kernel, are the pixel coordinates of the output feature map, c is the output channel index, , is the number of output channels; is the bias term, initialized to 0; is the boundary processing after padding; S3.1-2. The input size is , the convolutional kernel size is k k, the stride s = 1, without padding i.e., p = 0, then the output size is as follows: ; ; In the formula, and are the length and width of the output without filling; The deep feature extraction module in step S3 is stacked by multiple feature fusion blocks, and the module uses a two-way information interaction mechanism to enhance the feature expression ability, including a source information branch based on depth convolution and an information branch based on an attention mechanism, specifically: S3.2-1. The deep feature fusion block generates a channel attention map from the features of the convolutional branch through global average pooling, a 1×1 convolutional layer, and a Sigmoid activation function to enhance the feature representation of the self-attention branch. Among them, global average pooling can capture global context information, avoid local bias, calculate the global mean for each channel, and obtain a channel descriptor : ; Among them, is the c-th channel value of the input feature X at the position (i, j), is all channels; S3.2-2. Enhance the feature representation through a channel attention mechanism, that is: use the convolutional branch to extract local features, the self-attention branch models global dependencies, and the channel attention map can dynamically adjust the feature weights, and fuse the advantages of both. Among them, the number of channels is compressed through 1×1 convolution to reduce the dimension to C / r, where r is the compression ratio; ; Wherein, , s is the compressed feature vector, and ; in addition, an activation function is adopted in the channel attention map, and the formula is as follows: ; ReLU in the formula is a non-linear transformation; S3.2-3. Another information branch of the two-way information interaction mechanism comes from the spatial interaction module. The features of the self-attention branch are used to generate a spatial attention map through a 1×1 convolutional layer and a Sigmoid activation function, which is used to enhance the feature representation of the convolutional branch. When generating the spatial attention weights, a fully connected layer and a Sigmoid activation function are mainly used to generate the spatial attention weights: ; ; In the formula, is the weight of the fully connected layer, belongs to the real number range and is the bias term; is the sigmoid function to ensure the weight value ; The deep feature fusion block adopts a dual-branch parallel structure, including a convolutional branch and a self-attention branch, and realizes feature enhancement in the channel and spatial dimensions through a two-way information interaction mechanism. The specific steps are as follows: S3.3-1. The convolutional branch adopts a dual-path depth convolution structure. The first path uses a 3×3 depth convolution, and the second path uses a 5×5 depth convolution. The outputs of the two paths are spliced after dimensionality reduction through a 1×1 convolution to form the final convolutional branch features. The feature outputs of the two paths are calculated by the following formula: ; ; where X belongs to is the input feature, H×W is the spatial resolution, C is the number of channels, K represents the convolutional kernel, and b is the optional bias term for this convolutional kernel size; the output feature maps of the two branches are directly concatenated in the channel dimension to form the final fused feature. The formula for fusing the dual-path features is as follows: ; In the formula, and are the features of two path outputs, is a 1×1 dimensionality reduction convolution operation with a convolution kernel size of 1×1; S3.3-2. The self-attention branch adopts a window partitioning and multi-head self-attention mechanism for modeling global dependencies, that is, the input features are divided into multiple non-overlapping windows, self-attention is calculated within each window, the multi-head self-attention mechanism is adopted, the features are divided into multiple sub-spaces for parallel calculation, and a window offset strategy is adopted to enhance the information interaction between different windows. The attention weights are generated by calculating the correlation between features. The specific formula is as follows: ; Among them, Z is the feature representation after self-attention weighting, which contains global context information, D is the dimension of the feature, Softmax is the normalized exponential function, and Q, K, and V are the input feature X mapped to the query, key, and value respectively.
5. The feature integration method for interactive convolution and dynamic focusing of infrared images according to claim 1, characterized in that In step S4, the processed training set will be used to train the network, and the specific steps to optimize the loss function are as follows: S4.
1. To achieve the ability to distinguish enhanced feature representations, the model is realized by minimizing the loss function. The expression of the pixel-level loss function is: ; where L is the high-resolution image and the reconstructed image pixel difference of refers to the pixel value of the high-resolution image with coordinates i, h, w; S4.
2. Based on the attention mechanism, the model is fine-tuned through a supervised segmentation task, the segmentation loss function is calculated, and the network weights are optimized through backpropagation. The cross-entropy loss function is used to calculate the segmentation error: ; where the given input is \(x\) and the corresponding label , where is the true label of the \(i\)-th pixel, is the probability that the \(i\)-th pixel predicted by the model belongs to a certain category; S4.
3. After data processing and feature extraction, the model needs to generate new samples through a learning reconstruction process. The training objective is to minimize the reconstruction error, and the following loss function is used: ; Among them, is the result of the network reconstruction image, and the loss function is the error between the predicted result and the true original high-resolution data therebetween.
6. The feature integration method for interactive convolution and dynamic focusing of infrared images according to claim 1, wherein Step S5 is specifically as follows: S5.
1. The long skip connection can directly transmit shallow features and enhance the gradient backpropagation. The main task in this step is channel dimension splicing, and the formula is as follows: ; In the formula, is the feature map of the l-th layer, is the feature map of the k-th layer, and concat is the fusion method; S5.
2. Pixel reorganization operation is used for upsampling. The core idea is to expand the spatial dimensions H and W of the low-resolution feature map while reducing the number of channels C, so as to achieve efficient upsampling. The calculation formula is as follows: ; where r is the adopted factor, , , are the number of channels, height, and width after recombination; Y and X are the output and input features respectively; finally, after pixel recombination and a 3×3 convolutional layer, the image reconstruction module restores the feature map to a high-resolution image.
Citation Information
Patent Citations
Infrared image super-resolution reconstruction method combining two-way convolution and self-attention
CN117274047A
Multi-angle panchromatic and multispectral image progressive fusion method
CN119942285A
Medical image segmentation method based on u-net
US20220309674A1
Cited By
Infrared camera hybrid focusing system and method based on MDP decision
CN120568200A
Infrared image super-resolution method and device based on state space model and medium
CN120746839A
Industrial image super-resolution reconstruction method based on physical consistency self-supervision mechanism
CN120876229A
Image super-resolution method and device based on bidirectional focusing enhancement
CN120912438A
Infrared image restoration model training method adaptive to embedded platform
CN120931535A