Infrared image super-resolution method based on global context channel attention
By introducing a global context channel attention network in infrared image super-resolution reconstruction, the problem of poor infrared image super-resolution reconstruction performance in the prior art is solved, and a higher quality and practical infrared image super-resolution reconstruction effect is achieved.
Patent Information
- Application Number
- CN202411872359.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-18
AI Technical Summary
The existing deep learning super-resolution reconstruction algorithm is mainly aimed at visible light images, and it fails to effectively optimize the signal-to-noise ratio, contrast, edge structure richness and inhomogeneity of infrared images, resulting in poor super-resolution reconstruction performance of infrared images.
A super-resolution method based on global context channel attention is proposed. By constructing a global context channel attention network (GCCAN), the global context channel attention module is introduced in the deep feature extraction stage to enhance the network's understanding and reconstruction ability of infrared features.
Through the global context channel attention mechanism, different levels of features are effectively captured and utilized, and the overall quality and practicality of infrared image super-resolution reconstruction is improved, and it shows higher stability and performance excellence compared with traditional methods.
Smart Images

Figure CN119941507A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of image processing, and in particular relates to an infrared image super-resolution algorithm based on global context channel attention. Background Art
[0002] Due to its unique advantages, infrared imaging has become a branch of imaging that is difficult to replace with visible light, and is widely used in security monitoring, military reconnaissance, and medical imaging. All objects will generate infrared thermal radiation above absolute zero. Due to its special imaging method and longer wavelength than visible light imaging, infrared imaging technology can also present clear images under low light or haze conditions. As a result, it plays an important role in night navigation, border security, military reconnaissance, and biomedicine. However, due to the manufacturing process and materials of infrared imaging technology itself, the number of pixels cannot reach the level of visible light imaging. In addition, increasing the resolution by improving the performance of imaging equipment is also accompanied by higher costs. These factors have jointly promoted the demand for infrared image super-resolution technology.
[0003] Image super-resolution reconstruction technology is a basic task in computer vision, which aims to reconstruct the input low-resolution image (Low-resolution, LR) into a clearer high-resolution image (High-resolution, HR). With the remarkable success of AlexNet in target classification tasks, deep learning has become a focus of attention in the field of image processing, and various visual tasks have tried to use it as a new method to replace traditional methods. The same is true in the field of image super-resolution, and various image super-resolution reconstruction algorithms based on deep learning have emerged. However, many deep learning super-resolution reconstruction algorithms are currently mainly oriented to visible light images, and are not specifically optimized for the signal-to-noise ratio, contrast, edge structure richness and non-uniformity of infrared images. Therefore, it is particularly necessary to design an improved super-resolution reconstruction network algorithm for efficient infrared images to improve the super-resolution reconstruction performance of infrared images. Summary of the invention
[0004] In view of this, the main object of the present invention is to provide an infrared image super-resolution method based on global context channel attention.
[0005] To achieve the above object, the technical solution of the present invention is achieved as follows:
[0006] An embodiment of the present invention provides an infrared image super-resolution method based on global context channel attention, the method comprising the following steps:
[0007] Step 1: Get the training data set;
[0008] Step 2: Expand the training data set;
[0009] Step 3: Construct a global context channel attention network (GCCAN). The GCCAN network is based on the RCAN network architecture. While keeping the overall architecture of the RCAN network unchanged, it improves the residual group structure in the deep feature extraction stage. In the residual group structure, m global context channel attention blocks (GCCAB) connected in series are integrated;
[0010] Step 4: Construct structural loss function;
[0011] Step 5: Input the expanded training data set obtained in step 2 into the global context channel attention network constructed in step 3, and use the structural loss function constructed in step 4 for training as the optimization target for training. Use the ADAM optimizer to update the model parameters until the loss no longer decreases, and obtain a trained infrared image super-resolution reconstruction model.
[0012] Step 6: Use the infrared image super-resolution reconstruction model trained in step 5 to perform super-resolution reconstruction on the low-resolution images in the test set to obtain infrared super-resolution images.
[0013] In the above scheme, the acquisition of the training data set specifically includes: obtaining original high-resolution images from a public database or personal collection, and obtaining the high-resolution images through I LR =D w B l I HR +n o The image degradation model described is degraded to obtain, where I HR represents the original high-resolution infrared image, B l represents fuzzy interference, D w represents the downsampling process of high-resolution images, n o represents the noise in the imaging process, I LR For the collected low-resolution infrared images, a Gaussian blur kernel with a standard deviation of 1 is used in the degradation model, Bicubic interpolation downsampling is used, and additive Gaussian noise is used for noise. Finally, these one-to-one corresponding sets of high- and low-resolution image pairs are used as training data sets.
[0014] In the above scheme, the expanding the training data set specifically includes: adopting a data enhancement strategy for the acquired training data set, that is, randomly horizontally flipping or rotating the data set images by 90°, 180° and 270°.
[0015] In the above scheme, the global context channel attention network (GCCAN) specifically includes: the GCCAN includes three parts: shallow feature extraction, deep feature extraction and image reconstruction. In the shallow feature extraction stage, the model uses a convolution operation to process a single-channel infrared image with an input size of H×W×1, so that it is mapped into a feature map with a deeper channel; in the GCCAN, the channel depth after mapping is set to 64, and then the deep feature extraction stage includes n residual groups connected in series and a convolution operation at the tail end, and the image reconstruction part operation is sub-pixel convolution upsampling, which is used to increase the extracted feature map to the target resolution;
[0016] Among them, the GCCAN network is based on the RCAN network architecture. While keeping the overall architecture of the RCAN network unchanged, it improves the residual group structure in the deep feature extraction stage. In the residual group structure, m interconnected global context channel attention blocks (GCCAB) are integrated. In the GCCAN model, n and m are both adjustable hyperparameters, and n=m=10 is set.
[0017] In the above scheme, the global context channel attention module (GCCAB) specifically includes: GCCAB is based on the residual channel attention module of the RCAN algorithm, optimizes the mechanism of calculating channel attention, introduces global context channel attention, and embeds parameter-free channel shift operations to expand the receptive field of the network.
[0018] In the above scheme, the global context channel attention specifically includes: for a feature map F with an input size of H×W×C, the global context channel attention uses a 1×1 convolution to generate a spatial self-attention map shared by all channels Then the spatial self-attention map Deform the size from H×W×1 to 1×1×HW. Generate spatial self-attention map of probability distribution Among them, a represents The original value in Spatial self-attention map representing the generated probability distribution Elements in
[0019] Get the spatial self-attention map of the probability distribution After that, the feature map F is transformed into HW×C size and multiplied by its matrix, that is, according to Calculate the channel descriptor tensor Z, where F′ represents the deformed feature map, which is a tensor of size HW×C. is the spatial self-attention map of the probability distribution, with a size of 1×1×HW, and Z is the final channel descriptor tensor, with a size of 1×1×C;
[0020] After extracting the descriptors of each channel, the interdependencies between channels are captured.
[0021] In the above scheme, the channel shift operation specifically includes: inputting the feature map F of size H×W×C through Move the four partial channels up, down, left, and right by a unit distance in the width and height dimensions of the space, where γ is the ratio of the moved channels. It is the feature map after channel shift.
[0022] In the above scheme, the construction of the structural loss function specifically includes: using the L1 loss function: And the structural loss function: The weighted sum is used as the loss function when training the network: Loss = L L1 (I HR ,I G )+λL structure (I HR ,I G );
[0023] Where W and H represent the width and height of the image respectively, x ij Represents the pixel value of the reconstructed image, y ij Represents the pixel value of the target image, i and j are the indexes of the image, I HR and I G They represent the reconstructed high-resolution infrared image and the real high-resolution infrared image respectively, E represents the edge information image, i and j represent the coordinate index, W and H represent the width and height of the image respectively, and L structure (·) is the calculated structural loss function.
[0024] Compared with the prior art, the present invention has the following beneficial effects:
[0025] The present invention adopts a global context channel attention based infrared image super-resolution algorithm GCCAN. From the two key aspects of enhancing the network's understanding and reconstruction capabilities of infrared features, and efficiently fusing and processing global and local information, it helps to solve the limitations of current super-resolution technology in reconstructing infrared image details and characteristics, so as to improve the overall quality and practicality of the reconstructed image. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is the overall structure diagram of the GCCAN network in the present invention;
[0027] Figure 2It is a schematic diagram of GCCAB in the present invention;
[0028] Figure 3 It is the global context channel attention structure in the present invention;
[0029] Figure 4 Schematic diagram of channel shift in the present invention;
[0030] Figure 5 This is the PSNR change curve during training DETAILED DESCRIPTION
[0031] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0032] The embodiment of the present invention provides an infrared image super-resolution method based on global context channel attention, such as Figure 1-4 As shown, the method comprises the following steps:
[0033] Step 1: Get the training data set;
[0034] Specifically, the original high-resolution image is obtained from a public database or personal collection, and the obtained high-resolution image is obtained through I LR =D w B l I HR +n o The image degradation model described is degraded to obtain, where I HR represents the original high-resolution infrared image, B l represents fuzzy interference, D w represents the downsampling process of high-resolution images, n o represents the noise in the imaging process, I LR For the collected low-resolution infrared images, a Gaussian blur kernel with a standard deviation of 1 is used in the degradation model, Bicubic interpolation downsampling is used, and additive Gaussian noise is used for noise. Finally, these one-to-one corresponding sets of high- and low-resolution image pairs are used as training data sets.
[0035] Step 2: Expand the training data set;
[0036] Specifically, a data augmentation strategy is adopted for the acquired training dataset, that is, the dataset images are randomly horizontally flipped or rotated by 90°, 180°, and 270°.
[0037] Step 3: Construct a global context channel attention network (GCCAN). The GCCAN network is based on the RCAN network architecture. While keeping the overall architecture of the RCAN network unchanged, it improves the residual group structure in the deep feature extraction stage. In the residual group structure, m global context channel attention blocks (GCCAB) connected in series are integrated;
[0038] Specifically, the GCCAN includes three parts: shallow feature extraction, deep feature extraction and image reconstruction. In the shallow feature extraction stage, the model uses a convolution operation to process a single-channel infrared image with an input size of H×W×1, so that it is mapped into a feature map with a deeper channel; in GCCAN, the channel depth after mapping is set to 64, and then the deep feature extraction stage includes n residual groups connected in series and a convolution operation at the tail end. The image reconstruction part operation is sub-pixel convolution upsampling, which is used to increase the extracted feature map to the target resolution.
[0039] Among them, the GCCAN network is based on the RCAN network architecture. While keeping the overall architecture of the RCAN network unchanged, it improves the residual group structure in the deep feature extraction stage. In the residual group structure, m interconnected global context channel attention blocks (GCCAB) are integrated. In the GCCAN model, n and m are both adjustable hyperparameters, and n=m=10 is set.
[0040] GCCAB is based on the residual channel attention module of the RCAN algorithm, optimizes the mechanism of calculating channel attention, introduces global context channel attention, and embeds parameter-free channel shift operations to expand the receptive field of the network.
[0041] For a feature map F with an input size of H×W×C, the global context channel attention uses a 1×1 convolution to generate a spatial self-attention map shared by all channels. Then the spatial self-attention map Deform the size from H×W×1 to 1×1×HW. Generate spatial self-attention map of probability distribution Among them, a represents The original value in Spatial self-attention map representing the generated probability distribution Elements in
[0042] Get the spatial self-attention map of the probability distribution After that, the feature map F is transformed into HW×C size and multiplied by its matrix, that is, according to Calculate the channel descriptor tensor Z, where F′ represents the deformed feature map, which is a tensor of size HW×C. is the spatial self-attention map of the probability distribution, with a size of 1×1×HW, and Z is the final channel descriptor tensor, with a size of 1×1×C;
[0043] For a feature map F with an input size of H×W×C, the global context channel attention uses a 1×1 convolution to generate a spatial self-attention map shared by all channels. The spatial self-attention map The size is transformed from H×W×1 to 1×1×HW;
[0044] Generate spatial self-attention map of probability distribution through Softmax function
[0045] in, Represents the deformed spatial self-attention map of the input Softmax function, a represents The original value in Spatial self-attention map representing the generated probability distribution Elements in
[0046] Deform the feature map F into HW×C size and multiply it with its matrix to determine the channel descriptor tensor Z. Among them, F′ represents the deformed feature map, which is a tensor of size HW×C. is the spatial self-attention map of the probability distribution, with a size of 1×1×HW.
[0047] After extracting the descriptors of each channel, the interdependencies between channels are captured.
[0048] The channel shift operation specifically includes: shifting the input feature map F of size H×W×C by Move the four partial channels up, down, left, and right by a unit distance in the width and height dimensions of the space, where γ is the ratio of the moved channels. It is the feature map after channel shift.
[0049] Step 4: Construct structural loss function;
[0050] Specifically, the L1 loss function is used: And the structural loss function: The weighted sum is used as the loss function when training the network: Loss = L L1 (I HR ,I G)+λL structure (I HR ,I G );
[0051] Where W and H represent the width and height of the image respectively, x ij Represents the pixel value of the reconstructed image, y ij Represents the pixel value of the target image, i and j are the indexes of the image, I HR and I G They represent the reconstructed high-resolution infrared image and the real high-resolution infrared image respectively, E represents the edge information image, i and j represent the coordinate index, W and H represent the width and height of the image respectively, and L structure (·) is the calculated structural loss function. The weight factor λ of the structural loss function is set to 0.2 in the present invention.
[0052] According to g x (i,j)=(I*K x )(i,j),g y (i,j)=(I*K y )(i,j) to obtain edge information in the horizontal and vertical directions, where I represents the high-resolution image or super-resolution reconstructed image after Gaussian filtering, and K x and K y It is divided into horizontal and vertical detection operators, i and j are the coordinates of the pixels of the high-resolution image or super-resolution reconstructed image, g x and g y Respectively represent the extracted edge information in the horizontal and vertical directions.
[0053] The final key edge information E is obtained according to the edge information in the horizontal direction and the vertical direction,
[0054] Step 5: Input the expanded training data set obtained in step 2 into the global context channel attention network constructed in step 3, and use the structural loss function constructed in step 4 for training as the optimization target for training. Use the ADAM optimizer to update the model parameters until the loss no longer decreases, and obtain a trained infrared image super-resolution reconstruction model.
[0055] Step 6: Use the infrared image super-resolution reconstruction model trained in step 5 to perform super-resolution reconstruction on the low-resolution images in the test set to obtain infrared super-resolution images.
[0056] The test examples of the present invention are:
[0057] When training the GCCAN algorithm network with ×2 magnification, the PSNR evaluation test is performed on the obtained test set every thousand iterations. At the same time, the test results of the GCCAN algorithm are compared with those of the RCAN algorithm. The test and comparison results are shown in the figure. Figure 5 shown.
[0058] from Figure 5 It can be seen that as the training progresses, the test results of the GCCAN and RCAN algorithms gradually increase and tend to stabilize. GCCAN finally obtained a higher Peak Signal-to-Noise Ratio (PSNR) result, indicating that GCCAN has stronger learning ability. In addition, the GCCAN algorithm model showed high stability during the iteration process because it adopted a global context channel attention mechanism, which can effectively capture and utilize features at different levels, thereby providing more accurate detail recovery in super-resolution tasks. The above results show that compared with the RCAN model, the GCCAN model exhibits more outstanding performance during the training process, verifying the effectiveness and potential of the algorithm.
[0059] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention.
Claims
1. A method for infrared image super-resolution based on global contextual channel attention, characterized in that: The method comprises the following steps: Step 1: Get the training data set; Step 2: Expand the training data set; Step 3: Construct a global context channel attention network (GCCAN). The GCCAN network is based on the RCAN network architecture. While keeping the overall architecture of the RCAN network unchanged, it improves the residual group structure in the deep feature extraction stage. In the residual group structure, m global context channel attention blocks (GCCAB) connected in series are integrated; Step 4: Construct structural loss function; Step 5: Input the expanded training data set obtained in step 2 into the global context channel attention network constructed in step 3, and use the structural loss function constructed in step 4 for training as the optimization target for training. Use the ADAM optimizer to update the model parameters until the loss no longer decreases, and obtain a trained infrared image super-resolution reconstruction model. Step 6: Use the infrared image super-resolution reconstruction model trained in step 5 to perform super-resolution reconstruction on the low-resolution images in the test set to obtain infrared super-resolution images.
2. The infrared image super-resolution method based on global context channel attention according to claim 1, characterized in that: The acquisition of the training data set specifically includes: obtaining original high-resolution images from a public database or personal collection, and obtaining the high-resolution images through I LR =D w B l I HR +n o The image degradation model described is degraded to obtain, where I HR represents the original high-resolution infrared image, B l represents fuzzy interference, D w represents the downsampling process of high-resolution images, n o represents the noise in the imaging process, I LR For the collected low-resolution infrared images, a Gaussian blur kernel with a standard deviation of 1 is used in the degradation model, Bicubic interpolation downsampling is used, and additive Gaussian noise is used for noise. Finally, these one-to-one corresponding sets of high- and low-resolution image pairs are used as training data sets.
3. The infrared image super-resolution method based on global context channel attention according to claim 1 or 2, characterized in that: The expanding the training data set specifically includes: adopting a data enhancement strategy for the acquired training data set, that is, randomly horizontally flipping or rotating the data set images by 90°, 180°, and 270°.
4. The infrared image super-resolution method based on global context channel attention according to claim 3 is characterized in that: The global context channel attention network (GCCAN) specifically includes: the GCCAN includes three parts: shallow feature extraction, deep feature extraction and image reconstruction. In the shallow feature extraction stage, the model uses a convolution operation to process a single-channel infrared image with an input size of H×W×1, so that it is mapped into a feature map with a deeper channel; in the GCCAN, the channel depth after mapping is set to 64, and then the deep feature extraction stage includes n residual groups connected in series and a convolution operation at the tail end, and the image reconstruction part operation is sub-pixel convolution upsampling, which is used to increase the extracted feature map to the target resolution; Among them, the GCCAN network is based on the RCAN network architecture. While keeping the overall architecture of the RCAN network unchanged, it improves the residual group structure in the deep feature extraction stage. In the residual group structure, m interconnected global context channel attention blocks (GCCAB) are integrated. In the GCCAN model, n and m are both adjustable hyperparameters, and n=m=10 is set.
5. The infrared image super-resolution method based on global context channel attention according to claim 4, characterized in that: The global context channel attention module (GCCAB) specifically includes: GCCAB is based on the residual channel attention module of the RCAN algorithm, optimizes the mechanism of calculating channel attention, introduces global context channel attention, and embeds parameter-free channel shift operations to expand the receptive field of the network.
6. The infrared image super-resolution method based on global context channel attention according to claim 5, characterized in that: The global context channel attention specifically includes: for a feature map F with an input size of H×W×C, the global context channel attention uses a 1×1 convolution to generate a spatial self-attention map S shared by all channels, and then deforms the spatial self-attention map S from H×W×1 to 1×1×HW, and then, through Generate spatial self-attention map of probability distribution Among them, a represents the original value in S, The elements in the spatial self-attention map S representing the generated probability distribution; Get the spatial self-attention map of the probability distribution After that, the feature map F is transformed into HW×C size and multiplied by its matrix, that is, according to Calculate the channel descriptor tensor Z, where F′ represents the deformed feature map, which is a tensor of size HW×C. is the spatial self-attention map of the probability distribution, with a size of 1×1×HW, and Z is the final channel descriptor tensor, with a size of 1×1×C; After extracting the descriptors of each channel, the interdependencies between channels are captured.
7. The infrared image super-resolution method based on global context channel attention according to claim 6 is characterized in that: The channel shift operation specifically includes: shifting the input feature map F of size H×W×C by Move the four partial channels up, down, left, and right by a unit distance in the width and height dimensions of the space, where γ is the ratio of the moved channels. It is the feature map after channel shift.
8. The infrared image super-resolution method based on global context channel attention according to claim 7, characterized in that: The structural loss function is constructed, specifically including: using the L1 loss function: And the structural loss function: The weighted sum is used as the loss function when training the network: Loss = L L1 (I HR ,I G )+λL structure (I HR ,I G ); Where W and H represent the width and height of the image respectively, x ij Represents the pixel value of the reconstructed image, y ij Represents the pixel value of the target image, i and j are the indexes of the image, I HR and I G They represent the reconstructed high-resolution infrared image and the real high-resolution infrared image respectively, E represents the edge information image, i and j represent the coordinate index, W and H represent the width and height of the image respectively, and L structure (·) is the calculated structural loss function.
Citation Information
Patent Citations
Medical image super-resolution reconstruction method based on multi-attention residual feature fusion
CN113298717A
Super-resolution image reconstruction method based on multi-parallax attention module combination
CN113538243A
Light-weight multi-scale infrared image super-resolution reconstruction method
CN114092330A
Lightweight image super-resolution reconstruction method based on double attention mechanism
CN115496658A
Lightweight infrared super-resolution adaptive reconstruction method
CN116402679A