Infrared image super-resolution method based on generative adversarial network
Through the infrared image super-resolution method based on the generative adversarial network, combined with deep feature extraction and global residual learning, the problem of unclear details of infrared images is solved, and high-resolution reconstruction of infrared images and the improvement of visual effects is achieved.
Patent Information
- Application Number
- CN202510039784.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-01-10
AI Technical Summary
The spatial resolution of infrared images is usually not as good as that of visible light imaging systems, resulting in unclear details of infrared images and high-precision infrared imaging equipment, which limits its application.
The infrared image super-resolution method based on the generative adversarial network is adopted, and the deep feature extraction module and global residual learning are combined with the dual cubic interpolation technology to optimize the residual module to retain the low-frequency information of the infrared image to realize the reconstruction of high-resolution images.
A good balance is achieved in texture details and edge structure recovery. The generated infrared images are rich in texture, contain more information, and significantly improve the average peak signal-to-noise ratio (PSNR), achieving better visual effects.
Smart Images

Figure CN120107067A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of super-resolution in the field of computer vision, and in particular to an infrared image super-resolution method using a generative adversarial network. Background Art
[0002] Infrared imaging technology is an advanced imaging method that can capture infrared radiation signals that cannot be directly perceived by the human eye. These signals are emitted by objects and are usually related to the temperature of the object. Through special sensors, these infrared radiation signals are converted into photoelectric signals that can be processed by electronic equipment. Subsequently, these signals are processed by complex algorithms and converted into images that our naked eyes can recognize. The core advantage of this technology is its excellent anti-interference ability, which can work under harsh lighting conditions and provide clear images even in completely dark environments. In addition, infrared imaging technology also has the ability to penetrate certain materials, such as smoke and mist, and the ability to detect targets at a longer distance.
[0003] In the military field, infrared imaging technology is used in night vision equipment, thermal imaging reconnaissance, missile guidance and target tracking. It can help soldiers conduct combat and reconnaissance at night or in low visibility conditions, and improve battlefield situation awareness. In the civilian field, this technology is also widely used in security monitoring, industrial testing, medical imaging and environmental monitoring. For example, in industrial testing, infrared imaging can help detect equipment overheating and prevent malfunctions and accidents; in the medical field, it can be used to observe the temperature distribution of the human body and assist in the diagnosis of certain diseases.
[0004] Despite significant advances in infrared imaging, it still faces some challenges. The main problem is that the spatial resolution of infrared images is usually not as good as that of visible light imaging systems. This means that the details of infrared images are not as clear as those of visible light images. In addition, high-precision infrared imaging equipment is often expensive, which limits its application in certain fields.
[0005] Therefore, it is particularly urgent to improve the resolution of infrared images and develop infrared image super-resolution reconstruction technology. The goal of this technology is to improve the resolution of infrared images so that they can provide more detailed information. Through algorithm optimization and hardware improvement, infrared image super-resolution reconstruction technology is expected to significantly improve the performance of infrared imaging, reduce costs, and expand its application in various fields. This is of great significance for improving the application value of infrared imaging technology in military and civilian fields, and also opens up new possibilities for future technological development. Summary of the invention
[0006] To solve the above problems, the present invention proposes an infrared image super-resolution method based on a generative adversarial network, which deeply mines the original data in the low-resolution infrared image, optimizes the residual module, and combines it with the bicubic interpolation technology to retain the low-frequency information in the infrared image, thereby achieving a balance between clear visual effects and obtaining a higher peak signal-to-noise ratio (PSNR).
[0007] Technical solution, in order to achieve the above-mentioned purpose, the present invention proposes an infrared image super-resolution method based on a generative adversarial network, the method comprising the following steps:
[0008] S1, input the infrared low-resolution image into the model, and use the bicubic interpolation method to upsample the input image by four times; the low-resolution image pixel is 480×270;
[0009] S2, the input image enters the shallow feature extraction module of the generative network, and the low-resolution image features are initially extracted by expanding the number of channels;
[0010] S3, the extracted shallow image features are fed into the model deep feature extraction module, and the deep image features are output;
[0011] S4, adding the extracted shallow image features to the deep image features pixel by pixel to perform global residual learning;
[0012] S5, sending the final learned features to the upsampling module, enlarging and reducing the feature resolution to four channels, and obtaining the output image of the generated model;
[0013] S6, adding the image features obtained in S5 to the image obtained by upsampling in S1 pixel by pixel to obtain the final output result of the generative network;
[0014] S7. Input the original high-resolution image and the output image of the generative model into the discriminant network to determine whether the two images are consistent. If not, continue to adjust the parameters of the generative model and continue to iterate; if they are consistent, the model ends the iteration, and the final output is the output image of the generative model, and the high-resolution image is 1920×1080.
[0015] Furthermore, the implementation steps of S3 include:
[0016] The deep feature extraction module consists of 16 residual dense modules, each of which has the same structure and consists of three residual modules.
[0017] The input feature x1 first enters the first residual module, performs multi-scale feature extraction on the input, adds the image features extracted at different scales pixel by pixel, and then performs multiple feature extractions before sending them to the channel attention mechanism module to obtain the output x2. Finally, x2 is weighted and added to x2 to form the output rb1 of the first residual module.
[0018] Input rb1 to the second residual module, use local feature fusion and local residual learning to extract deep features, and fuse residuals at different levels to obtain better super-resolution effects. At the same time, the training process can be made more stable. Finally, the same weighted superposition method as the first residual module is used to obtain the output rb2 of the second residual module.
[0019] Input rb2 to the third residual module to obtain output rb3, the structure of this residual module is the same as the second residual module structure;
[0020] The weighted rb3 is superimposed with x1 to obtain the final output rbn1 of the residual dense module;
[0021] According to the above method, rbn1 is input into the second residual dense module to obtain rbn2, and so on to obtain rbn3 to rbn16. rbn1~rbn16 are merged according to the second dimension to obtain the output of the deep feature extraction module.
[0022] Furthermore, the discriminant network implementation steps involved in S5 include:
[0023] The number of image feature channels after global residual learning is 64. The sub-pixel convolution function PixelShuffle provided by pytorch is used for quadruple upsampling to reduce the number of channels to 4. PixelShuffle is a method of obtaining a high-resolution feature map by convolution and multi-channel recombination of low-resolution feature maps. The number of channels of the input feature map is the square of the upsampling multiple. Its core operation is to divide the channels of the input feature map into multiple groups, each containing r 2 channels, where r is the upsampling factor. The channels in each group are rearranged into a high-resolution feature map, where each pixel is upsampled from the original r 2 channels, and the input feature map size is H×W×(r 2 ×C), the output feature map size is (H×r)×(W×r)×C. Where H, W are the height and width of the feature map, respectively, and C is the number of feature channels.
[0024] Furthermore, the implementation steps of S6 include:
[0025] The image features output by S5 are input into a convolution layer with a convolution kernel size of 3×3, and the number of feature channels is adjusted to 3 to make it the same as the number of super-resolution image channels obtained by S1. Then, pixel-by-pixel superposition is performed to obtain the final output result of the generative network.
[0026] It can improve the PSNR value of super-resolution images compared to real high-resolution images, compensate for the low-frequency information of the image, and obtain better visual effects.
[0027] Furthermore, the discriminant network implementation steps involved in S7 include:
[0028] After the image is input into the discriminant network, it first passes through a convolutional layer of size 3x3, which has 64 channels. Then, the image passes through a LeakyReLU activation layer, which can prevent the occurrence of maximum pooling. After that, it passes through seven convolutional layers in sequence. The number of filters in each layer is 64, 128, 128, 256, 256, 512, 512. After processing by these 8 convolutional layers, 512 feature maps are extracted from the image. The classification probability is output through the sigmoid function.
[0029] Furthermore, the multi-scale feature extraction in the residual dense module is defined as:
[0030] The input features are input into two convolutional layers with convolution kernel sizes of 3×3 and 5×5 respectively, and then pass through the LeakyReLU activation layer respectively, and then add pixel by pixel to obtain the required multi-scale features for the subsequent residual convolution layer.
[0031] Furthermore, the channel attention mechanism in the residual dense module is defined as:
[0032] The input feature x is averaged and pooled to obtain a feature map with a resolution of 1×1, and then a convolution layer is used to compress the number of channels to the original number of channels. After the ReLU activation layer, the number of feature channels is expanded to the initial number of channels. The weights y of different channels of x learned by the channel attention mechanism are obtained. By multiplying x and y, the feature x with the channel attention weight can be obtained. ′ .
[0033] Beneficial effects: Compared with the prior art, the present invention has the following beneficial effects:
[0034] (1) The infrared image super-resolution algorithm of the present invention achieves a good balance between texture detail and edge structure restoration, and demonstrates excellent infrared image reconstruction effect;
[0035] (2) The present invention increases the receptive field of the convolution kernel to better capture feature information. When processing the infrared image super-resolution problem, the average peak signal-to-noise ratio (PSNR) exceeds that of other super-resolution technologies, and the generated image has richer texture and contains more information.
[0036] (3) The generation network part of the present invention combines the convolutional neural network with the traditional bicubic interpolation algorithm, fully utilizes the information of the low-frequency area of the infrared image, and adds the mean square error loss to achieve a balance between clear visual effects and obtaining higher objective evaluation indicators. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments are briefly introduced below.
[0038] Figure 1 A schematic diagram of a generated network model according to an embodiment of the present invention.
[0039] Figure 2 Schematic diagram of the discriminant network model.
[0040] Figure 3 Schematic diagram of the dense residual module. DETAILED DESCRIPTION
[0041] In order to more clearly illustrate the technical solution of the present invention, the present invention is further described below in conjunction with the accompanying drawings and embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several variations and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0042] Infrared imaging technology can convert infrared radiation signals emitted by objects that cannot be directly observed by the human eye into photoelectric signals, and further process them into images visible to our naked eyes. This technology has strong anti-interference, penetration ability and long working distance, and has been widely used in military and civilian fields. Although infrared imaging technology has made rapid progress, the spatial resolution of its images is usually not as good as that of visible light imaging systems, and high-precision infrared imaging equipment is expensive. Therefore, it has become particularly urgent to improve the resolution of infrared images and develop infrared image super-resolution reconstruction technology.
[0043] In the field of visible light image enhancement, many super-resolution (SR) technologies have been developed. Early traditional technologies include bilinear interpolation and bicubic interpolation based on interpolation. In recent years, with the rise of deep learning, neural networks have received great attention in the field of image processing due to their excellent data fitting ability. Generative adversarial networks (GANs) achieve excellent reconstruction effects through adversarial learning between generators and discriminators. Compared with visible light, the edges of infrared images are fuzzy, the texture features are fuzzy, and the noise is more obvious.
[0044] Current super-resolution methods all have problems such as insufficient performance and poor reconstruction capabilities. An improved infrared super-resolution method based on a generative adversarial network is proposed. Based on the above ideas, the present invention provides an infrared image super-resolution method based on a generative adversarial network, which uses a dense residual network to deepen the network structure, and introduces a multi-scale feature extraction module, which is combined with an interpolation upsampling method to achieve super-resolution reconstruction of infrared images. The super-resolution effect is significantly improved in terms of objective evaluation indicators, and it also shows good results in dealing with problems such as image noise, blur and artifacts.
[0045] See also Figure 1 and Figure 2 , Figure 1 A flowchart of the generation network part of the infrared image super-resolution method based on a generative adversarial network provided by the present invention, Figure 2 The flowchart of the discriminant network part of the infrared image super-resolution method based on the generative adversarial network provided by the present invention. Figure 1 and Figure 2 As shown, a super-resolution method for infrared images based on a generative adversarial network comprises the following steps:
[0046] Step 1: Input the infrared low-resolution image into the model and enlarge it four times using the bicubic interpolation method.
[0047] Step 2: The infrared low-resolution image is then fed into the shallow module of the generative network to preliminarily extract image features by increasing the number of channels;
[0048] Step 3: Perform deep feature extraction, send the features extracted from the shallow layer to the deep module, and further extract deep image features;
[0049] The deep feature extraction module consists of 16 residual dense modules with the same structure, and each network contains three residual modules. In the first residual module, the initial feature x1 is processed by multi-scale feature extraction and channel attention mechanism to obtain x2, and then the output rb1 is formed by weighted summation. The second residual module receives rb1, extracts deep features through local feature fusion and residual learning, optimizes the super-resolution effect and stabilizes the training process, and finally obtains the output rb2 by the same weighted superposition method. The third residual module receives rb2 and outputs rb3, which is weighted superposed with x1 to obtain the final output rbn1 of the residual dense module. rbn1 is then input into the next residual dense module to obtain rbn2 to rbn16 in sequence, and finally rbn1 to rbn16 are merged along the second dimension to form the final output of the deep feature extraction module. The multi-scale feature extraction module processes input features through two convolutional layers of different sizes (3×3 and 5×5), followed by a LeakyReLU activation layer. The feature map after pixel-by-pixel addition provides a rich feature basis for the subsequent residual convolution layer. The channel attention mechanism reduces the number of channels to 1 / 16 through average pooling and convolution layers, and then expands back to the original number through a ReLU activation layer. The weight y of each channel is calculated, and finally the original feature x is multiplied by the weight y to obtain the weighted feature x that incorporates the channel attention. ′ .
[0050] Step 4: Add shallow and deep features pixel by pixel to achieve global residual learning;
[0051] Step 5: Upsample the features, amplify the learned features through the upsampling module, and convert them into four channels to generate a super-resolution image;
[0052] After global residual learning, the number of channels of image features is expanded to 64. Subsequently, PyTorch's PixelShuffle function is used to perform quadruple upsampling and reduce the number of channels to 4. PixelShuffle converts low-resolution feature maps into high-resolution feature maps through convolution and channel reorganization techniques, where the number of channels of the input feature map is the square of the upsampling factor. Specifically, it divides the channels of the input feature map into several groups, each containing channels squared by the upsampling factor r, and then rearranges these channels to form a high-resolution feature map, where each pixel is converted from the original r to the high-resolution feature map. 2 channels. For example, if the size of the input feature map is H×W×(r 2 ×C), then the size of the output feature map becomes (H×r)×(W×r)×C.
[0053] Step 6: Perform image feature fusion, add the upsampled image features to the previously upsampled image pixel by pixel, and obtain the final output of the generated network;
[0054] The image features output in step 5 are first passed through a 3×3 convolutional layer to adjust the number of channels to 3 to match the number of channels of the super-resolution image obtained in step 1. Then the adjusted features are added pixel by pixel to the image in step 1. The purpose is to improve the PSNR value of the super-resolution image and supplement the low-frequency information of the image to obtain better visual effects.
[0055] Step 7: Input the original high-resolution image and the output image of the generative model into the discriminant network to determine whether the two images are consistent. If not, continue to adjust the parameters of the generative model and continue to iterate; if they are consistent, the model ends the iteration, and the final output is the output image of the generative model, and the high-resolution image is 1920×1080.
[0056] The image is first preprocessed through a 3x3 convolutional layer with 64 channels. Subsequently, a LeakyReLU activation layer is used to prevent the gradient from vanishing. The entire model consists of 8 convolutional layers, with the number of filters starting from 64 and gradually increasing to 512, namely 64, 64, 128, 128, 256, 256, 512, 512. After processing through these layers, the image is decomposed into 512 feature maps. Finally, the classification probability of the image is output through the sigmoid function.
[0057] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art may make various modifications or variations within the scope of the claims, which do not affect the essence of the present invention.
Claims
1. A super-resolution method for infrared images based on a generative adversarial network, characterized in that: The method comprises the following steps: S1, input the infrared low-resolution image into the model, and use the bicubic interpolation method to upsample the input image by four times; the low-resolution image pixel is 480×270; S2, the input image enters the shallow feature extraction module of the generative network, and the low-resolution image features are initially extracted by expanding the number of channels; S3, the extracted shallow image features are fed into the model deep feature extraction module, and the deep image features are output; S4, adding the extracted shallow image features to the deep image features pixel by pixel to perform global residual learning; S5, sending the final learned features to the upsampling module, enlarging and reducing the feature resolution to four channels, and obtaining the output image of the generated model; S6, adding the image features obtained in S5 to the image obtained by upsampling in S1 pixel by pixel to obtain the final output result of the generative network; S7. Input the original high-resolution image and the output image of the generative model into the discriminant network to determine whether the two images are consistent. If not, continue to adjust the parameters of the generative model and continue to iterate; if they are consistent, the model ends the iteration, and the final output is the output image of the generative model, and the high-resolution image is 1920×1080.
2. The infrared image super-resolution method using a generative adversarial network according to claim 1, characterized in that: The implementation steps of S3 include: The deep feature extraction module consists of 16 residual dense modules, each of which has the same structure and consists of three residual modules; The input feature x1 first enters the first residual module, performs multi-scale feature extraction on the input, adds the image features extracted at different scales pixel by pixel, and then performs multiple feature extractions before sending them to the channel attention mechanism module to obtain the output x2. Finally, x2 is weighted and added to x2 to form the output rb1 of the first residual module. Input rb1 to the second residual module, use local feature fusion and local residual learning to extract deep features, fuse the residuals at different levels, and use the same weighted superposition method as the first residual module to obtain the output rb2 of the second residual module; Input rb2 to the third residual module to obtain the output rb3, the structure of this residual module is the same as the second residual module; rb3 is weighted and superimposed with x1 to obtain the final output rbn1 of the residual dense module; According to the above method, rbn1 is input into the second residual dense module to obtain rbn2, and so on to obtain rbn3 to rbn16. rbn1~rbn16 are merged according to the second dimension to obtain the output of the deep feature extraction module.
3. The infrared image super-resolution method using a generative adversarial network according to claim 1, characterized in that: The discriminant network implementation steps involved in S5 include: The number of image feature channels after global residual learning is 64. The sub-pixel convolution function PixelShuffle provided by pytorch is used for quadruple upsampling to reduce the number of channels to 4. PixelShuffle is a method of obtaining a high-resolution feature map by convolution and multi-channel recombination of low-resolution feature maps. The number of channels of the input feature map is the square of the upsampling multiple. It divides the channels of the input feature map into multiple groups, each containing r 2 channels, where r is the upsampling factor. The channels in each group are rearranged into a high-resolution feature map, where each pixel is upsampled from the original r 2 channels, and the input feature map size is H×W×(r 2 ×C), the output feature map size is (H×r)×(W×r)×C, where H and W are the height and width of the feature map respectively, and C is the number of feature channels.
4. The infrared image super-resolution method using a generative adversarial network according to claim 1, characterized in that: The implementation steps of S6 include: The image features output by S5 are input into a convolution layer with a convolution kernel size of 3×3, and the number of feature channels is adjusted to 3 to make it the same as the number of super-resolution image channels obtained by S1. Then, pixel-by-pixel superposition is performed to obtain the final output result of the generative network.
5. The infrared image super-resolution method using a generative adversarial network according to claim 1, characterized in that: The steps of implementing the discriminant network involved in S7 include: After the image is input into the discriminant network, it first passes through a convolutional layer of size 3x3, which has 64 channels. Then, the image passes through a LeakyReLU activation layer; then it passes through seven convolutional layers in sequence, and the number of filters in each layer is: 64, 128, 128, 256, 256, 512, 512. After processing by the above 8 convolutional layers, 512 feature maps are extracted from the image; the classification probability is output through the sigmoid function.
6. The infrared image super-resolution method using a generative adversarial network according to claim 2, characterized in that: The multi-scale feature extraction in the residual dense module is defined as: The input features are input into two convolutional layers with convolution kernel sizes of 3×3 and 5×5 respectively, and then pass through the LeakyReLU activation layer respectively, and then add pixel by pixel to obtain the required multi-scale features for the subsequent residual convolution layer.
7. The infrared image super-resolution method using a generative adversarial network according to claim 2, characterized in that: The channel attention mechanism in the residual dense module is defined as: The input feature x is averaged and pooled to obtain a feature map with a resolution of 1×1, and then a convolution layer is used to compress the number of channels to the original number of channels. After the ReLU activation layer, the number of feature channels is expanded to the initial number of channels. At this point, the weights y of different channels of x learned by the channel attention mechanism are obtained. Multiply x and y to obtain the feature x with the channel attention weight added. ′ .
Citation Information
Patent Citations
Image super-resolution reconstruction method based on multi-scale pyramid network
CN111402128A
Infrared image super-resolution reconstruction method
CN112561799A
Image super-resolution reconstruction method based on attention mechanism and dual-channel network
CN113362223A
Thermal infrared image super-resolution algorithm of generative adversarial network based on multi-structure fusion
CN117372254A
Image super-resolution reconstruction method, terminal equipment and storage medium
CN117575915A
Cited By
Infrared image super-resolution reconstruction method and system based on convolutional neural network
CN121073798A
Infrared Image Super-Resolution Reconstruction Method and System Based on Convolutional Neural Networks
CN121073798B