An infrared image super-resolution reconstruction system and method integrating edge information

By generating adversarial network models to fuse edge information, the problems of low image resolution and unclear edges of infrared imaging equipment are solved, and high-resolution reconstruction of infrared images and improved visual effects are achieved.

CN114782254BActive Publication Date: 2025-07-29JIANGXI NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210534264.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-17
Publication Date
2025-07-29
Estimated Expiration
2042-05-17

AI Technical Summary

Technical Problem

The images captured by existing infrared imaging devices have low resolution, poor visual effects after reconstruction, poor edges, and lack sharp edge details.

Method used

The generative adversarial network model is adopted, combining edge detection networks and image super-resolution reconstruction networks, and the edge information is fused through the edge feature processing module and the deep feature extraction module. The image super-resolution reconstruction is performed using the generative adversarial network model, adding the fusion feature map of edge detection results and deep feature extraction results, and information integration is carried out through sub-pixel convolutional layer and convolutional layer.

Benefits of technology

The reconstructed infrared image has sharper edge details, better visual effects, and improves image resolution and clarity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114782254B_ABST
    Figure CN114782254B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of image processing, and relates to an infrared image super-resolution reconstruction system and method that fuses edge information. The system includes a generative adversarial network model, which is composed of an edge detection network, several edge feature processing modules, and an image super-resolution reconstruction network. The edge detection network includes several convolutional stages; the image super-resolution reconstruction network includes several hierarchical depth feature extraction modules. The present invention uses the edge detection network to extract the edge features of the image, and integrates the trained edge detection network into the super-resolution reconstruction network to jointly form the generative adversarial network model. The present invention enables the reconstructed high-resolution infrared image to have sharper edge detail information and better visual effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and in particular relates to an infrared image super-resolution reconstruction system and method that fuses edge information. Background Art

[0002] Infrared imaging systems require no auxiliary lighting to complete their images, functioning even in dim light sources or in complete darkness. Furthermore, the wavelength of visible light ranges from approximately 380nm to 760nm, while that of infrared light ranges from approximately 760nm to 1000nm. Longer wavelengths increase diffraction resistance, making infrared light more resistant to interference than visible light. Infrared imaging systems can still effectively detect objects in adverse weather conditions such as rain, snow, fog, and haze, compared to visible light. Due to these advantages, infrared imaging technology has been widely used in numerous fields, such as the military, where it enables precise target location in complex scenarios; in the power sector, where it can be used to detect line faults; and in security monitoring, where it plays a vital role in ensuring the safety of key departments. However, the images captured by infrared imaging devices generally have a low resolution, which greatly hinders their future application and development.

[0003] Currently, the problem of image super-resolution reconstruction has received widespread attention. In particular, image super-resolution reconstruction algorithms based on deep learning have been widely studied in recent years due to their powerful expressive power and have achieved good results. Convolutional neural networks have been widely studied and applied to the problem of image super-resolution reconstruction due to their advantages such as retaining spatial local information and powerful feature extraction capabilities. Convolutional neural networks usually increase the depth of the network model by stacking multiple convolutional layers to extract deeper feature information from the image, thereby increasing the model's expressive power. Although image super-resolution reconstruction algorithms based on convolutional neural networks have achieved good results in terms of objective evaluation indicators, the reconstructed high-resolution images have problems such as poor visual effects, overly smooth images, and unclear edges.

[0004] In an image, if the gradient difference of pixel values between adjacent points is large, it means that the gray value changes violently in this part, and then this area is considered as the edge part of the image; the image edge represents the boundary points where the changes in various parts of the figure are very obvious, which contains a large amount of structural information and has a great impact on the visual perception of the human eye. Therefore, edge repair is often carried out in image super-resolution reconstruction. For example, CN112288632A, a single-image super-resolution method based on a streamlined ESRGAN, includes the following steps: Step S1: Obtain the low-resolution image to be processed and preprocess it; Step S2: According to the preprocessed image, generate a super-resolution image through the generator module in the improved single-image super-resolution generative adversarial network. If the model is in the training stage, go to Step S3, otherwise go to Step S4; Step S3: Construct a discriminator and use the discriminator to judge whether the super-resolution image is a real high-resolution image. According to the result obtained by the discriminator, perform backpropagation to optimize the generator and repeat Step S2; Step S4: Perform edge repair processing on the obtained super-resolution image to obtain the final super-resolution image. CN111062872A discloses a method and system for image super-resolution reconstruction based on edge detection. Edge extraction uses the method of converting the image to the YCbCr color space to extract the image on the Y component, which retains the image texture information while reducing the size of the image, reducing the computational amount, and using the Sobel operator to perform edge extraction on the image on the Y component, enabling the super-resolution reconstruction network to effectively learn edge detail information. Summary of the Invention

[0005] In order to make the reconstructed image have richer edge detail information and better visual effects, the present invention provides a system and method for infrared image super-resolution reconstruction that integrates edge information, using edge information to assist in the super-resolution reconstruction of the image, so that the reconstructed image has better visual effects.

[0006] An infrared image super-resolution reconstruction system integrating edge information includes a generative adversarial network model. The generative adversarial network model consists of an edge detection network, several edge feature processing modules, and an image super-resolution reconstruction network. The edge detection network includes several convolutional stages; the image super-resolution reconstruction network includes several hierarchical depth feature extraction modules. The image edge feature maps output by each convolutional stage of the edge detection network are respectively processed by the edge feature processing modules and then fused with the feature maps extracted by the corresponding hierarchical depth feature extraction modules to obtain fused feature maps. The image edge feature maps of all convolutional stages of the edge detection network are concatenated together, and after convolutional processing, an image edge detection result is obtained. The image edge detection result is processed by the edge feature processing module and then superimposed on the feature map extracted by the depth feature extraction module of the last layer to obtain the last fused feature map. After the fused feature maps are successively convolved and integrated in reverse order, they are then processed by a sub-pixel convolutional layer and a convolutional layer and output.

[0007] The edge detection network includes five convolutional stages: a first convolutional stage, a second convolutional stage, a third convolutional stage, a fourth convolutional stage, and a fifth convolutional stage.

[0008] The image super-resolution reconstruction network is composed of a convolutional layer at the input end, a first depth feature extraction module, a second depth feature extraction module, a third depth feature extraction module, a fourth depth feature extraction module, a fifth depth feature extraction module, a sixth depth feature extraction module, multiple convolutional layers for successively implementing feature information integration, a sub-pixel convolutional layer, and a convolutional layer at the output end, which are arranged in sequence.

[0009] The edge feature processing module is composed of multiple convolution-activation modules, and the activation function used in each is ReLU.

[0010] The depth feature extraction module contains several dense residual modules.

[0011] In the dense residual module, several dense modules are stacked. The feature information of the input and output is fused through residual connections before and after each dense module, and finally, a residual connection that fuses the initial input and the final output of the dense module is added.

[0012] Each dense module includes multiple convolutional layers. Except for the last convolutional layer, each convolutional layer is followed by an activation layer, and the activation function used in each is Leaky ReLU. Then, using the idea of a dense network, the output of each layer is used as additional input for all subsequent layers through dense connections, thereby improving the utilization rate of the output features of each convolutional layer.

[0013] The image edge feature map output by the first convolution stage is processed by the edge feature processing module and then superimposed with the feature map extracted by the first depth feature extraction module to obtain the fused feature map X1. The image edge feature map output by the second convolution stage is processed by the edge feature processing module and then superimposed with the feature map extracted by the second depth feature extraction module to obtain the fused feature map X2. The image edge feature map output by the third convolution stage is processed by the edge feature processing module and then superimposed with the feature map extracted by the third depth feature extraction module to obtain the fused feature map X3. The image edge feature map output by the fourth convolution stage is processed by the edge feature processing module and then superimposed with the feature map extracted by the fourth depth feature extraction module to obtain the fused feature map X4. The image edge feature map output by the fifth convolution stage is processed by the edge feature processing module and then superimposed with the feature map extracted by the fifth depth feature extraction module to obtain the fused feature map X5. The image edge feature maps output by the five convolution stages are concatenated together, and finally, after a 1×1 convolution processing, the image edge detection result is obtained. The image edge detection result is processed by the edge feature processing module and then superimposed with the feature map extracted by the sixth depth feature extraction module to obtain the fused feature map X6. The fused feature map X6, fused feature map X5, fused feature map X4, fused feature map X4, fused feature map X3, fused feature map X2, and fused feature map X1 are successively convolutionally integrated through 6 convolutional layers for sequentially implementing feature information integration.

[0014] An infrared image super-resolution reconstruction method for fusing edge information, the steps are as follows:

[0015] Step 1: The infrared image needs to be preprocessed to construct an infrared image training dataset;

[0016] Step 2: In order to incorporate image edge information into the super-resolution reconstruction process, first, an edge detection network needs to be separately trained, the edge detection network and the image super-resolution reconstruction network are mixed, and then a generative adversarial network model is formed to achieve the super-resolution reconstruction of the edge information-assisted image;

[0017] Step 3: The high-resolution infrared image output by the generative adversarial network model is input into the discriminator. The discriminator performs a series of feature extraction operations, and finally outputs a probability according to the learned image features to describe the authenticity of the input image;

[0018] Step 4: Update the parameters of the generative adversarial network model; the update of the parameters of the entire generative adversarial network model is carried out alternately. First, fix the network parameters of the generative adversarial network model and only update the network parameters of the discriminator, and then fix the network parameters of the discriminator and update the generative adversarial network model. The two are alternately trained in sequence;

[0019] Step 5: Input a low-resolution image into the trained generative adversarial network model, and then output the reconstructed high-resolution infrared image.

[0020] Based on the generative adversarial network model, the present invention uses image edge information to assist in the super-resolution reconstruction of infrared images. In order to obtain clear image edge information, the present invention uses an edge detection network to extract the edge features of the image, and integrates the trained edge detection network into the super-resolution reconstruction network to jointly form the generative adversarial network model. In the generative adversarial network model, the image edge features output by five convolutional stages in the edge detection network and the finally output image edge detection results are intercepted. After passing through the edge feature processing module, they are successively superimposed on the feature maps obtained in the super-resolution reconstruction network, and then embedded into the subsequent part of the super-resolution reconstruction network for processing. In addition, in order to make more effective and full use of the rich edge information, a deep residual connection is added behind the feature map fused with the edge information. Combining the edge information from coarser to finer, it is successively input into the convolutional layer at the backend of the super-resolution reconstruction network for information integration, so that the reconstructed high-resolution infrared image has sharper edge detail information and better visual effects. Brief Description of the Drawings

[0021] Figure 1 is the structural diagram of the generative adversarial network model proposed by the present invention.

[0022] Figure 2 is the structural diagram of the edge detection network in the present invention.

[0023] Figure 3 is the structural diagram of the dense residual module unit in the super-resolution reconstruction network model of the present invention.

[0024] Figure 4 is the structural diagram of the discriminator network in the present invention.

[0025] Figure 5 is the image effect diagram obtained by doubling the super-resolution of the present invention.

[0026] Figure 6 is the image effect diagram obtained by quadrupling the super-resolution of the present invention. Detailed Embodiments

[0027] In order to understand the present invention more clearly, the following further detailed description will be given in conjunction with the accompanying drawings and specific embodiments.

[0028] Refer to Figure 1, an infrared image super-resolution reconstruction system integrating edge information, includes a generative adversarial network model. The generative adversarial network model consists of an edge detection network, 6 edge feature processing modules, and an image super-resolution reconstruction network. The edge detection network includes five convolutional stages: the first convolutional stage, the second convolutional stage, the third convolutional stage, the fourth convolutional stage, and the fifth convolutional stage. The image super-resolution reconstruction network is composed of a convolutional layer at one input end, six deep feature extraction modules (deep feature extraction module 1, deep feature extraction module 2, deep feature extraction module 3, deep feature extraction module 4, deep feature extraction module 5, deep feature extraction module 6), 6 convolutional layers for sequentially integrating feature information, a sub-pixel convolutional layer, and convolutional layers at two output ends. The image edge feature map output by the first convolutional stage is processed by the edge feature processing module 1 and then superimposed on the feature map extracted by the deep feature extraction module 1 to obtain the fused feature map X1. The image edge feature map output by the second convolutional stage is processed by the edge feature processing module 2 and then superimposed on the feature map extracted by the deep feature extraction module 2 to obtain the fused feature map X2. The image edge feature map output by the third convolutional stage is processed by the edge feature processing module 3 and then superimposed on the feature map extracted by the deep feature extraction module 3 to obtain the fused feature map X3. The image edge feature map output by the fourth convolutional stage is processed by the edge feature processing module 4 and then superimposed on the feature map extracted by the deep feature extraction module 4 to obtain the fused feature map X4. The image edge feature map output by the fifth convolutional stage is processed by the edge feature processing module 5 and then superimposed on the feature map extracted by the deep feature extraction module 5 to obtain the fused feature map X5. The image edge feature maps output by the five convolutional stages are concatenated together, and finally, after a 1×1 convolution processing, the image edge detection result is obtained. The image edge detection result is processed by the edge feature processing module 6 and then superimposed on the feature map extracted by the deep feature extraction module 6 to obtain the fused feature map X6. The fused feature map X6, fused feature map X5, fused feature map X4, fused feature map X4, fused feature map X3, fused feature map X2, and fused feature map X1 are sequentially convolutionally integrated through 6 convolutional layers for sequentially integrating feature information, and then processed by the sub-pixel convolutional layer and convolutional layers at two output ends and output.

[0029] In a preferred embodiment of the present invention, the edge detection network is as Figure 2As shown, it includes five convolutional stages. The first convolutional stage consists of two 3×3-64 convolutional layers, two 1×1-21 convolutional layers, and one 1×1 convolutional layer. The 3×3-64 convolutional layer is a convolutional layer with 64 convolutional kernels and a size of 3×3. The 1×1-21 convolutional layer is a convolutional layer with 21 convolutional kernels and a size of 1×1. The outputs of the two 3×3-64 convolutional layers respectively enter the two 1×1-21 convolutional layers for dimension increasing operations. After the features of the outputs of the two 1×1-21 convolutional layers are superimposed, they enter the 1×1 convolutional layer for dimension decreasing processing, thereby obtaining the image edge feature map output by the first convolutional stage. The feature maps obtained by convolving the two 3×3-64 convolutional layers are input into the second convolutional stage.

[0030] The second convolutional stage consists of two 3×3-128 convolutional layers, two 1×1-21 convolutional layers, one 1×1 convolutional layer, and one deconvolutional layer. The 3×3-128 convolutional layer is a convolutional layer with 128 convolutional kernels and a size of 3×3. The outputs of the two 3×3-128 convolutional layers respectively enter the two 1×1-21 convolutional layers for dimension increasing operations. After the features of the outputs of the two 1×1-21 convolutional layers are superimposed, they enter the 1×1 convolutional layer for dimension decreasing processing, and then enter the deconvolutional layer. The deconvolutional layer maps the image size back to the original size, and finally outputs the image edge feature map of the second convolutional stage.

[0031] The third convolutional stage consists of three 3×3-256 convolutional layers, three 1×1-21 convolutional layers, one 1×1 convolutional layer, and one deconvolutional layer. The 3×3-256 convolutional layer is a convolutional layer with 256 convolutional kernels and a size of 3×3. The outputs of the three 3×3-256 convolutional layers respectively enter the three 1×1-21 convolutional layers for dimension increasing operations. After the features of the outputs of the three 1×1-21 convolutional layers are superimposed, they enter the 1×1 convolutional layer for dimension decreasing processing, and then enter the deconvolutional layer. The deconvolutional layer maps the image size back to the original size, and finally outputs the image edge feature map of the third convolutional stage.

[0032] The fourth convolutional stage consists of three 3×3-512 convolutional layers, three 1×1-21 convolutional layers, one 1×1 convolutional layer, and one deconvolutional layer. The 3×3-512 convolutional layer is a convolutional layer with 512 convolutional kernels and a size of 3×3. The outputs of the three 3×3-512 convolutional layers respectively enter the three 1×1-21 convolutional layers for dimension increasing operations. After the features of the outputs of the three 1×1-21 convolutional layers are superimposed, they enter the 1×1 convolutional layer for dimension decreasing processing, and then enter the deconvolutional layer. The deconvolutional layer maps the image size back to the original size, and finally outputs the image edge feature map of the fourth convolutional stage.

[0033] The fifth convolutional stage consists of a first group of three 3×3-512 convolutional layers, a second group of three 3×3-512 convolutional layers, a 1×1 convolutional layer, and a transposed convolutional layer. The outputs of the first group of three 3×3-512 convolutional layers respectively enter the second group of three 3×3-512 convolutional layers for dimension increasing operations. After the features of the outputs of the second group of three 3×3-512 convolutional layers are stacked, they enter the 1×1 convolutional layer for dimension decreasing processing, and then enter the transposed convolutional layer. The transposed convolutional layer maps the image size back to the original size, and finally the image edge feature map of the fifth convolutional stage is obtained as the output.

[0034] The image edge feature maps output by the five convolutional stages are concatenated together, and finally the image edge detection result is obtained after a 1×1 convolutional processing.

[0035] In a preferred embodiment of the present invention, the edge feature processing module is composed of 4 convolutional-activation modules. The activation function used is ReLU in all cases, mainly used to adjust the number of channels of the image edge features, so as to fuse the edge information into the process of super-resolution reconstruction.

[0036] In a preferred embodiment of the present invention, the deep feature extraction module includes 4 dense residual modules. The structure of the dense residual module is as Figure 3 shown. In the dense residual module, 3 dense modules are stacked. Before and after each dense module, the feature information of the input and output is fused through residual connections, and finally a residual connection that fuses the initial input and the final output of the dense module is added. Each dense module is mainly composed of 5 convolutional layers. Except for the last convolutional layer, each convolutional layer is followed by an activation layer. The activation function used is Leaky ReLU in all cases. Then, using the idea of a dense network, the output of each layer is used as additional input for all subsequent layers through dense connections, thereby improving the utilization rate of the output features of each convolutional layer.

[0037] In the generative adversarial network model of the present invention, the input image passes through an image super-resolution reconstruction network and an edge detection network respectively. In the image super-resolution reconstruction network, the input image first passes through a convolutional layer and a deep feature extraction module 1, and then intercepts the output features of the input image in the first stage of the edge detection network. After passing through an edge feature processing module 1, the edge feature information extracted by the edge feature processing module 1 is fused with the feature map output by the deep feature extraction module 1 in the image super-resolution reconstruction network part, and a residual connection is added after the fused feature map, and then it is input into the convolutional layer part at the back end of the network for information integration. Subsequently, the edge feature information output by the edge feature processing modules 2 to 6 is successively added to the feature maps output by the deep feature extraction modules 2 to 6, and the fused feature maps are input into the convolutional layer part at the back end of the network through residual connections for information integration. In this way, rich image edge information is introduced into the image super-resolution reconstruction process through this hybrid network model.

[0038] In the generative adversarial network model of the present invention, since rich edge information is introduced in the super-resolution reconstruction process, it also increases the difficulty of model training to a certain extent. Therefore, in order to ensure the stability of model training, a deep residual connection is added in the middle of the super-resolution reconstruction network, and the features that incorporate edge information and have passed through the edge feature processing module are supplemented to the latter half of the network, and then the integration of feature information is achieved successively through 6 convolutional layers. At the end of the super-resolution reconstruction network, the present invention uses sub-pixel convolution to increase the image resolution. The feature map integrated through convolution is first passed through a convolutional layer to increase the number of feature maps, and then input into the sub-pixel convolutional layer. Through the way of channel shuffling, the pixels on all feature maps are rearranged, so as to increase the image resolution. Finally, two convolutional layers are passed through to control the number of channels of the output image and ensure the consistency of the number of channels of the input image and the output image.

[0039] Further preferably, the infrared image super-resolution reconstruction system integrating edge information of the present invention further includes a discriminator. The discriminator is mainly used to distinguish whether the input image comes from real data or simulated data output by the generative adversarial network model. The network structure of the discriminator is as Figure 4 shown. The input of the discriminator network is an image. Through a series of operations of feature extraction, and finally a probability is output according to the learned image features to describe the authenticity of the input image.

[0040] This embodiment provides an infrared image super-resolution reconstruction method integrating edge information, and the steps are as follows:

[0041] Step 1, the infrared image needs to be preprocessed to construct an infrared image training dataset.

[0042] In this embodiment, two infrared image datasets with resolution sizes of 640×512 and 640×480 respectively are adopted. Then, by mixing the two infrared image datasets, the training set, validation set, and test set are re-divided. Then, the high-resolution infrared image is downsampled by bicubic interpolation to obtain the corresponding low-resolution infrared image. Next, the training data is augmented by methods such as horizontal flipping, vertical flipping, and rotation, and the infrared image is cropped into sub-images of size k×k for input into the method of the present invention for training. It is recommended that k be 192.

[0043] Step 2: In order to incorporate image edge information into the super-resolution reconstruction process, first, an edge detection network needs to be trained separately. The edge detection network and the image super-resolution reconstruction network are mixed, and then a generative adversarial network model is formed to achieve edge information-assisted super-resolution reconstruction of the image.

[0044] For the input low-resolution infrared image l, on the one hand, after passing through a convolutional layer and the first deep feature extraction module of the image super-resolution reconstruction network, the feature map b1 is obtained. On the other hand, the image edge feature map r1 is output at the first convolutional stage of the edge detection network:

[0045] b1 = T1(f(l)) (1)

[0046] r1 = S1(l) (2)

[0047] X1 = b1 + P(r1) (3)

[0048] Where T1 and S1 respectively represent the operation sets of the first deep feature extraction module and the first convolutional stage of the edge detection network, and P represents the edge feature processing module for processing the output feature information of the edge detection network. The image edge feature map r1 after feature transformation is superimposed on the feature map r1 output by the deep feature extraction module 1 to obtain the fused feature map X1. Then, the fused feature map X1 is used as the input of the second deep feature extraction module. Subsequently, the subsequent processing process is the same as above, and the image edge feature maps of each stage are incorporated into the results of the deep feature extraction module. The obtained fused feature maps X1~X6 contain rich image edge information. Subsequently, the information is integrated through the six convolutional layers in the latter half of the image super-resolution reconstruction network in turn:

[0049] Y1 = f(X6) (4)

[0050] The above formula means that the feature map X6 is input into the convolutional layer for processing to obtain Y1, and then X5 to X1 are sequentially passed into the convolutional layer from back to front for feature transformation:

[0051] Y i= f(Y i-1 + X 7-i ) (5)

[0052] where the value range of i is from 2 to 6. Since the image super-resolution reconstruction network contains 6 depth feature extraction modules, 6 convolutional layers are also used to achieve information integration in the second half of the depth feature extraction module. Finally, the high-resolution infrared image h is obtained through the model reconstruction part:

[0053] h = f(f(UP(Y6))) (6)

[0054] where Y6 represents the output of the sixth convolutional layer, and UP represents the upsampling operation.

[0055] Step 3: Input the high-resolution infrared image output by the generative adversarial network model into the discriminator. The discriminator performs a series of feature extraction operations, and finally outputs a probability based on the learned image features to describe the authenticity of the input image. The main process is as follows:

[0056] The input image x first passes through two stacked convolutional layer-activation layer processes to obtain the feature map d:

[0057] d = F(F(x)) (7)

[0058] where F represents the set of convolutional and activation operations, and the Leaky ReLU activation function is used in all activation layers. The middle part of the discriminator network consists of 8 convolutional modules. The structure of each convolutional module is a convolutional layer - BN layer - activation layer, and the Leaky ReLU activation function is also used in this activation layer. Then, the obtained feature map is unfolded into a one-dimensional space and input into the fully connected layer, and then passes through an activation layer and a fully connected layer to output the probability value:

[0059] g = H8(H7(...H1(d))) (8)

[0060] p = C(σ(C(flatten(g)))) (9)

[0061] where flatten represents the operation of flattening the feature map, g represents the feature map extracted through a series of convolutional modules, p represents the probability finally output by the discriminator, H represents the processing process of the convolutional module, the subscript of H is the number of the convolutional module, C represents the fully connected layer, and σ represents the activation function, and the Leaky ReLU activation function is used for this activation function.

[0062] Step 4: Update the parameters of the generative adversarial network model.

[0063] The parameter update of the entire generative adversarial network model is carried out alternately. First, the network parameters of the generative adversarial network model are fixed, and only the discriminator network parameters are updated. Then, the discriminator network parameters are fixed and the generative adversarial network model is updated. The two are alternately trained in turn. The loss function adopted in the training of the present invention is as follows:

[0064] L = L percep + λL G + ηL1 (10)

[0065] L G = -E t [log(1 - D(t, f))] - E f [log(D(f, t))] (11)

[0066]

[0067] Among them, L is the loss function of the generative adversarial network model. This loss contains three parts, namely the perceptual loss L percep , the adversarial loss L G and the pixel loss L1. λ and η are the weight coefficients of the adversarial loss and the pixel loss in the total loss respectively. In the present invention, it is recommended that λ takes the value of 0.005, η takes the value of 0.01, E t represents taking the average value of the output results of this batch after inputting a batch of real images into the discriminator, E f represents taking the average value of the output results of this batch after inputting a batch of generated images into the discriminator. D(t, f) represents the probability that the real image is more real than the generated image, D(f, t) represents the probability that the generated image is less likely to be judged as real than the real image, t represents the real high-resolution infrared image, f represents the high-resolution infrared image reconstructed by the generator, i represents a certain current sample data, m represents the number of sample data, y i represents the real high-resolution infrared image, represents the reconstructed high-resolution infrared image. L percep is to extract features using the VGG network, and then calculate the difference between the feature maps of the generated image and the real image. L G is the loss generated by the discriminator network, which can be used to guide the generator to output more real images. The L1 loss is used to calculate the pixel difference between the network output image and the real image. The generative adversarial network adopts the Adam optimizer to realize parameter update.

[0068] Step five, input the low-resolution image into the trained generative adversarial network model, and then output the reconstructed high-resolution infrared image.

[0069] To prove the effectiveness of the present invention, a variety of super-resolution reconstruction methods were selected for comparison. To ensure the fairness of the comparative experiment results, all methods were trained and tested in the same environment. The present invention combines evaluation metrics and reconstruction effect diagrams to comprehensively illustrate the superiority of the method of the present invention, among which the evaluation metrics NIQE and LPIPS that are more in line with the subjective visual effect were selected.

[0070] NIQE is a no-reference image quality evaluation metric. It uses a set of quality-aware features to calculate the distance between the input image and the features extracted from the original image library, thereby judging the quality of the image. The smaller the value of NIQE, the better the image quality.

[0071] LPIPS calculates the perceptual distance between the real image and the generated image through a trained model:

[0072] where φ(·) l represents the features extracted by the model at the l-th layer, w l represents the parameters of the l-th layer, r represents the real high-resolution infrared image, s represents the reconstructed high-resolution infrared image, h and w are used to normalize the feature map extracted from the current l-th layer, H l and W l represent the feature sizes, used to calculate the average value in the spatial dimension, and finally sum over all channels. The larger the LPIPS value, the greater the gap between the reconstructed image and the real image. On the contrary, the smaller the LPIPS value, the better the quality of the reconstructed image.

[0073] Comparison of the NIQE and LPIPS evaluation metric results of each method for 2x super-resolution:

[0074]

[0075]

[0076] From the experimental results of each method in the above table, it can be seen that among the control methods, the present invention has the lowest values in both NIQE and LPIPS, which means that the images reconstructed by the present invention have the best quality. To intuitively illustrate that the images reconstructed by the present invention have a better visual effect, reconstruction effect diagrams were selected from the experimental results for a direct visual effect comparison. Figure 5 and Figure 6 are the reconstruction effect diagrams of each method shown. It can be seen from Figure 5 that only the images reconstructed by the method of the present invention can more accurately restore the texture information of the leaf part in the image, and the other comparison methods are prone to the situation where the leaves are intertwined with each other. InFigure 6 It can be seen that the image reconstructed by the method of the present invention can effectively restore the trunk part of the small tree.

Claims

1. An infrared image super-resolution reconstruction system that fuses edge information, including a generative adversarial network model, characterized in that, The generative adversarial network model is composed of an edge detection network, several edge feature processing modules, and an image super-resolution reconstruction network. The edge detection network includes several convolutional stages; the image super-resolution reconstruction network includes several hierarchical deep feature extraction modules. The image edge feature maps output by each convolutional stage of the edge detection network are respectively processed by the edge feature processing modules and then fused with the feature maps extracted by the corresponding hierarchical deep feature extraction modules to obtain fused feature maps. The image edge feature maps of all convolutional stages of the edge detection network are concatenated together, and after convolutional processing, an image edge detection result is obtained. The image edge detection result is processed by the edge feature processing module and then superimposed on the feature map extracted by the last hierarchical deep feature extraction module to obtain the last fused feature map. After the fused feature maps are sequentially convolved and integrated in reverse order, they are then processed by a sub-pixel convolutional layer and a convolutional layer and output; The edge detection network includes five convolutional stages: the first convolutional stage, the second convolutional stage, the third convolutional stage, the fourth convolutional stage, and the fifth convolutional stage. The image edge feature map output by the first convolutional stage is processed by the edge feature processing module and then superimposed on the feature map extracted by the first deep feature extraction module to obtain the fused feature map X1. The image edge feature map output by the second convolutional stage is processed by the edge feature processing module and then superimposed on the feature map extracted by the second deep feature extraction module to obtain the fused feature map X2. The image edge feature map output by the third convolutional stage is processed by the edge feature processing module and then superimposed on the feature map extracted by the third deep feature extraction module to obtain the fused feature map X3. The image edge feature map output by the fourth convolutional stage is processed by the edge feature processing module and then superimposed on the feature map extracted by the fourth deep feature extraction module to obtain the fused feature map X4. The image edge feature map output by the fifth convolutional stage is processed by the edge feature processing module and then superimposed on the feature map extracted by the fifth deep feature extraction module to obtain the fused feature map X5. The image edge feature maps output by the five convolutional stages are concatenated together, and finally, after a 1×1 convolution processing, an image edge detection result is obtained. The image edge detection result is processed by the edge feature processing module and then superimposed on the feature map extracted by the sixth deep feature extraction module to obtain the fused feature map X6. The fused feature map X6, the fused feature map X5, the fused feature map X4, the fused feature map X4, the fused feature map X3, the fused feature map X2, and the fused feature map X1 are sequentially convolved and integrated through 6 convolutional layers for sequentially implementing feature information integration.

2. The infrared image super-resolution reconstruction system integrating edge information according to claim 1, wherein The image super-resolution reconstruction network is composed of a convolutional layer at the input end, a first deep feature extraction module, a second deep feature extraction module, a third deep feature extraction module, a fourth deep feature extraction module, a fifth deep feature extraction module, a sixth deep feature extraction module, multiple convolutional layers for sequentially implementing feature information integration, a sub-pixel convolutional layer, and a convolutional layer at the output end, which are arranged in sequence.

3. An infrared image super-resolution reconstruction system integrating edge information according to claim 1, characterized in that, The edge feature processing module is composed of multiple convolutional-activation modules, and the ReLU activation function is adopted for all of them.

4. An infrared image super-resolution reconstruction system integrating edge information according to claim 1, characterized in that, The deep feature extraction module contains several dense residual modules.

5. The infrared image super-resolution reconstruction system integrating edge information according to claim 4, characterized in that, A number of dense modules are stacked in the dense residual module. Before and after each dense module, the feature information of the input and output is fused through residual connections. Additionally, a residual connection that fuses the initial input and the final output of the dense module is added at the end.

6. The infrared image super-resolution reconstruction system integrating edge information according to claim 5, characterized in that, Each dense module includes multiple convolutional layers. Except for the last convolutional layer, an activation layer follows each convolutional layer. The activation function used is Leaky ReLU. Then, following the idea of a dense network, the output of each layer is used as additional input for all subsequent layers through dense connections, thereby improving the utilization rate of the output features of each convolutional layer.

7. An infrared image super-resolution reconstruction method that fuses edge information, characterized in that, The steps are as follows: Step 1: The infrared image needs to be preprocessed to construct an infrared image training dataset. Step 2: In order to incorporate image edge information into the super-resolution reconstruction process, first, an edge detection network needs to be trained separately. The edge detection network and the image super-resolution reconstruction network are mixed, and then a generative adversarial network model is formed to achieve edge information-assisted super-resolution reconstruction of the image. Step 3: The high-resolution infrared image output by the generative adversarial network model is input into the discriminator. The discriminator performs a series of feature extraction operations and finally outputs a probability based on the learned image features to describe the authenticity of the input image. Step 4: Update the parameters of the generative adversarial network model; the update of the parameters of the entire generative adversarial network model is carried out alternately. First, the parameters of the generative adversarial network model are fixed, and only the parameters of the discriminator network are updated. Then, the parameters of the discriminator are fixed, and the generative adversarial network model is updated. The two are alternately trained in sequence. Step 5: Input the low-resolution image into the trained generative adversarial network model, and then output the reconstructed high-resolution infrared image. For the input low-resolution infrared image On the one hand, after passing through a convolutional layer and the first deep feature extraction module of the image super-resolution reconstruction network, a feature map is obtained On the other hand, the image edge feature map is output in the first convolutional stage of the edge detection network : (1); (2); (3); Among them and respectively represent the operation sets of the first depth feature extraction module and the first convolutional stage of the edge detection network. represents the edge feature processing module used to process the output feature information of the edge detection network, and fuses the image edge feature map after feature transformation with the feature map output by the depth feature extraction module 1 to obtain a fused feature map ; then the fused feature map is used as the input of the second depth feature extraction module, and the subsequent processing process is the same as above. The image edge feature maps of each stage are integrated into the results of the depth feature extraction module, and the obtained fused feature map contains rich image edge information; then it is successively passed through six convolutional layers in the second half of the image super-resolution reconstruction network for information integration: (4); Formula (4) indicates that the feature map is input into the convolutional layer and then is obtained. Then, starting from the back and proceeding one by one, to are passed into the convolutional layer for feature transformation: (5); Among them ranges from 2 to 6. Since the image super-resolution reconstruction network contains 6 deep feature extraction modules, 6 convolutional layers are also used to achieve information integration in the second half of the deep feature extraction module; finally, a high-resolution infrared image is obtained through the model reconstruction part : (6) Among them represents the output of the sixth convolutional layer, represents the upsampling operation.

Citation Information

Patent Citations

  • Image super-resolution reconstruction method and system based on edge detection

    CN111062872A

  • Image super-resolution reconstruction method for generating antagonistic network based on feature fusion

    CN109509152A