Image reconstruction method and system based on improved generative adversarial network

By introducing Res2net into the generative adversarial network to improve the multi-scale feature expression capability of residual blocks, building an SRGAN network, solving the problems of poor results and limited performance in the infrared image reconstruction process, and achieving higher-quality infrared image super-resolution reconstruction.

CN119991443APending Publication Date: 2025-05-13KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510144560.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art has poor results and limited performance in the infrared image reconstruction process, making it difficult to effectively restore high-frequency detailed information, resulting in blurred image and complex calculations, poor real-time performance, especially in case of large amplification factors (such as ×4 and ×8).

Method used

Using the improved generative adversarial network (SRGAN) based on Res2net, the multi-scale feature expression capability of residual blocks is improved in Res2net, and the sub-pixel convolution layer is used as the upsampling module to improve the effect of image super-resolution reconstruction.

Benefits of technology

The super-resolution reconstruction effect of infrared images is significantly improved, and the generated images are clearer and have higher resolution, solving the problems of poor results and limited performance in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991443A_ABST
    Figure CN119991443A_ABST
Patent Text Reader

Abstract

The invention discloses an image reconstruction method and system based on an improved generative adversarial network, and is applied to the technical field of image processing, and the method comprises the steps: obtaining an original infrared image data set, and carrying out the preprocessing, and obtaining an image data set; constructing an improved generative adversarial network based on Res2net, and training the generative adversarial network by adopting the image data set; and based on the trained generative adversarial network, performing reconstruction generation on a low-resolution infrared image needing to be reconstructed, and outputting a super-resolution infrared image. In this way, the training method of the SRGAN network is improved, the improved SRGAN network is obtained in combination with the idea of extracting multiple scales through the res2net, and therefore the quality of the super-resolution infrared image reconstruction result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology, and in particular to an image reconstruction method and system based on an improved generative adversarial network. Background Art

[0002] Image super-resolution reconstruction (abbreviated as image super-resolution) technology is a technology that improves image resolution through hardware or software methods. It aims to obtain high-resolution images based on low-resolution images. Without increasing hardware costs, it significantly improves image quality and application effects, and has broad practical value and far-reaching social impact. This technology is very important in many fields, such as digital imaging technology, video coding communication, deep space satellite remote sensing, target recognition analysis, and medical image analysis.

[0003] In the field of infrared imaging technology, image super-resolution technology is particularly important due to the following reasons that limit the existence of infrared technology: ① According to the Rayleigh resolution criterion, the longer wavelength of infrared optical fiber will lead to a decrease in the spatial resolution of the imaging system; ② In the manufacturing process of infrared detectors, it is necessary to balance the sensitivity of the detector and the spatial resolution. A larger pixel size will reduce the number of pixels per unit area, thereby reducing the spatial resolution; ③ In the manufacturing process of infrared detectors, technical problems such as non-uniformity, material defects, and the success rate of indium column interconnection may be encountered. These problems may lead to a decrease in the performance of the imaging system in terms of spatial resolution. Image super-resolution applications can help improve imaging quality and obtain clearer details in imaging results at the same manufacturing cost and detector volume, thereby improving the usability of images and the accuracy of analysis in specific application scenarios.

[0004] Traditional image super-resolution algorithms include reconstruction-based and example-based learning methods. These methods are generally difficult to restore high-frequency detail information, resulting in blurred reconstructed images, complex calculations, and low real-time performance. They are not suitable for large magnification factors (such as ×4 and ×8). When faced with complex real scenes and images with rich structural textures, these methods are often affected by factors such as noise, blur, and degradation, and their performance is limited. To solve these problems, deep learning has been used in image SR in recent years. Today, image SR based on deep learning has gradually become a mainstream method.

[0005] The method based on Generative Adversarial Network (GAN) is a very effective method to achieve image super-resolution reconstruction. It can generate high-quality images with rich details and textures in image super-resolution tasks. This method is particularly suitable for scenarios where high-resolution information needs to be restored from a single low-resolution image. GAN consists of two parts: a generator and a discriminator. Through adversarial training between the generator and the discriminator, the generator generates data that is as realistic as possible, and the discriminator distinguishes between fake data generated by the generator and real data. In the continuous confrontation between the two neural networks, GAN learns to generate data distribution. In the field of image super-resolution, since the data dimension of low-resolution images is much smaller than the reconstructed high-resolution images, the generator in GAN needs to solve the ill-posed problem when applying GAN. Building an efficient GAN generator is the basis for GAN to complete image super-resolution tasks. In the usual GAN ​​structure, feature information extraction and fitting are generally completed directly through multi-layer neural networks, and the multi-scale features of the input image are rarely considered to provide a variety of rich information.

[0006] Therefore, there is an urgent need for a technical solution to improve the resolution of infrared images based on generative adversarial networks. Summary of the invention

[0007] The present invention discloses an image reconstruction method and system based on an improved generative adversarial network. By utilizing the multi-scale feature expression capability of the improved residual block in Res2net, an improved SRGAN network is obtained, which at least solves the technical problems of poor effect and limited performance in the infrared image reconstruction process of the prior art.

[0008] According to a first aspect of the present disclosure, there is provided an image reconstruction method based on an improved generative adversarial network, comprising the following steps:

[0009] Acquire the original infrared image data set and perform preprocessing to obtain an image data set;

[0010] Constructing a generative adversarial network based on the improvement of Res2net, and using the image data set to train the generative adversarial network; wherein the generative adversarial network based on the improvement of Res2net includes a generator and a discriminator, the generator uses Res2net blocks and residual blocks to form a feature extraction module, and uses a sub-pixel convolution layer as an upsampling module; the discriminator includes 8 convolution layers, 2 dense layers and a sigmoid activation function;

[0011] Based on the trained generative adversarial network, the low-resolution infrared image that needs to be reconstructed is reconstructed and generated, and a super-resolution infrared image is output.

[0012] According to the above aspects and any possible implementation manner, an implementation manner is further provided, wherein the process of obtaining the original infrared image data set and preprocessing to obtain the image data set is:

[0013] An original infrared image data set is obtained, and a magnification factor is determined according to a target task. A downsampling operation is performed on the original infrared image data set according to the magnification factor to obtain different low-resolution infrared image data sets corresponding to different magnification factors.

[0014] As described above and any possible implementation method, an implementation method is further provided, wherein the generator includes convolution layer I, activation function PreLU I, Res2net block I, Res2net block II, Res2net block III, residual block I, residual block II, residual block III, pixel reconstruction layer I, and convolution layer II.

[0015] According to the above aspects and any possible implementation, an implementation is further provided, wherein the Res2net block includes a 1×1 convolution, a feature segmentation layer, a feature subset processing layer, a feature splicing layer, and a 1×1 convolution layer;

[0016] Among them, the feature segmentation layer divides the input features into s feature subsets, and the number of channels in each feature subset is 1 / s of the number of input feature channels;

[0017] The feature subset processing layer processes the segmented feature subsets, and directly outputs the first feature subset to the next level; for the second feature subset, the feature subset processing layer includes a 3×3 convolution layer, and then outputs to the next level; for the third to s feature subsets, the feature subset processing layer includes an addition layer and a 3×3 convolution layer with the output results of the s-1 feature subset processing layer, and then outputs to the next level;

[0018] The feature concatenation layer concatenates the s outputs obtained by the feature subset processing layer of the previous level into new features;

[0019] The residual block includes a first convolution layer, a pooling layer, an activation function, a second convolution layer, a pooling layer, and an addition layer; the feature information processed by the first convolution layer is weighted and retained as an input to a subsequent neural network convolution block, and is added to the feature information processed by the second convolution layer to strengthen the call of feature information in the processing process. At the same time, the weight of the intermediate residual feature information in the subsequent processing link is adjusted through the weighting parameter, so that the generator network can reasonably use the intermediate feature information as needed.

[0020] According to the aspects and any possible implementation methods described above, an implementation method is further provided, wherein the discriminator includes convolution layer I, activation function LeakyReLU I, convolution layer II, batch normalization layer BN I, activation function LeakyReLU II, convolution layer III, batch normalization layer BN II, activation function LeakyReLU III, convolution layer IV, batch normalization layer BN III, activation function LeakyReLU IV, convolution layer V, batch normalization layer BN IV, activation function LeakyReLU V, convolution layer VI, batch normalization layer BN V, activation function LeakyReLU VI, convolution layer VII, batch normalization layer BN VI, activation function LeakyReLU VII, convolution layer VIII, batch normalization layer BN VII, activation function LeakyReLU VIII, dense layer I, activation function LeakyReLU VIII, dense layer II, and activation function sigmoid I.

[0021] According to the above aspects and any possible implementation manner, an implementation manner is further provided, wherein the process of training the generative adversarial network to obtain the trained network is:

[0022] Dividing the image data set into a training set, a validation set and a test set, inputting the training set into the generator, reconstructing a super-resolution image, calculating the loss of the generator, and updating the generator weight;

[0023] Inputting the training set into the discriminator, classifying the reconstructed super-resolution image and the original high-resolution image through the discriminator network, calculating the discriminator loss, performing back propagation, and updating the generator weights;

[0024] After reaching the training times, the obtained parameters of the generative adversarial network are saved.

[0025] According to the above aspects and any possible implementation, a further implementation is provided, using the Adam optimization algorithm to update the weight w of the generator G in descending order. G :

[0026]

[0027] in, represents the generator weight w G The gradient of descent, z m Represents the super-resolution reconstructed image I SR The value of the mth pixel in , m = 1, 2, ..., M, M represents the number of pixels, D(G(z m) ) represents the discriminator D judging the super-resolution reconstructed image I SR The mth pixel in the image is the high-resolution image I HR; α represents the learning rate, β1 represents the exponential decay rate of the first-order moment estimate, and β2 represents the exponential decay rate of the second-order moment estimate;

[0028] According to the above aspects and any possible implementation, a further implementation is provided, in which the weight w of the discriminator D is updated in descending order using the Adam optimization algorithm. D :

[0029]

[0030] in, represents the discriminator weight w D The gradient of descent, x m Represents a high-resolution image I HR The value of the mth pixel, D(x m ) represents the discriminator D judging the high-resolution image I HR The mth pixel is the high-resolution image I HR The probability of a pixel in express The descending gradient, Indicates that the discriminator D judges it as a high-resolution image I HR The probability of a pixel in .

[0031] According to the above aspects and any possible implementation, an implementation is further provided, wherein the generator loss of the generative adversarial network is composed of content loss and adversarial loss, that is, the loss of the generator is l SR for:

[0032]

[0033] in, For content loss, is the generation loss; the content loss It consists of image loss, prediction loss and regularization loss weighted, where the image loss is calculated by MSE loss and the prediction loss is defined using VGG loss to describe the reconstructed image I SR and reference image I HR The Euclidean distance between the feature representations of ; the adversarial loss refers to adding the generative component of GAN to the perceptual loss, which is generated by the generative loss Defined by the probability of the discriminator on all training samples;

[0034] The loss of the discriminator is:

[0035]

[0036] in, For the discriminator D, the reference image I HRThe probability of being correctly identified as a reference image, For the discriminator D, the generator generates an image The probability of being identified as a reference image, the discriminator loss l D The overall description of the GAN discriminator's ability to correctly distinguish between reference images and generated images.

[0037] According to a second aspect of the present disclosure, there is provided an image reconstruction system based on an improved generative adversarial network, which is used to implement the image reconstruction method based on the improved generative adversarial network as described in the first aspect, comprising: an image acquisition module, an improved generative adversarial network construction module, and an image reconstruction module;

[0038] The image acquisition module is used to acquire the original infrared image data set and perform preprocessing to obtain the image data set;

[0039] The improved generative adversarial network construction module is used to construct a generative adversarial network based on the improved Res2net, and use the image data set to train the generative adversarial network; wherein the improved generative adversarial network construction module includes a generator generation module and a discriminator generation module;

[0040] The image reconstruction module is used to reconstruct a low-resolution infrared image that needs to be reconstructed based on a trained generative adversarial network, and output a super-resolution infrared image.

[0041] Compared with the prior art, the present invention has the following technical effects:

[0042] In view of the problem of incomplete utilization of multi-scale information features of images in generative adversarial networks, the present invention proposes a generative adversarial network based on the improved Res2net, utilizes the multi-scale feature expression ability of the improved residual block in Res2net, performs block structure replacement in the GAN generator, and utilizes the combined multi-scale features in the generator training process, so that the network can fully learn the multi-scale feature information in the input low-resolution image, thereby enhancing the full utilization of effective information in the process of super-resolution reconstruction of the input image by the network, and improving the effect of finally reconstructing the high-resolution image.

[0043] It should be understood that the contents described in the summary of the invention are not intended to limit the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, among which:

[0045] Figure 1 A schematic diagram of a process of an image reconstruction method based on an improved generative adversarial network according to an embodiment of the present disclosure is shown;

[0046] Figure 2 A schematic diagram of the block structure of a generator Res2net of an image reconstruction method based on an improved generative adversarial network according to an embodiment of the present disclosure is shown;

[0047] Figure 3 A schematic diagram of a residual feature block structure of a generator of an image reconstruction method based on an improved generative adversarial network according to an embodiment of the present disclosure is shown;

[0048] Figure 4 A schematic diagram of the structure of a generator of an image reconstruction method based on an improved generative adversarial network according to an embodiment of the present disclosure is shown;

[0049] Figure 5 A schematic diagram of a discriminator structure of an image reconstruction method based on an improved generative adversarial network according to an embodiment of the present disclosure is shown;

[0050] Figure 6 A schematic diagram of the structure of an image reconstruction system based on an improved generative adversarial network according to an embodiment of the present disclosure is shown;

[0051] Figure 7 A schematic diagram comparing an image reconstruction method based on an improved generative adversarial network according to an embodiment of the present disclosure with image reconstruction results in the prior art is shown. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.

[0053] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0054] Reference Figure 1 As shown, this embodiment provides an image reconstruction method based on an improved generative adversarial network, comprising the following steps:

[0055] S101, collecting infrared images, and preprocessing the infrared images to obtain a processed infrared image data set.

[0056] Specifically, in this embodiment, given the original infrared image data set I HR , a downsampling operation is performed on it, and the downsampling parameters a1, a2, a3... correspond to the super-resolution magnification factors A1, A2, A3... that the network needs to achieve, and different low-resolution infrared image data sets ILR1, ILR2, ILR3 corresponding to different magnification factors are obtained to divide them into training set, verification set and test set.

[0057] S102. Construct an improved generative adversarial network based on Res2net, and use an image dataset to train the generative adversarial network, wherein the generator includes a Res2net block and a residual block, and uses a sub-pixel convolutional layer as an upsampling module, and the discriminator includes 8 convolutional layers, 2 dense layers, and a sigmoid activation function.

[0058] The main framework of the model constructed in this embodiment is a GAN structure. The generator G is responsible for the main function of infrared super-resolution reconstruction, and the discriminator D is responsible for judging the relative authenticity of super-resolution SR images and high-resolution HR images. The discriminator network D and the generator network G are alternately optimized to solve the adversarial minimax problem. The specific training optimization objectives are:

[0059]

[0060] in, For the discriminator D, the reference image I HR The probability of being correctly identified as a reference image, For the discriminator D, the generator generates an image The probability of being identified as a reference image, the discriminator loss l D The overall description of the GAN discriminator's ability to correctly distinguish between reference images and generated images.

[0061] Specifically, in this embodiment, the improved generative adversarial network based on Res2net includes a generator G part and a discriminator D part. The generator uses Res2net blocks and residual blocks to form a feature extraction module, and uses a sub-pixel convolution layer as an upsampling module; Figure 2As shown in the figure, the Res2net block includes a 1×1 convolution, a feature segmentation layer, a feature subset processing layer, a feature splicing layer, and a 1×1 convolution layer; wherein the feature segmentation layer divides the input feature into s feature subsets, and the number of channels of each feature subset is 1 / s of the number of input feature channels; the feature subset processing layer processes the segmented feature subsets, and directly outputs the first feature subset to the next level; for the second feature subset, the feature subset processing layer includes a 3×3 convolution layer, and then outputs it to the next level; for the third to s feature subsets, the feature subset processing layer includes an addition layer and a 3×3 convolution layer with the output results of the s-1 feature subset processing layer, and then outputs it to the next level; the feature splicing layer re-splices the s outputs obtained by the feature subset processing layer of the previous level into new features; as shown in the figure, Figure 3 As shown in the figure, the residual block includes the first convolution layer, the pooling layer, the activation function, the second convolution layer, the pooling layer, and the addition layer; the feature information processed by the first convolution layer is weighted and retained as input to the subsequent neural network convolution block, and is added with the feature information processed by the second convolution layer to strengthen the call of the feature information in the processing process. At the same time, the weight of the intermediate residual feature information in the subsequent processing link is adjusted through the weighting parameter, so that the generator network can reasonably use the intermediate feature information as needed.

[0062] In this embodiment, if Figure 4 As shown, the generator includes convolution layer I, activation function PreLU I, Res2net block I, Res2net block II, Res2net block III, residual block I, residual block II, residual block III, pixel reconstruction layer I, and convolution layer II.

[0063] Furthermore, in this embodiment, the Res2net module is specifically described as follows: the input is subjected to a 1×1 convolution to obtain a feature x0, which is then evenly divided into s feature map subsets, denoted as x i , where i∈{1, 2, .., s}; each feature subset x i It has the same spatial size as the input feature map, but the number of channels is 1 / s of the input feature map; except for x1, each x corresponds to a 3×3 convolution, denoted by K i (),x i After convolution, the output y i =K i (); feature subset x i With K i-1 () outputs are added and input to K i () in; i Written as:

[0064]

[0065] Among them, x i is the feature map subset, Ki () is convolution, y i is the convolution output.

[0066] The discriminator contains 8 convolutional layers with increasing number of 3×3 convolution kernels, using the same architecture as the VGG network, increasing from 64 to 512 kernels; every time the number of features doubles, serial convolution is used to reduce the resolution of the image; the resulting 512 feature maps are sequentially passed through two dense layers and a final sigmoid activation function to obtain a probability for sample classification.

[0067] Specifically, in this embodiment, if Figure 5 As shown in the figure, the discriminator network includes convolution layer I, activation function LeakyReLU I, convolution layer II, batch standardization layer BN I, activation function LeakyReLU II, convolution layer III, batch standardization layer BN II, activation function LeakyReLU III, convolution layer IV, batch standardization layer BN III, activation function LeakyReLU IV, convolution layer V, batch standardization layer BN IV, activation function LeakyReLU V, convolution layer VI, batch standardization layer BN V, activation function LeakyReLU VI, convolution layer VII, batch standardization layer BN VI, activation function LeakyReLU VII, convolution layer VIII, batch standardization layer BN VII, activation function LeakyReLU VIII, dense layer I, activation function LeakyReLU VIII, dense layer II, activation function sigmoid I. Finally, a probability for sample classification is obtained.

[0068] Specifically, in this embodiment, the training process of the network is as follows: during the training process, the generator and the discriminator are trained in turn in each training round, wherein the generator training requires the input of the low-resolution image I in the training set. LR As input data, generate the reconstructed super-resolution image I SR , calculate the generator loss and update the generator weights; then complete the discriminator training, requiring the use of the discriminator network to reconstruct the super-resolution image I SR and high resolution image I HR Perform classification, calculate the discriminator loss, perform backpropagation, and update the generator weights; after completing the required training rounds, save the relevant parameters of the complete GAN network.

[0069] Furthermore, in this embodiment, the Adam optimization algorithm is used to optimize the objective function of the generator G and the discriminator to realize the training process. The specific method is: using the Adam optimization algorithm, the weight w of the generator G is updated in descending order. G :

[0070]

[0071] in, represents the generator weight w G The gradient of descent, z m Represents the super-resolution reconstructed image I SR The value of the mth pixel in , m = 1, 2, ..., M, M represents the number of pixels, D(G(z m )) indicates that the discriminator D judges the super-resolution reconstructed image I SR The mth pixel in the image is the high-resolution image I HR The probability of the pixel in , α represents the learning rate, β1 represents the exponential decay rate of the first-order moment estimate, and β2 represents the exponential decay rate of the second-order moment estimate;

[0072] Use the Adam optimization algorithm to update the weight w of the discriminator D in descending order D :

[0073]

[0074] in, Denotes the discriminator weight w D The gradient of descent, x m Represents a high-resolution image I HR The value of the mth pixel, D(x m ) represents the discriminator D judging the high-resolution image I HR The mth pixel is the high-resolution image I HR The probability of a pixel in express The descending gradient, Indicates that the discriminator D judges it as a high-resolution image I HR The probability of a pixel in .

[0075] Furthermore, in this embodiment, the generator loss of the network consists of content loss and adversarial loss, specifically:

[0076]

[0077] in, For content loss, To generate losses.

[0078] In this embodiment, the content loss It includes: image loss, prediction loss and regularization loss weighted composition, where the image loss is calculated by MSE loss, and the prediction loss is defined using VGG loss to describe the reconstructed image I SR and reference image I HR The Euclidean distance between the feature representations is:

[0079]

[0080] Among them, limage is the image loss, l perception To predict the loss, l TV is the regularization loss;

[0081]

[0082] Among them, l MSE is the MSE loss, W is the number of pixels in the width direction of the reference image and the reconstructed image, H is the number of pixels in the height direction of the reference image and the reconstructed image, and r is the scaling factor between the reference image and the reconstructed image, that is, the value of the magnification factor;

[0083]

[0084] Among them, l VGG / i,j is the VGG loss, where W i,j and H i,j Respectively represent the dimensions of each feature map in the VGG network, with φ i,j represents the feature map obtained by the jth convolution (after activation) before the i-th maximum pooling layer in the VGG19 network,

[0085]

[0086] Among them, l TV The regularization loss representing the smoothness of the output image is calculated for the total variation loss. i,j represents the i,jth pixel value in the output image, x i,j-1 and x i,j-1 It represents the element value adjacent to the i,j-th pixel value in the output image, and β is the smoothing order.

[0087] In this embodiment, the adversarial loss refers to adding the generative component of GAN to the perceptual loss. Defined by the probability of the discriminator on all training samples:

[0088]

[0089] in, Generate images for the generator is the probability of the reference image, using the minimized Make the loss function have better gradient performance.

[0090] Furthermore, in this embodiment, the loss of the discriminator network is defined as:

[0091]

[0092] in, For the discriminator D, the reference image I HRThe probability of being correctly identified as a reference image, For the discriminator D, the generator generates an image The probability of being identified as a reference image.

[0093] S103, reconstructing the low-resolution infrared image that needs to be reconstructed based on the trained generative adversarial network, and outputting a super-resolution infrared image.

[0094] like Figure 6 As shown, this embodiment also provides an image reconstruction system based on an improved generative adversarial network, including: an image acquisition module 1, an improved generative adversarial network construction module 2 and an image reconstruction module 3;

[0095] The image acquisition module 1 is used to obtain the original infrared image data set and perform preprocessing to obtain the image data set;

[0096] The improved generative adversarial network construction module 2 is used to construct a generative adversarial network based on the improved Res2net, and use the image data set to train the generative adversarial network; wherein the improved generative adversarial network construction module 2 includes a generator generation module 21 and a discriminator generation module 22;

[0097] The image reconstruction module 3 is used to reconstruct the low-resolution infrared image that needs to be reconstructed based on the trained generative adversarial network, and output a super-resolution infrared image.

[0098] Example

[0099] In order to verify the effectiveness of the present invention, this embodiment uses the DIV2K dataset and the infrared image dataset commonly used in the field of image super-resolution to form a training dataset, wherein the DIV2K dataset is a dataset containing 800 training images and 100 verification images. The images in the dataset all have 2K resolution, have extremely high clarity and details, and are very suitable for training and evaluating super-resolution algorithms; in order to achieve the purpose of super-resolution of infrared images by the method, the M3FD, KAIST, and IRSTD-1k datasets are used as infrared image datasets to be added to the training dataset; the infrared image dataset is preprocessed, and the amplification factor is determined according to the target task, and the training dataset is downsampled according to the amplification factor to obtain a low-resolution image dataset for the training task, and the dataset is divided to prepare the corresponding training set and test set;

[0100] In this embodiment, a ×4 magnification factor is used to downsample the training set and test set data, and the image of the original size of h×w is downsampled to 1 / 4*h×1 / 3w. To avoid overfitting, data enhancement operations such as horizontal flipping, translation and scaling are used on the image pairs.

[0101] The method effect is tested using the test set data. The low-resolution image prepared in the test set is input into the generative adversarial network generator to reconstruct the super-resolution image. It is compared with the original high-resolution image, and the PSNR and SSIM between the generated image and the original image are calculated as evaluation indicators to evaluate the quality of the generated super-resolution image. The calculation formulas for PSNR and SSIM are as follows:

[0102]

[0103] Among them, the prediction of the MSE calculation model The degree of closeness to the true label Y, n is the number of elements; MAXI in PSNR is the maximum value representing the color of the image point; in SSIM, μ x and μ y represent the average values ​​of X and Y respectively; and Denote the variance of X and Y respectively; δ XY Represents the covariance of X and Y; the SSIM value range is [0,1]. The larger the value, the smaller the difference between the output image and the undistorted image, that is, the better the image quality. C1 and C2 are two constants, C1=(k1L) 2 ,C1=(kL) 2 , k1 and k2 are commonly used default values ​​of 0.01 and 0.03, and L represents the range of image pixel values.

[0104] In this experimental verification, Bicubic algorithm, SRCNN, VDSR, and SRGAN are selected as comparison methods. Random images in SET5, SET14, BSD100, and infrared data sets are used as test samples. First, the test samples are downsampled to obtain low-resolution input images, and then the low-resolution images are reconstructed using the method of the present invention and the comparison method for super-resolution images, and the reconstruction results are compared. Figure 7 For the test image and the reconstructed image obtained by the comparison method, the reconstructed image is represented by a local area to facilitate the observation of the reconstruction result. The verification result is as follows Figure 7 shown.

[0105] It can be seen that the present invention improves the feature extraction module in the generator of the GAN network in the framework of SRGAN, and based on the idea of ​​multi-scale feature extraction of input components in Res2net, extracts and utilizes multiple scale features of the input image, and obtains the reconstruction result of the expected magnification factor in the process of image super-resolution. Compared with the prior art, the image reconstructed by the technology of the present invention is clearer and has a higher resolution. Therefore, the present invention can improve the super-resolution effect of infrared images.

[0106] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present disclosure is not limited by the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present disclosure.

[0107] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0108] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. An image reconstruction method based on an improved generative adversarial network, characterized in that: The following steps are involved: Acquire the original infrared image data set and perform preprocessing to obtain an image data set; Constructing a generative adversarial network based on the improvement of Res2net, and using the image data set to train the generative adversarial network; wherein the generative adversarial network based on the improvement of Res2net includes a generator and a discriminator, the generator uses Res2net blocks and residual blocks to form a feature extraction module, and uses a sub-pixel convolution layer as an upsampling module; the discriminator includes 8 convolution layers, 2 dense layers and a sigmoid activation function; Based on the trained generative adversarial network, the low-resolution infrared image that needs to be reconstructed is reconstructed and generated, and a super-resolution infrared image is output.

2. The image reconstruction method based on the improved generative adversarial network according to claim 1, characterized in that: The process of obtaining the original infrared image data set and preprocessing to obtain the image data set is as follows: An original infrared image data set is obtained, and a magnification factor is determined according to a target task. A downsampling operation is performed on the original infrared image data set according to the magnification factor to obtain different low-resolution infrared image data sets corresponding to different magnification factors.

3. The image reconstruction method based on the improved generative adversarial network according to claim 1, characterized in that: The generator includes convolution layer I, activation function PreLU I, Res2net block I, Res2net block II, Res2net block III, residual block I, residual block II, residual block III, pixel reconstruction layer I, and convolution layer II.

4. The image reconstruction method based on the improved generative adversarial network according to claim 3 is characterized in that: The Res2net block includes a 1×1 convolution, a feature segmentation layer, a feature subset processing layer, a feature splicing layer, and a 1×1 convolution layer; Among them, the feature segmentation layer divides the input features into s feature subsets, and the number of channels in each feature subset is 1 / s of the number of input feature channels; The feature subset processing layer processes the segmented feature subsets, and directly outputs the first feature subset to the next level; for the second feature subset, the feature subset processing layer includes a 3×3 convolution layer, and then outputs to the next level; for the third to s feature subsets, the feature subset processing layer includes an addition layer and a 3×3 convolution layer with the output results of the s-1 feature subset processing layer, and then outputs to the next level; The feature concatenation layer concatenates the s outputs obtained by the feature subset processing layer of the previous level into new features; The residual block includes a first convolution layer, a pooling layer, an activation function, a second convolution layer, a pooling layer, and an addition layer; the feature information processed by the first convolution layer is weighted and retained as an input to a subsequent neural network convolution block, and is added to the feature information processed by the second convolution layer to strengthen the call of feature information in the processing process. At the same time, the weight of the intermediate residual feature information in the subsequent processing link is adjusted through the weighting parameter, so that the generator network can reasonably use the intermediate feature information as needed.

5. The image reconstruction method based on the improved generative adversarial network according to claim 4, characterized in that: The discriminator includes convolution layer I, activation function LeakyReLU I, convolution layer II, batch standardization layer BN I, activation function LeakyReLU II, convolution layer III, batch standardization layer BN II, activation function LeakyReLU III, convolution layer IV, batch standardization layer BN III, activation function LeakyReLU IV, convolution layer V, batch standardization layer BN IV, activation function LeakyReLU V, convolution layer VI, batch standardization layer BN V, activation function LeakyReLU VI, convolution layer VII, batch standardization layer BN VI, activation function LeakyReLU VII, convolution layer VIII, batch standardization layer BN VII, activation function LeakyReLU VIII, dense layer I, activation function LeakyReLU VIII, dense layer II, activation function sigmoid I.

6. The image reconstruction method based on improved generative adversarial network according to claim 1, characterized in that: The process of training the generative adversarial network to obtain the trained network is as follows: Dividing the image data set into a training set, a validation set and a test set, inputting the training set into the generator, reconstructing a super-resolution image, calculating the loss of the generator, and updating the generator weight; Inputting the training set into the discriminator, classifying the reconstructed super-resolution image and the original high-resolution image through the discriminator network, calculating the discriminator loss, performing back propagation, and updating the generator weights; After reaching the training times, the obtained parameters of the generative adversarial network are saved.

7. The image reconstruction method based on improved generative adversarial network according to claim 6, characterized in that: Using the Adam optimization algorithm, update the weight w of the generator G in descending order G : in, represents the generator weight w G The gradient of descent, z m Represents the super-resolution reconstructed image I SR The value of the mth pixel in , m = 1, 2, ..., M, M represents the number of pixels, D(G(z m )) indicates that the discriminator D judges the super-resolution reconstructed image I SR The mth pixel in the image is the high-resolution image I HR ; α represents the learning rate, β1 represents the exponential decay rate of the first-order moment estimate, and β2 represents the exponential decay rate of the second-order moment estimate.

8. The image reconstruction method based on improved generative adversarial network according to claim 6, characterized in that: Using the Adam optimization algorithm, update the weight w of the discriminator D in descending order D : in, Denotes the discriminator weight w D The gradient of descent, x m Represents a high-resolution image I HR The value of the mth pixel, D(x m ) represents the discriminator D judging the high-resolution image I HR The mth pixel is the high-resolution image I HR The probability of a pixel in express The descending gradient, Indicates that the discriminator D judges it as a high-resolution image I HR The probability of a pixel in .

9. The image reconstruction method based on improved generative adversarial network according to claim 6, characterized in that: The generator loss of the generative adversarial network consists of content loss and adversarial loss, that is, the loss of the generator l SR for: in, For content loss, is the generation loss; the content loss It consists of image loss, prediction loss and regularization loss weighted, where the image loss is calculated by MSE loss and the prediction loss is defined using VGG loss to describe the reconstructed image I SR and reference image I HR The Euclidean distance between the feature representations of ; the adversarial loss refers to adding the generative component of GAN to the perceptual loss, which is generated by the generative loss Defined by the probability of the discriminator on all training samples; The discriminator loss l D for: in, For the discriminator D, the reference image I HR The probability of being correctly identified as a reference image, For the discriminator D, the generator generates an image The probability of being identified as a reference image, the discriminator loss l D The overall description of the GAN discriminator's ability to correctly distinguish between reference images and generated images.

10. An image reconstruction system based on an improved generative adversarial network, used to implement the image reconstruction method based on an improved generative adversarial network as claimed in any one of claims 1 to 9, characterized in that: include: An image acquisition module (1), an improved generative adversarial network construction module (2), and an image reconstruction module (3); The image acquisition module (1) is used to acquire an original infrared image data set and perform preprocessing to obtain an image data set; The improved generative adversarial network construction module (2) is used to construct a generative adversarial network improved based on Res2net, and use the image data set to train the generative adversarial network; wherein the improved generative adversarial network construction module (2) includes a generator generation module (21) and a discriminator generation module (22); The image reconstruction module (3) is used to reconstruct a low-resolution infrared image that needs to be reconstructed based on a trained generative adversarial network, and output a super-resolution infrared image.

Citation Information

Cited By

  • Infrared image progressive super-resolution method based on gradient guidance

    CN122453613A