Infrared image super-resolution reconstruction method based on frequency division guided iterative diffusion
By using the frequency-dividing iterative diffusion method in infrared image super-resolution reconstruction, the image is decomposed into high-frequency and low-frequency information, and the image quality is improved through the iterative feedback mechanism, the problems of insufficient high-frequency detail recovery and poor low-frequency information enhancement effect in the prior art are solved, and higher quality image reconstruction is achieved.
Patent Information
- Application Number
- CN202510095537.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The existing infrared image super-resolution reconstruction methods are insufficient in high-frequency detail recovery, and the low-frequency information enhancement effect is poor, resulting in low-quality image quality.
Using a frequency-dividing guided iterative diffusion method, the iterative feedback mechanism is implemented to improve the image detail and clarity by decomposing the input image into high-frequency and low-frequency information, and using the high-frequency guided information encoder and conditional mapping layer to improve the reverse process of the pre-trained stable diffusion model.
Improves the high-frequency detail recovery ability and low-frequency information enhancement effect of infrared images, and generates higher quality super-resolution images.
Smart Images

Figure CN120070177A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of infrared image super-resolution reconstruction, and specifically to an infrared image super-resolution reconstruction method based on frequency division guided iterative diffusion. Background Art
[0002] Infrared images are widely used in fields such as thermal imaging, video surveillance, medical diagnosis, and remote sensing. However, due to the limitations of optical devices and detector sizes, infrared images often suffer from blurring and low resolution problems, far inferior to visible light images. To improve image quality, super-resolution technology restores details by reconstructing high-resolution images and has shown potential in fields such as military, security, and surveillance. However, infrared images are often accompanied by low contrast, blurring, and noise, which limit their practical applications. Most existing infrared super-resolution techniques are designed based on fixed degradation assumptions, resulting in a decline in reconstruction performance when the actual degradation does not match, and problems such as insufficient accuracy and poor adaptability. Therefore, we propose an infrared image super-resolution method based on frequency division guided iterative diffusion to solve the above problems.
[0003] Chinese authorized announcement number "CN117391938B", titled "An Infrared Image Reconstruction Method, System, Device, and Terminal". This method first simulates the low-resolution state of real images through various degradation strategies, then constructs a super-resolution network, which includes a convolutional layer for primary feature extraction, three branches for deep feature extraction, and pre-trains the network with a visible light dataset, and then optimizes the network with an infrared dataset. Finally, a high-resolution image is obtained through sub-pixel convolution and three interpolation upsampling operations. This method does not clearly distinguish between the processing of low-frequency and high-frequency information during the feature extraction and feature restoration processes, resulting in insufficient recovery of high-frequency details, and this method relies on the feature migration of the visible light dataset and cannot fully reflect the characteristics of infrared images. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention proposes an infrared image super-resolution reconstruction method based on frequency division guided iterative diffusion, which solves the problems of insufficient restoration of high-frequency details and poor enhancement effect of low-frequency information in existing infrared image super-resolution reconstruction methods.
[0005] The present invention specifically adopts the following technical solutions to achieve the above objectives:
[0006] An infrared image super-resolution reconstruction method based on frequency division guided iterative diffusion, comprising the following steps.
[0007] Step 1, construct a network model: The entire super-resolution network consists of an upsampling layer, a Gaussian filter, a high-frequency guidance information encoder, a conditional mapping layer, and a pre-trained stable diffusion model.
[0008] Step 2, Prepare the dataset: The present invention selects the LLVIP infrared dataset and performs selection and preprocessing on the dataset.
[0009] Step 3, Select the loss function and evaluation metrics of this super-resolution method: Construct the overall loss function of the super-resolution network to minimize the difference between the generated super-resolution image and the real high-resolution image, and select the evaluation metrics of this super-resolution method to evaluate the performance of the overall network model.
[0010] Step 4, Train the network model: Input the training set into the network model for training. The entire training process follows the frequency-divided guidance iterative training strategy. At the same time, to ensure the effect of the pre-trained stable diffusion model, fix the weights of this model and only optimize the parameters of the high-frequency guidance information encoder and the conditional mapping layer until the optimal training weights are obtained.
[0011] Step 5, Save the model: After the training is completed, solidify the network parameters, store the parameters, weights, code, etc. in the memory, and determine the final infrared image super-resolution reconstruction model.
[0012] Furthermore, the high-frequency information guidance encoder in Step 1 is used to encode the input high-frequency information into guidance information for guiding the pre-trained stable diffusion model. This encoder is divided into two branches to extract and enhance the high-frequency information in the spatial domain and frequency domain respectively, and then splice the outputs of the two to obtain four-level guidance information.
[0013] Furthermore, the conditional mapping layer in Step 1 takes the high-frequency guidance information encoder as input, obtains the latent representation of the guidance information to adapt to the latent diffusion of the stable diffusion model, and transmits the guidance information to the decoder of its U-net network under the condition of not affecting the weight parameters of the pre-trained stable diffusion model to achieve the role of guiding the generation of its reverse process.
[0014] Furthermore, the pre-trained stable diffusion model in Step 1 selects the V1.5 version, which has rich prior knowledge and can help the network generate higher-quality super-resolution images.
[0015] Furthermore, the overall loss in Step 3 consists of pixel loss, perceptual loss, KL divergence loss, control loss, and high-frequency loss, and the evaluation metrics include peak signal-to-noise ratio and structural similarity.
[0016] Furthermore, in the frequency division guided iterative training strategy in step 4, the input infrared image is decomposed into high-frequency information and low-frequency information. The high-frequency information contains details, textures, etc. and is used as the input of the high-frequency guided information encoder, while the low-frequency information contains the overall structure of the image and is used as the input of the pre-trained stable diffusion model. The high-frequency information is encoded into conditional information to guide the generation of the pre-trained stable diffusion model. During the reverse process of the pre-trained stable diffusion model, the low-frequency information at each time step is combined with the high-frequency conditional information to generate results. The results generated at the nth time step are extracted and used as the input for the (n + 1)th round of training and fed back to the entire network for iteration, and the above process is repeated, thereby gradually improving the details and overall clarity of the image and finally generating high-quality super-resolution images.
[0017] Compared with the prior art, the present invention provides an infrared image super-resolution reconstruction method based on frequency division guided iterative diffusion, which has the following beneficial effects:
[0018] The present invention provides a frequency division guided iterative training strategy. This strategy decomposes the input image into high-frequency and low-frequency information, which are respectively used as the input of the high-frequency guided information encoder and the pre-trained stable diffusion model, and improves the reverse process of the pre-trained stable diffusion model to implement an iterative feedback mechanism, and gradually optimizes the details and clarity of the image, and finally obtains high-quality super-resolution images.
[0019] The present invention provides a high-frequency perception three-dimensional attention module. By introducing a wavelet extraction layer, the high-frequency components in the input features are obtained, and the three-dimensional attention map obtained by the attention mechanism is used to adaptively adjust the high-frequency components, helping the network to better focus on the high-frequency information.
[0020] The present invention provides a conditional guided diffusion network model. The conditional information generated by the high-frequency guided information encoder is used to constrain the pre-trained stable diffusion model, solving the problem of inconsistent generation results of the pre-trained stable diffusion model and realizing more realistic super-resolution image reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a flowchart of an infrared image super-resolution reconstruction method based on frequency division guided iterative diffusion provided by the present invention;
[0022] Figure 2 It is a schematic structural diagram of a high-frequency conditional guided diffusion network model provided by the present invention;
[0023] Figure 3 It is a schematic structural diagram of a high-frequency guided information encoder provided by the present invention;
[0024] Figure 4 It is a schematic structural diagram of a frequency domain extraction module provided by the present invention.
[0025] Figure 5 Schematic diagram of the airspace extraction module provided by the present invention.
[0026] Figure 6 Schematic diagram of the expansion and fusion module provided by the present invention
[0027] Figure 7 Schematic diagram of the high-frequency perception three-dimensional attention module provided by the present invention.
[0028] Figure 8 Schematic diagram of the conditional mapping layer provided by the present invention.
[0029] Figure 9 Schematic diagram of the comparison of evaluation indexes between the method provided by the present invention and the prior art Specific implementation manners
[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0031] Embodiment 1
[0032] As Figure 1 shown, the flowchart of a method for infrared image super-resolution reconstruction based on frequency division-guided iterative diffusion provided by an embodiment of the present invention. The method specifically includes the following steps:
[0033] Step 1, construct a network model: As Figure 2As shown, it is the structural diagram of the high-frequency conditional guidance diffusion network model, which specifically includes: an upsampling layer, a Gaussian filter, a high-frequency guidance information encoder, a conditional mapping layer, and a pre-trained stable diffusion model; the high-frequency guidance information encoder extracts and enhances the features of the input high-frequency image in two dimensions, namely the frequency domain and the spatial domain. The frequency domain feature extraction branch is composed of four frequency domain extraction blocks with local residual connections; the spatial domain feature extraction branch is also composed of four spatial domain extraction blocks with local residual connections. Each frequency domain and spatial domain extraction block corresponds one by one and is divided into four levels. The features processed at each level are concatenated to obtain four-level guidance information. Finally, a 1×1 convolutional layer is used to process the concatenated features of the last level to obtain the processed high-frequency feature map; the conditional mapping layer consists of four 1×1 convolutional layers, a VAE encoder, and four 0 convolutional layers. Among them, the 1×1 convolutional layer is used to adjust the number of channels of each level of conditional information, and the VAE encoder is used to map the conditional information to the latent space to adapt to the latent diffusion process of the stable diffusion model. Subsequently, the conditional information is mapped to the U-net decoder of the stable diffusion model through the 0 convolutional layer to guide the generation of high-quality images; the selected version of the pre-trained stable diffusion model is V1.5, and its rich prior knowledge is used to generate high-quality super-resolution images;
[0034] Step 2, Prepare the dataset: This method uses the LLVIP dataset and selects and preprocesses the dataset;
[0035] Step 3, Select the loss function and evaluation metrics for this super-resolution method: According to the network model in Step 1, select a loss function to minimize the difference between the super-resolution image reconstructed by the network and the real high-resolution image, and select evaluation metrics to evaluate the performance of the network; The present invention uses pixel loss, perceptual loss, KL divergence loss, control loss, and high-frequency loss; Pixel loss is used to ensure that the super-resolution image is as close as possible to the real high-resolution image at the pixel level; Perceptual loss is used to improve the visual quality of the super-resolution image, making the super-resolution image more perceptually close to the real image; KL divergence loss is used to regularize the latent distribution of the VAE encoder to make the latent space structure more reasonable; Control loss is used to ensure the conditional generation of the pre-trained stable diffusion model; High-frequency loss is used to guide and strengthen the high-frequency guidance encoder to make it closer to the real image in terms of texture and contour. During the training process, peak signal-to-noise ratio and structural similarity are used as evaluation metrics;
[0036] Step 4, Train the model: Input the training set in the dataset constructed in Step 2 into the network model to train the network model. The training process of the entire model follows the frequency division guidance iterative training strategy, as Figure 2As shown in the figure, the training process of the frequency-divided guided iterative training strategy is presented. The frequency-divided guided iterative training strategy decomposes the input infrared image into high-frequency information and low-frequency information, where the high-frequency information contains details, textures, etc., and the low-frequency part contains the overall structure of the image. The high-frequency information is encoded into conditional information to guide the generation of the pre-trained stable diffusion model. During the reverse process of the stable diffusion model, the low-frequency features at each time step will combine with the high-frequency conditional information to generate results. The results generated at the nth time are extracted and used as the input for the (n + 1)th round of training and fed back into the entire network for iteration until the training threshold is reached, thereby gradually enhancing the details and overall clarity of the image and finally generating high-quality super-resolution images. To ensure the effect of the stable diffusion model, its weight parameters are kept fixed, and only the parameters of the high-frequency guided information encoder and the conditional mapping layer are optimized until the optimal training weights are obtained;
[0037] Step 5, save the model. After the training is completed, solidify the network parameters, store the parameters, weights, code, etc. in the memory, and determine the final infrared image super-resolution reconstruction model;
[0038] Embodiment 2
[0039] As Figure 1 shown, the flowchart of a method for infrared image super-resolution reconstruction based on frequency-divided guided iterative diffusion proposed in an embodiment of the present invention is as follows. The method specifically includes the following steps:
[0040] Step 1, construct a network model: As Figure 2 shown, the high-frequency conditional guided diffusion network model constructed by the present invention specifically includes: an upsampling layer, a Gaussian filter, a high-frequency guided information encoder, a conditional mapping layer, and a pre-trained stable diffusion model;
[0041] The structure of the high-frequency guided information encoder is as Figure 3 shown. The high-frequency guided information encoder uses both the spatial domain and the frequency domain to perform multi-level feature extraction and enhancement on the input high-frequency features to ensure that the generated conditional information retains rich high-frequency details. In the encoder, both the spatial domain and the frequency domain feature extraction branches consist of four feature extraction blocks with local residual connections. Each frequency domain extraction block corresponds to a spatial domain extraction block one by one. Therefore, the entire encoder can be divided into four levels, and the spatial domain and frequency domain features are concatenated at each level to generate the fused four-level conditional information to guide the stable diffusion model to generate high-quality images. Finally, the features fused at the fourth level are processed by a 1×1 convolutional layer, and the output is the enhanced high-frequency feature map;
[0042] The structure of the frequency domain extraction block is as Figure 4As shown, first, the input high-frequency features are subjected to inverse Fourier transform to generate a spectrogram. Subsequently, the high-frequency and low-frequency information in the spectrogram is separated using high-low frequency masks. Then, the high-low frequency information is extracted through 1×1 convolution and L-type activation function. For the spectrogram S(f x ,f y ) after low-frequency masking, the adaptive moving weighted average algorithm is used to enhance the high-frequency information therein;
[0043] The spectrogram is divided into N×M blocks, and each block is denoted as B i,j (f x ,f y ), where i and j represent the indices of the block and column respectively. An enhancement coefficient A i,j (f x ,f y ) needs to be set for each block to reasonably constrain the intensity of high-frequency information enhancement. The setting of this enhancement coefficient is based on the distance between each block and the center frequency (f cx ,f cy ). The enhancement coefficient A i,j (f x ,f y ) is expressed as:
[0044]
[0045] where α represents the coefficient controlling the enhancement intensity, and d i.j (f x ,f y ) represents the normalized distance from the frequency component (f x ,f y ) within the block to the center frequency (f cx ,f cy ), and σ i,j (f x ,f y ) is the local variance within the block, which is used to adaptively control the enhancement coefficient.
[0046] After calculating the adaptive enhancement coefficient A i,j (f x ,f y ) for each block, the enhancement results of each block are weighted to ensure a smooth transition between each block, and finally the enhanced high-frequency spectral component (f x ,f y ) is obtained, and its expression form is:
[0047]
[0048] where k i,j (f x ,f y) represents the weighting coefficient between each block, S i,j (f x , f y ) represents the high-frequency spectral components of the original spectrogram within block B i,j (f x , f y ).
[0049] Multiply the enhanced high-frequency components by the original high-frequency spectral components to obtain enhanced high-frequency information. Add it to the low-frequency information extracted through 1×1 convolution and L-type activation function after high-frequency masking to generate an enhanced spectrogram. Finally, convert this spectrogram from the frequency domain back to the spatial domain through inverse Fourier transform to obtain the output of the frequency domain extraction block.
[0050] The structure of the spatial domain extraction block is as shown in Figure 5 . For the input high-frequency features , after two downsamplings, we get and , and perform feature extraction on the high-frequency features of three different dimensions respectively. Each feature extraction branch includes a multi-scale dilation fusion block, a wavelet three-dimensional attention module, and a 1×1 convolutional layer;
[0051] The structure of the multi-scale dilation fusion block is as shown in Figure 6 . It consists of three dilated convolutional layers and a concatenation operation. Among them, the three dilated convolutional layers are respectively composed of dilated convolutions with a convolution kernel size of 3×3 and dilation rates of 2, 3, and 4, and an L-type activation function to achieve feature extraction at different scales. Finally, the features extracted at three different scales are concatenated as the output of the dilation fusion block;
[0052] The structure of the high-frequency perception three-dimensional attention module is as shown in Figure 7 . The structure includes spatial attention, channel attention, and a wavelet extraction layer. The input feature map is passed into these three parts in parallel. Spatial attention: Perform global max pooling and global average pooling on the input feature map to generate a spatial attention map Channel attention: Perform global max pooling and global average pooling on the input feature map in the channel dimension. The results are passed into a fully connected layer to generate a channel average feature map and a channel max feature map Wavelet extraction layer: Use discrete wavelet transform to decompose the input feature map into low-frequency and high-frequency components. Integrate the high-frequency components through Convolution Layer 1 (convolution kernel size of 3×3 with an R-type activation function) to obtain a high-frequency feature map, and convert the high-frequency feature map into a high-frequency spatial feature map through Convolution Layer 2 (convolution kernel size of 3×3 with an R-type activation function) Convert the high-frequency feature map into a high-frequency channel feature map through an average pooling layer The high-frequency spatial feature map and the spatial attention map are concatenated along the channels and processed by Convolution Layer 4 (with a convolution kernel size of 3×3 and an R-type activation function) to generate a high-frequency spatial attention map. The average and maximum features of the high-frequency channel feature map and the channel attention are subjected to pixel-level addition operations to generate a high-frequency channel attention map. Using the broadcasting mechanism, the high-frequency spatial attention map and the high-frequency channel attention map are multiplied to generate a three-dimensional attention map. The Sigmoid activation function is used to obtain a three-dimensional attention weight map, which is multiplied by the input features through a residual connection to obtain the output of the high-frequency perception three-dimensional attention module.
[0053] The outputs of the high-frequency perception three-dimensional attention modules of the three branches all pass through a 1×1 convolution layer to obtain three features, which are respectively and For Deconvolution operations are used and added to the feature through pixel-level addition to obtain the feature For the feature Deconvolution operations are used and added to the feature through pixel-level addition to obtain the feature Finally, the feature is added to through pixel-level addition to obtain the output of the spatial domain extraction block.
[0054] The structure of the conditional mapping layer is as Figure 8 shown, including four 1×1 convolution layers (with a convolution kernel size of 1×1 and an R-type activation function), a VAE encoder, and four 0 convolution layers (with a convolution kernel size of 1×1 and weights initialized to 0). The four 1×1 convolution layers respectively receive the four-level guidance information generated by the high-frequency guidance information encoder. After adjusting the number of channels, they are used as the input of the VAE encoder to generate the latent representation of the four-level guidance information. Each level of latent representation is sequentially used as the input of the four 0 convolution layers to achieve the injection of conditional information, while keeping the parameters of the pre-trained stable diffusion model unaffected, so that the conditional information can act on the four upsampling layers of the model encoder.
[0055] The present invention uses a pre-trained stable diffusion model of version V1.5. Since it has been trained with a large amount of data, it has quite rich prior knowledge to help the network generate clearer and more realistic high-resolution images in the infrared image super-resolution reconstruction task.
[0056] To ensure the robustness of the network, introduce more non - linear factors, and retain more image information, the present invention adopts four activation functions, namely: R - type activation function, L - type activation function, and S - type activation function, which are defined as follows:
[0057]
[0058] Step 2, Prepare the dataset: The LLVIP dataset contains 17,088 infrared images, covering 10 types of scenarios. We selected 4,690 infrared images with a size of 1080 pixels × 720 pixels, performed image augmentation, and expanded the dataset to obtain 14,570 infrared images with a size of 256×256. Using a ratio of 6:2:2, the dataset was divided into a training set, a validation set, and a test set, and the training set was subjected to adaptive histogram equalization operation to improve the image quality of the training set.
[0059] Step 3, Select the loss function and evaluation metrics for this super - resolution method: According to the network model in Step 1, select the loss function and evaluation metrics. The high - frequency guidance information encoder uses the high - frequency loss, the pre - trained StableDiffusion model uses the control loss, the VAE encoder uses the KL - divergence loss, and the entire super - resolution network model uses pixel loss and perceptual loss;
[0060] The high - frequency guidance information encoder uses the high - frequency loss to minimize the difference between the generated high - frequency features I H and the high - frequency information extracted from the real high - resolution image I HR by the Laplacian operator Δ. The specific mathematical expression is:
[0061]
[0062] The pre - trained StableDiffusion model uses the control loss to generate high - quality images under the influence of the guidance information c at each time step t. The specific mathematical expression is:
[0063]
[0064] where ε represents the noise added to the input low - frequency information in the forward process at each time step, and ε θ (x t ,t,c) represents the noise predicted by the model under the control of the guidance information at each time step.
[0065] The VAE encoder uses the KL - divergence loss To measure the difference between the potential distribution q(z|x) of the output and the prior distribution p(z), the potential distribution q(z|x) is the distribution of the latent variable z under the condition of the guiding information x, and the prior distribution p(z) is a hypothetical distribution used to define the overall structure of the latent space. Therefore, the divergence loss has the following mathematical expression:
[0066]
[0067] The pixel loss is used to measure the difference in pixels between the generated super-resolution I SR and the real high-resolution image I HR The specific mathematical expression is as follows:
[0068]
[0069] The perceptual loss passes through a pre-trained convolutional neural network to extract high-level features from the generated super-resolution image I SR and the real high-resolution image I HR respectively, and evaluates the similarity between the generated image I SR and the target image I HR through the difference, making the image generated by the network more in line with human perception. The specific mathematical expression is as follows:
[0070]
[0071] where represents the i-th feature layer in the pre-trained network and N represents the number of selected feature layers.
[0072] The present invention combines the above losses, and the overall loss is as follows:
[0073]
[0074] where λ 1 、λ 2 、λ 3 、λ 4 and λ 5 respectively represent the weights controlling different losses in the overall loss. By optimizing the overall loss, it helps the network improve the reconstruction accuracy and generate more realistic super-resolution images.
[0075] The peak signal-to-noise ratio (RNSR) and the similarity contrast (SSIM) are selected as the evaluation indicators of the present invention, and are defined as follows respectively:
[0076]
[0077] Among them, RNSR represents the peak signal-to-noise ratio, MAX I represents the maximum possible value of the image pixels, MSE represents the mean square error used to measure the difference between the original image and the compressed or reconstructed image, σ x represents the contrast of image x, σ y represents the contrast of image y, C 2 is a constant to prevent the denominator from being zero.
[0078] Step 4, training the network model: The training process of the entire network model follows the frequency division guided iterative training strategy proposed by the present invention. As Figure 2 shown, the training process of the frequency division guided iterative training strategy is presented. First, the input infrared image passes through a Gaussian filter to extract the low-frequency features of the image, and then subtracts the original image to obtain the high-frequency features. This process can be expressed as:
[0079]
[0080] I H = I - I L , where I represents the input infrared image, I H represents the high-frequency features of the image, including details and textures, I L represents the low-frequency features of the image, including the overall structure of the image, represents the Gaussian filter operation. For the parameter settings of the filter, the standard deviation of the filter is set to 2, and the filter kernel size is 11×11. Such an operation smooths the low-frequency structure while ensuring that the high-frequency details can be better separated.
[0081] The extracted high-frequency features and low-frequency features are respectively used as the input of the high-frequency guidance information encoder and the stable diffusion model. The high-frequency guidance information encoder extracts and strengthens the key features in the high-frequency information and encodes them into conditional information. This process can be expressed as:
[0082] C H = ε(I H ),
[0083] where C H represents the conditional information generated by the high-frequency features extracted through the high-frequency guidance information encoder, and ε(·) represents the high-frequency guidance information encoding operation.
[0084] The low-frequency features are used as the input of the pre-trained stable diffusion model to realize image generation through the reverse diffusion process. At each time step t, the input low-frequency features are subjected to feature extraction through the encoder in the U-net network of the stable diffusion model. In the decoding stage, these feature representations are reconstructed into the generated image at this time step under the guidance of the conditional information. This process follows a normal distribution and can be expressed as:
[0085]
[0086] Among them, I S,t+1 represents the image generated by the reverse process of the Stable Diffusion model at time step t+1, represents the reverse diffusion process of the Stable Diffusion model, I L,t represents the low-frequency features input into the model at time step t.
[0087] After the Stable Diffusion model completes the reverse process of a single time step each time, the result generated at this moment is extracted and used as the new input of the entire super-resolution network to iterate the above training process. Therefore, the high- and low-frequency feature extraction process can be expressed as:
[0088]
[0089] Among them represents the image generated by the Stable Diffusion model in the nth iteration of the training process, represents that when the network uses as the new input, the low-frequency features extracted in the (n+1)th iteration, represents the high-frequency features extracted in the (n+1)th iteration.
[0090] The high-frequency guidance information encoding process in the (n+1)th iteration can be expressed as:
[0091]
[0092] Among them represents the high-frequency features extracted by the network in the (n+1)th iteration The conditional information generated by the high-frequency guidance encoder.
[0093] The reverse process of the Stable Diffusion model in the (n+1)th iteration can be expressed as:
[0094]
[0095] Among them represents the result generated by the reverse process of the Stable Diffusion model in the (n+1)th iteration.
[0096] In summary, the entire training process of this strategy can be expressed as:
[0097]
[0098] Among them represents the super-resolution image generated by the entire network after N iterations.
[0099] Training is completed when the loss of training reaches the set threshold. During training, the Adam optimizer is used to continuously optimize the network parameters of the high-frequency guidance information encoder and the conditional mapping layer. The parameters of the pre-trained Stable Diffusion model are fixed and do not participate in the update. The learning rate is set to 1×10 -4 . After training is completed, a test set is selected to test the model, and it is evaluated by the evaluation metrics in step 3.
[0100] Step 5, save the model: After the network model training is completed, all the trained parameters are solidified to determine the final super-resolution network. When performing the infrared image super-resolution task, the low-resolution infrared image can be directly input into the trained end-to-end network model to obtain the final super-resolution image.
[0101] Among them, the implementations of convolution, deconvolution, activation function, concatenation operation, upsampling, and Gaussian filtering are algorithms well-known to those skilled in the art, and the specific processes and methods can be found in the corresponding textbooks or technical literatures.
[0102] The present invention designs a method for infrared image super-resolution based on frequency division guided iterative diffusion, and applies it to the super-resolution task of infrared images, realizing the end-to-end network input and output, and preferably solving the problem of detail loss in the infrared image super-resolution effect for a long time; making the effect of the infrared super-resolution task more excellent; under the same conditions, the relevant indicators of the super-resolution images obtained by the existing methods further verify the feasibility and superiority of this method.
[0103] A schematic diagram comparing the evaluation metrics of the existing technology and the method proposed by the present invention is shown in Figure 9. It can be seen from the figure that when the reconstruction multiples are 2 and 4, the method proposed by the present invention has higher peak signal-to-noise ratio and structural similarity than the existing methods. These metrics indicate that the method proposed by the present invention has a better effect than the existing methods, and the reconstructed high-resolution images are clearer.
[0104] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for super-resolution reconstruction of infrared images based on frequency division guided iterative diffusion, characterized in that: The specific steps are: Step 1, build the network model: the entire super-resolution network consists of an upsampling layer, a Gaussian filter, a high-frequency guidance information encoder, a conditional mapping layer, and a pre-trained stable diffusion model; Step 2, prepare the data set: use the LLVIP data set, and select and preprocess the data set; Step 3, selecting a loss function and an evaluation index of the super-resolution method: selecting a loss function according to the network model in step 1 to minimize the difference between the super-resolution image reconstructed by the network and the real high-resolution image, and selecting an evaluation index to evaluate the performance of the network; the present invention uses pixel loss, perceptual loss, KL divergence loss, control loss and high-frequency loss; and selecting peak signal-to-noise ratio and structural similarity as evaluation indexes of the super-resolution method; Step 4, training the network model: input the training set in the data set constructed in step 2 into the network model to train the network model. The entire training process follows the frequency division guided iterative training strategy. At the same time, in order to ensure the effect of the pre-trained stable diffusion model, its weight parameters remain fixed, and only the parameters of the high-frequency guided information encoder and the conditional mapping layer are optimized until the optimal training weight is obtained; Step 5, save the model. After the training is completed, solidify the network parameters, store the parameters, weights, and codes in the memory, and determine the final infrared image super-resolution reconstruction model.
2. The infrared image super-resolution reconstruction method based on frequency division guided iterative diffusion according to claim 1 is characterized in that: The high-frequency guidance information encoder in step 1 extracts and enhances the input high-frequency information in two dimensions: frequency domain and spatial domain; The frequency domain feature extraction branch consists of four frequency domain extraction blocks with local residual connections; The spatial feature extraction branch is also composed of four spatial extraction blocks with local residual connections; each frequency domain and spatial extraction block corresponds one to one, and is divided into four levels. The features processed at each level are spliced to obtain four levels of guidance information.
3. The infrared image super-resolution reconstruction method based on frequency division guided iterative diffusion according to claim 1, characterized in that: The spatial domain extraction block in step 1 is composed of a downsampling operation, three multi-scale expansion fusion blocks, three high-frequency perception three-dimensional attention modules, three 1×1 convolutional layers, and three deconvolutional layers.
4. The infrared image super-resolution reconstruction method based on frequency division guided iterative diffusion according to claim 1, characterized in that: The high-frequency perception three-dimensional attention module includes spatial attention, channel attention and wavelet extraction layers; spatial attention generates a spatial attention map through global pooling; channel attention generates a channel feature map through global pooling and a fully connected layer; The wavelet extraction layer uses discrete wavelet transform to decompose the feature map, and the high-frequency components are processed through convolution and average pooling layers to generate high-frequency spatial and high-frequency channel feature maps.
5. The infrared image super-resolution reconstruction method based on frequency division guided iterative diffusion according to claim 1, characterized in that: High-frequency spatial features are fused with spatial attention to generate a high-frequency spatial attention map, and high-frequency channel features are fused with channel features to generate a high-frequency channel attention map; finally, a three-dimensional attention map is generated by multiplying them through a broadcast mechanism, and the result is output through an activation function and residual connection.
6. The infrared image super-resolution reconstruction method based on frequency division guided iterative diffusion according to claim 5, characterized in that: The frequency division guided iterative training strategy in step 4 first decomposes the input infrared image into high-frequency information and low-frequency information, wherein the high-frequency information contains details and texture information, and the low-frequency part contains the overall structure of the image.
7. The infrared image super-resolution reconstruction method based on frequency division guided iterative diffusion according to claim 6, characterized in that: The high-frequency information is used as the input of the high-frequency guidance information encoder, and the low-frequency information is used as the input of the pre-trained stable diffusion model. The high-frequency information is encoded into conditional information to guide the generation of the pre-trained stable diffusion model. In the reverse process of the stable diffusion model, the low-frequency features at each time step will be combined with the high-frequency conditional information to generate results. The result generated at the nth time step is extracted as the input of the n+1th round of training and fed back to the entire network for iteration until the training threshold is reached.
Citation Information
Patent Citations
Infrared image super-resolution reconstruction method, system, device and terminal
CN117391938B
Infrared image super-resolution reconstruction method based on transfer learning
CN114708148A
Infrared image super-resolution reconstruction method, system, device and terminal
CN117391938A
Conditional diffusion probability model training method and omnidirectional image generation method
CN118096532A
Multi-scale infrared image super-resolution reconstruction method and system and storage medium
CN118644394A
Cited By
MRI image reconstruction method and device, electronic device and storage medium
CN120388095A
Power transmission and transformation equipment inspection infrared image super-resolution reconstruction method and system
CN120976022A
Virtual reality data acquisition and restoration method
CN121707869A
Model training method, image task processing method and related device
CN122134571A
A model training method, an image task processing method and related devices
CN122134571B