Infrared image super-resolution reconstruction method based on noise decoupling
Through the infrared image super-resolution reconstruction method based on noise decoupling, the noise components are decoupled and processed, combined with reverse projection guidance and multi-scale feature extraction, the problem of poor recovery effect of infrared images in the prior art in different environments is solved, and higher quality and stable super-resolution reconstruction is achieved.
Patent Information
- Application Number
- CN202510489706.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The existing infrared image super-resolution methods have poor recovery effects in different sensors or high noise environments, and are not robust enough to adapt to specific equipment or noise conditions.
The super-resolution reconstruction method of infrared images based on noise decoupling is adopted. The noise components of low-resolution images are decoupled through noise decoupling and dynamic split injection strategy, and consistency correction is performed through the reverse projection guidance module. The super-resolution model is constructed by combining multi-scale feature extraction and high-frequency cross-domain fusion.
The stability and generalization ability of the model in different noise environments and sensors is improved, and the detail recovery ability and image quality of infrared images are significantly improved, solving the problems of blurred details and artifacts in traditional methods.
Smart Images

Figure CN120198293A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of super-resolution reconstruction, and specifically relates to an infrared image super-resolution reconstruction method based on noise decoupling. Background Art
[0002] Infrared image super-resolution reconstruction has wide applications in remote sensing detection, night monitoring, medical imaging and other fields. However, due to the limitations of the physical characteristics of sensors and environmental factors in infrared imaging devices, it is often difficult to obtain high-resolution infrared data. In addition, traditional super-resolution methods usually rely on large-scale high-low resolution paired data for supervised learning, while the data acquisition cost of infrared images is high and annotation is difficult. When existing visible-light-based methods (such as DDRM, DiffPIR) are directly migrated to infrared images, due to ignoring the non-Gaussian characteristics of thermal noise and the band response differences, artifacts (such as stripe noise) and spectral aliasing will appear in the reconstructed images.
[0003] The Chinese patent publication number is "CN114913069A", and the name is "An Infrared Image Super-Resolution Reconstruction Method Based on Deep Neural Network". This method effectively enhances the data set and the non-linear fitting ability of the network by using various data enhancement methods for low-resolution infrared images, and realizes multi-path learning through the parallel connection of convolution kernels of different sizes in the residual block, improves the learning ability of the network, enables the network to learn local and global features simultaneously, but the designed super-resolution network model is difficult to restore the details of infrared images, has weak generalization ability in various scenarios, and the model performance drops significantly without a large-scale low-resolution and high-resolution paired data set. Summary of the Invention
[0004] (1) Technical Problems to be Solved
[0005] In view of the above problems existing in the prior art, the present invention provides an infrared image super-resolution reconstruction method based on noise decoupling, which solves the problems that existing infrared image super-resolution methods need to be specially trained for specific devices or noise conditions, have limited generalization ability, resulting in poor restoration effect and insufficient robustness in different sensors or high-noise environments.
[0006] (2) Technical Solutions
[0007] The technical solution of the present invention to solve the above problems is to provide an infrared image super-resolution reconstruction method based on noise decoupling, including the following steps:
[0008] S1: Input data and preprocessing: Using the FLIR infrared data set, downsample the high-resolution image to a low-resolution input of 64×64 or 32×32 by bicubic interpolation, and normalize the pixel values to [0,1] to adapt to the low-resolution infrared image I of the neural networkLR As subsequent input, and for this I LR Perform noise decoupling;
[0009] S2: Pretrained model loading: Load the pretrained consistency model as the denoising network. Through feature extraction, residual block groups, noise-conditioned embedding, and back-projection guidance modules, calculate the noise mean μ of the input noise n and standard deviation σ n , and use a multi-layer perceptron (MLP) to encode the normalized noise statistical information to generate a noise condition and add it to the feature map channel by channel. Finally, iteratively output the infrared high-frequency feature component F through back-projection HR to be used as subsequent conditional input. When loading the model, freeze the underlying feature extraction layer and only fine-tune the high-level residual blocks and noise embedding layer to adapt to the brightness distribution and noise characteristics of infrared images;
[0010] S3: Construct a super-resolution model: Based on a multi-scale feature extraction module, a heterogeneity high-frequency information fusion module, and sub-pixel convolution, construct a super-resolution model to enhance the detail recovery ability of infrared images; The low-resolution image I LR generates super-resolution features F through the multi-scale feature extraction module SR , and performs cross-domain fusion with the infrared high-frequency features F generated by denoising with the consistency model HR . Finally, generate the super-resolution image I through sub-pixel convolution SR ;
[0011] S4: Model training: The low-resolution image obtained in S1 undergoes noise decoupling to generate the initial noise N LR as conditional input and is input into the denoising network model in S2 for pre-training, and perform back-projection correction on it to extract the infrared high-frequency feature vector F HR ; Subsequently, use the multi-scale feature extraction module to decompose the features of the low-resolution image in S1 to obtain the super-resolution feature component F SR ; Perform heterogeneity characteristic information fusion and sub-pixel convolution on F HR and F SR to finally obtain the super-resolution reconstructed image I SR ;
[0012] S5: Model optimization and loss function design: The composite loss function includes the mean square error loss function (L MSE ), perceptual loss function (L LPIPS ), consistency loss function (L Con ), and gradient regularization loss function (L Reg ), and perform joint optimization to ensure that the finally output super-resolution infrared image reaches the optimal quality;
[0013] Furthermore, the low-resolution infrared image I in S1 LR undergoes three-layer discrete wavelet transform (DWT) with the basis function Symlet4, and is decomposed into a low-frequency sub-band {LL3} and high-frequency sub-bands; the adaptive hard threshold function is used to extract the noise components in the high-frequency sub-bands; subsequently, the noise components z are reconstructed through inverse wavelet transform (IDWT) LR to achieve noise decoupling.
[0014] Furthermore, in S2, the low-resolution noise image z LR is input into the denoising network, which includes a feature extraction, a residual block group, a noise condition embedding, and a back-projection guidance module; the feature extraction uses 3×3 convolution combined with ReLU activation for preliminary feature extraction, the residual block group contains 4 residual units, each residual unit consists of 3×3 convolution, ReLU activation, and skip connection, and the skip connection enhances the high-frequency detail recovery ability; the noise condition embedding layer encodes the noise level (1 + δ)τ n and fuses it with the feature map to achieve adaptive denoising, specifically expressed as:[[]]
[0015] The noise condition embedding: separates and independently controls the noise level (1 + δ)τ used in the denoising process n from the noise level τ injected into the image n and splits the noise into anti-correlated estimated noise and random noise to achieve fine modeling of the noise, specifically expressed as:[[]]
[0016]
[0017] Ensures that the noise intervention not only retains historical information but also has the ability of random perturbation, balancing detail recovery and avoiding local optima;
[0018] The back-projection guidance module: based on data consistency correction, matches the original observation through downsampling to ensure that the reconstructed image has both rich details and maintains the same degradation relationship with the input image, specifically expressed as:[[]]
[0019]
[0020] Furthermore, in S3, the multi-scale feature extraction module uses 3×3, 5×5, and 7×7 convolution kernels to extract multi-scale features in parallel and fuses different-scale information through Softmax normalization weights;
[0021] The heterogeneous high-frequency information fusion module aligns the feature domains through linear transformation and calculates the fused features based on the channel attention mechanism Finally, the fusion weight γ = Sigmoid(FC([F HR ,F SR )) is calculated for F final =γ·F fusion +(1 - γ)·F SR 。
[0022] Furthermore, the mean square error loss function (L MSE ) in S5: Constrains the super-resolution image to approximate the real image at the pixel level, ensuring the consistency of the basic structure and brightness;
[0023] The perceptual loss function (L LPIPS ):By performing similarity calculations in the feature space, it improves the subjective perceptual quality of the reconstructed image, captures texture and high-order semantic information, making the reconstructed image superior in perceptual quality to the results trained only using MSE;
[0024] The consistency loss function (L Con ):Through consistency distillation, it ensures that the outputs of the model at different time steps remain stable, reducing the impact of noise on super-resolution reconstruction;
[0025] The gradient regularization loss function (L Reg ):Smooths the gradient during the optimization process to prevent gradient explosion or model overfitting;
[0026] The overall composite loss function is represented as follows:
[0027] L Total =λ1L MSE +λ2L LPIPS +λ3L Con +λ4L Reg
[0028] where the loss weights are set as λ1 = 1.0, λ2 = 0.1, λ3 = 0.5, λ4 = 10 -4 。
[0029] (III) Beneficial Effects
[0030] Compared with the prior art, the present invention provides an infrared image super-resolution reconstruction method based on noise decoupling, having the following beneficial effects:
[0031] 1. The present invention proposes a noise decoupling and dynamic splitting injection strategy. By decoupling the noise component of the input low-resolution image from the image content and dynamically separating it into a historical noise estimation term and a random perturbation term, it solves the problems of detail blurring and artifacts caused by noise coupling in traditional super-resolution methods, enabling the model to make full use of noise statistical information and improving the stability and generalization ability of the denoising process.
[0032] 2. The present invention provides a noise consistency correction strategy combined with back-projection guidance, which corrects the noise output by the consistency model through back-projection, calculates the error gradient of the low-resolution noise, and guides the model to maintain statistical consistency with the input noise during the denoising process. This solution solves the problem of the noise distribution shift after the model denoises, enabling the model to accurately retain the statistical characteristics of the real noise and improving the authenticity and consistency of the denoised image.
[0033] 3. Through the end-to-end collaborative optimization of multi-scale feature extraction, noise decoupling guidance, and high-frequency cross-domain fusion, the present invention solves the problems of difficult alignment of heterogeneous features and serious loss of details in infrared image super-resolution, enabling the model to adapt to different scenarios and different noise intensities, and significantly improving the image quality and structural information integrity of super-resolution reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0035] Figure 1 It is a flowchart of the infrared image super-resolution reconstruction method based on noise decoupling according to the present invention;
[0036] Figure 2 It is a framework diagram of the super-resolution model based on consistency denoising and multi-scale feature fusion according to the present invention;
[0037] Figure 3 It is a structural block diagram of the consistency denoising network according to the present invention;
[0038] Figure 4 It is a structural block diagram of the super-resolution according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all of them. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the protection scope of the present invention.
[0040] Embodiment
[0041] As Figures 1-4 shown, an example of the present invention provides an infrared image super-resolution reconstruction method based on noise decoupling, including the following steps:
[0042] S1: Input data and preprocessing: Prepare the FLIR infrared dataset for model testing. Use bicubic interpolation to downsample the high-resolution images in the dataset to generate low-resolution inputs of 64×64 or 32×32, scale the pixel value range to [0,1] to adapt to the neural network input, and decouple the noise from the low-resolution degraded images;
[0043] In the noise decoupling stage, input the low-resolution infrared image I LR , and use 3-layer discrete wavelet transform (DWT) with the Symlet4 basis function to decompose it into the low-frequency subband {LL3} and high-frequency subbands First, extract the noise components from the high-frequency subbands through an adaptive threshold function. The hard threshold function is defined as:
[0044]
[0045] where the threshold of the k-th layer σ k is the noise standard deviation of the k-th layer high-frequency subband, and N is the number of subband pixels. Then reconstruct the noise, only retain the noise components of the high-frequency subbands, and generate the decoupled noise z LR :
[0046] z LR = IDWT({0,Γ hard (LH1),Γ hard (HL1),Γ hard (HH1)},...,{0,Γ hard (LH3),Γ hard (HL3),Γ hard (HH3)}) Through the above decoupling strategy, the structural information of the infrared image is retained, and at the same time, the noise components are effectively separated, improving the noise modeling ability and providing higher-quality input data for subsequent super-resolution reconstruction.
[0047] S2: Loading the pre-trained model: Construct a denoising network f θ based on the consistency models (CMs), and use the pre-trained consistency model as the denoising network to improve the stability and robustness in the infrared super-resolution reconstruction process;
[0048] Input the low-resolution noise z of the infrared image extracted after noise decoupling in S1 LR into the CMs denoising network for processing. The CMs network is as Figure 3 shown, adopting a hierarchical denoising structure, including multiple key modules such as feature extraction, residual block groups, noise-conditioned embedding, and reconstruction layers.
[0049] The feature extraction part uses a 3×3 convolution combined with a ReLU activation function for preliminary feature extraction to enhance the network's ability to understand infrared noise; the feature extraction layer captures the low-level features of infrared noise through convolution while retaining global context information, providing a basis for subsequent noise modeling and high-resolution detail restoration.
[0050] During the noise condition embedding process, first calculate the mean (μ n ) and standard deviation (σ n ) of the input noise to statistically model the input noise:
[0051] μ n , σ n = f noise (z LR )
[0052] Subsequently, the residual block group consists of 4 residual units, each containing a 3×3 convolution, a ReLU activation, and a skip connection to ensure the effective transmission of information and improve the ability to restore high-frequency details of the image;
[0053] Noise condition encoding: Encode the noise level (1 + δ)τ n through a fully connected layer, and normalize the noise statistical information (μ n , σ n ) of the infrared image and input it into a multi-layer perceptron (MLP) to calculate the final noise condition, making it adapt to the infrared noise distribution under different scenarios:
[0054]
[0055] The calculated noise condition is used as additional conditional information and added to the feature map channel by channel to achieve adaptive denoising for different noise intensities.
[0056] Adopt a noise condition injection strategy, divide the injected noise into anti-correlated estimated noise and random noise z to ensure that the noise intervention retains both historical information and random perturbation ability;
[0057] In each iterative optimization step, first calculate the current noise estimate:
[0058]
[0059] where is the denoising noise of the current iterative state, and z LR is the original noise term. This noise term reflects the update trend of the image in the current iterative state.
[0060] Then, independently sample standard normal distribution noise And use the hyperparameter η for weighted combination to generate the injected noise:
[0061]
[0062] Among them, η is used to balance the influence of historical noise information and random noise. Through this mechanism, the present invention can effectively utilize the existing noise estimation information, and at the same time introduce an appropriate amount of random perturbation to prevent the model from falling into local optimum and improve the ability to restore image details. The feature map after injecting noise is used for the next back-projection correction.
[0063] Back-projection guidance module: To ensure that the super-resolution image generated by the model is consistent with the original low-resolution observation data in the degradation relationship and reduce information loss, the present invention further designs a back-projection guidance module. First, use the degradation operator A to downsample the noise output by the consistency model to simulate the real image noise degradation process:
[0064]
[0065] Among them, z LR,est represents the downscaled estimated noise. Then, by calculating the error between the estimated low-resolution noise z LR,est and the input low-resolution noise component z LR , the gradient information is obtained:
[0066]
[0067] Among them, is the pseudo-inverse operator, and this term is used to correct the statistical consistency of the noise generated by the model. In infrared image processing, the noise is related to temperature radiation, so the noise is further corrected by thermal radiation physics to ensure the physical consistency of the infrared noise and the physical correlation with the missing high-frequency components:
[0068]
[0069] Among them, Blackbody(·) corrects the pixel radiation value based on the Stefan-Boltzmann law.
[0070] In each iteration, the super-resolution high-frequency information is updated through this gradient information:
[0071]
[0072] Among them, μ n is the dynamically adjusted step size parameter, and an adaptive step size adjustment strategy is adopted to optimize according to the current gradient amplitude. τ n-1 controls the intensity of the noise perturbation, and z comb is the combined noise provided by the noise splitting module, which is used to prevent falling into local optimum.
[0073] After the above processing, the denoised infrared high-frequency feature vector F is output HR , to ensure the consistency of the image value range and avoid information loss and reconstruction distortion;
[0074] In addition, when loading the pre-trained model, the underlying feature extraction layer is frozen to retain the global feature representation, and only the high-level residual blocks and noise embedding layer are fine-tuned to adapt to the brightness distribution and noise characteristics of the infrared image;
[0075] Through the above model design and pre-training process, this consistency model can efficiently extract the structural information of infrared images, complete the denoising process within a small number of iterations, and at the same time adapt to different noise intensities and scenarios, providing high-quality input for subsequent super-resolution reconstruction.
[0076] S3: Construct a super-resolution model: As Figure 4 shown, the model is composed of multiple functional modules, aiming to improve the detail recovery ability of low-resolution images through iterative optimization while ensuring data consistency. The model specifically includes: The low-resolution image I LR generates the super-resolution feature component F through the multi-scale feature extraction module SR , and performs heterogeneous high-frequency component information fusion with the infrared high-frequency feature component F generated by denoising with the consistency model in S2 HR , and finally generates the super-resolution image I through sub-pixel convolution SR ;
[0077] Multi-scale feature extraction module: In order to effectively restore the texture details of different scales in infrared images and solve the problem that it is difficult for single-scale convolution to extract cross-scale features, the present invention introduces a multi-scale feature fusion method to enhance the ability to restore details of different scales. This module first uses 3×3, 5×5, and 7×7 convolutional kernels in parallel to perform parallel feature extraction on the original low-resolution image I LR , generates multi-scale feature maps, after splicing the feature maps of different scales, a fully connected layer is introduced to calculate weights, and weight normalization is achieved through the Softmax function, and the fused multi-scale features are input into multiple residual blocks to further enhance the detail information of the image. The residual block consists of a 3×3 convolution, a ReLU activation function, and a skip connection, which can effectively extract high-frequency details and retain the integrity of feature transmission. Excluding the infrared feature F output in S2 HR to obtain the enhanced feature F SR as the super-resolution feature component to enter the next heterogeneous information fusion module.
[0078] Heterogeneous high-frequency information fusion module: The high-resolution feature F output by consistency denoising in S2 HR is fused with the multi-scale super-resolution feature F SRPerform cross - domain fusion. Since F HR and F SR are heterogeneous, it is necessary to align the feature domains of the two. A linear transformation is used for feature mapping to align them in the same feature space:
[0079]
[0080] Calculate the statistical information of the transformed features, including the mean μ and variance σ:
[0081]
[0082] Subsequently, based on the attention mechanism, calculate the channel attention weights of the two feature maps to enhance the high - frequency information expression ability, perform a one - dimensional convolution operation on the transformed features, and normalize them through the Softmax function to obtain the channel attention weights:
[0083]
[0084] Use the above - mentioned channel attention weights to perform weighted fusion on the features to obtain heterogeneous high - frequency weighted features:
[0085]
[0086] Adopt the channel attention mechanism to calculate the final fusion weight to dynamically adjust the contribution ratio of the infrared high - frequency features and the super - resolution features:
[0087] γ = Sigmoid(FC([F HR ,F SR ))
[0088] Finally, calculate the final fusion features according to the fusion weight:
[0089] F final = γ·F fusion +(1 - γ)·F SR
[0090] Through the above steps, the deep fusion of infrared high - frequency information and super - resolution features is achieved, thereby enhancing the detail performance ability of the super - resolution reconstructed image. Finally, use sub - pixel convolution for upsampling to generate the final super - resolution image I SR :
[0091] I SR = SubpixelConv(F fusion )
[0092] Through the collaborative action of the above modules, the clarity of the image is gradually improved during the iterative optimization process, while maintaining data consistency and structural integrity, thus effectively solving the problems of detail loss and instability existing in traditional super-resolution methods.
[0093] S4: Model training: The overall training framework of the model is as Figure 2 shown. The low-resolution image obtained in S1 is decoupled from noise to generate the initial noise N LR as the conditional input, and is input into the denoising network model in S2 for pre-training, and is corrected by back-projection to extract the infrared high-frequency feature vector F HR . Subsequently, the multi-scale feature extraction module is used to decompose the features of the low-resolution image in S1 to obtain the super-resolution feature component F SR . For F HR and F SR , heterogeneous feature information fusion and sub-pixel convolution are performed, and finally the super-resolution reconstructed image I SR is obtained.
[0094] During the optimization process, the AdamW optimizer (β1 = 0.9, β2 = 0.999) is adopted, the initial learning rate is set to 2×10 -4 , and the cosine annealing strategy is adopted for gradual attenuation:
[0095]
[0096] The minimum learning rate is set to 1×10 -5 , the training cycle is 50 rounds, and the weight decay parameter is set to 10 -4 to prevent parameter overfitting. The model completes super-resolution reconstruction within no more than 4 neural network computations (NFE), and the loss weight coefficients λ1 = 1.0, λ2 = 0.1, λ3 = 0.5, λ4 = 10 -4 are set. The weight coefficients are optimized through cross-validation and experiments to ensure the balance of the loss function, which not only improves the pixel-level reconstruction accuracy but also enhances the perceptual quality. During the training process, the hyperparameters are dynamically adjusted according to the peak signal-to-noise ratio (PSNR) evaluation index. To ensure the stability of model training, an early stopping strategy is adopted for training termination control: if the PSNR of the validation set does not improve in 5 consecutive rounds of training, the training is terminated to ensure the final output of high-quality infrared super-resolution images.
[0097] S5: Model optimization and loss function design: In order to ensure high-quality infrared image super-resolution reconstruction in a few steps of iteration, this method designs a weighted joint loss function, that is, a composite loss function;
[0098] The mean squared error loss function (L MSE), constrain the super-resolution image to approximate the real image at the pixel level, ensuring the consistency of the basic structure and brightness:
[0099]
[0100] where I HR,i is the real high-resolution image, I SR,i is the final output of the super-resolution network, and N is the number of training samples.
[0101] Introduce the perceptual loss function (L LPIPS ), and the L LPIPS loss can better capture texture and high-order semantic information, making the reconstructed image superior to the result trained only with MSE in terms of perceptual quality:
[0102]
[0103] where VGG(·) represents the feature representation extracted by the pre-trained VGG network, and this loss calculates the Euclidean distance between the super-resolution image and the real image in the VGG feature space.
[0104] Introduce the consistency loss function (L Con ), and through consistency distillation, ensure that the outputs of the model at different time steps remain stable, reducing the impact of noise on super-resolution reconstruction:
[0105]
[0106] where f θ is the consistency model, and I t and are the denoising results at different time steps t and t′ respectively.
[0107] Introduce the gradient regularization loss function (L Reg ), to avoid unstable optimization caused by excessive gradient updates during training:
[0108]
[0109] where represents the gradient of the loss function with respect to the model parameters.
[0110] Considering all the above loss functions, the overall composite loss function is expressed as follows:
[0111] L Total = λ1L MSE + λ2L LPIPS + λ3L Con + λ4L Reg
[0112] Among them, the loss weights are set as λ1 = 1.0, λ2 = 0.1, λ3 = 0.5, and λ4 = 10 -4 , so as to ensure the balance between the reconstruction quality and the perceptual effect. Through the above loss function design and optimization strategy, this patent realizes the efficient reconstruction of infrared image super-resolution, taking into account accuracy, speed, and robustness.
[0113] Through the efficient infrared image super-resolution reconstruction based on noise decoupling, combined with the back-projection guidance and the optimized noise injection strategy, high-quality reconstruction is achieved with extremely few neural network calculations. Compared with traditional super-resolution methods, this method has the advantages of low computational cost, high recovery quality, and strong adaptability, and is applicable to infrared image application fields such as remote sensing, night vision, and medical imaging.
[0114] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. The infrared image super-resolution reconstruction method based on noise decoupling is characterized by: The steps include: S1: Input data and preprocessing: Using the FLIR infrared dataset, the high-resolution image is downsampled to a low-resolution input of 64×64 or 32×32 through bicubic interpolation, and the pixel values are normalized to [0,1] to adapt the low-resolution infrared image I of the neural network. LR As a subsequent input, and for this I LR Perform noise decoupling; S2: Pre-trained model loading: Load the pre-trained consistency model as the denoising network, calculate the noise mean μ of the input noise through feature extraction, residual block group, noise condition embedding and back-projection guidance module n and standard deviation σ n , a multi-layer perceptron (MLP) is used to normalize the noise statistics Code generation noise conditions And add it to the feature map channel by channel, and finally output the infrared high-frequency feature component F through back projection iteration. HR Used as subsequent conditional input, the underlying feature extraction layer is frozen when the model is loaded, and only the high-level residual block and noise embedding layer are fine-tuned to adapt to the brightness distribution and noise characteristics of the infrared image; S3: Building a super-resolution model: Based on the multi-scale feature extraction module, the heterogeneous high-frequency information fusion module and the sub-pixel convolution, a super-resolution model is built to enhance the ability to restore infrared image details; low-resolution image I LR The super-resolution feature F is generated by the multi-scale feature extraction module SR , and the infrared high-frequency feature F generated by denoising with the consistency model HR Perform cross-domain fusion and finally generate a super-resolution image I through sub-pixel convolution SR ; S4: Model training: The low-resolution image obtained in S1 is decoupled by noise to generate the initial noise N LR As the conditional input, it is input into the denoising network model in S2 for pre-training, and then back-projected and corrected to extract the infrared high-frequency feature vector F HR ; Then, the multi-scale feature extraction module is used to decompose the low-resolution image in S1 to obtain the super-resolution feature component F SR ; for F HR and F SR Perform heterogeneous feature information fusion and sub-pixel convolution to finally obtain the super-resolution reconstructed image I SR ; S5: Model optimization and loss function design: The composite loss function includes the mean square error loss function (L MSE ), the perceptual loss function (L PIPS ), consistency loss function (L Con ) and the gradient regularization loss function (L Reg ), perform joint optimization to ensure that the final output super-resolution infrared image reaches the best quality.
2. The infrared image super-resolution reconstruction method based on noise decoupling according to claim 1 is characterized in that: The low-resolution infrared image I in S1 LR After three layers of discrete wavelet transform (DWT), the basis function is Symlet4, which is decomposed into low-frequency subband {LL3} and high-frequency subband; using the adaptive hard threshold function Extracting high frequency sub-band noise components; Then, the noise component z is reconstructed by inverse wavelet transform (IDWT) LR , to achieve noise decoupling.
3. The infrared image super-resolution reconstruction method based on noise decoupling according to claim 1 is characterized in that: In S2, the low-resolution noise image z LR Input denoising network, which includes feature extraction, residual block group, noise conditional embedding and back-projection guidance module; The feature extraction uses 3×3 convolution combined with ReLU activation for preliminary feature extraction. The residual block group contains 4 residual units, each of which is composed of 3×3 convolution, ReLU activation and jump connection. The jump connection enhances the ability to restore high-frequency details. The noise conditional embedding layer has an effect on the noise level (1+δ)τ n Encode and fuse with the feature map to achieve adaptive denoising, which is specifically expressed as: The noise condition embedding: The noise level used in the denoising process is (1+δ)τ n The noise level τ injected into the image n Separate and independently control the noise into anti-correlated estimated noise and random noise Achieve fine modeling of noise, specifically expressed as: Ensure that noise intervention not only retains historical information but also has the ability of random perturbation, balancing detail recovery and avoiding local optimality; The back-projection guidance module: based on data consistency correction, matches the original observation by downsampling to ensure that the reconstructed image has rich details and maintains the same degradation relationship with the input image, which is specifically expressed as:
4. The infrared image super-resolution reconstruction method based on noise decoupling according to claim 1 is characterized in that: The multi-scale feature extraction module in S3 uses 3×3, 5×5, and 7×7 convolution kernels to extract multi-scale features in parallel, and fuses information of different scales through Softmax normalization weights; The heterogeneous high-frequency information fusion module is transformed by linear Align feature domains and calculate fusion features based on channel attention mechanism Finally, according to the fusion weight γ = Sigmoid (FC ([F HR ,F SR ]))Calculate F final =γ·F fusion +(1-γ)·F SR .
5. The infrared image super-resolution reconstruction method based on noise decoupling according to claim 1 is characterized in that: The composite loss function in S5 is generally expressed as follows: L Total =λ1L MSE +λ2L LPIPS +λ3L Con +λ4L Reg The loss weights are set as λ1=1.0, λ2=0.1, λ3=0.5, λ4=10 -4 .
Citation Information
Patent Citations
Image super-resolution reconstruction method of stacking attention mechanism coding and decoding unit
CN111681166A
Infrared image super-resolution reconstruction method
CN112561799A
Real scene-oriented infrared image super-resolution reconstruction method based on deep learning
CN117422620A
Infrared video super-resolution reconstruction method, device and equipment
CN118134766A
Image super-resolution reconstruction method based on generative adversarial network
CN118396853A
Cited By
Optical image noise reduction method and system of composite imaging assembly
CN120525759A
Optical image noise reduction method and system for composite imaging assembly
CN120525759B
Infrared image denoising method and system based on artificial intelligence
CN120852210A
Artificial Intelligence-Based Infrared Image Denoising Method and System
CN120852210B
Image processing method and device for infrared imaging and storage medium
CN121660896A