Dual guided infrared image super-resolution method and device based on diffusion model

By employing a dual-guided approach of global distribution and local structure for infrared image super-resolution, this method addresses the issues of weakened generation capabilities and difficulty in maintaining the realism of local structures in existing technologies. It achieves efficient and accurate infrared image super-resolution reconstruction, improving the quality of reconstructed images and the performance of downstream tasks.

CN122472984APending Publication Date: 2026-07-28UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UNIV OF SCI & TECH BEIJING
Filing Date
2026-05-09
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

Existing infrared image super-resolution techniques tend to weaken the generation capability of pre-trained diffusion models during the adaptation process, making it difficult to effectively suppress prior interference from visible light, resulting in global distribution deviations in the reconstruction results, and making it difficult to maintain the authenticity of local structures during diffusion sampling.

Method used

A dual-guided infrared image super-resolution method is adopted, which uses visible light images and high-resolution infrared images of the same scene for training through global distribution modulation and local structure refinement mechanism. Global representation modulation and local structure guidance are introduced to optimize the diffusion model to maintain the generation capability and reconstruction quality.

Benefits of technology

It improves the overall distribution consistency and local structure consistency of infrared image super-resolution results, reduces the distribution shift of visible light priors, enhances the ability to preserve target contours, edges and local geometry, and improves the quality of reconstructed images and the performance of downstream tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122472984A_ABST
    Figure CN122472984A_ABST
Patent Text Reader

Abstract

The application discloses a dual-guided infrared image super-resolution method and device based on a diffusion model, and relates to the technical field of image super-resolution. The method comprises the following steps: based on a preset Gaussian noise, denoising diffusion training is performed on a dual-guided super-resolution framework according to a time step, a low-resolution infrared image, a visible light image and a high-resolution infrared true value image; loss function calculation is performed according to infrared global representation, visible light global representation, modal classification probability, the preset Gaussian noise and predicted Gaussian noise; the dual-guided super-resolution framework is subjected to alternating step optimization of partial parameter freezing according to the classification loss and the comprehensive loss; a to-be-processed low-resolution infrared image is acquired; and image reconstruction is performed on the to-be-processed low-resolution infrared image by using the optimized dual-guided super-resolution framework. The application is an efficient and accurate infrared image super-resolution method based on dual guidance of global distribution and local structure based on a diffusion model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image super-resolution technology, and in particular to a dual-guided infrared image super-resolution method and apparatus based on a diffusion model. Background Technology

[0002] Infrared images can characterize the thermal radiation information of a target and still possess good perception capabilities in low-light, no-light, and complex environments, thus having significant application value in fields such as intelligent driving, target detection, and night vision surveillance. Due to limitations in infrared imaging principles and equipment, infrared images typically suffer from low resolution. Therefore, infrared image super-resolution reconstruction has become an important technical means to improve infrared imaging quality and the performance of subsequent tasks. Infrared image super-resolution not only requires visual quality of the generated image but also demands that the overall distribution consistency and local structural authenticity of the infrared image be maintained as much as possible during the reconstruction process, avoiding the generation of false details that do not conform to the laws of infrared imaging.

[0003] Existing infrared image super-resolution techniques generally follow the development path of general image super-resolution, with adaptive improvements tailored to the characteristics of infrared imaging. Early typical schemes mainly included super-resolution methods based on convolutional neural networks, which achieve pixel-level reconstruction by learning the mapping relationship between low-resolution and high-resolution images end-to-end. Generative adversarial networks were introduced into super-resolution tasks to enhance high-frequency details and perceptual quality. Transformer-based methods improved the recovery capability of complex structures and global information by modeling long-range dependencies. In recent years, diffusion models, due to their strong distribution modeling capabilities and high generation quality, have also begun to be applied to image inpainting and super-resolution tasks. In particular, the latent space diffusion framework, by first mapping the image to the latent space and then performing progressive denoising and reconstruction in the latent space, reduces computational complexity while maintaining generation quality, and has become one of the typical technical routes in generative super-resolution.

[0004] Diffusion models typically achieve high-quality image generation through progressive denoising, while latent space diffusion frameworks further enhance the feasibility of such methods in practical applications. These methods generally include a latent space encoding module, a diffusion denoising module, and an image decoding module. An encoder maps a low-resolution infrared image to a low-dimensional latent space, and then, combined with time-step information, inputs it into a diffusion denoising network. Through progressive sampling, high-resolution latent variable recovery is achieved, and finally, a decoder outputs a super-resolution image. The training process typically uses reconstruction or noise prediction loss as constraints. Because latent space diffusion strikes a good balance between generation quality and computational efficiency, it has become one of the important technical approaches in generative super-resolution of infrared images.

[0005] Building upon this foundation, some studies have made adaptive improvements to suit the characteristics of infrared imaging. Existing methods generally employ techniques such as infrared domain fine-tuning, conditional guidance, distribution constraints, or structural enhancement to make pre-trained diffusion models more suitable for infrared image reconstruction tasks. Specifically, some methods adapt the pre-trained model to the infrared domain to leverage existing generative priors and improve reconstruction capabilities; others enhance the overall distribution consistency between the reconstructed results and the real infrared image by adding distribution constraints or feature alignment; furthermore, current methods, while introducing structural priors such as edge information, contour information, or high-frequency information to improve structural consistency through loss functions applied to the model training objective, have already addressed this issue.

[0006] However, existing technologies still have some unresolved issues. Some methods adapt to infrared image super-resolution tasks by fine-tuning or additionally guiding pre-trained models, but this process can easily affect the model's original generative capabilities, making it difficult to balance the preservation of generative priors with infrared task adaptation. Although some methods have begun to focus on the overall distributional consistency between super-resolution results and real infrared images, and have attempted to improve infrared adaptability through distribution constraints and feature alignment, the lack of continuous and effective constraints on the generation process means that the model is still susceptible to interference from visible light priors during iterative reconstruction, causing the reconstruction results to deviate from the real infrared images in terms of grayscale distribution and thermal feature representation. Current methods mostly rely on edge information, contour information, or high-frequency information for static constraints to improve structural consistency, with insufficient intervention in the intermediate states during the progressive sampling process of the diffusion model. Therefore, there is still room for further improvement in preserving local contours, boundary details, and structural realism.

[0007] In the existing technology, there is a lack of an efficient and accurate infrared image super-resolution method based on the dual guidance of global distribution and local structure of diffusion model. Summary of the Invention

[0008] To address the technical problems of existing technologies, such as the easy weakening of the original generative capacity of the pre-trained diffusion model during adaptation; the difficulty in effectively suppressing visible light prior interference in the pre-trained diffusion model, leading to global distribution bias in the reconstruction results; and the difficulty in maintaining the authenticity of local structures during diffusion sampling, this invention provides a dual-guided infrared image super-resolution method and apparatus based on a diffusion model. The technical solution is as follows: On the one hand, a dual-guided infrared image super-resolution method based on a diffusion model is provided. This method is implemented by a dual-guided infrared image super-resolution device and includes: Acquire low-resolution infrared images, visible light images of the same scene, and high-resolution infrared ground truth images corresponding to the low-resolution infrared images; Obtain the time step; based on the preset Gaussian noise, the dual-guided super-resolution framework is denoised and diffused trained according to the time step, low-resolution infrared image, visible light image and high-resolution infrared ground image to obtain infrared global representation, visible light global representation, modality classification probability and prediction Gaussian noise. The loss function is calculated based on the infrared global representation, the visible light global representation, the modality classification probability, the preset Gaussian noise, and the predicted Gaussian noise to obtain the classification loss and the comprehensive loss. Based on the classification loss and the comprehensive loss, the dual-guided super-resolution framework is optimized by alternating step-by-step optimization with partial parameter freezing to obtain the optimized dual-guided super-resolution framework. Acquire the low-resolution infrared image to be processed; based on the low-resolution infrared image to be processed, perform image reconstruction using an optimized dual-guided super-resolution framework.

[0009] On the other hand, a dual-guided infrared image super-resolution device based on a diffusion model is provided. This device is applied to a dual-guided infrared image super-resolution method based on a diffusion model. The device includes: The data acquisition module is used to acquire low-resolution infrared images, visible light images of the same scene, and high-resolution infrared true images corresponding to the low-resolution infrared images. The denoising training module is used to obtain the time step; based on the preset Gaussian noise, the dual-guided super-resolution framework is denoised and diffused to obtain the infrared global representation, visible light global representation, modality classification probability and prediction Gaussian noise according to the time step, low-resolution infrared image, visible light image and high-resolution infrared ground truth image. The loss calculation module is used to calculate the loss function based on the infrared global representation, the visible light global representation, the modal classification probability, the preset Gaussian noise, and the predicted Gaussian noise, so as to obtain the classification loss and the comprehensive loss. The alternating optimization module is used to perform alternating step-by-step optimization of the dual-guided super-resolution framework by freezing some parameters based on the classification loss and the comprehensive loss, so as to obtain the optimized dual-guided super-resolution framework. The image reconstruction module is used to acquire the low-resolution infrared image to be processed; based on the low-resolution infrared image to be processed, image reconstruction is performed using an optimized dual-guided super-resolution framework.

[0010] On the other hand, a dual-guided infrared image super-resolution device is provided, the dual-guided infrared image super-resolution device comprising: a processor; a memory storing computer-readable instructions, wherein when the computer-readable instructions are executed by the processor, any one of the methods in the above-described dual-guided infrared image super-resolution method based on a diffusion model is implemented.

[0011] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored therein, the at least one instruction being loaded and executed by a processor to implement any of the above-described methods of the dual-guided infrared image super-resolution method based on the diffusion model.

[0012] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: This invention proposes a dual-guided infrared image super-resolution method based on a diffusion model. While making full use of the pre-trained diffusion model to generate priors, it minimizes the modification of its backbone network and avoids significant damage to the original generation capability of the model during the infrared task adaptation process. This allows the model to inherit the high-quality generation capability of the pre-trained model and meet the requirements of infrared image super-resolution reconstruction. Based on the effective guidance modulation mechanism of global modal distribution of infrared images, global representation modulation information oriented towards infrared modes is introduced in the diffusion reconstruction process to intervene in the model generation direction, thereby reducing the distribution offset caused by visible light prior and improving the consistency between the reconstruction results and the real infrared images in terms of overall distribution. Based on a stepwise constraint and refinement structure guidance mechanism for local structure in infrared images, this invention introduces structure guidance information during the iterative sampling process of the diffusion model to continuously adjust intermediate reconstruction states, thereby enhancing the preservation of target contours, edges, and local geometric structures and improving the structural consistency of the super-resolution results. This invention is an efficient and accurate infrared image super-resolution method based on dual guidance from global distribution and local structure in a diffusion model. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 This is a flowchart of a dual-guided infrared image super-resolution method based on a diffusion model provided in an embodiment of the present invention; Figure 2 This is a block diagram of a dual-guided infrared image super-resolution device based on a diffusion model provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a dual-guided infrared image super-resolution device provided in an embodiment of the present invention. Detailed Implementation

[0015] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0016] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0017] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0018] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0019] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0020] This invention provides a dual-guided infrared image super-resolution method based on a diffusion model. This method can be implemented using a dual-guided infrared image super-resolution device, which can be a terminal or a server. Figure 1 The flowchart shown is for a dual-guided infrared image super-resolution method based on a diffusion model. The processing flow of this method may include the following steps: S1. Acquire a low-resolution infrared image, a visible light image of the same scene, and a high-resolution infrared ground truth image corresponding to the low-resolution infrared image. In one feasible implementation, the low-resolution infrared image and the visible light image in this invention are infrared-visible light image pairs registered in the same scene, which can be obtained synchronously by an infrared thermal imaging device and a visible light imaging device; the high-resolution infrared ground truth image is a high-resolution infrared reference image corresponding to the low-resolution infrared image, used to provide a supervision signal for super-resolution reconstruction. After preprocessing, the above image data is input into the encoder of the pre-trained latent space diffusion model, and mapped to the low-dimensional latent space to obtain low-resolution infrared latent space features, visible light latent space features, and high-resolution infrared latent space ground truth, providing a data foundation for subsequent global distribution guidance and local structure guidance.

[0021] S2. Obtain the time step; Based on the preset Gaussian noise, the dual-guided super-resolution framework is denoised and diffused trained according to the time step, low-resolution infrared image, visible light image and high-resolution infrared ground image to obtain infrared global representation, visible light global representation, modality classification probability and prediction Gaussian noise. The dual-guided super-resolution framework includes a pre-trained latent space diffusion model, a global representation modulation module, and a local structure refinement module. The pre-trained latent space diffusion model includes a latent space coding submodule, a diffusion denoising submodule, and an image decoding submodule; The global representation modulation module includes a feature extractor, a classifier, a projection layer, and a gated modulation unit; The local structure refinement module includes the Sobel operator and the time-varying modulation function.

[0022] In one feasible implementation, the dual-guided super-resolution framework proposed in this invention includes a training latent space diffusion model, a global representation modulation module, and a local structure refinement module.

[0023] The pre-trained latent space diffusion model adopts a latent space diffusion framework, in which the latent space encoding submodule and the image decoding submodule are built based on variational autoencoders to realize bidirectional mapping between images and latent space variables. The network parameters of the pre-trained latent space diffusion model are frozen during training to preserve the original generative capabilities of the pre-trained model.

[0024] The Global Representation Modulation (GRM) module primarily extracts infrared modal representation information from the latent space features of infrared images and injects this information into the temporal embedding of the diffusion model, thereby continuously guiding the generation process at each stage of diffusion sampling. The GRM module mainly includes a feature extractor, a classifier, a projection layer, and a gated modulation unit. The feature extractor extracts infrared modal representations from the input latent space features. The classifier enhances the representational capability of the feature extractor by discriminating the modalities of the features extracted by the feature extractor. The projection layer is a linear layer used to map the infrared modal representations to the feature dimensions matching the temporal embedding of the diffusion model. The gated modulation unit controls the intensity of the modal representation information injected into the diffusion model.

[0025] The feature extractor is built upon convolutional layers. This encoder consists of an initial convolutional layer and three downsampled convolutional layers, each composed of a convolution with a stride of 2, group normalization, and an activation function. Successive downsampling expands the receptive field and improves global representation capabilities. To enhance the feature extractor's ability to extract infrared features, visible light image features corresponding to the same scene are introduced as reference information during the training phase.

[0026] To address the problem that existing infrared image super-resolution methods struggle to maintain the authenticity of local structures during diffusion sampling, leading to blurred edges, missing contours, and distortion of local geometry, this invention proposes a Local Structure Refinement (LSR) module. This module extracts local structure priors from the latent space features of low-resolution infrared images and continuously refines the intermediate latent variable states during the stepwise sampling process of the diffusion model, thereby enhancing the edge, contour, and local structure preservation capabilities of the reconstructed results.

[0027] Optionally, based on preset Gaussian noise, the dual-guided super-resolution framework is trained by denoising and diffusion according to the time step, low-resolution infrared image, visible light image, and high-resolution infrared ground truth image to obtain infrared global representation, visible light global representation, modality classification probability, and prediction Gaussian noise, including: Global guidance is constructed based on low-resolution infrared images, visible light images, and high-resolution infrared ground truth images to obtain low-resolution infrared latent space features, high-resolution infrared latent space features, infrared global representation, visible light global representation, and modality classification probability. Based on the global distributed modulation function, time embedding modulation is generated according to the infrared global representation and time step to obtain the modulation time embedding. Based on the time step, forward noise processing is performed according to high-resolution infrared latent space features and preset Gaussian noise to obtain noisy intermediate latent features. Based on the local structure modulation function, local structure-guided modulation is performed according to the low-resolution infrared latent space characteristics and time step to obtain the structure correction amount. Based on the structural correction amount, adaptive structural prior injection is performed on the noisy intermediate latent features to obtain the corrected intermediate latent features; Based on the time step, noise prediction is performed according to the modified intermediate latent features and the modulation time embedding to obtain the predicted Gaussian noise.

[0028] In one feasible implementation, infrared super-resolution technology can improve the application limitations caused by low spatial resolution. Existing methods attempt to maintain the global distribution and structural consistency of the results while improving sharpness. However, these methods either fail to address the problem thoroughly or require significant modifications that lead to a decline in reconstruction quality; this shortcoming is more pronounced with diffusion models. To address these issues, this invention proposes an infrared image super-resolution method guided by both global distribution and local structure based on diffusion models. The aim of this method is to improve the consistency of infrared super-resolution results while preserving the generative capabilities of diffusion models.

[0029] Optionally, global guidance is constructed based on low-resolution infrared images, visible light images, and high-resolution infrared ground truth images to obtain low-resolution infrared latent space features, high-resolution infrared latent space features, infrared global representation, visible light global representation, and modality classification probabilities, including: Latent space coding is performed based on low-resolution infrared images, visible light images, and high-resolution infrared ground truth images to obtain low-resolution infrared latent space features, visible light latent space features, and high-resolution infrared latent space features. Global representations are extracted based on low-resolution infrared latent space features and visible light latent space features to obtain infrared global representations and visible light global representations. Modal classification supervision is performed based on the infrared global representation and the visible light global representation to obtain the modal classification probability.

[0030] In one feasible implementation, the latent space features corresponding to the low-resolution infrared image are: The latent space features corresponding to the visible light images of the same scene are Through feature extractor Encoding the two allows us to obtain the infrared global representation. and visible light global representation The process is as follows (1): (1); Classifier Supervised infrared and visible light representations are employed to enhance the modal characteristics of the global representation information. The classifier utilizes a convolutional structure, consisting of convolutional layers, nonlinear activation functions, adaptive pooling layers, Flatten functions, and fully connected layers. Each convolutional layer includes an initial convolutional layer and two downsampling convolutional layers. LeakyReLU is used as the nonlinear activation function between the convolutional layers. Following the convolutional layers, an adaptive average pooling layer further extracts important information. The Flatten function reduces the feature dimension to one dimension. Finally, a linear layer module outputs the probability that the input feature belongs to either the infrared or visible light modality.

[0031] Optionally, based on the globally distributed modulation function, time-embedded modulation is generated according to the global infrared representation and time steps to obtain the modulation time embedding, including: Based on linear projection and gating control mechanism, modulation mapping is performed according to the infrared global representation to obtain the infrared prior modulation amount; The time encoding mechanism based on the diffusion model performs time embedding encoding according to the time step to obtain the original time embedding. Based on the globally distributed modulation function, the modulation time embedding is obtained by weighted fusion of the infrared prior modulation amount and the original time embedding.

[0032] In one feasible implementation, based on the infrared global representation This invention utilizes a linear projection layer Mapping to the same dimension as the temporal embedding in the diffusion model. A learnable gating is also introduced. Controlling the early injection intensity of information improves the stability of the training process; gating parameters... The initial value is -2.0. This represents the Sigmoid function.

[0033] Infrared prior modulation amount It can be expressed as follows (2): (2); The above infrared prior modulation amount Temporal embeddings fused into the diffusion model In, and through a globally distributed modulation function that varies with time step. This controls the guidance intensity at different diffusion stages. The modulated time embedding can be expressed as follows (3): (3); Since the diffusion model uses temporal embedding information at every time step, the above method can continuously apply infrared modal priors throughout the entire denoising process, rather than being limited to the input or output. The model is guided by infrared distribution characteristics at each sampling step, thereby gradually correcting the generation bias caused by visible light pre-training priors.

[0034] Optionally, based on the local structure modulation function, local structure-guided modulation is performed according to the low-resolution infrared latent space characteristics and time step to obtain the structure correction amount, including: Edge detection calculations are performed based on low-resolution infrared latent space features to obtain edge structure features; The edge gradient features are normalized to obtain normalized structural features; Based on the local structure modulation function, weighted modulation is performed according to the normalized structure characteristics and time step to obtain the structure correction amount.

[0035] In one feasible implementation, latent space features corresponding to low-resolution images are... The Sobel operator was used to extract local gradient information. The gradient responses of features in the horizontal and vertical directions are calculated to obtain edge information. To mitigate the problem of large differences in contrast and gradient responses between different images, the extracted structural features are normalized. The normalized structural features... It can be expressed as follows (4): (4); in, This indicates that the Sobel operator acts on low-resolution latent space features. The resulting gradient response is the edge structure feature. This represents the mean of the absolute values ​​of the gradient response. To prevent small positive numbers with a denominator of zero, we take... By performing the normalization process described above, the structural priors obtained from different input images can maintain a relatively stable scale range, thereby improving the robustness and numerical stability of subsequent structure-guided computation.

[0036] Considering the different requirements for structure guidance at different stages of diffusion sampling, this invention further designs a local structure modulation function. This is used to adaptively control the prior injection intensity of the structure. The corresponding structural correction amount... It can be expressed as follows (5): (5); During the training phase, this structural correction is incorporated into the intermediate latent variables of the diffusion model's forward degradation process, enabling the model to consider the influence of local edge and contour information on the reconstruction results when learning high-resolution reconstruction. During the sampling phase, this structural correction is added to the current intermediate state, allowing the model to be continuously guided by local structural information while gradually recovering image content. This structural injection is performed continuously in each sampling step, thus more effectively adjusting the intermediate reconstruction results of the diffusion model and reducing problems such as edge blurring, contour breakage, and local structural distortion.

[0037] S3. Calculate the loss function based on the infrared global representation, the visible light global representation, the modal classification probability, the preset Gaussian noise, and the predicted Gaussian noise to obtain the classification loss and the comprehensive loss. Optionally, a loss function is calculated based on the infrared global representation, the visible light global representation, the modality classification probability, the preset Gaussian noise, and the predicted Gaussian noise to obtain the classification loss and the comprehensive loss, including: The classification loss is obtained by calculating the binary cross-entropy loss function based on the infrared global representation, the visible light global representation, and the modality classification probability. The reconstruction loss is obtained by calculating the mean square error based on the preset Gaussian noise and the predicted Gaussian noise. The comprehensive loss is obtained by weighting the classification loss and reconstruction loss.

[0038] In one feasible implementation, during the training process, the classifier is optimized using a binary cross-entropy loss function, and the classification loss is calculated as follows (6): (6); The feature extractor and projection layer are updated using the comprehensive loss function, as shown in equation (7): (7); in, For the reconstruction loss of the diffusion model, , which is a weighting coefficient used to balance generation fidelity and modal constraint strength.

[0039] S4. Based on the classification loss and the comprehensive loss, the dual-guided super-resolution framework is optimized by alternating step-by-step freezing of some parameters to obtain the optimized dual-guided super-resolution framework. In one feasible implementation, the present invention employs a two-stage alternating update strategy. First, the feature extractor, projection layer, and gated modulation unit in the global representation modulation module are frozen, and the classifier parameters are updated only using the classification loss to enhance the modal discriminative power of the infrared global representation. Second, the classifier parameters are frozen, and the feature extractor, projection layer, and gated modulation unit are jointly updated using the comprehensive loss.

[0040] All parameters of the latent space coding submodule, diffusion denoising submodule, and image decoding submodule of the pre-trained latent space diffusion model are kept frozen to avoid damaging the model's original generative capabilities during infrared task adaptation. The above alternating update process is executed iteratively until convergence, resulting in the optimized dual-guided super-resolution framework.

[0041] S5. Acquire the low-resolution infrared image to be processed; based on the low-resolution infrared image to be processed, use an optimized dual-guided super-resolution framework to reconstruct the image.

[0042] In one feasible implementation, extensive experiments have shown that this method effectively improves the consistency of distribution and structure while maintaining competitive super-resolution performance, and performs well in downstream tasks.

[0043] To verify the effectiveness of the method of the present invention, a quantitative comparative analysis was conducted using subsets Set5 (test set 1), Set15 (test set 2), and Set20 (test set 3) from the Multi-scenario Multi-Modality Fusion Dataset (M3FD). Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index Measure (SSIM), and Learned Perceptual Image Patch Similarity (LPIPS) were selected as evaluation metrics.

[0044] The methods used include Residual Shifting Diffusion Model for Image Super-Resolution (ResShift), Single-Step Diffusion Model for Image Super-Resolution (SinSR), Binarized Diffusion Model for Image Super-Resolution (Bi-DiffSR), Adaptive Token Dictionary Super-Resolution (ATD), Mamba Image Restoration (MambaIR), Contourlet Residual for Prompt Learning Enhanced Infrared Image Super-Resolution (CoRPLE), Infrared Feature Fusion Network (InfraFFN), and Diffusion Model with Gradient Guidance for Infrared Image Super-Resolution. Super-Resolution (DifIISR) results are shown in Table 1 (Comparison Table of Super-Resolution Results Indicators).

[0045] Table 1

[0046] Compared with the comparative methods, the method of the present invention achieved the best SSIM and LPIPS on all three test subsets, indicating that the method of the present invention has a stable advantage in improving the structural consistency and visual perception quality of the reconstruction results; the method of the present invention also maintained a high level in terms of PSNR, and its overall performance is highly competitive.

[0047] From the perspective of distribution consistency, the method of this invention also compared the gray-level histogram distribution between the reconstructed result and the real infrared image, and evaluated them using Bach's distance, cosine similarity, and Manhattan distance. The results are shown in Table 2 (Comparison of Super-resolution Results and Ground Values). The method of this invention outperforms the compared ResShift and DifIISR methods in all three metrics. This indicates that the super-resolution result generated by the method of this invention is closer to the real infrared image in terms of overall gray-level distribution and can better maintain the modal distribution characteristics of the infrared image.

[0048] Table 2

[0049] The super-resolution image generated by the method of this invention is closer to the real infrared image overall, exhibiting better stability in brightness distribution, target contour, and local detail recovery. Especially in areas with slender structures, complex boundaries, and small targets, the method of this invention can improve image clarity while better preserving the original structural features of the infrared image, avoiding significant mode shift. Compared with models such as ATD, BI-DiffSR, SinSR, SinSRv2, ResShift, MambaIR, MambaIRv2, InfraFFN, DifIISR, and CoRPLE, the results obtained by the method of this invention are visually closer to the corresponding high-resolution ground truth images. While existing methods can generate relatively sharp textures in local areas, they are prone to introducing false details inconsistent with the infrared imaging mechanism, or deviating from the real infrared image in overall distribution. Comparing the normalized grayscale histograms of the reconstructed image and the real image, the distribution curve of the method of this invention matches the ground truth distribution to a higher degree, while ResShift and DifIISR still show more significant deviations from the ground truth.

[0050] The practical application value of this method was further validated in downstream tasks. In the object detection task, an object detection model (You Only Look Once version 5 small, YOLOv5s) was used as the detector, fine-tuned on the M3FD training set, and tested on infrared images reconstructed by different super-resolution methods; the BI-DiffSR, SinSR, ResShift, MambaIRv2, DifIISR, and CoRPLE models were compared. The results show that other methods all have missed detections, and the ResShift and SinSR models also have false detections. This method can better recover the thermal distribution features, contour boundaries, and semantic information related to object recognition, thereby reducing missed detections and false detections.

[0051] In the semantic segmentation task, the Deep Laboratory version 3+ (DeepLabv3+) augmented model was adopted as the segmentation model and tested on the MSRS dataset after fine-tuning. The model was compared with BI-DiffSR, SinSR, ResShift, MambaIRv2, DifIISR, and CoRPLE. The results show that, compared with other methods, this method can more accurately preserve semantic region boundaries and local structural details, reducing class confusion and inaccurate boundary issues.

[0052] The above results demonstrate that the method of the present invention can not only improve the quality of infrared image super-resolution reconstruction, but also significantly enhance the usability of the reconstructed image in practical downstream vision tasks such as target detection and semantic segmentation.

[0053] This invention proposes a dual-guided infrared image super-resolution method based on a diffusion model. While making full use of the pre-trained diffusion model to generate priors, it minimizes the modification of its backbone network and avoids significant damage to the original generation capability of the model during the infrared task adaptation process. This allows the model to inherit the high-quality generation capability of the pre-trained model and meet the requirements of infrared image super-resolution reconstruction. Based on the effective guidance modulation mechanism of global modal distribution of infrared images, global representation modulation information oriented towards infrared modes is introduced in the diffusion reconstruction process to intervene in the model generation direction, thereby reducing the distribution offset caused by visible light prior and improving the consistency between the reconstruction results and the real infrared images in terms of overall distribution. Based on a stepwise constraint and refinement structure guidance mechanism for local structure in infrared images, this invention introduces structure guidance information during the iterative sampling process of the diffusion model to continuously adjust intermediate reconstruction states, thereby enhancing the preservation of target contours, edges, and local geometric structures and improving the structural consistency of the super-resolution results. This invention is an efficient and accurate infrared image super-resolution method based on dual guidance from global distribution and local structure in a diffusion model.

[0054] Figure 2 This is a block diagram of a dual-guided infrared image super-resolution device based on a diffusion model, provided in an embodiment of the present invention. This device is used in a dual-guided infrared image super-resolution method based on a diffusion model. (Refer to...) Figure 2 The device includes a data acquisition module 210, a denoising training module 220, a loss calculation module 230, an alternating optimization module 240, and an image reconstruction module 250. Wherein: The data acquisition module 210 is used to acquire low-resolution infrared images, visible light images of the same scene, and high-resolution infrared true images corresponding to the low-resolution infrared images. The denoising training module 220 is used to obtain the time step; based on the preset Gaussian noise, the dual-guided super-resolution framework is denoised and diffused trained according to the time step, low-resolution infrared image, visible light image and high-resolution infrared ground image to obtain infrared global representation, visible light global representation, modality classification probability and prediction Gaussian noise. The loss calculation module 230 is used to calculate the loss function based on the infrared global representation, the visible light global representation, the modal classification probability, the preset Gaussian noise and the predicted Gaussian noise, so as to obtain the classification loss and the comprehensive loss. Alternating optimization module 240 is used to perform alternating step-by-step optimization of the dual-guided super-resolution framework by partially freezing some parameters based on classification loss and comprehensive loss, so as to obtain an optimized dual-guided super-resolution framework. Image reconstruction module 250 is used to acquire a low-resolution infrared image to be processed; and to perform image reconstruction using an optimized dual-guided super-resolution framework based on the low-resolution infrared image to be processed.

[0055] The dual-guided super-resolution framework includes a pre-trained latent space diffusion model, a global representation modulation module, and a local structure refinement module. The pre-trained latent space diffusion model includes a latent space coding submodule, a diffusion denoising submodule, and an image decoding submodule; The global representation modulation module includes a feature extractor, a classifier, a projection layer, and a gated modulation unit; The local structure refinement module includes the Sobel operator and the time-varying modulation function.

[0056] Optionally, the denoising training module 220 is further used for: Global guidance is constructed based on low-resolution infrared images, visible light images, and high-resolution infrared ground truth images to obtain low-resolution infrared latent space features, high-resolution infrared latent space features, infrared global representation, visible light global representation, and modality classification probability. Based on the global distributed modulation function, time embedding modulation is generated according to the infrared global representation and time step to obtain the modulation time embedding. Based on the time step, forward noise processing is performed according to high-resolution infrared latent space features and preset Gaussian noise to obtain noisy intermediate latent features. Based on the local structure modulation function, local structure-guided modulation is performed according to the low-resolution infrared latent space characteristics and time step to obtain the structure correction amount. Based on the structural correction amount, adaptive structural prior injection is performed on the noisy intermediate latent features to obtain the corrected intermediate latent features; Based on the time step, noise prediction is performed according to the modified intermediate latent features and the modulation time embedding to obtain the predicted Gaussian noise.

[0057] Optionally, the denoising training module 220 is further used for: Latent space coding is performed based on low-resolution infrared images, visible light images, and high-resolution infrared ground truth images to obtain low-resolution infrared latent space features, visible light latent space features, and high-resolution infrared latent space features. Global representations are extracted based on low-resolution infrared latent space features and visible light latent space features to obtain infrared global representations and visible light global representations. Modal classification supervision is performed based on the infrared global representation and the visible light global representation to obtain the modal classification probability.

[0058] Optionally, the denoising training module 220 is further used for: Based on linear projection and gating control mechanism, modulation mapping is performed according to the infrared global representation to obtain the infrared prior modulation amount; The time encoding mechanism based on the diffusion model performs time embedding encoding according to the time step to obtain the original time embedding. Based on the globally distributed modulation function, the modulation time embedding is obtained by weighted fusion of the infrared prior modulation amount and the original time embedding.

[0059] Optionally, the denoising training module 220 is further used for: Edge detection calculations are performed based on low-resolution infrared latent space features to obtain edge structure features; The edge gradient features are normalized to obtain normalized structural features; Based on the local structure modulation function, weighted modulation is performed according to the normalized structure characteristics and time step to obtain the structure correction amount.

[0060] Optionally, the loss calculation module 230 is further used for: The classification loss is obtained by calculating the binary cross-entropy loss function based on the infrared global representation, the visible light global representation, and the modality classification probability. The reconstruction loss is obtained by calculating the mean square error based on the preset Gaussian noise and the predicted Gaussian noise. The comprehensive loss is obtained by weighting the classification loss and reconstruction loss.

[0061] This invention proposes a dual-guided infrared image super-resolution method based on a diffusion model. While making full use of the pre-trained diffusion model to generate priors, it minimizes the modification of its backbone network and avoids significant damage to the original generation capability of the model during the infrared task adaptation process. This allows the model to inherit the high-quality generation capability of the pre-trained model and meet the requirements of infrared image super-resolution reconstruction. Based on the effective guidance modulation mechanism of global modal distribution of infrared images, global representation modulation information oriented towards infrared modes is introduced in the diffusion reconstruction process to intervene in the model generation direction, thereby reducing the distribution offset caused by visible light prior and improving the consistency between the reconstruction results and the real infrared images in terms of overall distribution. Based on a stepwise constraint and refinement structure guidance mechanism for local structure in infrared images, this invention introduces structure guidance information during the iterative sampling process of the diffusion model to continuously adjust intermediate reconstruction states, thereby enhancing the preservation of target contours, edges, and local geometric structures and improving the structural consistency of the super-resolution results. This invention is an efficient and accurate infrared image super-resolution method based on dual guidance from global distribution and local structure in a diffusion model.

[0062] Figure 3 This is a schematic diagram of the structure of a dual-guided infrared image super-resolution device provided in an embodiment of the present invention, as shown below. Figure 3 As shown, the dual-guided infrared imaging super-resolution device may include the above-mentioned Figure 2 The illustrated dual-guided infrared image super-resolution device is based on a diffusion model. Optionally, the dual-guided infrared image super-resolution device 310 may include a first processor 2001.

[0063] Optionally, the dual-guided infrared imaging super-resolution device 310 may also include a memory 2002 and a transceiver 2003.

[0064] The first processor 2001, memory 2002, and transceiver 2003 can be connected via a communication bus.

[0065] The following is combined with Figure 3 A detailed introduction to each component of the dual-guided infrared imaging super-resolution device 310 is provided below: The first processor 2001 is the control center of the dual-guided infrared image super-resolution device 310. It can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 can be one or more central processing units (CPUs), application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs).

[0066] Optionally, the first processor 2001 can perform various functions of the dual-guided infrared image super-resolution device 310 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.

[0067] In a specific implementation, as one example, the first processor 2001 may include one or more CPUs, for example... Figure 3 CPU0 and CPU1 are shown in the diagram.

[0068] In a specific implementation, as one example, the dual-guided infrared image super-resolution device 310 may also include multiple processors, such as... Figure 3 The first processor 2001 and the second processor 2004 are shown in the diagram. Each of these processors can be a single-core processor or a multi-core processor. Here, a processor can refer to one or more devices, circuits, and / or processing cores used to process data (such as computer program instructions).

[0069] The memory 2002 is used to store the software program that executes the present invention, and is controlled by the first processor 2001 to execute it. The specific implementation method can be referred to the above method embodiment, and will not be repeated here.

[0070] Optionally, the memory 2002 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently, and may be connected via the interface circuit of the dual-boot infrared image super-resolution device 310. Figure 3 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.

[0071] The transceiver 2003 is used to communicate with network devices or with terminal devices.

[0072] Alternatively, transceiver 2003 may include a receiver and a transmitter. Figure 3 (Not shown separately). The receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function.

[0073] Optionally, the transceiver 2003 can be integrated with the first processor 2001 or exist independently, and can be connected to the interface circuit of the dual-guided infrared image super-resolution device 310. Figure 3 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.

[0074] It should be noted that, Figure 3 The structure of the dual-guided infrared image super-resolution device 310 shown does not constitute a limitation on the router. Actual dual-guided infrared image super-resolution devices may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0075] Furthermore, the technical effect of the dual-guided infrared image super-resolution device 310 can be referred to the technical effect of the diffusion-based dual-guided infrared image super-resolution method described in the above method embodiments, and will not be repeated here.

[0076] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or it may be any conventional processor, etc.

[0077] It should also be understood that the memory in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DRRAM).

[0078] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0079] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.

[0080] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.

[0081] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0082] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0083] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0084] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0085] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0086] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0087] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0088] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A dual-guided infrared image super-resolution method based on a diffusion model, characterized in that, The method includes: Acquire low-resolution infrared images, visible light images of the same scene, and high-resolution infrared ground truth images corresponding to the low-resolution infrared images; Obtain the time step; based on the preset Gaussian noise, the dual-guided super-resolution framework is denoised and diffused trained according to the time step, low-resolution infrared image, visible light image and high-resolution infrared ground image to obtain infrared global representation, visible light global representation, modality classification probability and prediction Gaussian noise. The loss function is calculated based on the infrared global representation, the visible light global representation, the modality classification probability, the preset Gaussian noise, and the predicted Gaussian noise to obtain the classification loss and the comprehensive loss. Based on the classification loss and the comprehensive loss, the dual-guided super-resolution framework is optimized by alternating step-by-step optimization with partial parameter freezing to obtain the optimized dual-guided super-resolution framework. Acquire the low-resolution infrared image to be processed; based on the low-resolution infrared image to be processed, perform image reconstruction using an optimized dual-guided super-resolution framework.

2. The dual-guided infrared image super-resolution method based on a diffusion model according to claim 1, characterized in that, The dual-guided super-resolution framework includes a pre-trained latent space diffusion model, a global representation modulation module, and a local structure refinement module. The pre-trained latent space diffusion model includes a latent space coding submodule, a diffusion denoising submodule, and an image decoding submodule; The global representation modulation module includes a feature extractor, a classifier, a projection layer, and a gated modulation unit; The local structure refinement module includes the Sobel operator and the time-varying modulation function.

3. The dual-guided infrared image super-resolution method based on a diffusion model according to claim 1, characterized in that, The process involves training a dual-guided super-resolution framework using preset Gaussian noise, based on a time step, a low-resolution infrared image, a visible light image, and a high-resolution infrared ground truth image. This training yields infrared global representation, visible light global representation, modality classification probability, and predicted Gaussian noise, including: Global guidance is constructed based on low-resolution infrared images, visible light images, and high-resolution infrared ground truth images to obtain low-resolution infrared latent space features, high-resolution infrared latent space features, infrared global representation, visible light global representation, and modality classification probability. Based on the global distributed modulation function, time embedding modulation is generated according to the infrared global representation and time step to obtain the modulation time embedding. Based on the time step, forward noise processing is performed according to high-resolution infrared latent space features and preset Gaussian noise to obtain noisy intermediate latent features. Based on the local structure modulation function, local structure-guided modulation is performed according to the low-resolution infrared latent space characteristics and time step to obtain the structure correction amount. Based on the structural correction amount, adaptive structural prior injection is performed on the noisy intermediate latent features to obtain the corrected intermediate latent features; Based on the time step, noise prediction is performed according to the modified intermediate latent features and the modulation time embedding to obtain the predicted Gaussian noise.

4. The dual-guided infrared image super-resolution method based on a diffusion model according to claim 3, characterized in that, The process of constructing a global-guided model based on low-resolution infrared images, visible light images, and high-resolution infrared ground truth images to obtain low-resolution infrared latent space features, high-resolution infrared latent space features, infrared global representation, visible light global representation, and modality classification probability includes: Latent space coding is performed based on low-resolution infrared images, visible light images, and high-resolution infrared ground truth images to obtain low-resolution infrared latent space features, visible light latent space features, and high-resolution infrared latent space features. Global representations are extracted based on low-resolution infrared latent space features and visible light latent space features to obtain infrared global representations and visible light global representations. Modal classification supervision is performed based on the infrared global representation and the visible light global representation to obtain the modal classification probability.

5. The dual-guided infrared image super-resolution method based on a diffusion model according to claim 3, characterized in that, The process of generating a modulation time embedding based on a globally distributed modulation function, according to the global infrared representation and time steps, includes: Based on linear projection and gating control mechanism, modulation mapping is performed according to the infrared global representation to obtain the infrared prior modulation amount; The time encoding mechanism based on the diffusion model performs time embedding encoding according to the time step to obtain the original time embedding. Based on the globally distributed modulation function, the modulation time embedding is obtained by weighted fusion of the infrared prior modulation amount and the original time embedding.

6. The dual-guided infrared image super-resolution method based on a diffusion model according to claim 3, characterized in that, The method based on a local structure modulation function, which performs local structure-guided modulation according to low-resolution infrared latent space characteristics and time steps to obtain a structure correction amount, includes: Edge detection calculations are performed based on low-resolution infrared latent space features to obtain edge structure features; The edge gradient features are normalized to obtain normalized structural features; Based on the local structure modulation function, weighted modulation is performed according to the normalized structure characteristics and time step to obtain the structure correction amount.

7. The dual-guided infrared image super-resolution method based on a diffusion model according to claim 1, characterized in that, The loss function is calculated based on the infrared global representation, the visible light global representation, the modality classification probability, the preset Gaussian noise, and the predicted Gaussian noise to obtain the classification loss and the comprehensive loss, including: The classification loss is obtained by calculating the binary cross-entropy loss function based on the infrared global representation, the visible light global representation, and the modality classification probability. The reconstruction loss is obtained by calculating the mean square error based on the preset Gaussian noise and the predicted Gaussian noise. The comprehensive loss is obtained by weighting the classification loss and reconstruction loss.

8. A diffusion-model-based dual-guided infrared image super-resolution device, wherein the diffusion-model-based dual-guided infrared image super-resolution device is used to implement the diffusion-model-based dual-guided infrared image super-resolution method as described in any one of claims 1-7, characterized in that, The device includes: The data acquisition module is used to acquire low-resolution infrared images, visible light images of the same scene, and high-resolution infrared true images corresponding to the low-resolution infrared images. The denoising training module is used to obtain the time step; based on the preset Gaussian noise, the dual-guided super-resolution framework is denoised and diffused to obtain the infrared global representation, visible light global representation, modality classification probability and prediction Gaussian noise according to the time step, low-resolution infrared image, visible light image and high-resolution infrared ground truth image. The loss calculation module is used to calculate the loss function based on the infrared global representation, the visible light global representation, the modal classification probability, the preset Gaussian noise, and the predicted Gaussian noise, so as to obtain the classification loss and the comprehensive loss. The alternating optimization module is used to perform alternating step-by-step optimization of the dual-guided super-resolution framework by freezing some parameters based on the classification loss and the comprehensive loss, so as to obtain the optimized dual-guided super-resolution framework. The image reconstruction module is used to acquire the low-resolution infrared image to be processed; based on the low-resolution infrared image to be processed, image reconstruction is performed using an optimized dual-guided super-resolution framework.

9. A dual-guided infrared image super-resolution device, characterized in that, The dual-guided infrared imaging super-resolution device includes: processor; A memory storing computer-readable instructions that, when executed by the processor, implement the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains program code that can be invoked by a processor to execute the method as described in any one of claims 1 to 7.