An infrared image super-resolution method based on physical guidance visual autoregression

By constructing an infrared physical encoder and a scale-wise autoregressive generator, combined with a physical consistency verification layer and a four-stage training mechanism, the problem of physical inconsistency in infrared image super-resolution is solved, achieving high-quality infrared image reconstruction, which is suitable for precise infrared temperature measurement and quantitative analysis of thermal radiation.

CN122367744APending Publication Date: 2026-07-10CHINA UNIV OF MINING & TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA UNIV OF MINING & TECH
Filing Date
2026-06-05
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing infrared image super-resolution methods fail to effectively utilize the physical characteristics of infrared images, resulting in physical inconsistencies in the reconstruction results and failing to meet the needs of accurate infrared temperature measurement and quantitative analysis of thermal radiation.

Method used

By constructing an infrared physical encoder to extract multi-branch physical prior representations, designing a physical-guided scale-wise autoregressive generator, and establishing a physical consistency verification layer, we implement hard constraints on temperature range, soft constraints on radiation energy conservation, and thermal diffusion smoothing structural constraints. We also construct a four-stage training and physical annealing optimization mechanism to ensure the physical consistency and accuracy of the reconstruction results.

Benefits of technology

It significantly improves the temperature accuracy and energy conservation of infrared image reconstruction, enhances the precision of thermal target detection and the physical consistency of reconstructed images, and is suitable for resolution enhancement of infrared thermal imaging systems in fields such as industrial inspection, security monitoring, medical diagnosis and national defense.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122367744A_ABST
    Figure CN122367744A_ABST
Patent Text Reader

Abstract

This invention proposes a physically guided visual autoregressive method for infrared image super-resolution. An infrared physical encoder is constructed to extract temperature, radiation intensity, thermal gradient, and thermal diffusion fields from low-resolution images, forming a multi-branch physical prior representation. A physically guided scale-wise autoregressive generator is designed, embedding the physical priors as conditions into the image encoder. Spatial relationships are constrained by temperature gradient modulation of rotational position encoding, and a physical conditional cross-attention mechanism is introduced to promote feature interaction in thermally similar regions. A physical consistency verification layer is established, implementing hard constraints on temperature range, soft constraints on radiation energy conservation, and thermal diffusion smoothing structure constraints. This method effectively solves the physical inconsistency problem caused by the visual autoregressive model ignoring infrared physical characteristics, significantly improving the temperature accuracy, energy conservation, and physical consistency of reconstructed images. It is applicable to scenarios such as precise infrared thermometry, quantitative analysis of thermal radiation, and thermal target detection and recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision and infrared thermal imaging technology, specifically relating to an infrared image super-resolution method based on physical-guided visual autoregression, which is particularly suitable for scenarios that require ensuring the physical consistency of reconstruction results, such as precise infrared temperature measurement, quantitative analysis of thermal radiation, and detection and identification of thermal targets. Background Technology

[0002] With the widespread application of infrared thermal imaging technology in industrial inspection, security monitoring, medical diagnosis, and national defense, infrared image super-resolution reconstruction technology has become an important means to improve the spatial resolution of infrared imaging systems. However, due to limitations in infrared sensor manufacturing capabilities, cost, and device size, existing infrared images often suffer from low resolution, low contrast, and blurred edge details, severely restricting their application in high-value scenarios such as precise temperature measurement and quantitative analysis of thermal radiation.

[0003] In the field of super-resolution reconstruction technology, early methods mainly employed traditional approaches such as bicubic interpolation. While these methods could amplify image resolution, their low-pass characteristics resulted in the loss of high-frequency components, leading to blurred reconstructed images. In recent years, deep learning-based super-resolution methods have made significant progress, continuously improving the visual quality of reconstructed images. However, most of these methods are designed for visible light images, treating infrared images simply as ordinary grayscale images and completely ignoring the inherent physical characteristics of infrared images.

[0004] Infrared images differ fundamentally from visible light images in their imaging principles and physical characteristics: the pixel values ​​of infrared images represent temperature or radiation energy, and their imaging is based on the thermal radiation intensity of an object's surface rather than the intensity of reflected light. Infrared images possess unique properties such as blurred edges but clear thermal boundaries, noise primarily consisting of fixed pattern noise and non-uniform noise, and target characteristics dependent on temperature differences. Existing methods ignore these physical characteristics, leading to the following key problems: coarse-scale predictions cannot effectively locate thermal targets; the physical relationship between temperature and the fourth power of radiation intensity is disrupted during magnification; and the radiation characteristics of thermal targets are distorted after reconstruction, failing to meet the quantitative application requirements such as precise infrared thermometry.

[0005] Visual autoregressive models, as an emerging image generation paradigm, demonstrate powerful modeling capabilities in visible light image generation and super-resolution tasks by predicting image lexical sequences at each scale. Standard visual autoregressive models use low-resolution images as prefix lexical inputs to a Transformer, achieving progressive resolution improvement through multi-scale autoregression. However, this model has significant limitations when applied to infrared images: its positional encoding only considers spatial geometric relationships, neglecting thermal diffusion direction constraints; its attention mechanism is based solely on content similarity calculations, failing to consider the physical intuition of thermal conduction in regions with similar temperatures; and its generation process lacks explicit constraints on physical laws such as temperature range, conservation of radiative energy, and smooth thermal diffusion. Directly applying standard visual autoregressive models to infrared image super-resolution will lead to physical inconsistencies in the reconstruction results, failing to guarantee temperature accuracy and energy conservation.

[0006] Therefore, there is an urgent need for an infrared image super-resolution method that can fully utilize infrared physical priors, implement scale-by-scale physical guidance for generation, and apply physical consistency verification within a visual autoregressive framework, so as to achieve synergistic optimization of reconstruction quality and physical consistency and meet the practical application needs of infrared thermal imaging in the field of quantitative analysis. Summary of the Invention

[0007] This invention provides a physical-guided visual autoregressive method for infrared image super-resolution. The method constructs an infrared physical encoder to extract multi-branch physical prior representations, designs a physical-guided scale-wise autoregressive generator to implement thermal gradient modulation and interaction with thermally similar region features, and establishes a physical consistency verification layer to implement hard constraints on temperature range, soft constraints on radiation energy conservation, and thermal diffusion smoothing structure constraints. This supports high-resolution image reconstruction with physical consistency, energy conservation, and accurate temperature in scenarios such as precise infrared thermometry, quantitative analysis of thermal radiation, and thermal target detection and recognition.

[0008] To achieve the above objectives, this invention provides a physical-guided visual autoregressive method for infrared image super-resolution, comprising the following steps:

[0009] S1. The input low-resolution infrared image is extracted by a multi-branch physical prior through an infrared physical encoder. The original digital quantization value is mapped into a temperature field, a radiation intensity field, a thermal gradient field, and a thermal diffusion field to form a physical prior representation vector.

[0010] S2. A physical-guided, scale-wise autoregressive generator is used for high-resolution image reconstruction. The physical priors at each scale are embedded as conditions into the image encoder. Spatial relationships are constrained by temperature gradient modulation of rotational position encoding. A physical conditional cross-attention mechanism is introduced to promote the interaction of thermally similar region features, thereby achieving stepwise prediction from coarse to fine scale.

[0011] S3. Establish a physical consistency verification layer to implement hard constraints on temperature range, soft constraints on radiation energy conservation, and thermal diffusion smoothing structure constraints on the high-resolution images generated by autoregression, so as to ensure that the reconstruction results meet the laws of infrared physics.

[0012] S4. Construct a four-stage training and physical annealing optimization mechanism, which sequentially performs physical encoder pre-training, basic generator training, joint training and end-to-end fine-tuning. By progressively introducing physical constraint weights, we prevent premature limitation of the model's learning ability and achieve synergistic optimization of reconstruction quality and physical consistency.

[0013] Furthermore, in S1, multi-branch physical prior extraction is performed on the input low-resolution infrared image to form a physical prior representation vector, specifically including the following:

[0014] S1.1 Receive the input low-resolution infrared image ,in and These are the image height and width, respectively. The number of channels; the pixel values ​​of the low-resolution infrared image are the original digital quantization values ​​of the sensor, which need to be converted into physically meaningful temperature values ​​through physical calibration;

[0015] S1.2 Perform multi-branch physical encoding on the input image to generate four physical feature fields:

[0016] Temperature field branch: Converts the input low-resolution infrared image into an absolute temperature map by using sensor calibration parameters. The conversion relationship satisfies the sensor calibration curve. ,in This is a nonlinear mapping function obtained based on the calibration of a blackbody radiation source, with the output unit being Kelvin (K). The original digital quantization value of the sensor;

[0017] Radiation intensity field branch: Calculation of radiation intensity map based on Stefan-Boltzmann law The calculation formula is as follows:

[0018] ;

[0019] in, Emissivity based on the gray-body approximation assumption; The Stefan-Boltzmann constant has a value of [value missing]. ;

[0020] Thermal gradient field branch: Calculating temperature gradient maps using the Sobel operator. Simultaneously encode gradient magnitude and gradient direction ,in:

[0021] ;

[0022] Thermal diffusion field branch: Calculating temperature Laplace plot Used to locate heat sources and heat sinks:

[0023] ;

[0024] In this context, positive values ​​represent heat sources, and negative values ​​represent heat sinks.

[0025] S1.3. Concatenate the four physical feature maps along the channel dimension to generate a physical prior vector. This serves as the conditional input for the subsequent scale-wise autoregressive generator.

[0026] Furthermore, the high-resolution image reconstruction based on the physically guided scale-wise autoregressive generator in S2 specifically includes the following:

[0027] S2.1, The physical prior vector is divided by the scaling decomposer. Analysis into a series of ordered scales of physical conditions , of which Each scale corresponds to a resolution Physical prior subgraphs, at each scale This corresponds to a generation subtask that can be independently guided;

[0028] S2.2, The autoregressive generator is used for each scale Call the physical condition embedding module to embed the physical priors at the current scale. Image prefix features Merging into a unified representation of generation conditions Its calculation process satisfies:

[0029] ;

[0030] in, For image feature embedding function, For physical condition projection function, Indicates feature concatenation operation;

[0031] S2.3. In each Transformer layer at each scale, a temperature gradient modulated rotational position code is introduced, with standard scale aligned rotational position code. After temperature gradient modulation, it becomes:

[0032] ;

[0033] in, To map the thermal gradient to the phase-shifted modulation function, specifically:

[0034] thermal gradient direction Rotation angle offset mapped to position encoding:

[0035] ;

[0036] thermal gradient amplitude Mapped to position-coded amplitude modulation:

[0037] ;

[0038] in, These are learnable modulation coefficients; at thermal boundaries, the modulated position encoding exhibits thermal boundary spatial separation characteristics, preventing the model from confusing high-temperature and low-temperature regions.

[0039] S2.4. In each scale of the Transformer layer, a physically-conditional cross-attention mechanism is introduced to promote the interaction of features in thermally similar regions. Specifically, the standard self-attention computation is extended as follows:

[0040] ;

[0041] in, Attention bias matrix generated for physical priors For sequence length, The physical guidance intensity coefficient; the bias matrix is ​​calculated as follows:

[0042] ;

[0043] in, For the magnitude of attention bias, For temperature similarity bandwidth, This is the temperature difference threshold; for regions with similar temperatures, It promotes focused attention and facilitates feature interaction and information fusion between thermally similar regions; for regions with large temperature differences... Suppress attentional interactions to prevent feature confusion between thermally heterogeneous regions;

[0044] S2.5 During the generation phase, a physical constraint mask is applied to the output image term set, allowing only physically valid image term sets retrieved from the dataset. Preset temperature range whitelist Set of terminators The output probability distribution of the inequality is obtained by decoding the inequality in the set of inequalities and satisfies the following:

[0045] ;

[0046] in, For the set of image terms that are allowed to be generated, Image lexical The unnormalized log probability value after classification layer mapping. Parameters for model training;

[0047] S2.6 The scale incrementer generates the results of the current scale. As a prefix input for the next scale, it enables the transition from... arrive The progressively increasing resolution ultimately outputs high-resolution infrared images. ,in This is the super-resolution factor.

[0048] Furthermore, the physical consistency verification layer established in S3 performs multi-constraint verification on the reconstruction results, specifically including the following:

[0049] S3.1 High-resolution images generated by autoregression Implement hard temperature range constraints to crop pixel values ​​that exceed the physically reasonable range to the valid range:

[0050] ;

[0051] in, and These are the minimum and maximum effective temperatures calibrated by the sensor, ensuring that the reconstructed temperature is not lower than absolute zero and does not exceed the sensor's range; simultaneously, a non-negative constraint on radiation intensity is applied.

[0052] ;

[0053] S3.2, the cropped image Implement soft constraints on radiation energy conservation and calculate energy conservation losses:

[0054] ;

[0055] in, and The images show the radio intensity maps for high-resolution and low-resolution images, respectively. This is the super-resolution factor; this loss ensures that the total amplified radiated energy is conserved from the low-resolution input, assuming a constant emissivity. Exceeding the energy conservation loss threshold This triggers the energy correction mechanism, for Perform global brightness scaling:

[0056] ;

[0057] S3.3, Image after energy correction Implement thermal diffusion smoothing structural constraints and calculate thermal diffusion smoothing losses:

[0058] ;

[0059] in, For pixels The neighborhood of , typically 4 or 8 neighbors; weight Determined by the physical thermal diffusivity:

[0060] ;

[0061] in, This reflects the thermal diffusion characteristics of different media: the air region allows for a larger temperature difference due to slow thermal diffusion, while the solid region is forced to have a smaller temperature difference due to faster thermal diffusion; if Exceeding the thermal diffusion smoothing loss threshold Then anisotropic thermal diffusion filtering is applied to the image:

[0062] = ;

[0063] in, For based on The constructed diffusion tensor, This is the diffusion step size, and its value is greater than 0.

[0064] S3.4, The verifier verifies the final output. Implement double check and cyclic correction:

[0065] First, set the maximum number of iterations. Initialize the iteration index Let the current image to be verified Start processing the current step image Perform double verification;

[0066] Physical rationality verification: Verify that the temperature values ​​of all pixels meet the requirements. And the radiation intensity satisfies ;

[0067] Energy conservation verification: ,in The allowable energy conservation error threshold;

[0068] If both checks pass, then the current step image will be... As the final output image;

[0069] If any check fails and the iterative index... Then for Re-perform temperature range cropping, global scaling of radiant energy, and thermal diffusion smoothing filtering, and record the corrected image as follows: ,make And re-perform the double check;

[0070] like If the double verification still fails, the pixel or region will be downgraded to a bicubic interpolation result and marked as physically untrusted.

[0071] Furthermore, the four-stage training and physical annealing optimization mechanism constructed in S4 specifically includes the following:

[0072] S4.1 Constructing a four-stage training process:

[0073] Phase 1 – Physical Encoder Pre-training: Freeze generator parameters and train only the infrared physical encoder, optimizing the physical feature reconstruction loss.

[0074] ;

[0075] Among them, subscript and These represent the predicted value and the actual value, respectively. Weight the loss for each branch;

[0076] Phase Two – Basic Generator Training: Freeze the physical encoder and train the basic VARSR generator with the following optimization objective:

[0077] ;

[0078] in, For pixel-level L1 reconstruction loss, For the perceptual loss based on the pre-trained VGG network, Weights for perceived loss;

[0079] Phase 3 – Joint Training: Unfreeze the physical encoder and generator, jointly train the physical consistency verification layer, with the optimization objective being:

[0080] ;

[0081] in, and The time-varying physical constraint weights are gradually increased using a physical annealing strategy.

[0082] Phase 4 – End-to-End Fine-Tuning: Performing end-to-end fine-tuning with the complete objective function:

[0083] ;

[0084] in, To balance the hyperparameters of reconstruction quality and physical constraints;

[0085] S4.2 In the third stage, a physical annealing training strategy is implemented, with physical constraint weights introduced gradually according to an exponential growth pattern:

[0086] ;

[0087] ;

[0088] in, This represents the current number of training steps. and The annealing time constant is and The maximum constraint weights are used; the physical annealing training strategy is used to avoid the strong physical constraints from prematurely limiting the model's learning ability, so that the model first masters the basic generation ability and then gradually adapts to the physical laws.

[0089] S4.3. In the incremental training phase, a continuous learning regularization term is introduced to prevent catastrophic forgetting of historical data distributions: based on the old model. Approximate calculation of the Fisher information matrix And use this matrix as the parameter importance matrix; construct the Elastic Weight Consolidation (EWC) loss:

[0090] ;

[0091] in, These are the current model parameters. These are the optimal parameters for the old model. As a diagonal element, it represents the first element. The importance of each parameter; simultaneously generating soft-label distributions using the old model. As the output of the teacher model, the new model output is aligned using knowledge distillation loss constraints:

[0092] ;

[0093] The final joint optimization objective is:

[0094] ;

[0095] in, This is a regular intensity coefficient; through this mechanism, while adapting to infrared data in new scenarios, backward compatibility with existing sensor models, temperature ranges, and physical laws is maintained.

[0096] S4.4 Record the following multi-dimensional feedback signals during training: reconstruction quality metrics, including peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and temperature error. Physical consistency indicators, including energy conservation error. and thermal diffusion smoothness Downstream task metrics, including mean accuracy (mAP) for thermal target detection; all data are grouped into a quintuple of input-physical prior-generated results-verification feedback-downstream metrics. Archived data for continuous model optimization and adaptive adjustment of physical constraint weights.

[0097] Beneficial Effects: This invention extracts multi-branch physical prior representations such as temperature field, radiation intensity field, thermal gradient field, and thermal diffusion field by constructing an infrared physical encoder, integrating the physical essence of infrared imaging into the generation process. This effectively solves the physical inconsistency problem caused by the visual autoregressive model ignoring infrared physical characteristics, significantly improving the temperature accuracy and physical consistency of the reconstructed image. By designing a physically guided scale-wise autoregressive generator, the thermal boundary spatial relationship is constrained by temperature gradient modulated rotational position encoding, and a physical conditional cross-attention mechanism is introduced to promote the interaction of features in thermally similar regions, effectively improving the accuracy of thermal target positioning and the ability to maintain thermal boundaries. By establishing a physical consistency verification layer, hard constraints on temperature range, soft constraints on radiation energy conservation, and thermal diffusion smoothing structure constraints are implemented to ensure that the reconstruction results strictly meet the laws of infrared physics, significantly improving energy conservation and radiation characteristic fidelity. Furthermore, through a four-stage training and physical annealing optimization mechanism, progressive physical constraints are introduced through a physical annealing training strategy, combined with elastic weight consolidation and knowledge distillation to prevent catastrophic forgetting, maintaining backward compatibility with existing sensor models and physical laws while adapting to new scene data. In summary, this invention effectively solves the key problems of existing visual autoregressive models in infrared image super-resolution, such as physical inconsistency, poor temperature accuracy, and insufficient energy conservation. It significantly improves the physical consistency, temperature accuracy, and energy conservation of reconstructed images, demonstrating stronger robustness and reliability in high-value scenarios such as precise infrared temperature measurement, quantitative analysis of thermal radiation, and detection and identification of thermal targets. It can be widely used in resolution enhancement of infrared thermal imaging systems in fields such as industrial inspection, security monitoring, medical diagnosis, and national defense. Attached Figure Description

[0098] Figure 1 This is a schematic diagram of the overall process of the present invention;

[0099] Figure 2 This is a flowchart of the multi-branch physical prior extraction process for infrared physical encoders;

[0100] Figure 3 This is a schematic diagram of a physical-guided scale-wise autoregressive generator structure;

[0101] Figure 4 This is a flowchart of the physical consistency verification layer constraints;

[0102] Figure 5This is a flowchart of the four-stage training and physical annealing optimization mechanism; Detailed Implementation

[0103] The invention will now be further described with reference to the accompanying drawings.

[0104] Example

[0105] Furthermore, such as Figure 1 As shown, an infrared image super-resolution method based on physical-guided visual autoregression includes the following steps:

[0106] S1. The input low-resolution infrared image is extracted by a multi-branch physical prior through an infrared physical encoder. The original digital quantization value is mapped into a temperature field, a radiation intensity field, a thermal gradient field, and a thermal diffusion field to form a physical prior representation vector.

[0107] S2. A physical-guided, scale-wise autoregressive generator is used for high-resolution image reconstruction. The physical priors at each scale are embedded as conditions into the image encoder. Spatial relationships are constrained by temperature gradient modulation of rotational position encoding. A physical conditional cross-attention mechanism is introduced to promote the interaction of thermally similar region features, thereby achieving stepwise prediction from coarse to fine scale.

[0108] S3. Establish a physical consistency verification layer to implement hard constraints on temperature range, soft constraints on radiation energy conservation, and thermal diffusion smoothing structure constraints on the high-resolution images generated by autoregression, so as to ensure that the reconstruction results meet the laws of infrared physics.

[0109] S4. Construct a four-stage training and physical annealing optimization mechanism, which sequentially performs physical encoder pre-training, basic generator training, joint training and end-to-end fine-tuning. By progressively introducing physical constraint weights, we prevent premature limitation of the model's learning ability and achieve synergistic optimization of reconstruction quality and physical consistency.

[0110] Furthermore, such as Figure 2 As shown, in step S1, multi-branch physical prior extraction is performed on the input low-resolution infrared image to form a physical prior representation vector. The specific steps are as follows:

[0111] S1.1 Receive the input low-resolution infrared image ,in and These are the image height and width, respectively. The number of channels; the pixel values ​​of the low-resolution infrared image are the sensor's raw digital quantization values, which need to be converted into physically meaningful temperature values ​​through physical calibration; the physical calibration parameters are obtained by fitting the sensor's raw digital quantization values ​​collected by a blackbody radiation source at multiple known temperature points, and the calibration temperature range covers... to To ensure temperature conversion accuracy across the entire range;

[0112] S1.2 Perform multi-branch physical encoding on the input image to generate four physical feature fields:

[0113] Temperature field branch: Converting raw digital quantized values ​​into an absolute temperature map using sensor calibration parameters. The conversion relationship satisfies the sensor calibration curve. ,in The calibration function is a nonlinear mapping function obtained based on the calibration of a blackbody radiation source, and the output unit is Kelvin (K). The calibration curve is implemented by piecewise polynomial fitting or table lookup. For uncooled infrared focal plane array sensors, the calibration function needs to consider the detector response nonlinearity and ambient temperature drift compensation.

[0114] Radiation intensity field branch: Calculation of radiation intensity map based on Stefan-Boltzmann law The calculation formula is as follows:

[0115] ;

[0116] in, Emissivity based on the gray-body approximation assumption, The Stefan-Boltzmann constant has a value of [value missing]. For different material surfaces, emissivity It can be dynamically adjusted based on prior knowledge or the material recognition module;

[0117] Thermal gradient field branch: Calculating temperature gradient maps using the Sobel operator. Simultaneously encode gradient magnitude and gradient direction ,in:

[0118] ;

[0119] The Sobel operator uses Convolution kernel, the horizontal kernel is The vertical core is The gradient magnitude is normalized to Gaussian smoothing. interval;

[0120] Thermal diffusion field branch: Calculating temperature Laplace plot Used to locate heat sources and heat sinks:

[0121] ;

[0122] In this context, positive regions represent heat sources, and negative regions represent heat sinks; the Laplace operator is implemented using a 5-point difference scheme, and a mirror-fill strategy is used at the boundaries.

[0123] S1.3. Concatenate the four physical feature maps along the channel dimension to generate a physical prior vector. This serves as the conditional input for the subsequent scale-wise autoregressive generator; the splicing order is... Each channel is scaled to the same numerical range after being processed independently in batches.

[0124] Furthermore, such as Figure 3 As shown, the high-resolution image reconstruction based on the physically guided scale-wise autoregressive generator in S2 specifically includes the following:

[0125] S2.1, The physical prior vector is divided by the scaling decomposer. Analysis into a series of ordered scales of physical conditions , of which Each scale corresponds to a resolution Physical prior subgraphs, at each scale This corresponds to an independently guided generation subtask; the scale decomposition is achieved using bilinear interpolation downsampling, with a total number of scales. Based on the target super-resolution factor Decision, satisfaction ;

[0126] S2.2, The autoregressive generator is used for each scale Call the physical condition embedding module to embed the physical priors at the current scale. Image prefix features Merging into a unified representation of generation conditions Its calculation process satisfies:

[0127] ;

[0128] in, For image feature embedding functions, a lightweight convolutional encoder is used to map the prefix image to the latent space; The physical condition projection function is used, and a two-layer MLP is employed to map the physical prior to the same dimensional space as the image features. This indicates a feature concatenation operation, specifically implemented as channel concatenation followed by... Convolution dimensionality reduction;

[0129] S2.3. In each Transformer layer at each scale, a temperature gradient modulated rotational position code is introduced, with standard scale aligned rotational position code. After temperature gradient modulation, it becomes:

[0130] ;

[0131] in, To map the thermal gradient to the phase-shifted modulation function, specifically:

[0132] thermal gradient direction Rotation angle offset mapped to position encoding:

[0133] ;

[0134] thermal gradient amplitude Mapped to position-coded amplitude modulation:

[0135] ;

[0136] in, The initial values ​​of the learnable modulation coefficients are set as follows: And 0.5, adaptively adjusted through end-to-end training; at the thermal boundary, the modulated position encoding presents the thermal boundary spatial separation characteristics to prevent the model from confusing high temperature and low temperature regions;

[0137] S2.4. In each scale of the Transformer layer, a physically-conditional cross-attention mechanism is introduced to promote the interaction of features in thermally similar regions. Specifically, the standard self-attention computation is extended as follows:

[0138] ;

[0139] in, Attention bias matrix generated for physical priors For sequence length, The physical guidance strength coefficient is initially set to 0.3 and can be dynamically adjusted during training; the bias matrix is ​​calculated as follows:

[0140] ;

[0141] in, The attention bias magnitude is set to 0.1; The bandwidth for temperature similarity is set to 5K. The temperature difference threshold is set to 10K; for regions with similar temperatures, To promote focused attention and facilitate feature interaction and information fusion between thermally similar regions; for regions with large temperature differences, Suppress attentional interactions to prevent feature confusion between thermally heterogeneous regions;

[0142] S2.5 During the generation phase, a physical constraint mask is applied to the output image term set, allowing only physically valid image term sets retrieved from the dataset. Preset temperature range whitelist Set of terminators The output probability distribution of the inequality is obtained by decoding the inequality in the set of inequalities and satisfies the following:

[0143] ;

[0144] in, For the set of image terms that are allowed to be generated, Image lexical The unnormalized log probability value after classification layer mapping. The model training parameters; the physically reasonable image lexical set The subset of image terms containing temperature values ​​within the valid range is obtained from statistics of the training set; the temperature range whitelist. Dynamically generated based on sensor calibration range, ensuring that the temperature values ​​corresponding to the generated image terms meet the requirements. ;

[0145] S2.6 The scale incrementer generates the results of the current scale. As a prefix input for the next scale, it enables the transition from... arrive The progressively increasing resolution ultimately outputs high-resolution infrared images. ,in The super-resolution factor is used; the scale progression is achieved by combining nearest neighbor upsampling with convolutional thinning, and detailed information is passed between adjacent scales through residual connections.

[0146] Furthermore, such as Figure 4 As shown, the physical consistency verification layer established in S3 performs multi-constraint verification on the reconstruction results, specifically including the following:

[0147] S3.1 High-resolution images generated by autoregression Implement hard temperature range constraints to crop pixel values ​​that exceed the physically reasonable range to the valid range:

[0148] ;

[0149] in, and These are the minimum and maximum effective temperatures calibrated by the sensor, ensuring that the reconstructed temperature is not lower than absolute zero and does not exceed the sensor's range; simultaneously, a non-negative constraint on radiation intensity is applied.

[0150] ;

[0151] The cropping operation is performed immediately after each scale is generated to prevent outliers from accumulating and spreading during the progressive scaling process.

[0152] S3.2, the cropped image Implement soft constraints on radiation energy conservation and calculate energy conservation losses:

[0153] ;

[0154] in, and The images show the radio intensity maps for high-resolution and low-resolution images, respectively. This is the super-resolution factor; this loss ensures that the total amplified radiated energy is conserved from the low-resolution input, assuming a constant emissivity. This triggers the energy correction mechanism, for Perform global brightness scaling:

[0155] ;

[0156] in, The energy conservation tolerance threshold is set to 5% of the total radiant energy of the low-resolution input.

[0157] S3.3, Image after energy correction Implement thermal diffusion smoothing structural constraints and calculate thermal diffusion smoothing losses:

[0158] ;

[0159] in, For pixels The neighborhood of the denoted 'x' is set to a size of 4; the weights are... Determined by the physical thermal diffusivity:

[0160] ;

[0161] in, This reflects the thermal diffusion characteristics of different media: the air region allows for a larger temperature difference due to slow thermal diffusion, while the solid region is forced to have a smaller temperature difference due to faster thermal diffusion; typical values ​​are... =1.0, =0.5, =0.1; The region classification is automatically determined based on a temperature gradient threshold, where the gradient magnitude is greater than 0.1. The pixels are marked as thermal boundaries; if Exceeding the thermal diffusion smoothing loss threshold Then anisotropic thermal diffusion filtering is applied to the image:

[0162] = ;

[0163] in, For based on The constructed diffusion tensor, The diffusion step size is set to 0.1;

[0164] S3.4, The verifier verifies the final output. Implement double check and cyclic correction:

[0165] First, set the maximum number of iterations. Initialize the iteration index Let the current image to be verified Start processing the current step image Perform double verification;

[0166] Physical rationality verification: Verify that the temperature values ​​of all pixels meet the requirements. And the radiation intensity satisfies ;

[0167] Energy conservation verification: ,in The allowable energy conservation error threshold is set to 0.05;

[0168] If both checks pass, then the current step image will be... As the final output image;

[0169] If any check fails and the iterative index... Then for Re-perform temperature range cropping, global scaling of radiant energy, and thermal diffusion smoothing filtering, and record the corrected image as follows: ,make And re-perform the double check;

[0170] like If the double verification still fails, the pixel or region is downgraded to a bicubic interpolation result and marked as physically untrusted. The maximum number of iterations in the cyclic correction process... The value is set based on a balance between real-time and accuracy requirements, with a typical value being [value missing]. =3. The dual verification and cyclic correction mechanism ensures that the system maintains the physical security and consistency of the output under sensor malfunctions or extreme scenarios.

[0171] Furthermore, such as Figure 5 As shown, the four-stage training and physical annealing optimization mechanism constructed in S4 specifically includes the following:

[0172] S4.1 Constructing a four-stage training process:

[0173] Phase 1 – Physical Encoder Pre-training: Freeze generator parameters and train only the infrared physical encoder, optimizing the physical feature reconstruction loss.

[0174] ;

[0175] Among them, subscript and These represent the predicted value and the actual value, respectively. The loss weights for each branch are typically set to a value of [value to be filled in]. =0.1, =0.5, =0.3; the number of training rounds in this stage is set to 50, and the learning rate is 0.3. It uses the Adam optimizer;

[0176] Phase Two – Basic Generator Training: Freeze the physical encoder and train the basic VARSR generator with the following optimization objective:

[0177] ;

[0178] in, For pixel-level L1 reconstruction loss, For the perceptual loss based on the pre-trained VGG network, The perceptual loss weight is set to 0.1; the number of training epochs in this phase is set to 100, and the learning rate is [missing information]. ;

[0179] Phase 3 – Joint Training: Unfreeze the physical encoder and generator, jointly train the physical consistency verification layer, with the optimization objective being:

[0180] ;

[0181] in, and The physical constraint weights are time-varying and progressively increased using a physical annealing strategy; the number of training epochs in this stage is set to 80, and the base learning rate is... ;

[0182] Phase 4 – End-to-End Fine-Tuning: Performing end-to-end fine-tuning with the complete objective function:

[0183] ;

[0184] in, To balance the reconstruction mass and physical constraints, the hyperparameters typically take the value of [value missing]. =0.5, =0.3, =0.1; the number of training rounds in this stage is set to 30, and the learning rate is 0.1. ;

[0185] S4.2 In the third stage, a physical annealing training strategy is implemented, with physical constraint weights introduced gradually according to an exponential growth pattern:

[0186] ;

[0187] ;

[0188] in, This represents the current number of training steps. and The annealing time constants are set to 5000 and 3000 steps respectively; and The maximum constraint weights are set to 1.0 and 0.8 respectively; the physical annealing training strategy avoids premature limitation of the model's learning ability by strong physical constraints, allowing the model to first master basic generative capabilities and then gradually adapt to physical laws.

[0189] S4.3. In the incremental training phase, a continuous learning regularization term is introduced to prevent catastrophic forgetting of historical data distributions: based on the old model. Approximate calculation of the Fisher information matrix And use this matrix as the parameter importance matrix; construct the Elastic Weight Consolidation (EWC) loss:

[0190] ;

[0191] in, These are the current model parameters. These are the optimal parameters for the old model. As a diagonal element, it represents the first element. The importance of each parameter; simultaneously generating soft-label distributions using the old model. As the output of the teacher model, the new model output is aligned using knowledge distillation loss constraints:

[0192] ;

[0193] The final joint optimization objective is:

[0194] ;

[0195] in, This is the regular intensity coefficient, typically taking the value of [value missing]. , This mechanism allows for adaptation to infrared data in new scenarios while maintaining backward compatibility with existing sensor models, temperature ranges, and physical laws.

[0196] S4.4 Record the following multi-dimensional feedback signals during training: reconstruction quality metrics, including peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and temperature error. Physical consistency indicators, including energy conservation error. and thermal diffusion smoothness Downstream task metrics, including mean accuracy (mAP) for thermal target detection; all data are grouped into a quintuple of input-physical prior-generated results-verification feedback-downstream metrics. Archived data is used for continuous model optimization and adaptive adjustment of physical constraint weights; the archived data is stored in the training warehouse after being de-identified, supporting sample resampling and hard sample mining during incremental training.

[0197] Furthermore, in this embodiment, M is used. 3 The FD dataset is used as the model pre-training dataset, with image patches of predetermined size as input. Batch training is performed at predetermined batch sizes, and the physically guided visual autoregressive network is trained end-to-end according to a four-stage training mechanism. Specifically, in the first stage, the generator parameters are frozen, and only the infrared physical encoder is trained (50 training epochs) using the Adam optimizer. In the second stage, the physical encoder is frozen, and the basic VARSR generator is trained (100 training epochs). In the third stage, the physical encoder and generator are unfrozen, and the physical consistency verification layer is jointly trained (80 training epochs), with a physical annealing strategy to progressively increase the physical constraint weights. In the fourth stage, end-to-end fine-tuning is performed using the complete objective function (30 training epochs). The learning rate for each stage is preset according to a progressive strategy, decreasing gradually as training progresses. This progressive introduction of physical constraint weights prevents premature limitation of the model's learning ability, achieving synergistic optimization of reconstruction quality and physical consistency.

[0198] During model testing, low-resolution infrared images from the test set were input into the trained physical-guided visual autoregressive network for prediction. Peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and natural image quality evaluation index (NIQE) were used as quantitative metrics to validate and compare the model's reconstruction performance on the test set. Considering that infrared image super-resolution reconstruction needs to simultaneously ensure pixel-level fidelity, structural consistency, and visual quality, PSNR, SSIM, and NIQE were used to comprehensively evaluate the model's performance. Details are as follows:

[0199] Peak signal-to-noise ratio (PSNR) is used to measure the fidelity of the reconstructed high-resolution image to the real high-resolution image at the pixel level. The larger the PSNR, the smaller the pixel-level mean square error of the reconstructed image and the higher the reconstruction accuracy.

[0200] The Structural Similarity Index (SSIM) measures the overall similarity between a reconstructed image and a real image in terms of brightness, contrast, and structure. The closer the SSIM is to 1, the higher the structural fidelity and visual consistency of the reconstructed image.

[0201] The Natural Image Quality Evaluation (NIQE) metric is used to measure the naturalness and perceived quality of reconstructed images without reference. The smaller the NIQE, the more the reconstructed image conforms to the statistical characteristics of a natural image, and the better the perceived quality.

[0202] Comparative experiments were conducted with eight representative image super-resolution reconstruction methods, including SwinIR, DifIISR, HAT, MambaIR, MambaIRv2, IRSRMamba, DCMamba, and VARSR. 2x and 4x super-resolution reconstructions were performed on the Result-A and Result-C infrared super-resolution benchmark test sets, respectively. The comparison results are shown in the table below.

[0203]

[0204] In the table, the units for PSNR, SSIM, and NIQE are dB, dimensionless, and dimensionless, respectively. The best values ​​in each column are indicated in bold, and the second-best values ​​are indicated by underline. In the Result-A 2x super-resolution task, our method achieves PSNR and SSIM of 39.32 and 0.9440, respectively, significantly outperforming existing methods (the highest PSNR and SSIM values ​​among the comparative methods are 39.21 and 0.9425, respectively), while reducing NIQE to 5.0845, achieving the best image naturalness and perceptual quality. In the Result-C 2x task, our method achieves PSNR, SSIM, and NIQE of 40.16, 0.9552, and 4.9630, respectively, also outperforming all comparative methods. In the Result-A 4x super-resolution task, our method achieved PSNR and SSIM of 34.71 and 0.8578, respectively, significantly outperforming existing methods (the highest PSNR and SSIM values ​​among the comparative methods were 34.54 and 0.8561, respectively), while reducing the NIQE to 6.7578. In the Result-C 4x task, our method achieved PSNR, SSIM, and NIQE of 35.31, 0.8746, and 6.8356, respectively, also outperforming all comparative methods. The combined results of the four sets of experiments show that, in terms of reconstruction fidelity metrics (PSNR and SSIM) and visual quality metrics (NIQE), our method exhibits superior overall performance and stronger robustness across all magnifications and test scenarios. This verifies the effectiveness of the physically guided visual autoregressive framework in infrared image super-resolution reconstruction, achieving better visual quality while maintaining reconstruction accuracy and physical consistency.

[0205] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. The scope of protection of the present invention should be determined by the scope of protection of the appended claims.

Claims

1. A physical-guided visual autoregressive method for infrared image super-resolution, characterized in that, Includes the following steps: S1. The input low-resolution infrared image is extracted by a multi-branch physical prior through an infrared physical encoder. The original digital quantization value is mapped into a temperature field, a radiation intensity field, a thermal gradient field, and a thermal diffusion field to form a physical prior representation vector. S2. A physical-guided, scale-wise autoregressive generator is used for high-resolution image reconstruction. The physical priors at each scale are embedded as conditions into the image encoder. Spatial relationships are constrained by temperature gradient modulation of rotational position encoding. A physical conditional cross-attention mechanism is introduced to promote the interaction of thermally similar region features, thereby achieving stepwise prediction from coarse to fine scale. S3. Establish a physical consistency verification layer to implement hard constraints on temperature range, soft constraints on radiation energy conservation, and thermal diffusion smoothing structure constraints on the high-resolution images generated by autoregression, so as to ensure that the reconstruction results meet the laws of infrared physics. S4. Construct a four-stage training and physical annealing optimization mechanism, which sequentially performs physical encoder pre-training, basic generator training, joint training and end-to-end fine-tuning. By progressively introducing physical constraint weights, we prevent premature limitation of the model's learning ability and achieve synergistic optimization of reconstruction quality and physical consistency.

2. The infrared image super-resolution method based on physically guided visual autoregression according to claim 1, characterized in that, S1 includes the following steps: S1.1 Receive the input low-resolution infrared image ,in and These are the image height and width, respectively. The number of channels; the pixel values ​​of the low-resolution infrared image are the original digital quantization values ​​of the sensor, which need to be converted into physically meaningful temperature values ​​through physical calibration; S1.2 Perform multi-branch physical encoding on the input image to generate four physical feature fields: Temperature field branch: Converts the input low-resolution infrared image into an absolute temperature map by using sensor calibration parameters. The conversion relationship satisfies the sensor calibration curve. ,in This is a nonlinear mapping function obtained based on the calibration of a blackbody radiation source, with the output unit being Kelvin; The original digital quantization value of the sensor; Radiation intensity field branch: Calculation of radiation intensity map based on Stefan-Boltzmann law The calculation formula is as follows: ; in, Emissivity based on the gray-body approximation assumption; The Stefan-Boltzmann constant has a value of [value missing]. ; Thermal gradient field branch: Calculating temperature gradient maps using the Sobel operator. Simultaneously encode gradient magnitude and gradient direction ,in: ; Thermal diffusion field branch: Calculating temperature Laplace plot Used to locate heat sources and heat sinks: ; In this context, positive values ​​represent heat sources, and negative values ​​represent heat sinks. S1.

3. Concatenate the four physical feature maps along the channel dimension to generate a physical prior vector. This serves as the conditional input for the subsequent scale-wise autoregressive generator.

3. The infrared image super-resolution method based on physically guided visual autoregression according to claim 1, characterized in that, S2 includes the following steps: S2.1, The physical prior vector is divided by the scaling decomposer. Analysis into a series of ordered scales of physical conditions , of which Each scale corresponds to a resolution Physical prior subgraphs, at each scale This corresponds to a generation subtask that can be independently guided; S2.2, The autoregressive generator is used for each scale Call the physical condition embedding module to embed the physical priors at the current scale. Image prefix features Merging into a unified representation of generation conditions Its calculation process satisfies: ; in, For image feature embedding function, For physical condition projection function, Indicates feature concatenation operation; S2.

3. In each Transformer layer at each scale, a temperature gradient modulated rotational position code is introduced, with standard scale aligned rotational position code. After temperature gradient modulation, it becomes: ; in, To map the thermal gradient to the phase-shifted modulation function, specifically: thermal gradient direction Rotation angle offset mapped to position encoding: ; thermal gradient amplitude Mapped to position-coded amplitude modulation: ; in, These are learnable modulation coefficients; at thermal boundaries, the modulated position encoding exhibits thermal boundary spatial separation characteristics, preventing the model from confusing high-temperature and low-temperature regions. S2.

4. In each scale of the Transformer layer, a physically-conditional cross-attention mechanism is introduced to promote the interaction of features in thermally similar regions; specifically, the standard self-attention computation is extended to: ; in, Attention bias matrix generated for physical priors For sequence length, The physical guidance intensity coefficient; the bias matrix is ​​calculated as follows: ; in, For the magnitude of attention bias, For temperature similarity bandwidth, This is the temperature difference threshold; for regions with similar temperatures, To promote focused attention and facilitate feature interaction and information fusion between thermally similar regions; for regions with large temperature differences, Suppress attentional interactions to prevent feature confusion between thermally heterogeneous regions; S2.5 During the generation phase, a physical constraint mask is applied to the output image term set, allowing only physically valid image term sets retrieved from the dataset. Preset temperature range whitelist Set of terminators The output probability distribution of the inequality is obtained by decoding the inequality in the set of inequalities and satisfies the following: ; in, For the set of image terms that are allowed to be generated, Image lexical The unnormalized log probability value after classification layer mapping. Parameters for model training; S2.6 The scale incrementer generates the results of the current scale. As a prefix input for the next scale, it enables the transition from... arrive The progressively increasing resolution ultimately outputs high-resolution infrared images. ,in This is the super-resolution factor.

4. The infrared image super-resolution method based on physically guided visual autoregression according to claim 1, characterized in that, S3 includes the following steps: S3.1 High-resolution images generated by autoregression Implement hard temperature range constraints to crop pixel values ​​that exceed the physically reasonable range to the valid range: ; in, and These are the minimum and maximum effective temperatures calibrated by the sensor, ensuring that the reconstructed temperature is not lower than absolute zero and does not exceed the sensor's range; simultaneously, a non-negative constraint on radiation intensity is applied. ; S3.2, the cropped image Implement soft constraints on radiation energy conservation and calculate energy conservation losses: ; in, and The images show the radio intensity maps for high-resolution and low-resolution images, respectively. This is the super-resolution factor; this loss ensures that the total amplified radiated energy is conserved from the low-resolution input, assuming a constant emissivity. Exceeding the energy conservation loss threshold This triggers the energy correction mechanism, for Perform global brightness scaling: ; S3.3, Image after energy correction Implement thermal diffusion smoothing structural constraints and calculate thermal diffusion smoothing losses: ; in, For pixels The neighborhood of , typically 4 or 8 neighbors; weight Determined by the physical thermal diffusivity: ; in, This reflects the thermal diffusion characteristics of different media: the air region allows for a larger temperature difference due to slow thermal diffusion, while the solid region is forced to have a smaller temperature difference due to faster thermal diffusion; if Exceeding the thermal diffusion smoothing loss threshold Then anisotropic thermal diffusion filtering is applied to the image: = ; in, For based on The constructed diffusion tensor; This is the diffusion step size, and its value is greater than 0. S3.4, The verifier verifies the final output. Implement double check and cyclic correction: First, set the maximum number of iterations. Initialize the iteration index Let the current image to be verified Start processing the current step image Perform double verification; Physical rationality verification: Verify that the temperature values ​​of all pixels meet the requirements. And the radiation intensity satisfies ; Energy conservation verification: ,in The allowable energy conservation error threshold; If both checks pass, then the current step image will be... As the final output image; If any check fails and the iterative index... Then for Re-perform temperature range cropping, global scaling of radiant energy, and thermal diffusion smoothing filtering, and record the corrected image as follows: ,make And re-perform the double check; like If the double verification still fails, the pixel or region will be downgraded to a bicubic interpolation result and marked as physically untrusted.

5. The infrared image super-resolution method based on physically guided visual autoregression according to claim 1, characterized in that, S4 includes the following steps: S4.1 Constructing a four-stage training process: Phase 1 – Physical Encoder Pre-training: Freeze generator parameters and train only the infrared physical encoder, optimizing the physical feature reconstruction loss. ; Among them, subscript and These represent the predicted value and the actual value, respectively. Weight the loss for each branch; Phase Two – Basic Generator Training: Freeze the physical encoder and train the basic VARSR generator with the following optimization objective: ; in, For pixel-level L1 reconstruction loss, For the perceptual loss based on the pre-trained VGG network, Weights for perceived loss; Phase 3 – Joint Training: Unfreeze the physical encoder and generator, jointly train the physical consistency verification layer, with the optimization objective being: ; in, and The time-varying physical constraint weights are gradually increased using a physical annealing strategy. Phase 4 – End-to-End Fine-Tuning: Performing end-to-end fine-tuning with the complete objective function: ; in, To balance the hyperparameters of reconstruction quality and physical constraints; S4.2 In the third stage, a physical annealing training strategy is implemented, with physical constraint weights introduced gradually according to an exponential growth pattern: ; ; in, This represents the current number of training steps. and The annealing time constant is and The maximum constraint weights are used; the physical annealing training strategy is used to avoid the strong physical constraints from limiting the model's learning ability too early, so that the model can first master the basic generation ability and then gradually adapt to the physical consistency constraints. S4.

3. In the incremental training phase, a continuous learning regularization term is introduced to prevent catastrophic forgetting of historical data distributions: based on the old model. Approximate calculation of the Fisher information matrix This matrix is ​​then used as the parameter importance matrix; an elastic weighted consolidation loss is constructed. as follows: ; in, These are the current model parameters. These are the optimal parameters for the old model. As a diagonal element, it represents the first element. The importance of each parameter; simultaneously generating soft-label distributions using the old model. As the output of the teacher model, the new model output is aligned using knowledge distillation loss constraints: ; The final joint optimization objective is: ; in, This is a regular intensity coefficient; through this mechanism, while adapting to infrared data in new scenarios, backward compatibility with existing sensor models, temperature ranges, and physical laws is maintained. S4.4 Record the following multi-dimensional feedback signals during training: reconstruction quality metrics, including peak signal-to-noise ratio, structural similarity index, and temperature error. Physical consistency indicators, including energy conservation error. and thermal diffusion smoothness Downstream task metrics, including average accuracy of thermal target detection; all data are grouped into a five-tuple: input - physical prior - generated result - verification feedback - downstream metrics. Archived data for continuous model optimization and adaptive adjustment of physical constraint weights.