Unsupervised infrared polarimetric super-resolution reconstruction method based on gating and physical constraints

CN122550362APending Publication Date: 2026-08-11CHANGCHUN UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-01
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

这种“合成退化”无法模拟真实红外成像中复杂的光学衍射模糊、非均匀性噪声及探测器响应

Benefits of technology

(1)本发明引入短波红外图像作为可见光图像与长波红外图像之间的结构一致性参考,并通过门控机制对可见光高频特征进行选择性筛选。与直接融合可见光信息的方法相比,本发明能够在保留有效结构细节的同时,抑制可见光中与长波红外热辐射分布不一致的纹理信息,降低跨模态融合过程中的模态污染,从而提高红外偏振超分辨率重建结果的真实性和可靠性,有效抑制跨模态误引导造成的非物理纹理注入。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122550362A_ABST
    Figure CN122550362A_ABST
Patent Text Reader

Abstract

An Unsupervised Infrared Polarization Super-Resolution Reconstruction Method Based on Gated and Physical Constraints. This method relates to the fields of photoelectric imaging and remote sensing imaging technology, specifically an unsupervised infrared polarization super-resolution reconstruction method based on gated and physical constraints. The method includes the following steps: acquiring images of four different dimensions and spectral bands within the same field of view using an imaging device; performing cross-modal geometric registration on the four images to construct guiding data in a unified coordinate system; performing SWIR-gated cross-modal filtering; constructing a two-stream asymmetric polarization reconstruction structure, guiding polarization component reconstruction through high-level semantic features of the intensity component; constructing a composite target total loss function using physical period consistency loss, cross-modal structure loss, polarization correlation loss, and auxiliary regularization terms; iteratively training the two-stream asymmetric polarization reconstruction structure to obtain the final reconstructed image, and calculating the degree of linear polarization and polarization angle based on the final reconstructed image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of photoelectric imaging and remote sensing imaging technology, specifically to an unsupervised infrared polarization super-resolution reconstruction method based on gating and physical constraints. Background Technology

[0002] With the development of infrared stealth coatings and thermal camouflage technologies, the temperature contrast between the target and the background has been significantly compressed. Traditional infrared thermal imagers, which rely solely on intensity information, have seen a marked decline in detection capabilities under conditions of low temperature differences and complex backgrounds. To compensate for this deficiency, infrared polarization imaging has gradually become an important supplement to traditional thermal imaging. Unlike images that only reflect the intensity of thermal radiation, polarization information is also related to the material, roughness, and observation geometry of the object's surface, providing clues to structural and surface properties that are not easily revealed in intensity images. Therefore, infrared polarization imaging technology can improve target recognition and scene understanding capabilities in complex scenarios.

[0003] Despite the significant advantages of infrared polarization imaging, its practical applications are still limited by detector resolution and image quality. Currently, real-time infrared polarization imaging primarily employs division of focal plane (DoFP) polarization detectors in the infrared band. Due to their compact structure and real-time imaging capabilities, they have become the mainstream detection method. These detectors typically integrate multiple micro-polarizers with different polarization directions within the same focal plane, enabling the simultaneous acquisition of intensity images in multiple polarization directions under a single exposure, thus achieving instantaneous detection of polarization information.

[0004] However, DoFP polarization detectors acquire polarization states in different directions by integrating micro-polarization arrays on pixels, essentially at the expense of spatial sampling density. The effective resolution of a single polarization channel is typically only one-quarter that of the original detector, while instantaneous field-of-view (IFOV) errors between different polarization pixels and detector non-uniform noise further degrade image quality and polarization solution stability.

[0005] Chinese invention patent application "A Super-Resolution Reconstruction Method and System for Visible Light Feature Transfer to Infrared" (CN119107233A) uses an image registration module to register infrared and visible light images. A feature extraction module converts the registered visible light and infrared images into two high-dimensional feature vectors. A feature transfer module performs feature fusion processing on the high-dimensional feature vectors of the registered visible light and infrared images. A super-resolution reconstruction module extracts the residual of the high-dimensional mixed feature vector after visible light feature transfer. Based on the extracted residual, super-resolution reconstruction is performed. The super-resolution reconstruction result effectively restores weak texture areas and increases visible light texture feature details.

[0006] Chinese invention patent application "Super-resolution method for infrared polarization images based on cross-attention dual-branch network" (CN120876234A) utilizes two infrared polarization cameras with different image resolutions to acquire two infrared polarization images of different resolutions in the same scene; matches the two infrared polarization images of different resolutions to obtain multiple image datasets; constructs a dual-branch super-resolution network model, trains the dual-branch super-resolution network model based on the multiple image datasets until a robust fitting state is reached, and obtains the trained dual-branch super-resolution network model; and uses the dual-branch super-resolution network model to obtain a high-resolution infrared polarization image from the low-resolution infrared polarization image to be super-resolution.

[0007] However, existing technologies still have problems, mainly in the following two aspects: Infrared polarization super-resolution techniques guided by high-resolution visible light images often employ a "forced fusion" mechanism, blindly injecting high-frequency details from visible light into the infrared network. This easily leads to the erroneous transfer of non-physical textures unique to visible light to the infrared thermal image, causing severe "modal contamination." This visual "pseudo-clarity" directly destroys the numerical accuracy of the Stokes vector, causing subsequent polarization-based inversion algorithms to completely fail.

[0008] On the other hand, most current methods rely on supervised learning frameworks, and their performance is highly dependent on paired "low-resolution-high-resolution" ground truths. However, obtaining high-resolution ground truths for dynamic scenes in the infrared polarization domain is extremely difficult, with challenges such as data acquisition and calibration difficulties, limited data volume or high costs, and limited super-resolution magnification. Existing research often synthesizes training data by manually downsampling static images (e.g., bicubic downsampling). This "synthetic degradation" cannot simulate the complex optical diffraction blur, non-uniform noise, and detector response in real infrared imaging. Therefore, models trained on synthetic data face significant domain differences when applied to real low-quality infrared data, resulting in a severe decrease in generalization ability. There is an urgent need for a real-domain reconstruction method that does not rely on paired ground truths and conforms to physical laws. Summary of the Invention

[0009] To address the aforementioned problems, the purpose of this invention is to propose an unsupervised infrared polarization super-resolution reconstruction method based on gating and physical constraints. This method does not rely on high-resolution polarization ground truth and can utilize multimodal information for guidance, but it does not inject false non-polarized textures, while ensuring the physical consistency of polarization.

[0010] The method includes the following steps: S1. Acquire low-resolution long-wave infrared polarization images within the same field of view. Long-wave infrared intensity reference image High-resolution visible light images and shortwave infrared images ; S2. For the images obtained in step S1, perform cross-modal geometric registration to construct a multimodal guided image group in the same coordinate system. ; S3, based on and Constructing a gated mask and will Acting on and The final guide image is obtained. ; S4, based on A dual-stream asymmetric polarization reconstruction structure is constructed. By using high-level semantic features of the intensity component, the polarization component is reconstructed to obtain the reconstructed intensity component. and the reconstructed polarization components and ; S5, based on , and A composite objective total loss function is constructed using physical period consistency loss, cross-modal structure loss, polarization correlation loss, and auxiliary regularization term; S6. Iteratively train the dual-stream asymmetric polarization reconstruction structure to obtain the final reconstructed image. ; according to , and Calculate the degree of linear polarization DoLP and the polarization angle AoP.

[0011] Furthermore, gated masks The construction method is as follows: ,in, It is the Sigmoid activation function. for Convolutional layer For channel splicing operations; Final boot image The formula for calculation is: ,in, This is element-wise multiplication.

[0012] Furthermore, the dual-flow asymmetric polarization reconstruction structure includes: an intensity flow branch and a polarization flow branch; by The original resolution is used as the input for the two-stream asymmetric polarization reconstruction structure. , For the original intensity component, and All are original polarization components; The intensity flow branch includes an intensity flow branch backbone network module, which outputs high-level semantic features of the intensity components. , and After residual stitching, the reconstructed intensity components are output. ; The polarization flow branch includes a polarization flow branch backbone network module and a polarization refinement module. The polarization flow branch backbone network module outputs high-level semantic features of the polarization components. and , and and and The polarization thinning module is input together to obtain polarization thinning features. , respectively with and After residual stitching, the reconstructed polarization components are output. and .

[0013] Furthermore, the calculation formula for the intensity flow branch backbone network module is as follows: ,in, include: , and , For spatial feature transformation function, As the input to the backbone network module, it is generated in the intensity flow branch by... as well as Calculations show that in the polarized flow branch, by , , as well as Calculations show that Scaling factor The translation factor; The structure of the intensity flow branch backbone network module is the same as that of the polarization flow branch backbone network module.

[0014] Furthermore, the formula for calculating the physical cycle consistency loss is:

[0015] in, for L 1-norm, For degenerate operators, For the output of the intensity flow branch, For the input of the intensity flow branch, For the output of the polarization flow branch, This is the input for the polarization flow branch.

[0016] Furthermore, cross-modal structural loss The formula for calculation is:

[0017] in, For Sobel edge operators, for High-frequency structural information, and They represent and The square of, for High-frequency structural information.

[0018] Furthermore, polarization correlation loss The formula for calculation is:

[0019] in, The index value of the pixel. for The pixels are spatial weights constructed using DoLP. It is a constant. In order to be in At the pixel point, The result of the degradation operation. for L 2-norm.

[0020] Furthermore, auxiliary regularization terms The formula for calculation is:

[0021] in, For total variation loss, To perceive style loss, and This represents the loss coefficient for the corresponding parameter.

[0022] Furthermore, the composite objective total loss function The formula for calculation is:

[0023] in, , , and This represents the loss coefficient for the corresponding parameter.

[0024] Furthermore, the formula for calculating the degree of linear polarization (DoLP) is: ,in, For the final intensity component, and For the final polarization component; The formula for calculating the polarization angle AoP is: .

[0025] The beneficial effects of the method described in this invention are as follows: (1) This invention introduces short-wave infrared images as a structural consistency reference between visible light images and long-wave infrared images, and selectively filters high-frequency features of visible light through a gating mechanism. Compared with methods that directly fuse visible light information, this invention can suppress texture information in visible light that is inconsistent with the thermal radiation distribution of long-wave infrared while preserving effective structural details, reducing modal contamination during cross-modal fusion, thereby improving the authenticity and reliability of infrared polarization super-resolution reconstruction results and effectively suppressing non-physical texture injection caused by cross-modal misguidance.

[0026] (2) This invention constructs a physical degradation closed loop that includes point spread function convolution, downsampling, and noise modeling, remapping the high-resolution prediction results output by the network to the low-resolution observation space and constraining consistency with the real low-resolution input. Therefore, this invention does not rely on paired low-resolution / high-resolution training samples and can directly use real-collected low-resolution infrared polarization data for training, reducing the problem of reduced generalization performance caused by inconsistency between artificially synthesized degradation and real imaging degradation.

[0027] (3) To address the polarization distortion problem that easily occurs in infrared polarization super-resolution reconstruction, this invention introduces polarization correlation constraints and polarization degree spatial continuity constraints to limit the directional relationship, amplitude relationship, and spatial distribution between polarization components in the Stokes parameters. This reduces problems such as energy inconsistency, polarity reversal, and polarization angle jumps that occur during the reconstruction process, making the reconstructed DoLP and AoP results more stable and improving the reliability of subsequent target recognition, camouflage detection, and material property analysis. Attached Figure Description

[0028] Figure 1 This is a schematic diagram of the four coaxial image acquisition systems with different dimensions and spectral bands described in this invention; Figure 2 This is a flowchart of the method described in this invention; Figure 3 This is a schematic diagram of the network structure of the method described in this invention; Figure 4The following are schematic diagrams of the super-resolution results and local magnified details of the method described in this invention: (a) is a schematic diagram of the super-resolution results and local magnified details of a geometric building (roof); (b) is a schematic diagram of the super-resolution results and local magnified details of a thermal radiation target (human figure). Figure 5 This is a schematic diagram illustrating the super-resolution reconstruction effect of the method described in this invention and the comparative method in a vehicle scene; Figure 6 This is a schematic diagram illustrating the super-resolution reconstruction effect of the method described in this invention and the comparative method in a building facade scene. Detailed Implementation

[0029] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0030] Example 1 This embodiment provides an unsupervised infrared polarization super-resolution reconstruction method based on gating and physical constraints. The flowchart of the method is as follows: Figure 1 As shown, the method includes the following steps: S1. Acquire low-resolution long-wave infrared polarization images within the same field of view. Long-wave infrared intensity reference image High-resolution visible light images and shortwave infrared images ; The relevant operations in step S1 will be introduced with specific examples: like Figure 2 As shown, a multi-band aperture parallel optical axis acquisition system was constructed. The system is integrated into a unified high-precision optical platform and achieves coaxial alignment and synchronous exposure to acquire images of four different dimensions and spectral bands within the same field of view. The system includes the following four independent camera modules: 1) LWIR-Pol (Long-wave Infrared DoFP Polarization Imaging Camera): The detector resolution is 640×512, and after polarization calculation, a 320×256 polarization image is obtained, which serves as the actual low-resolution input for the network. ; 2) LWIR-Base (Long-wave Infrared Intensity Imaging Camera): Operating wavelength 8–14 μm, 640×512 resolution, provides a thermal radiation reference for acquiring long-wave infrared intensity reference (LWIR-Base) images. ; 3) VIS (Visible Light Imaging Camera): 2448×2048 resolution, providing high-resolution geometric edge and texture priors for acquiring high-resolution visible light (VIS) images. ; 4) SWIR (Shortwave Infrared Imaging Camera): Operating wavelength 0.9–1.7μm, 640×512 resolution, auxiliary gating to filter cross-modal high-frequency features, used to acquire shortwave infrared (SWIR) images. .

[0031] S2. For the images obtained in step S1, perform cross-modal geometric registration to construct a multimodal guided image group in the same coordinate system. ; The relevant operations in step S2 will be introduced with specific examples: To address the disparity and resolution differences in multimodal systems, this embodiment uses low resolution. Using the original field of view as a reference (without spatial interpolation or resampling to preserve true blur, noise, and sampling degradation characteristics), , and Mapping to a unified coordinate system to construct a multimodal guided image group Specifically, distortion correction and baseline adjustment are first performed on the images of each modality; then, stable edge and corner features are extracted using phase consistency, and sub-pixel refinement of control points is performed using mutual information; finally, the cross-modal homography matrix is ​​estimated using the least squares method. Under a 4×10⁻⁶ super-resolution setting, at a resolution of 320×256... Construct a high-resolution virtual canvas of 1280×1024 for the anchor point, and... , and Mapped into this unified coordinate system ,and Keep the original resolution as network input , .

[0032] After completing the above registration and mapping, a scale-synchronized sliding window strategy is adopted. and Image patches were synchronously cropped at corresponding positions and combined with data augmentation operations such as random flipping and rotation to construct approximately 1200 training sample pairs; each sample pair consisted of a low-resolution polarization input patch (taken from...). ) and a spatially aligned multimodal guide image patch (taken from )composition.

[0033] S3, based on and Constructing a gated mask and will Acting on and The final guide image is obtained. ; The relevant operations in step S3 will be introduced with specific examples: To improve the reliability of cross-modal guidance, this embodiment introduces... As an intermediate reference mode, an SGM (short-wave infrared gating module) was designed to spatially filter high-resolution visible light features. When and Consistent structure, retain Details; Inconsistency, Suppression Texture.

[0034] like Figure 3 As shown, the input to the SGM includes a high-resolution visible light image. and shortwave infrared images Spatial gated masks are generated through convolutional mapping and sigmoid activation. :

[0035] Subsequently, the mask was applied to Image, and injected into LWIR baseline features in the form of residuals. In the process, the final guide image is obtained. :

[0036] in, This represents the Sigmoid activation function. This represents element-wise multiplication. This module represents the channel concatenation operation. It learns a spatial gating mask by concatenating two modal features and feeding them into a convolutional layer. Furthermore, the visible light characteristics Perform feature space filtering: when and When the structures are consistent, the mask is preserved. Details; mask suppression when the two structures are inconsistent. Redundant textures; for Convolutional layer. It is the intensity benchmark for long-wave infrared radiation, which contains the extremely important distribution of "absolute temperature / radiant energy". It acts as a semantic filter, cross-modal features in the image and When a local area exhibits high consistency, the gated mask tends to retain more... Structural details; conversely, suppressing inconsistent textures that may cause modal shifts.

[0037] S4, based on A dual-stream asymmetric polarization reconstruction structure is constructed. By using high-level semantic features of the intensity component, the polarization component is reconstructed to obtain the reconstructed intensity component. and the reconstructed polarization components and ; The relevant operations in step S4 will be introduced with specific examples: like Figure 3 As shown, the dual-flow asymmetric polarization reconstruction structure includes an intensity flow branch and a polarization flow branch; The intensity flow branch uses the original intensity components at low resolution. As input, the focus is on recovering radiation intensity and object contours. This branch receives guided features from the SGM through an SFT layer (backbone network module). By using affine transformation to inject guiding textures into infrared features, the reconstructed intensity components are obtained. ; like Figure 3 As shown, the backbone network employs four residual Swin Transformer (RSTB) blocks based on the Swin Transformer. To effectively inject guiding features into low-resolution infrared features, each branch utilizes a Spatial Feature Transform (SFT) layer. The SFT layer generates affine modulation parameters based on the guiding features, thereby enhancing texture details while preserving the original infrared distribution.

[0038] The formula for calculating the backbone network module is: ,in, The inputs are for the backbone network modules, including the inputs for the intensity flow branch backbone module and the polarization flow backbone module; In the intensity flow branch, The acquisition process is as follows: in the intensity flow branch, The signal passes sequentially through the upsampling module (using cubic splines to achieve 4x upsampling) and the first branch of the intensity flow. After processing by the convolutional layer, intensity features are obtained. , and exist In the module, after channel splicing, it then passes through 5 channel units and a Hadamard Product module in sequence to obtain the result. ; and The SFT layer is based on the guiding features. The generated scaling factor and translation factor The output of the backbone network module, Including the output of the intensity flow branch backbone network module and the output of the polarization flow branch backbone network module and , High-level semantic features representing intensity components, and High-level semantic features representing polarization components.

[0039] After the second branch of intensity flow After processing by the convolutional layer, and The output of the intensity flow branch is obtained through residual connection: the reconstructed intensity component. .

[0040] Each channel unit passes through the following sequence from input to output: Convolutional layers and channel attention mechanism module (Squeeze-and-Excitation, SE).

[0041] Polarization flow branch with low-resolution polarization components As input, the focus is on recovering subtle polarization differences. To enhance the robustness of polarization features, this branch employs a cascaded fusion strategy: not only receiving... It also receives high-level semantic features from the intensity flow branch through cross-branch connections. By utilizing structural priors based on intensity information to assist polarization reconstruction, the reconstructed polarization components are obtained. and .

[0042] Considering the characteristics of polarization signals, such as low signal-to-noise ratio and polarity sensitivity, this embodiment introduces a polarization refinement block (PRB) at the end of the polarization flow branch to suppress local abnormal polarization response and enhance the consistency of polarization edge.

[0043] like Figure 3 As shown, in the polarization flow branch, and After upsampling module and polarization flow branch first After processing by the convolutional layer, polarization features are obtained. , , and Passing by together The module processes the data to obtain the input of the polarization flow branch backbone network module. The input to the polarization flow branch backbone network module is processed by the polarization flow branch backbone network module. The processed data from the polarization flow branch backbone network module is then compared with... Inputting both into the PRB module yields polarization refinement features. .

[0044] After the second polarization flow branch After processing by the convolutional layer, respectively with and The output of the polarization flow branch is obtained through residual connection: the reconstructed polarization components. and .

[0045] Module structure and The modules are consistent.

[0046] S5, based on , and A composite objective total loss function is constructed using physical period consistency loss, cross-modal structure loss, polarization correlation loss, and auxiliary regularization term; The relevant operations in step S5 will be introduced with specific examples: To simultaneously ensure the consistency of content, structural clarity, and polarization physical reliability of the reconstruction results, a composite objective total loss function is adopted, consisting of physical period consistency loss, cross-modal structure loss, polarization correlation loss, and auxiliary regularization term. . Defined as:

[0047] in, For physical cycle consistency loss, For cross-modal structural loss, For polarization-dependent loss, To assist regularization terms, , , and For the loss coefficients of the corresponding parameters, in this embodiment , , , .

[0048] Physical cycle consistency loss To constrain the physical consistency between network outputs and input observations, a cycle consistency loss (physical cycle consistency loss) is defined:

[0049] in, This represents the reconstructed polarization components. include: and , For the output of the intensity flow branch ( ), Input for the intensity flow branch ( ), For the output of the polarization flow branch ( and ), Input for polarization flow branch ( and ), For degenerate operators, for L 1-norm.

[0050] Explicit physical degradation operator Its purpose is to construct an unsupervised training loop, which will convert the predicted images generated by the network into a closed loop. ( and Projected back into the low-resolution observation space:

[0051] Where * denotes the point spread function (PSF). k Convolution operation, The scale is represented as s Downsampling (in this embodiment, s =4), n For simulated thermal noise ( n =0.43DN). This design enables the network training process to explicitly consider blur, sampling, and noise degradation in LWIR imaging, through constraints. With the original input The consistency ensures the closed-loop constraints of the reconstruction results' content, energy distribution, and physical interpretability.

[0052] Cross-modal structural loss: This embodiment constructs a cross-modal structural loss in the gradient domain:

[0053] in, This represents the Sobel edge operator. The loss function has two independent constraints: intensity and polarization. The first term influences the infrared intensity... Absorption Multimodal Guidance Map High-frequency structural information .in for The high-frequency characteristics after Fourier transform. The second term represents the amplitude of linearly polarized radiation. The gradient distribution is constrained by the radiation reference. High-frequency structural information On L, and They represent and The square of.

[0054] Polarization dependence loss: Considering polarization components and The sign direction directly affects the computational stability of AoP. This embodiment constructs a polarization correlation loss based on local cosine similarity at the LR scale, for pixels... The polarization vector is defined as:

[0055] The corresponding correlation loss is:

[0056] in In order to be in At the pixel point, Perform degeneracy operations ( The result of ) This indicates low-resolution input. include and , To prevent division by zero errors, numerically stable terms, For pixels The space weights are constructed using DoLP. , 2 is L 2-norm.

[0057] Auxiliary regular expression terms: This embodiment introduces an auxiliary loss term. This term is composed of the Total Variation Loss (TVS). L TV ) and Perceptual Style Loss Composition: ,in, and In this embodiment, the loss coefficients for the corresponding parameters are... , , L TVApplying to the DoLP plot to constrain the local smoothness of the polarization degree distribution on the surface of the same object:

[0058] Furthermore, this embodiment introduces texture statistical constraints based on the Gram matrix:

[0059] in This represents the pre-trained VGG feature extraction network used for loss function calculation. This represents the Gram matrix, which is used to statistically constrain the local texture distribution of the reconstruction results to help improve the naturalness of the structure and reduce artifacts.

[0060] S6. Iteratively train the dual-stream asymmetric polarization reconstruction structure to obtain the final reconstructed image. ; according to , and Calculate the degree of linear polarization DoLP and the polarization angle AoP; The relevant operations in step S6 will be introduced with specific examples: The 1200 datasets constructed in step S2 are randomly divided proportionally into a training set (1000), a validation set (100), and a test set (100). This embodiment is implemented based on PyTorch and trained on a single RTX 4090 GPU. The AdamW optimizer is used for training, with β1=0.9, β2=0.999, and weight decay of 0.01. The initial learning rate is set to 1×10⁻⁶. -4 And a cosine annealing strategy is used to decay it to 1×10. -7 The batch size is set to 4, and the total number of training rounds is 200.

[0061] During iterative training, based on the composite objective total loss function Optimize shortwave infrared gating module, Module, Module, backbone network module, second intensity flow branch Convolutional layer and polarization flow branch second Convolutional layer.

[0062] After 200 iterations, the high-frequency final intensity component is output. and high-frequency final polarization components and .

[0063] according to , and Calculate the degree of linear polarization DoLP and the polarization angle AoP.

[0064]

[0065] Example 2 This embodiment is a further limitation of Embodiment 1.

[0066] To fully verify the performance of the method described in Example 1, six representative algorithms were selected for comparison: Traditional interpolation and physical reconstruction: Bicubic interpolation and Newton polynomial interpolation commonly used in DoFP desmosaic. Infrared polarization single-frame super-resolution (SISR): The representative network SwinIPISR was selected; Multimodal Guided Super-Resolution (RefSR / GuidedSR): Includes the classic feature matching reference super-resolution model MASA-SR, as well as the state-of-the-art deep learning frameworks SwinFuSR and GuidedSR designed for multimodal fusion.

[0067] In this embodiment, the training strategy of the method of the present invention is the same as that of Embodiment 1. The training strategy of the comparative method is as follows: Since existing supervised infrared polarization super-resolution methods rely on training with pairs of high- and low-resolution samples, and it is difficult to obtain high-resolution polarization ground truth in real-world scenarios, this embodiment constructs pseudo-paired training data for the supervised baseline: a real 128×128 LWIR polarization image patch is used as a pseudo-high-resolution target, and a corresponding low-resolution input is generated through Gaussian blur, 4× bicubic downsampling, and additive Gaussian noise.

[0068] During the testing phase, all methods uniformly input real LWIR low-resolution images that have not been artificially degraded, and uniformly output 4× super-resolution results for blind evaluation.

[0069] (1) Multimodal guidance and verification of real-domain super-resolution effect To visually demonstrate the reconstruction capability of the method described in this invention by fusing multimodal features in real-world complex scenarios, Figure 4 The super-resolution results and magnified details of the model on representative geometric buildings (roof skeleton, Fig. 4(a)) and thermal radiation targets (human figures, Fig. 4(b)) are shown.

[0070] In the rooftop scene, limited by the spatial sampling rate and the point spread function of the imaging system, the low-resolution input exhibits obvious edge blurring in the S0 image, while jagged edges and blocky discontinuities exist in the DoLP and AoP images. After reconstruction using the method described in this invention, the edge contours in the S0 image are clearer, and the continuity of geometric lines is significantly improved. In the portrait scene in Figure 4(b), the human contour of the low-resolution input, especially the thermal radiation boundary at the knee bend, exhibits obvious blurring and spatial aliasing. The reconstruction result, guided by multimodal priors, restores a smoother and more natural boundary transition, without obvious overshoot or pseudo-contour phenomena. At the same time, the model restores clearer polarization feature differences at the junction of the leg surface and the step, and reduces phase breaks in the original low-resolution image.

[0071] Overall, the method described in this invention exhibits smoother boundary transitions and more stable polarization distributions in both the roof structure and the human body edge region, indicating that the proposed physical closed-loop and gating mechanism can effectively improve the structural restoration quality and maintain polarization consistency.

[0072] (2) Visual comparison with mainstream super-resolution algorithms To visually demonstrate the super-resolution reconstruction effects of various methods in real-world scenarios, Figures 5 (vehicle scene) and 6 (building facade scene) show a 4× super-resolution visual comparison between the method described in this invention and six mainstream baseline algorithms under real physical degradation. To clearly reveal texture restoration and polarization polarity, key structural areas (such as wheel hubs, window edges, and building textures) are magnified for display.

[0073] As can be seen from the comparison images, traditional interpolation methods (Bicubic and Newton Polynomial) mainly rely on local pixel weighting for reconstruction, failing to introduce new high-frequency structural information. In the intensity map (S0), this manifests as overall blurring with relatively smooth edge transitions. In DoLP and AoP images, noticeable mosaic effects and jagged edges exist in local areas, with limited structural continuity. For single-frame infrared polarization super-resolution (SwinIPISR), overall contrast and structural sharpness are improved compared to interpolation methods. However, due to its reliance on only a single low-resolution input frame, its recovery of high-frequency details under realistic degradation conditions remains limited. A certain degree of structural smoothing can be observed in the wheel hub area in Figure 5 and the building texture area in Figure 6, indicating limited recovery of detail levels.

[0074] In contrast, the method described in this invention exhibits relatively stable reconstruction results in both scenarios. In the wheel hub area of ​​Figure 5, the edge structure of the hub in the intensity map is more continuous; no obvious over-sharpening or texture drift is observed. In the DoLP image, the polarization distribution shows a reasonable variation trend at the structural boundaries, maintaining a relatively smooth transition within homogeneous areas. In the AoP image, the pseudo-color distribution in the vehicle body and building wall areas is more continuous, without large areas of abnormal color blocks or drastic polarity jumps. Figure 6 In complex architectural facade scenes, the differences between the methods become more pronounced. Traditional interpolation methods struggle to recover high-frequency detail structures. SwinIPISR exhibits a degree of structural smoothing in DoLP images. Some reference-guided methods produce noticeable high-frequency color oscillations in AoP images. GuidedSR reveals structural discontinuities in detail areas. This method demonstrates greater continuity at window frame edges, particularly in DoLP and AoP images, where the texture structure in magnified areas is clearer; the polarization degree and polarization angle are also more stably distributed within homogeneous regions.

[0075] As discussed above, traditional interpolation methods, limited by mathematical interpolation mechanisms, struggle to recover lost high-frequency structural information. Single-frame super-resolution methods can improve overall sharpness, but their ability to recover details under realistic physical degradation conditions is limited. Reference graph-guided methods, while enhancing the structure, may introduce high-frequency information inconsistent with the target modal statistical characteristics. Our proposed method achieves a relatively balanced reconstruction effect between structural enhancement and polarization physical consistency through cross-modal gating mechanisms and physical closed-loop constraints.

Claims

1. A method for unsupervised infrared polarimetric super-resolution reconstruction based on gating and physical constraints, characterized in that, The method includes the following steps: S1, acquiring a low-resolution long-wave infrared polarization image within a range in a same field of view , a long-wave infrared intensity reference image , a high-resolution visible light image , and a short-wave infrared image ; S2. For the images obtained in step S1, perform cross-modal geometric registration to construct a multimodal guided image group in the same coordinate system. ; S3, based on and Constructing a gated mask and will Acting on and The final guide image is obtained. ; S4, based on A dual-stream asymmetric polarization reconstruction structure is constructed. By using high-level semantic features of the intensity component, the polarization component is reconstructed to obtain the reconstructed intensity component. and the reconstructed polarization components and ; S5, based on , and A composite objective total loss function is constructed using physical period consistency loss, cross-modal structure loss, polarization correlation loss, and auxiliary regularization term; S6. Iteratively train the dual-stream asymmetric polarization reconstruction structure to obtain the final reconstructed image. ; according to , and Calculate the degree of linear polarization DoLP and the polarization angle AoP.

2. The unsupervised infrared polarization super-resolution reconstruction method based on gating and physical constraints according to claim 1, characterized in that, Gated Mask The construction method is as follows: ,in, It is the Sigmoid activation function. for Convolutional layer For channel splicing operations; Final boot image The formula for calculation is: ,in, This is element-wise multiplication.

3. The unsupervised infrared polarization super-resolution reconstruction method based on gating and physical constraints according to claim 2, characterized in that, The dual-flow asymmetric polarization reconstruction structure includes: an intensity flow branch and a polarization flow branch; by The original resolution is used as the input for the two-stream asymmetric polarization reconstruction structure. , For the original intensity component, and All are original polarization components; The intensity flow branch includes an intensity flow branch backbone network module, which outputs high-level semantic features of the intensity components. , and After residual stitching, the reconstructed intensity components are output. ; The polarization flow branch includes a polarization flow branch backbone network module and a polarization refinement module. The polarization flow branch backbone network module outputs high-level semantic features of the polarization components. and , and and and The polarization thinning module is input together to obtain polarization thinning features. , respectively with and After residual stitching, the reconstructed polarization components are output. and .

4. The unsupervised infrared polarization super-resolution reconstruction method based on gating and physical constraints according to claim 3, characterized in that, The calculation formula for the intensity flow branch backbone network module is: ,in, include: , and , For spatial feature transformation function, As the input to the backbone network module, it is generated in the intensity flow branch by... as well as Calculations show that in the polarized flow branch, by , , as well as Calculations show that Scaling factor The translation factor; The structure of the intensity flow branch backbone network module is the same as that of the polarization flow branch backbone network module.

5. The unsupervised infrared polarization super-resolution reconstruction method based on gating and physical constraints according to claim 4, characterized in that, The formula for calculating physical cycle consistency loss is: in, for L 1-norm, For degenerate operators, For the output of the intensity flow branch, For the input of the intensity flow branch, For the output of the polarization flow branch, This is the input for the polarization flow branch.

6. The unsupervised infrared polarization super-resolution reconstruction method based on gating and physical constraints according to claim 5, characterized in that, Cross-modal structural loss The formula for calculation is: in, For Sobel edge operators, for High-frequency structural information, and They represent and The square of, for High-frequency structural information.

7. The unsupervised infrared polarization super-resolution reconstruction method based on gating and physical constraints according to claim 6, characterized in that, Polarization correlation loss The formula for calculation is: in, The index value of the pixel. for The pixels are spatial weights constructed using DoLP. It is a constant. In order to be in At the pixel, The result of the degradation operation. for L 2-norm.

8. The unsupervised infrared polarization super-resolution reconstruction method based on gating and physical constraints according to claim 7, characterized in that, Auxiliary regularization terms The formula for calculation is: in, For total variation loss, To perceive style loss, and This represents the loss coefficient for the corresponding parameter.

9. The unsupervised infrared polarization super-resolution reconstruction method based on gating and physical constraints according to claim 8, characterized in that, Composite objective total loss function The formula for calculation is: in, , , and This represents the loss coefficient for the corresponding parameter.

10. The unsupervised infrared polarization super-resolution reconstruction method based on gating and physical constraints according to claim 9, characterized in that, The formula for calculating the linear polarization degree DoLP is: ,in, For the final intensity component, and For the final polarization component; The formula for calculating the polarization angle AoP is: .

Citation Information

Patent Citations

  • Super-resolution reconstruction method and system for transmitting visible light characteristics to infrared light

    CN119107233A

  • Infrared polarization image super-resolution method based on cross attention double-branch network

    CN120876234A