An Optically Guided Thermal Infrared Image Super-Resolution Method and System
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-08-14
AI Technical Summary
[0005]本申请提供一种光学引导的热红外图像超分辨率方法及系统,旨在解决光学引导热红外图像超分辨率重建中,光学纹理易干扰热红外重建结果、二维空间结构保持能力不足以及热红外成像物理特性约束不充分的技术问题
[0019]本申请实施例提供一种光学引导的热红外图像超分辨率方法及系统,该方法通过获取低分辨率热红外图像及其空间对齐的高分辨率光学图像,先从热红外图像中提取热红外浅层特征,并从光学图像中获得光学结构特征、光学边缘图和扩散系数图,使光学图像提供的结构信息不再作为无约束纹理直接注入热红外重建过程,而是先转化为表征热特征扩散能力的约束信息;进一步通过扩散系数图调制物理约束状态空间编码器中的状态空间传播参数,使热特征传播能够结合光学边界和热扩散约束形成物理约束热先验特征,从而降低热结构过度平滑和跨边界传播失真的风险;同时,基于光学边缘图选取光学结构枢轴并对物理约束热先验特征进行稀疏流形重构,可以利用较可靠的光学结构信息修复热特征中的空间结构缺失,再根据重构热结构特征与光学结构特征之间的结构一致性进行门控细化,使光学引导信息的引入受结构一致性约束,减少不匹配光学纹理对热红外结果的干扰。因此,本申请能够在提高热红外图像空间分辨率的同时,抑制伪边缘、纹理溢出、局部伪热点和热结构失真,解决光学引导热红外图像超分辨率重建中光学纹理干扰、二维空间结构保持不足以及热红外成像物理特性约束不充分的问题。
Smart Images

Figure CN122573708A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to an optically guided thermal infrared image super-resolution method and system. Background Technology
[0002] With the continuous development of applications such as drone inspection, nighttime security, disaster relief, border monitoring, environmental perception, and target search under complex weather conditions, thermal infrared imaging, due to its ability to perceive the thermal radiation information of targets, has significant application value in low-light, nighttime, foggy, and visible-light-limited environments. Thermal infrared images can provide effective evidence for target detection, scene understanding, and anomaly identification. However, due to factors such as the size of the thermal infrared detector array, the aperture of the optical system, long-distance imaging conditions, platform motion, and imaging noise, the acquired thermal infrared images often suffer from low spatial resolution, blurred target boundaries, insufficient local texture, and loss of high-frequency details, affecting the accuracy of subsequent target recognition, monitoring analysis, and intelligent perception tasks.
[0003] In related technologies, to improve the spatial resolution of thermal infrared images, single-mode thermal infrared image super-resolution methods are commonly used, or high-resolution optical images spatially aligned with the thermal infrared image are introduced for auxiliary reconstruction. High-resolution optical images typically contain rich information on edges, textures, geometric contours, and spatial structures, providing a certain reference for thermal infrared image reconstruction. However, optical images mainly reflect the visible light reflection characteristics of the target surface, while thermal infrared images mainly reflect the thermal radiation distribution of the target; the two differ significantly in their imaging mechanisms and physical meanings. Directly fusing optical textures with thermal infrared features can easily introduce visible light textures unrelated to thermal radiation into the thermal infrared reconstruction results, leading to false edges, texture overflow, local false hotspots, and thermal structure distortion. Furthermore, while state-space models used in image restoration and super-resolution tasks in recent years have good long-range dependency modeling capabilities and computational efficiency, they may weaken the original two-dimensional spatial adjacency relationships of the image when converting two-dimensional image features into sequences for processing. This is especially true in complex boundaries, slender targets, and local high-frequency regions, easily leading to decreased structural continuity or unstable local detail recovery. Furthermore, existing reconstruction methods do not adequately consider the physical characteristics of thermal infrared imaging, making it difficult to improve image clarity while simultaneously ensuring the rationality of thermal structure distribution.
[0004] Therefore, in optically guided super-resolution reconstruction of thermal infrared images, there are significant differences in the imaging mechanisms between optical and thermal infrared images. Optical textures can easily interfere with the thermal infrared reconstruction results. The state-space model still has insufficient ability to preserve the spatial structure when dealing with complex boundaries and slender structures in two-dimensional images. The existing reconstruction process does not adequately constrain the physical properties of thermal infrared imaging, which has become an urgent problem to be solved. Summary of the Invention
[0005] This application provides an optically guided thermal infrared image super-resolution method and system, aiming to solve the technical problems in optically guided thermal infrared image super-resolution reconstruction, such as the easy interference of optical textures with thermal infrared reconstruction results, insufficient ability to maintain two-dimensional spatial structure, and insufficient constraints of the physical characteristics of thermal infrared imaging.
[0006] In a first aspect, this application provides an optically guided thermal infrared image super-resolution method, the method comprising: Acquire a low-resolution thermal infrared image to be processed and a high-resolution optical image spatially aligned with the low-resolution thermal infrared image; Thermal features are extracted from the low-resolution thermal infrared image to obtain shallow thermal infrared features; Optical structural features and optical edge maps are generated based on the high-resolution optical image, and a diffusion coefficient map is determined based on the optical edge map; wherein, the diffusion coefficient map is used to characterize the thermal feature diffusion capability at different spatial locations; The thermal infrared shallow layer features and the diffusion coefficient map are input into the physical constraint state space encoder, and the state space propagation parameters in the physical constraint state space encoder are modulated based on the diffusion coefficient map to obtain the physical constraint thermal prior features. Based on the optical edge map, an optical structure pivot is selected from the optical structure features, and the physical constraint thermal prior features are reconstructed using the optical structure pivot to obtain the reconstructed thermal structure features. Based on the structural consistency between the reconstructed thermal structural features and the optical structural features, the reconstructed thermal structural features are optically guided and gated to obtain refined thermal features. Super-resolution reconstruction is performed based on the refined thermal features to obtain a high-resolution thermal infrared image.
[0007] In one possible design, generating optical structural features and an optical edge map based on the high-resolution optical image, and determining a diffusion coefficient map based on the optical edge map, includes: The high-resolution optical image is scale-aligned to obtain a scale-aligned optical image corresponding to the thermal infrared shallow layer features. Structural features are extracted from the scale-aligned optical image to obtain the optical structural features; Gradient feature extraction and edge fusion processing are performed on the scale-aligned optical image to obtain the optical edge map; The diffusion coefficient map is determined based on the edge response intensity at different spatial locations in the optical edge map.
[0008] In one possible design, the modulation of the state-space propagation parameters in the physical constraint state-space encoder based on the diffusion coefficient map to obtain physical constraint thermal prior features includes: The thermal infrared shallow layer features and the diffusion coefficient map are respectively serialized to obtain thermal feature sequences and diffusion coefficient sequences. Based on the thermal feature sequence, determine the data-driven modulation information; Based on the diffusion coefficient sequence, determine the physical diffusion modulation information; The state space propagation parameters are determined based on the data-driven modulation information and the physical diffusion modulation information; wherein the state space propagation parameters include discrete time step and / or input injection parameters; Based on the state-space propagation parameters, the thermal feature sequence is propagated in the state space to obtain the physical constraint thermal prior features.
[0009] In one possible design, the physical constraint state space encoder includes multiple cascaded residual physical diffusion coding groups; Each of the residual physical diffusion coding groups is used to propagate the state space features of the input thermal features based on the diffusion coefficient map, and to perform local convolution correction on the thermal features after propagation of the state space features, so as to obtain the output thermal features of the corresponding residual physical diffusion coding group. Among them, the output thermal features of the last residual physical diffusion coding group are used to determine the physical constraint thermal prior features.
[0010] In one possible design, the step of selecting an optical structure pivot from the optical structure features based on the optical edge map, and using the optical structure pivot to perform sparse manifold reconstruction of the physically constrained thermal prior features to obtain reconstructed thermal structure features includes: Based on the optical edge map, edge saliency information is determined; Based on the edge saliency information, the position of the optical structure feature that satisfies the preset saliency condition is selected from the optical structure features as the optical structure pivot; Based on the edge saliency information, the location of the thermal feature to be reconstructed in the physical constraint thermal prior features is determined; The thermal features corresponding to the positions of the thermal features to be reconstructed are masked, and sparse cross-attention processing is performed based on the masked thermal features and the optical structure pivot to obtain the reconstructed thermal structure features.
[0011] In one possible design, the optically guided gating refinement of the reconstructed thermal structure features based on the structural consistency between the reconstructed thermal structure features and the optical structure features to obtain refined thermal features includes: The reconstructed thermal structure features are mapped to a shared comparison space to obtain thermal structure comparison features; The optical structure features are mapped to the shared comparison space to obtain optical structure comparison features; Based on the thermal structure comparison features and the optical structure comparison features, the structural consistency between the reconstructed thermal structure features and the optical structure features is determined; A trust gate is generated based on the aforementioned structural consistency; The reconstructed thermal structure features are optically guided and refined based on the trust gate to obtain the refined thermal features.
[0012] In one possible design, the optically guided gating refinement of the reconstructed thermal structure features based on the trust gate to obtain the refined thermal features includes: Based on the reconstructed thermal structure features, thermal infrared manifold preservation features are generated; Based on the aforementioned optical structural features, optical compensation features are generated; The injection intensity of the optical compensation feature relative to the thermal infrared manifold preservation feature is determined based on the trust gate; The refined thermal feature is obtained by fusing the thermal infrared manifold preservation feature and the optical compensation feature according to the injection intensity.
[0013] In one possible design, the super-resolution reconstruction based on the refined thermal features to obtain a high-resolution thermal infrared image includes: The refined thermal features are then subjected to feature aggregation to obtain aggregated thermal features; The polymerization thermal features are subjected to nonlinear feature mapping to obtain the reconstructed thermal features; The reconstructed thermal features are upsampled to obtain the high-resolution thermal infrared image.
[0014] In one possible design, before performing super-resolution reconstruction based on the refined thermal features to obtain a high-resolution thermal infrared image, the method further includes: Obtain a training sample set; wherein the training sample set includes low-resolution thermal infrared sample images, high-resolution optical sample images spatially aligned with the low-resolution thermal infrared sample images, and high-resolution thermal infrared ground truth images; The low-resolution thermal infrared sample image and the high-resolution optical sample image are input into the thermal infrared image super-resolution network to be trained to obtain the training output image. Based on the training output image and the high-resolution thermal infrared ground truth image, a composite loss is determined; wherein, the composite loss includes pixel reconstruction loss, perceptual structure loss and edge supervision loss; The thermal infrared image super-resolution network to be trained is trained based on the composite loss to obtain the trained thermal infrared image super-resolution network.
[0015] Secondly, this application provides an optically guided thermal infrared imaging super-resolution system, the system comprising: An image acquisition module is used to acquire a low-resolution thermal infrared image to be processed and a high-resolution optical image spatially aligned with the low-resolution thermal infrared image. A thermal feature extraction module is used to extract thermal features from the low-resolution thermal infrared image to obtain shallow thermal infrared features. An optical prior generation module is used to generate optical structural features and an optical edge map based on the high-resolution optical image, and to determine a diffusion coefficient map based on the optical edge map; wherein the diffusion coefficient map is used to characterize the thermal feature diffusion capability at different spatial locations; The physical constraint coding module is used to input the thermal infrared shallow layer features and the diffusion coefficient map into the physical constraint state space encoder, and modulate the state space propagation parameters in the physical constraint state space encoder based on the diffusion coefficient map to obtain the physical constraint thermal prior features. The manifold reconstruction module is used to select optical structure pivots from the optical structure features based on the optical edge map, and to perform sparse manifold reconstruction on the physical constraint thermal prior features using the optical structure pivots to obtain reconstructed thermal structure features. The gated refinement module is used to perform optically guided gated refinement of the reconstructed thermal structure features based on the structural consistency between the reconstructed thermal structure features and the optical structure features, so as to obtain refined thermal features; The reconstruction output module is used to perform super-resolution reconstruction based on the refined thermal features to obtain a high-resolution thermal infrared image.
[0016] Thirdly, this application provides an electronic device, including: a memory and at least one processor; The memory stores computer-executed instructions; The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the method described in the first aspect or various possible designs of the first aspect.
[0017] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed, implement the method described in the first aspect or various possible designs of the first aspect.
[0018] Fifthly, this application provides a computer program product, which includes computer program code that, when run on a computer, causes the computer to implement the method described in the first aspect or various possible designs of the first aspect.
[0019] This application provides an optically guided thermal infrared image super-resolution method and system. The method acquires a low-resolution thermal infrared image and its spatially aligned high-resolution optical image. First, it extracts shallow thermal infrared features from the thermal infrared image and obtains optical structural features, optical edge maps, and diffusion coefficient maps from the optical image. This transforms the structural information provided by the optical image from unconstrained textures directly into constrained information characterizing the thermal feature diffusion capability. Furthermore, it modulates the state-space propagation parameters in the physical constraint state-space encoder using the diffusion coefficient map, enabling thermal feature propagation to combine optical boundaries and thermal diffusion constraints to form physically constrained thermal prior features. This reduces the risk of excessive smoothing of the thermal structure and cross-boundary propagation distortion. Simultaneously, by selecting an optical structural pivot based on the optical edge map and reconstructing the physically constrained thermal prior features using sparse manifolds, it can repair spatial structural deficiencies in thermal features using reliable optical structural information. Then, it performs gating refinement based on the structural consistency between the reconstructed thermal structural features and the optical structural features, ensuring that the introduction of optical guidance information is constrained by structural consistency and reducing interference from mismatched optical textures on the thermal infrared results. Therefore, this application can improve the spatial resolution of thermal infrared images while suppressing false edges, texture overflow, local false hot spots and thermal structure distortion, thus solving the problems of optical texture interference, insufficient preservation of two-dimensional spatial structure and insufficient constraints of physical properties in optically guided thermal infrared image super-resolution reconstruction. Attached Figure Description
[0020] Figure 1 A flowchart illustrating an optically guided thermal infrared image super-resolution method provided in this application embodiment; Figure 2 A schematic diagram of the overall structure of an optically guided thermal infrared image super-resolution network provided in this application embodiment; Figure 3 A schematic diagram of the structure of a physical constraint state space encoder provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of a sparse manifold reconstruction module provided in an embodiment of this application; Figure 5 A schematic diagram of the layered workflow of an optically guided thermal infrared image super-resolution method provided in this application embodiment; Figure 6 A schematic diagram of an optically guided thermal infrared imaging super-resolution system provided in this application embodiment; Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims and drawings of this application are intended to cover non-exclusive inclusion.
[0023] The term "embodiment" as used herein means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of the phrase "embodiment" in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0024] In this article, the term "and / or" simply describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists, A and B can exist simultaneously, and B exists. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0025] Furthermore, the terms "first," "second," etc., in the specification and claims of this application or in the aforementioned drawings are used to distinguish different objects rather than to describe a specific order, and may explicitly or implicitly include one or more of the features.
[0026] In the description of this application, unless otherwise stated, "multiple" and "at least two" mean two or more (including two), and similarly, "multiple groups" and "at least two groups" mean two or more (including two groups).
[0027] In the description of this application, it should be noted that, unless otherwise explicitly specified and limited, the terms "connected" and "linked" should be interpreted broadly. For example, "connected" or "linked" can refer not only to a physical connection, but also to an electrical connection or a signal connection. For instance, it can be a direct connection, i.e., a physical connection, or an indirect connection through at least one intermediate component, as long as the circuit is connected. It can also refer to the internal connection between two components. A signal connection can refer not only to a signal connection through a circuit, but also to a signal connection through a medium, such as radio waves. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0028] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. It should be noted that, unless otherwise specified, different technical features in this application can be combined with each other.
[0029] With the development of applications such as drone inspection, nighttime security, disaster relief, border monitoring, environmental perception, and target search under complex weather conditions, thermal infrared imaging technology has high application value in scenarios where visible light imaging is limited, such as low light, nighttime, and fog / haze, due to its ability to perceive the thermal radiation information of targets. Thermal infrared images can reflect the temperature distribution and thermal radiation differences in targets or scenes, and can be used for tasks such as target identification, anomaly detection, scene perception, and security monitoring. However, thermal infrared sensors are usually limited by factors such as detector array size, optical system aperture, long-distance imaging conditions, platform motion, and imaging noise. The thermal infrared images actually acquired often suffer from problems such as low spatial resolution, blurred target boundaries, insufficient local texture, and loss of high-frequency details, which in turn affect the accuracy of subsequent visual analysis tasks.
[0030] In related technologies, to improve the spatial resolution of thermal infrared images, single-mode thermal infrared image super-resolution reconstruction is commonly employed, or a high-resolution optical image spatially aligned with the thermal infrared image is introduced for auxiliary reconstruction. High-resolution optical images typically contain rich information on edges, textures, geometric contours, and spatial structures, providing a reference for detail recovery from thermal infrared images. However, optical images primarily reflect the visible light reflectance characteristics of the target surface, while thermal infrared images primarily reflect the thermal radiation distribution of the target; the two differ in their imaging mechanisms and physical meanings. Directly fusing optical textures with thermal infrared features can easily introduce visible light textures that do not correspond to the thermal radiation distribution into the thermal infrared reconstruction results, leading to problems such as false edges, texture overflow, localized false hotspots, or thermal structure distortion.
[0031] In recent years, State Space Models (SSMs) have been increasingly applied to image restoration and super-resolution tasks due to their superior ability to model long-range dependencies. Among them, the Mamba model, as a model structure based on a selective state space mechanism, can model long-range feature dependencies with relatively low computational complexity, making it suitable for processing large-size image features. However, in vision tasks, the Mamba model typically requires converting two-dimensional image features into a one-dimensional sequence according to a pre-defined scanning method. This processing method, while modeling long-range dependencies, may weaken the original two-dimensional spatial adjacency relationships of the image. Especially with slender targets, complex boundaries, and local high-frequency regions, the one-dimensional scanning sequence may lead to a decrease in local structural continuity, perturbation of topological relationships, or unstable feature propagation.
[0032] Furthermore, some related techniques constrain the thermal infrared image reconstruction results through physical priors, such as imposing gradient, edge, or structural consistency constraints on the reconstructed image at the loss function level or the final output level. However, these constraints typically apply to the training objective or output result level, and have a weaker constraint on the feature propagation process within the network, making it difficult to fully reflect the heat conduction or heat diffusion laws during the feature evolution stage. Therefore, when introducing optically assisted information for thermal infrared image super-resolution reconstruction, problems such as overly smoothed thermal boundaries, unreasonable cross-boundary feature propagation, and inconsistent thermal structure distribution may still occur.
[0033] Therefore, in the process of optically guided super-resolution reconstruction of thermal infrared images, how to utilize the structural information of optical images while avoiding interference of optical textures on thermal infrared modes, how to maintain the two-dimensional spatial structural relationship when the state space model processes image features, and how to make the reconstruction process more fully conform to the physical characteristics of thermal infrared imaging are issues that need further improvement in related technologies.
[0034] To address the aforementioned issues, this application provides an optically guided super-resolution method and system for thermal infrared images. The method first generates optical structural features and optical edge maps using a high-resolution optical image, and further determines a diffusion coefficient map to characterize the thermal feature diffusion capability, transforming the structural boundaries in the optical image into prior information constraining thermal feature propagation. Subsequently, the shallow thermal infrared features and diffusion coefficient map corresponding to the low-resolution thermal infrared image are input into a physically constrained state-space encoder. The diffusion coefficient map modulates the state-space propagation parameters, subjecting the thermal feature propagation process to the combined constraints of optical boundaries and thermal diffusion characteristics. Further, based on the optical edge map, an optical structural pivot is selected from the optical structural features to reconstruct the physically constrained thermal prior features using sparse manifolds, improving the two-dimensional spatial structure preservation effect. Finally, based on the structural consistency between the reconstructed thermal structural features and the optical structural features, optically guided gating refinement is applied to the reconstructed thermal structural features, thereby enhancing the details and edge clarity of the thermal infrared image while reducing optical texture errors, false edges, local false hotspots, and thermal structural distortion.
[0035] Figure 1 This is a schematic flowchart illustrating an optically guided thermal infrared image super-resolution method provided in an embodiment of this application. Figure 1 As shown, the method includes S101 to S107, and S101 to S107 are described in detail below.
[0036] It should be noted that this method can be executed by an optically guided thermal infrared image super-resolution system, or by electronic devices, servers, edge computing devices, UAV-borne processing devices, vehicle-mounted multimodal imaging devices, or fixed monitoring and processing devices with image processing capabilities.
[0037] S101. Acquire the low-resolution thermal infrared image to be processed and the high-resolution optical image spatially aligned with the low-resolution thermal infrared image.
[0038] Among them, the low-resolution thermal infrared image is the thermal imaging image to be reconstructed in super-resolution, and the high-resolution optical image is the visible light image or other optical image used to provide structural guidance information.
[0039] Low-resolution thermal infrared images and high-resolution optical images can be acquired by dual-light cameras on drones, vehicle-mounted multimodal imaging equipment, fixed monitoring equipment, or other optical-thermal infrared synchronous acquisition equipment.
[0040] It should be noted that spatial alignment refers to the correspondence between low-resolution thermal infrared images and high-resolution optical images in terms of pixel positions, local regions, or target regions. Based on this spatial correspondence, the edge structures, geometric contours, and local spatial topology in the high-resolution optical image can serve as structural references in the super-resolution reconstruction process of the low-resolution thermal infrared image, and provide a data foundation for the subsequent generation of optical structural features, optical edge maps, and diffusion coefficient maps.
[0041] S102. Extract thermal features from the low-resolution thermal infrared image to obtain shallow thermal infrared features.
[0042] Specifically, a low-resolution thermal infrared image can be input into a shallow thermal feature extractor. The shallow thermal feature extractor extracts the initial thermal radiation distribution features, local grayscale change features, and basic thermal structure features from the low-resolution thermal infrared image to obtain shallow thermal infrared features.
[0043] A shallow thermal feature extractor can be composed of convolutional layers and nonlinear activation functions, or it can be composed of other feature extraction structures that can extract shallow features from thermal infrared images.
[0044] The shallow thermal infrared features serve as input features for subsequent physical constraint state space coding, preserving the thermal modal information of the low-resolution thermal infrared image itself. This allows the subsequent reconstruction process to be based on the thermal radiation expression of the thermal infrared image, rather than directly relying on the visible light texture in the high-resolution optical image to generate the thermal structure.
[0045] In one alternative implementation, a low-resolution thermal infrared image can be input into a shallow thermal feature extractor to obtain shallow thermal infrared features. This process can be represented as follows: ;in, Indicates shallow thermal infrared characteristics. This indicates a shallow thermal feature extractor. This indicates a low-resolution thermal infrared image.
[0046] S103. Generate optical structural features and optical edge maps based on high-resolution optical images, and determine the diffusion coefficient map based on the optical edge maps.
[0047] The diffusion coefficient diagram is used to characterize the thermal diffusion capacity at different spatial locations.
[0048] Specifically, optical structural features and optical edge maps can be extracted from high-resolution optical images. Optical structural features are used to characterize the edge structures, geometric contours, and local spatial topology in high-resolution optical images, providing a reference for structure reconstruction from subsequent thermal infrared images. Optical edge maps are used to characterize the strength of optical structural boundaries at different spatial locations.
[0049] When determining the diffusion coefficient map based on the optical edge map, locations with strong optical edge responses can correspond to lower thermal feature diffusion capabilities, while locations with weaker or relatively smooth optical edge responses can correspond to higher thermal feature diffusion capabilities. Thus, the diffusion coefficient map can transform structural boundaries in high-resolution optical images into diffusion constraints during thermal feature propagation. This suppresses subsequent thermal feature propagation at structural boundaries and maintains continuous propagation within relatively smooth regions, thereby avoiding the direct and unconstrained injection of optical textures into thermal infrared features.
[0050] S104. Input the thermal infrared shallow layer features and diffusion coefficient map into the physical constraint state space encoder, and modulate the state space propagation parameters in the physical constraint state space encoder based on the diffusion coefficient map to obtain the physical constraint thermal prior features.
[0051] Specifically, a physically constrained state-space encoder is used to propagate state-space features from shallow thermal infrared features to obtain thermal prior features containing long-range thermal feature dependencies. The physically constrained state-space encoder can employ a Mamba-based state-space coding structure, enabling feature updates of shallow thermal infrared features during state-space propagation.
[0052] During state-space feature propagation, the diffusion coefficient map is used to modulate the state-space propagation parameters. In other words, the diffusion coefficient map is not simply a concatenation of ordinary image features and shallow thermal infrared features, but rather it is used to influence the propagation capability of shallow thermal infrared features at different spatial locations.
[0053] When the diffusion coefficient at a certain spatial location in the diffusion coefficient diagram is low, it indicates that this spatial location is more likely to correspond to the optical structure boundary, thus suppressing the propagation of thermal features across the boundary. Conversely, when the diffusion coefficient at a certain spatial location in the diffusion coefficient diagram is high, it indicates that this spatial location is more likely to be located in a relatively smooth region, thus allowing the continuous propagation of thermal features within that region. Therefore, the physically constrained state-space encoder can introduce diffusion constraints corresponding to optical boundaries during the propagation of shallow thermal infrared features, obtaining physically constrained thermal prior features.
[0054] S105. Based on the optical edge map, select the optical structure pivot from the optical structure features, and use the optical structure pivot to reconstruct the sparse manifold of the physical constraint thermal prior features to obtain the reconstructed thermal structure features.
[0055] Specifically, an optical structure pivot is a structural reference position within an optical structure feature that can characterize a significant edge, geometric contour, or local spatial topology. Based on the edge salience at different spatial locations in the optical edge map, some feature positions with high structural reliability can be selected from the optical structure features as optical structure pivots.
[0056] Because physical constraint state-space encoders can convert two-dimensional image features into sequential features for propagation when processing image features, while the sequential propagation process is beneficial for modeling long-range dependencies, it may also weaken the original spatial adjacency relationships of the two-dimensional image. Especially in complex boundaries, slender targets, or local high-frequency regions, thermal features may exhibit a decrease in local structural continuity.
[0057] Based on this, after obtaining the physically constrained thermal prior features, sparse manifold reconstruction can be performed on these features using optical structural pivots. Sparse manifold reconstruction uses optical structural pivots as structural anchors to compensate for and repair the spatial structure in the physically constrained thermal prior features, thereby obtaining reconstructed thermal structural features. Since this process utilizes only a portion of the optical structural pivots, rather than directly injecting all the texture from the high-resolution optical image into the thermal infrared features, it can improve the spatial structure representation of thermal features while reducing the interference of optical textures on the thermal infrared modes.
[0058] S106. Based on the structural consistency between the reconstructed thermal structural features and the optical structural features, optically guided gating refinement is performed on the reconstructed thermal structural features to obtain refined thermal features.
[0059] Specifically, structural consistency is used to characterize the degree of structural matching between reconstructed thermal structural features and optical structural features at corresponding spatial locations. Since high-resolution optical images reflect visible light reflection characteristics, while low-resolution thermal infrared images reflect thermal radiation distribution, and the two differ in imaging mechanisms and physical meanings, not all optical structures are suitable as compensation information for thermal infrared image reconstruction.
[0060] When performing gating refinement of optical guidance, the structural consistency between the reconstructed thermal structural features and the optical structural features can be determined first. Then, the intensity of the effect of optical guidance information on the reconstructed thermal structural features can be controlled based on the structural consistency. When the structural consistency is high, it indicates that the optical structure and thermal structure at the corresponding position have a high degree of matching, which can enhance the refinement effect of the optical structure on the reconstructed thermal structural features. When the structural consistency is low, it indicates that the optical texture at the corresponding position may not match the thermal radiation structure, which can weaken the influence of the optical structure on the reconstructed thermal structural features and allow the thermal infrared mode's own structure to be preferentially preserved.
[0061] By using the above-mentioned gating refinement process, while utilizing high-resolution optical images to provide structural references, it is possible to avoid excessive introduction of optical textures that do not correspond to the thermal radiation distribution into the thermal infrared reconstruction process, thereby obtaining refined thermal features.
[0062] S107. Super-resolution reconstruction is performed based on refined thermal features to obtain a high-resolution thermal infrared image.
[0063] Specifically, refined thermal features can be input into the reconstruction output module, which then maps these features into a high-resolution thermal infrared image at the target scale. The resulting high-resolution thermal infrared image, compared to a low-resolution thermal infrared image, possesses higher spatial resolution and better representation of target boundaries, local details, and thermal structure continuity. This high-resolution thermal infrared image can be used for UAV thermal infrared image enhancement, nighttime inspection, security monitoring, target recognition, disaster search and rescue, environmental monitoring, and other optically guided thermal imaging enhancement scenarios.
[0064] This application provides an optically guided thermal infrared image super-resolution method. By acquiring a low-resolution thermal infrared image and its spatially aligned high-resolution optical image, shallow thermal infrared features are first extracted from the thermal infrared image. Optical structural features, optical edge maps, and diffusion coefficient maps are then obtained from the optical image. This transforms the structural information provided by the optical image from unconstrained textures directly into constrained information characterizing the thermal feature diffusion capability. Furthermore, the diffusion coefficient map modulates the state-space propagation parameters in the physical constraint state-space encoder, enabling thermal feature propagation to combine optical boundaries and thermal diffusion constraints to form physically constrained thermal prior features. This reduces the risk of excessive smoothing of the thermal structure and cross-boundary propagation distortion. Simultaneously, by selecting an optical structural pivot based on the optical edge map and reconstructing the physically constrained thermal prior features using sparse manifolds, reliable optical structural information can be used to repair spatial structural deficiencies in the thermal features. Finally, gating refinement is performed based on the structural consistency between the reconstructed thermal structural features and the optical structural features, ensuring that the introduction of optical guidance information is constrained by structural consistency and reducing interference from mismatched optical textures on the thermal infrared results. Therefore, this application can improve the spatial resolution of thermal infrared images while suppressing false edges, texture overflow, local false hot spots, and thermal structure distortion.
[0065] Figure 2 This is a schematic diagram of the overall structure of an optically guided thermal infrared image super-resolution network provided in an embodiment of this application. Figure 2 As shown, the network may include a thermal feature extraction part, an optical prior generation part, a physical constraint state space encoding part, a sparse manifold reconstruction part, a consistency gating refinement part, and a reconstruction output part.
[0066] Specifically, the network takes a low-resolution thermal infrared image and a spatially aligned high-resolution optical image as input. For the low-resolution thermal infrared image, shallow thermal infrared features can be extracted using a shallow thermal feature extractor. For the high-resolution optical image, optical structural features, optical edge maps, and diffusion coefficient maps are obtained through processing such as scale alignment, gradient extraction, edge fusion, and structural feature extraction. The diffusion coefficient map is used to characterize the thermal feature diffusion capability at different spatial locations and serves as a physical prior for the subsequent propagation of physical constraint states in space.
[0067] Furthermore, the shallow thermal infrared features and diffusion coefficient map are input into the physically constrained state space coding section. This section may include multiple residual physical diffusion coding groups, each of which modulates the state space propagation process using the diffusion coefficient map, thereby suppressing the cross-boundary propagation of thermal features in strong boundary regions and maintaining continuous propagation in relatively smooth regions, thus obtaining physically constrained thermal prior features.
[0068] Furthermore, the physically constrained thermal prior features are input into the sparse manifold reconstruction part. This part selects the optical structure pivot based on the optical edge map and performs sparse manifold reconstruction on the physically constrained thermal prior features based on the optical structure pivot to repair the two-dimensional spatial structure damage that may be caused during the one-dimensional scanning of the state space model, thereby obtaining the reconstructed thermal structure features.
[0069] Furthermore, the reconstructed thermal and optical structural features are input into a consistency-gated refinement section. This section generates a trust gate by calculating the structural consistency between the reconstructed thermal and optical structural features, and controls the injection intensity of optical compensation information based on the trust gate, thereby obtaining refined thermal features. Finally, the reconstruction output section performs feature aggregation, nonlinear feature mapping, and upsampling processing on the refined thermal features to obtain a high-resolution thermal infrared image.
[0070] pass Figure 2 The overall network structure shown in this application embodiment can transform the structural information in the optical image into thermal feature propagation constraints, structural reconstruction references, and reliable compensation information, thereby improving the spatial resolution of the thermal infrared image while reducing false edges, texture overflow, and thermal structural distortion caused by unconstrained injection of optical textures.
[0071] In one possible embodiment, the method steps shown in S103 can be implemented by S1031 to S1034, which are described in detail below.
[0072] S1031. Scale-align the high-resolution optical image to obtain a scale-aligned optical image corresponding to the shallow thermal infrared features.
[0073] Since the spatial resolution of high-resolution optical images is usually higher than the feature scale corresponding to low-resolution thermal infrared images, and subsequent physical constraint state space coding and sparse manifold reconstruction mainly focus on processing shallow thermal infrared features, it is necessary to first make the high-resolution optical images and shallow thermal infrared features corresponding to each other at the processing scale.
[0074] Specifically, high-resolution optical images can be downsampled, scaled, or transformed in terms of feature scale to scale them down to a low-resolution physical domain corresponding to shallow thermal infrared features, thus obtaining a scale-aligned optical image. This process can be represented as: ,in, Represents high-resolution optical images. This indicates a downsampling or scale alignment operation. This indicates a scale-aligned optical image.
[0075] The scale-aligned optical image and the thermal infrared shallow layer features have a correspondence in spatial location, local region or feature unit, so that the optical structure information extracted from the scale-aligned optical image can be applied to the spatial location corresponding to the thermal infrared shallow layer features.
[0076] In this embodiment, scale alignment processing can avoid the scale mismatch problem caused by directly using high-resolution optical images. This allows the edge structures, geometric contours, and local spatial topology in the optical images to participate in subsequent processing in a form that is compatible with the shallow thermal infrared features, thereby providing a scale-consistent data foundation for the generation of optical structural features, optical edge maps, and diffusion coefficient maps.
[0077] S1032. Extract structural features from the scale-aligned optical image to obtain optical structural features.
[0078] Specifically, the scale-aligned optical image can be input into an optical structure feature extraction network or feature extraction module to extract features from the edge orientation, geometric contour, local spatial topology, and stable structural regions in the scale-aligned optical image, thereby obtaining optical structure features.
[0079] Optical structural features are used to characterize the structural information of high-resolution optical images at the corresponding processing scale. They are not used to directly replace the thermal radiation information in thermal infrared images, but rather to provide a reference for structural recovery of thermal infrared images in subsequent processing.
[0080] In this embodiment, optical structural features may include feature representations related to target edges, elongated structures, complex boundaries, local high-frequency regions, or spatial contours. Since high-resolution optical images typically have clearer geometric structures and edge information than low-resolution thermal infrared images, extracting optical structural features can provide a feature basis for subsequent optical structure pivot selection, sparse manifold reconstruction, and structural consistency judgment.
[0081] S1033. Perform gradient feature extraction and edge fusion processing on the scale-aligned optical image to obtain an optical edge map.
[0082] Specifically, horizontal and vertical gradient extraction can be performed on the scale-aligned optical image to obtain horizontal and vertical gradient features. This process can be represented as: , ;in, Represents the gradient characteristics in the horizontal direction. Represents the gradient characteristics in the vertical direction. This represents the gradient extraction operator in the horizontal direction. This represents the gradient extraction operator in the vertical direction.
[0083] Horizontal gradient features are used to characterize the brightness or structural changes in a scale-aligned optical image in the horizontal direction, while vertical gradient features are used to characterize the brightness or structural changes in a scale-aligned optical image in the vertical direction.
[0084] After obtaining the horizontal and vertical gradient features, an edge fusion head can be used to perform edge fusion processing on the horizontal and vertical gradient features, and an optical edge map can be generated using an activation function. This process can be represented as: ,in, Represents the optical edge diagram. Indicates edge blending head, This represents the Sigmoid activation function. The edge fusion head can fuse gradient responses from different directions into a unified edge response result; the Sigmoid activation function can map edge responses to a preset value range, enabling the optical edge map to characterize the edge response intensity at different spatial locations.
[0085] Optical edge maps are used to characterize the strength of optical structure boundaries at different spatial locations in scale-aligned optical images.
[0086] Among them, the higher the edge response intensity, the more likely the corresponding spatial location belongs to the boundary of a significant optical structure; the lower the edge response intensity, the more likely the corresponding spatial location belongs to a weak edge region or a relatively smooth region.
[0087] In this embodiment, based on the optical edge map, on the one hand, it can provide the basis for the edge saliency of the subsequent selection of the optical structure pivot, and on the other hand, it can provide the basis for the boundary strength of the diffusion coefficient map, so that the structural boundary in the optical image can be transformed into constraint information in the thermal feature propagation process.
[0088] S1034. Determine the diffusion coefficient diagram based on the edge response intensity at different spatial locations in the optical edge diagram.
[0089] Specifically, a diffusion coefficient map can be generated based on the optical edge map. The diffusion coefficient map is used to characterize the thermal diffusion capability at different spatial locations. This process can be expressed as: Where D represents the diffusion coefficient map and E represents the optical edge map. From the above relationship, it can be seen that the edge response intensity in the diffusion coefficient map and the optical edge map are negatively correlated.
[0090] In other words, the stronger the optical edge response, the smaller the corresponding diffusion coefficient, indicating a weaker thermal feature diffusion capability at that location, which helps suppress the propagation of thermal features across structural boundaries. Conversely, the weaker the optical edge response or the location being in a relatively smooth region, the larger the corresponding diffusion coefficient, indicating a stronger thermal feature diffusion capability at that location, which allows thermal information to propagate within a local area. In this way, the optical boundary can function similarly to an adiabatic boundary during the propagation of thermal features.
[0091] In this embodiment, after determining the diffusion coefficient map, the edge structure in the high-resolution optical image can be converted into a diffusion constraint during the thermal feature propagation process. This diffusion constraint does not directly inject optical texture into thermal infrared features, but rather adjusts the thermal feature diffusion capability at different spatial locations, so that thermal features are suppressed at structural boundaries and continue to propagate in relatively smooth regions, thereby providing boundary-aware physical priors for subsequent physical constraint state space encoding.
[0092] This application's embodiments can convert structural information in high-resolution optical images into optical structural features, optical edge maps, and diffusion coefficient maps. The optical structural features provide a geometric reference, the optical edge maps characterize the strength of optical structural boundaries, and the diffusion coefficient maps characterize the thermal feature diffusion capability at different spatial locations. Therefore, the structural information of the optical image can be utilized in subsequent thermal infrared image super-resolution reconstruction. Simultaneously, the diffusion coefficient maps allow the optical boundaries to function similarly to adiabatic boundaries during thermal feature propagation, preventing the unconstrained fusion of optical textures into thermal infrared features and helping to reduce the risks of false edges, texture overflow, and thermal structural distortion.
[0093] In one possible embodiment, the method steps of "modulating the state space propagation parameters in the physical constraint state space encoder based on the diffusion coefficient map to obtain the physical constraint thermal prior features" shown in S104 can be implemented by S1041 to S1045, and S1041 to S1045 are described in detail below.
[0094] S1041. The shallow thermal infrared features and diffusion coefficient map are serialized to obtain the thermal feature sequence and diffusion coefficient sequence, respectively.
[0095] Specifically, the physical constraint state-space encoder can be a physical information diffusion Mamba encoder, which is used to input thermal features and diffusion coefficient maps into the state-space propagation process to generate physical constraint thermal prior features. Since the state-space propagation process in the Mamba encoder is usually processed in a sequential manner, the two-dimensional thermal infrared shallow features and diffusion coefficient maps can be first unfolded into a one-dimensional sequence according to a preset scanning order.
[0096] In one alternative implementation, the serialization process of shallow thermal infrared features can be represented as: ;in, This represents the thermal feature input at position t. This indicates an operation that unfolds the sequence according to a preset scanning order. This indicates shallow thermal infrared characteristics.
[0097] The serialization process of the diffusion coefficient map can be represented as follows: ;in, Let represent the diffusion coefficient at position t, and D represent the diffusion coefficient diagram.
[0098] Furthermore, the thermal feature sequence can be represented as The diffusion coefficient sequence can be expressed as Among them, thermal feature sequences Includes multiple thermal feature inputs arranged in a preset scanning order. diffusion coefficient sequence Including multiple diffusion coefficients arranged in the same scanning order Due to the diffusion coefficient diagram D and the shallow thermal infrared characteristics... Having corresponding spatial relationships, therefore, after serialization, the thermal feature input... With diffusion coefficient They still correspond to the same spatial location or the same feature unit.
[0099] In this embodiment, the thermal infrared shallow features and diffusion coefficient map can be converted into a one-dimensional sequence form suitable for state space propagation, while maintaining the positional correspondence between the thermal feature input and the diffusion coefficient, thereby providing a basis for subsequent physical modulation of the state space propagation parameters using the diffusion coefficient.
[0100] S1042. Determine the data-driven modulation information based on the thermal feature sequence.
[0101] Specifically, after obtaining the thermal feature sequence, the thermal feature at position t in the thermal feature sequence can be used as input. Determine the data-driven modulation information. The data-driven modulation information is used to reflect the content distribution, local thermal radiation expression, and feature response of the shallow thermal infrared features themselves, so that the state-space propagation process can be modulated according to the data features of the thermal infrared image itself.
[0102] In one alternative implementation, data-driven modulation information can be obtained through data-driven mapping. The data-driven mapping process can be represented as: ,in, This indicates a data-driven mapping. This represents the data-driven term determined by the thermal feature input at position t. This data-driven term is used to participate in the construction of the subsequent physical modulation time step, enabling the state-space propagation parameters to adaptively adjust as the thermal infrared feature content changes.
[0103] In this embodiment, the physical constraint state space encoder can not only rely on fixed propagation parameters when performing state space propagation, but also generate corresponding data-driven modulation information by combining the thermal feature input of the current sequence position, thereby improving the adaptability of thermal feature propagation to changes in image content.
[0104] S1043. Determine the physical diffusion modulation information based on the diffusion coefficient sequence.
[0105] Specifically, after obtaining the diffusion coefficient sequence, the diffusion coefficient at position t in the diffusion coefficient sequence can be used as a reference. Determine the physical diffusion modulation information. The physical diffusion modulation information is used to characterize the thermal feature diffusion capability corresponding to this spatial location, so that the state space propagation process can be introduced with the thermal diffusion physical constraints determined by the optical edge map.
[0106] In one alternative implementation, the physical diffusion modulation information can be obtained from a physical diffusion mapping. The physical diffusion mapping process can be represented as: ,in, Represents the physical diffusion map. This represents the physical diffusion term determined by the diffusion coefficient at position t.
[0107] Due to diffusion coefficient Determined by the optical edge map, therefore, It can reflect the thermal characteristic diffusion capability at position t. When When the value is smaller, it indicates that the t-th position is more likely to be near the boundary of a strong optical structure, and the thermal feature diffusion capability at this position is weaker; when When the value is large, it indicates that the t-th position is more likely to be in a relatively smooth region, and the thermal feature diffusion capability at this position is stronger. Therefore, the physical diffusion modulation information can introduce the diffusion constraint corresponding to the optical boundary into the subsequent determination process of the state-space propagation parameters.
[0108] S1044. Determine the state-space propagation parameters based on the data-driven modulation information and the physical diffusion modulation information.
[0109] The state-space propagation parameters include the discrete time step and / or the input injection parameters.
[0110] Specifically, the physical modulation time step can be constructed by combining data-driven modulation information and physical diffusion modulation information. Unlike traditional state-space models where the time step is mainly determined by data-driven methods, this embodiment introduces the physical laws of heat conduction into the Mamba state-space propagation process, so that the time step is jointly modulated by data-driven terms and physical diffusion terms.
[0111] In one alternative implementation, the physical modulation time step can be expressed as: ;in, This represents the physical modulation time step corresponding to the t-th position. This represents the time step constraint mapping. This represents the physical modulation coefficient.
[0112] From the above relationship, it can be seen that the physical modulation time step Data-driven items and physical diffusion term The terms are jointly determined. The data-driven term reflects the content changes of the shallow thermal infrared features themselves, the physical diffusion term reflects the thermal feature diffusion capability at that location, and the physical modulation coefficient... Time step constraint mapping is used to adjust the influence of the physical diffusion term on the time step. Used to map the modulation result to a time step suitable for state-space discretization propagation.
[0113] Based on the modulated time step, discrete updates of the state-space model can be performed. The discrete state-space propagation parameters can be expressed as: , ;in, This represents the discrete state transition parameter corresponding to the t-th position. This represents the input injection parameter corresponding to the t-th position. Represents the state transition matrix. Indicates the input projection matrix. Represents the identity matrix. This indicates matrix exponentiation operations.
[0114] When the diffusion coefficient When the diffusion coefficient approaches zero, it indicates that the current position is near a strong structural boundary. At this point, the physical modulation time step and input injection capability are suppressed, thereby reducing the propagation of thermal features across the boundary; when the diffusion coefficient... A larger value indicates that the current position is in a relatively smooth region, allowing thermal features to spread sufficiently within the local area. Therefore, the Mamba encoder is no longer a one-dimensional sequence propagator without physical constraints, but rather becomes an anisotropic thermal diffusion feature propagation module with boundary awareness.
[0115] S1045. Based on the state-space propagation parameters, the thermal feature sequence is propagated in the state space to obtain the physical constraint thermal prior features.
[0116] Specifically, after obtaining the discrete state transition parameters corresponding to the t-th position... and input injection parameters Afterwards, it can be based on and Input thermal features from thermal feature sequences Perform state-space propagation. This process can be represented as: ;in, Indicates the current state. This represents the previous state corresponding to the (t-1)th position.
[0117] After completing the state space propagation at each position in the thermal feature sequence, the propagated state sequence can be restored to a two-dimensional thermal feature expression according to the spatial mapping relationship corresponding to the serialization process, thus obtaining the physical constraint thermal prior features.
[0118] In this embodiment, the physically constrained thermal prior features not only include the thermal radiation information and long-range dependencies of the low-resolution thermal infrared image itself, but also the boundary-aware diffusion constraints introduced by the diffusion coefficient map. Therefore, the physically constrained state-space encoder can suppress unreasonable cross-boundary thermal feature propagation at the optical structure boundaries and maintain continuous thermal feature propagation within relatively smooth regions, thereby improving physical consistency, boundary preservation capability, and thermal structure representation stability during the thermal infrared image super-resolution reconstruction process.
[0119] In one possible embodiment, the physical constraint state space encoder includes multiple cascaded residual physical diffusion coding groups.
[0120] Multiple residual physical diffusion coding groups are sequentially connected along the direction of thermal feature propagation. The output thermal feature of the previous residual physical diffusion coding group can be used as the input thermal feature of the next residual physical diffusion coding group. By cascading multiple residual physical diffusion coding groups, the input thermal feature can undergo multi-stage physical constraint state space propagation and local structure correction, thereby gradually enhancing the long-range thermal dependence expression and local spatial structure expression in the thermal infrared feature.
[0121] Each residual physical diffusion coding group is used to propagate the state space features of the input thermal features based on the diffusion coefficient map, and to perform local convolution correction on the thermal features after propagation of the state space features, so as to obtain the output thermal features of the corresponding residual physical diffusion coding group. Among them, the output thermal features of the last residual physical diffusion coding group are used to determine the physical constraint thermal prior features.
[0122] Specifically, each residual physical diffusion coding group may include at least one physical information Mamba module and a local convolution correction unit. The physical information Mamba module is used to propagate the input thermal features in state space based on the diffusion coefficient map, so that the input thermal features are physically constrained by the diffusion coefficient map during the propagation in state space. The local convolution correction unit is used to perform local convolution correction on the thermal features after the propagation of state space features, so as to compensate for the local neighborhood structure information that may be weakened during the propagation in state space.
[0123] In the g-th residual physical information diffusion group, the input thermal characteristics of this residual physical diffusion coding group can be denoted as: The diffusion coefficient map is then unfolded into a diffusion coefficient sequence according to the scanning order corresponding to the input thermal features. The g-th residual physical diffusion coding block can be represented as: ;in, This represents the output thermal characteristic obtained after processing by the g-th residual physical information diffusion group, and can be used as the input thermal characteristic of the next residual physical information diffusion group. Indicates the first Input features of a residual physical information diffusion group This represents the expanded diffusion coefficient sequence. Indicates the first A residual physical information diffusion group.
[0124] In the above formula, the residual connection term Residual diffusion features are used to preserve the thermal feature representation before entering the g-th residual physical diffusion coding group. Used to characterize the diffusion coefficient sequence The modulated state space propagation results and the local convolution correction results are combined. By adding the two together, long-range thermal dependence features with physical diffusion constraints and local spatial correction features can be introduced while preserving the original thermal feature information, thereby reducing the problems of thermal feature degradation or local structure loss during multi-layer coding.
[0125] Furthermore, when each residual physical diffusion coding group performs state-space feature propagation on the input thermal features based on the diffusion coefficient map, it can continue to use the aforementioned physical modulation time step and discrete state-space update process. This suppresses the cross-boundary propagation of thermal features near strong structural boundaries, allowing thermal features in relatively smooth regions to propagate continuously. The thermal features after state-space feature propagation are then input into the local convolution correction unit, which compensates for the local spatial structure within the neighborhood, combining the long-range dependency information obtained from state-space propagation with the local neighborhood information obtained from the convolution operation.
[0126] In the case of multiple residual physical diffusion coding groups cascaded together, the following can be obtained sequentially: , ... The physical constraint state space encoder employs multi-level output thermal features. Each level of residual physical diffusion coding can continue to propagate the state space using the diffusion coefficient map based on the output thermal features of the previous level, and correct the local spatial structure through local convolution correction units. Thus, the physical constraint state space encoder can gradually form a more stable physical constraint thermal feature representation during multi-level feature propagation.
[0127] The output thermal features of the last residual physical diffusion coding group are used to determine the thermal prior features of the physical constraints. In other words, after multiple residual physical diffusion coding groups have been cascaded, the output thermal features of the last stage can be used as the coding result of the physical constraint state space encoder, or the output thermal features of the last stage can be used as the thermal prior features of the physical constraints after adaptive feature mapping.
[0128] This application embodiment incorporates multiple cascaded residual physical diffusion coding groups within a physically constrained state-space encoder, enabling thermal features to undergo multi-stage boundary-aware state-space propagation and local convolutional correction. Compared to single-stage state-space propagation, multiple residual physical diffusion coding groups enhance long-range thermal dependency modeling capabilities and preserve the local spatial structure of the image through local convolutional correction. This suppresses cross-boundary thermal feature propagation while improving the stability of thermal structure representation in complex boundaries, slender targets, and local high-frequency regions, contributing to enhanced physical consistency and structural clarity in subsequent thermal infrared image super-resolution reconstruction.
[0129] Figure 3 This is a schematic diagram of the structure of a physical constraint state-space encoder provided in an embodiment of this application. Figure 3 As shown, the physical constraint state space encoder can adopt a physical information diffusion Mamba structure to input thermal features and diffusion coefficient maps into the state space propagation process and generate physical constraint thermal prior features.
[0130] Specifically, the shallow thermal infrared features can be expanded into a thermal feature sequence according to a preset scanning order, and the diffusion coefficient map can be expanded into a diffusion coefficient sequence according to the same scanning order. For the t-th position in the sequence, the thermal feature input and the corresponding diffusion coefficient jointly participate in the generation of the physical modulation time step. The physical modulation time step is further used to determine the discrete state transition parameters and input injection parameters, and the state space is updated based on the discrete state transition parameters and input injection parameters.
[0131] like Figure 3 As shown, when the diffusion coefficient is close to zero, it indicates that the current position is more likely to be near a strong structural boundary, and the time step and input injection capability in the state space propagation are suppressed, thereby reducing the propagation of thermal features across structural boundaries. When the diffusion coefficient is large, it indicates that the current position is more likely to be in a relatively smooth region, allowing thermal features to diffuse and propagate within the local region. Thus, the physically constrained state space encoder is no longer a one-dimensional sequence propagation structure without physical constraints, but rather an anisotropic thermal diffusion feature propagation structure capable of boundary awareness.
[0132] Furthermore, the physical constraint state space encoder may include multiple cascaded residual physical diffusion coding groups. Each residual physical diffusion coding group may include a physical information Mamba module and a local convolutional correction unit. The physical information Mamba module is used for long-range thermal dependency modeling constrained by the diffusion coefficient map, and the local convolutional correction unit is used to compensate for the weakening of local neighborhood structure during state space propagation. By cascading multiple residual physical diffusion coding groups, the ability to model long-range dependencies can be maintained while enhancing the expression of local spatial structure, resulting in more stable physical constraint thermal prior features.
[0133] In one possible embodiment, the method steps shown in S105 can be implemented by S1051 to S1054, which are described in detail below.
[0134] S1051. Determine edge saliency information based on optical edge maps.
[0135] Specifically, since the Mamba model typically uses a one-dimensional sequential scanning method when processing two-dimensional image features, it may weaken the original two-dimensional spatial adjacency relationship in complex boundaries, slender targets, and local high-frequency regions. Therefore, edge saliency information can be determined based on the optical edge map so that reliable structural locations in high-resolution optical images can be used to perform topological repair on the physically constrained thermal prior features.
[0136] In one alternative implementation, the optical edge map can be unfolded into a saliency vector, which can be represented as: ;in, Represents the optical edge diagram. This represents the saliency vector obtained from the expansion of the optical edge map. (Saliency vector) The elements in the matrix are used to characterize the edge response intensity of the corresponding spatial location. The higher the edge response intensity, the more likely the corresponding location is to belong to a significant optical structure boundary; the lower the edge response intensity, the more likely the corresponding location is to belong to a weak edge region or a relatively smooth region.
[0137] In this embodiment, the two-dimensional optical edge map can be converted into a saliency vector suitable for index selection, so that the subsequent selection of optical structure pivots and determination of the location of thermal features to be reconstructed can be based on unified edge saliency information.
[0138] S1052. Based on the edge saliency information, select the optical structure feature position that meets the preset saliency condition from the optical structure features as the optical structure pivot.
[0139] Specifically, it can be based on the significance vector. The edge response intensity at each spatial location is used to select locations with higher edge responses from the optical structural features as optical structural pivots. Optical structural pivots are structural reference locations within optical structural features that can characterize significant edges, geometric contours, or local spatial topology. They are used as geometric anchors to assist in repairing the spatial topology of thermal infrared features during subsequent sparse manifold reconstruction.
[0140] In one alternative implementation, it can be derived from the saliency vector. Selecting several positions with the highest response values yields the set of optical structure pivot indices. This process can be represented as: ;in, Represents the set of pivot indices for optical structures. This indicates the operation of selecting the first K positions based on the response value. Indicates the pivot ratio, Indicates the total number of spatial locations. This indicates the number of positions selected as the optical structure pivot.
[0141] Based on optical structure pivot index set It can be derived from optical structural feature sequences The optical structure pivot features at the corresponding positions are selected. It should be noted that the optical structure feature sequence... It can be a feature sequence of optical structural features obtained from high-resolution optical images after being unfolded in the spatial dimension. By selecting only a few salient locations as optical structural pivots, instead of directly inputting all optical textures into the thermal infrared feature reconstruction process, it is possible to utilize more reliable geometric boundary information in the optical images, while reducing the interference of irrelevant optical textures on thermal infrared features.
[0142] S1053. Based on the edge saliency information, determine the location of the thermal feature to be reconstructed in the physical constraint thermal prior features.
[0143] Specifically, after determining the marginal saliency information, it can be based on the saliency vector. Identify the locations of thermal features requiring structural reconstruction from the physical constraint thermal prior features. Since complex boundaries, slender targets, and local high-frequency regions are more susceptible to the effects of one-dimensional sequence scanning, locations with high edge saliency or spatial locations that still have high saliency after perturbation can be prioritized as the locations of thermal features to be reconstructed.
[0144] In one alternative implementation, during the i-th reconstruction stage, the saliency vector can be used as a basis. Determining the set of indices of locations to be recovered from the thermal features can be represented as follows: ;in, This represents the set of locations to be restored in the i-th reconstruction stage. Represents the random disturbance term. This represents the mask ratio for the i-th reconstruction stage. This represents the number of locations to be restored selected in the i-th reconstruction stage.
[0145] In the above process, a random disturbance term is introduced. This allows for differences in the locations to be restored at different reconstruction stages, avoiding the need to reconstruct only fixed edge locations, thereby enhancing the coverage of complex boundaries, slender structures, and local high-frequency regions by multiple reconstruction stages.
[0146] In this embodiment, the location in the physical constraint thermal prior features that requires more spatial structure repair can be determined based on the edge saliency represented by the optical edge map, providing a location basis for subsequent masking and sparse cross-attention processing.
[0147] S1054. Mask the thermal features corresponding to the positions of the thermal features to be reconstructed, and perform sparse cross-attention processing based on the masked thermal features and the optical structure pivot to obtain the reconstructed thermal structure features.
[0148] Specifically, in determining the set of locations to be restored in the i-th reconstruction stage Then, the thermal features corresponding to the positions to be recovered in the physical constraint thermal prior features can be masked.
[0149] Specifically, for the location to be restored, the corresponding thermal features can be replaced using a shared learnable mask vector; for unselected non-restored locations, the original thermal features can be retained. After masking, the input thermal features for the i-th reconstruction stage are obtained, denoted as . .
[0150] Subsequently, sparse cross-attention processing can be performed using the masked hot features as queries and the optical structure pivot features as keys and values.
[0151] In one alternative implementation, sparse cross-attention processing can be represented as: , , , ;in, This represents the input thermal features after masking in the i-th reconstruction stage. Represents the sequence of optical structural features. Represents the set of pivot indices for optical structures. Represents the sequence of optical structural features According to the set of pivot indices of optical structure Selected optical structure pivot features; , and These represent the query projection matrix, the key projection matrix, and the value projection matrix, respectively. This represents the input thermal characteristics after masking. After querying the projection matrix The query features obtained from the mapping Indicates the pivotal features of the optical structure Key projection matrix The key features obtained from the mapping Indicates the pivotal features of the optical structure Longitude projection matrix The characteristics of the values obtained from the mapping; Key features transpose, Indicates the single-head feature dimension. This represents the normalized exponential function, used to obtain query features. Key features Attention weights between them; Represents the input thermal features after masking. Optical structure feature sequence and the set of optical structure pivot indices The sparse cross-attention results obtained.
[0152] In the aforementioned sparse cross-attention processing, not all optical textures are directly transferred to the thermal infrared features. Instead, only a small number of reliable optical structural pivots are selected as geometric anchors, enabling the thermal infrared features to recover the two-dimensional topological structure with the assistance of these optical structural pivots. Since the optical structural pivots originate from locations with high edge saliency, they are more likely to correspond to real geometric boundaries or stable structural contours, thus providing an effective reference for the structural reconstruction of physically constrained thermal prior features.
[0153] Furthermore, the hot features after masking and the sparse cross-attention results can be normalized and feedforward mapped to obtain the output hot features of the i-th reconstruction stage. This process can be expressed as: ;in, Indicates a feedforward network. This indicates normalization processing.
[0154] In this embodiment, a small number of reliable optical structural pivots can be used to repair the structural positions of the thermal features to be reconstructed, thereby restoring the spatial topological relationships in the physically constrained thermal prior features that may have been damaged by sequential scanning. After one or more reconstruction stages, the final thermal features can be used as the reconstructed thermal structural features.
[0155] This application embodiment determines edge saliency information based on optical edge maps and selects optical structural pivots and the locations of thermal features to be reconstructed based on this edge saliency information. This concentrates the reconstruction process on locations prone to spatial structural damage, such as complex boundaries, slender targets, and local high-frequency regions. By masking the locations of the thermal features to be reconstructed and using sparse cross-attention processing with optical structural pivots, the two-dimensional topological structure of thermal infrared features can be recovered using reliable optical structures as geometric anchors without directly introducing all optical textures. Therefore, this application embodiment can improve the spatial structural continuity of physically constrained thermal prior features, reduce optical texture contamination, and improve the boundary sharpness and structural stability of subsequent thermal infrared image super-resolution reconstruction results.
[0156] Figure 4 This is a schematic diagram of a sparse manifold reconstruction module provided in an embodiment of this application. Figure 4 As shown, the sparse manifold reconstruction module is used to select optical structure pivots based on optical edge maps and to reconstruct sparse manifolds based on physical constraint thermal prior features using optical structure pivots.
[0157] Specifically, the optical edge map can be first expanded into a saliency vector, and the positions with higher response values can be selected from the saliency vector as the set of optical structure pivot indices. Based on the set of optical structure pivot indices, the corresponding optical structure pivot features can be extracted from the optical structure feature sequence. Since the optical structure pivots originate from salient edges or stable structure locations, they can serve as reliable geometric anchors in the thermal infrared feature space structure repair process.
[0158] Furthermore, the positions of the thermal features to be reconstructed in the physical constraint thermal prior features can be determined based on the saliency vector, and the thermal features corresponding to the positions to be reconstructed are masked. The masked thermal features are used as query features, and the optical structure pivot features are used as key and value features, respectively, and sparse cross-attention processing is performed. In this way, the sparse manifold reconstruction module does not directly transfer all optical textures to the thermal infrared features, but only uses a small number of reliable optical structure pivots to assist in recovering the two-dimensional spatial topological relationships in the thermal infrared features.
[0159] like Figure 4 As shown, the sparse manifold reconstruction module can include multiple reconstruction stages. In different reconstruction stages, different sets of locations to be restored can be determined based on edge saliency information and perturbation terms, and structural repair can be performed stage by stage on complex boundaries, slender targets, and local high-frequency regions in the thermal features. Through multi-stage sparse manifold reconstruction, the problems of weakened local adjacency relationships, spatial manifold distortion, and topological breaks caused by one-dimensional scanning of the state-space model can be alleviated, thereby obtaining the reconstructed thermal structural features.
[0160] pass Figure 4 The sparse manifold reconstruction module shown in this application embodiment can perform structural repair on thermal infrared features using optical structural pivots without directly introducing all optical textures, thereby improving the structural recovery capability of complex boundaries, slender targets and local high-frequency regions, and reducing the risk of optical textures contaminating thermal infrared reconstruction results.
[0161] In one possible embodiment, the method steps shown in S106 can be implemented by S1061 to S1065, which are described in detail below.
[0162] S1061. Map the reconstructed thermal structure features to the shared comparison space to obtain thermal structure comparison features.
[0163] Specifically, since the reconstructed thermal structural features belong to thermal infrared modal features and the optical structural features belong to optical modal features, their feature distributions, response intensities, and physical meanings differ. To compare their consistency in spatial structure, the reconstructed thermal structural features can be projected onto a shared comparison space to obtain comparative thermal structural features.
[0164] In one alternative implementation, the projection process for reconstructing the thermal structure features can be represented as: ;in, This indicates the reconstructed thermal structure characteristics. This represents the shared comparison space projection function corresponding to the thermal structural features. This represents the comparative features of the thermal structure. Through this projection process, the reconstructed thermal structure features can be converted into feature expressions suitable for consistent comparison with optical structure features.
[0165] S1062. Map the optical structure features to the shared comparison space to obtain the optical structure comparison features.
[0166] Specifically, optical structural features can be projected onto a shared comparison space identical to that of thermal structural comparison features to obtain optical structural comparison features. This shared comparison space is used to eliminate the differences between thermal infrared modes and optical modes in feature scale and channel representation, enabling them to perform structural consistency measurement within a unified feature space.
[0167] In one alternative implementation, the projection process of the optical structural features can be represented as: ;in, Indicates optical structural features. This represents the shared comparison space projection function corresponding to the optical structure features. This represents the comparative features of the optical structure. Through this projection process, the optical structural features can be converted into feature expressions comparable to the comparative features of the thermal structure.
[0168] S1063. Based on the thermal structure comparison characteristics and the optical structure comparison characteristics, determine the structural consistency between the reconstructed thermal structure characteristics and the optical structure characteristics.
[0169] Specifically, after obtaining the thermal structure comparison features and optical structure comparison features, channel normalization can be performed on both, and the similarity at each spatial location can be calculated to obtain structural consistency. Since there are modal differences between optical images and thermal infrared images, not all optical structures correspond to the true thermal radiation boundary. Therefore, structural consistency is needed to determine whether the optical structure at the corresponding spatial location has reliable reference value for thermal infrared reconstruction.
[0170] In one alternative implementation, spatial location can be calculated. The cosine similarity at a given point can be expressed as: ;in, Indicates spatial location Structural consistency at the location Indicates the spatial location of thermal structural comparison features The feature vector at that location, Indicates the spatial location of optical structural comparison features The feature vector at that location, Let denote the L2 norm, and 〈·, ·〉 denote the vector inner product. The larger the value, the higher the spatial location. The more consistent the thermal structure and optical structure are at that location; The smaller the value, the higher the spatial location. The optical structure at that location may not match the thermal radiation structure.
[0171] S1064. Generate trust gates based on structural consistency.
[0172] Specifically, a trust gate can be generated based on structural consistency. The trust gate is used to characterize the reliability of optical structural information at different spatial locations. Locations with high structural consistency have higher trust gate values, indicating that the optical structural information has high reference value for thermal infrared reconstruction. Locations with low structural consistency have lower trust gate values, indicating that the optical structural information may not match the thermal radiation structure, and the injection intensity of the optical information should be reduced.
[0173] In one alternative implementation, the trust gate can be represented as: ;in, This indicates a trust gate. This represents the convolution mapping function.
[0174] Through convolution mapping function It can perform spatial mapping and local context integration for structural consistency; through the Sigmoid activation function, the mapping result can be converted into gating weights within a preset range.
[0175] S1065. Based on the trust gate, optically guided gating refinement of the reconstructed thermal structure features is performed to obtain refined thermal features.
[0176] Specifically, after obtaining the trust gate, the influence of optical guidance information on the reconstructed thermal structure features can be controlled. When the trust gate value is high, the participation of optical structure information in the refinement of the thermal structure can be enhanced; when the trust gate value is low, the participation of optical structure information in the refinement of the thermal structure can be weakened, and the structural expression of the thermal infrared mode itself can be maintained first.
[0177] In this embodiment, by sharing the comparison spatial projection, structural consistency calculation, and trust gate generation, the reliability of the optical structural information can be determined before performing optical guided refinement. This avoids directly injecting optical textures inconsistent with the thermal radiation structure into the thermal infrared reconstruction process, thereby reducing the risks of false edges, optical texture overflow, and localized false hotspots.
[0178] This application's embodiments map reconstructed thermal and optical structural features to a shared comparison space, and generate a trust gate based on their structural consistency in spatial location. This transforms the introduction of optical guidance information from a fixed intensity or unconditional fusion into a dynamic control based on the consistency between the thermal and optical structures. Consequently, optical structural compensation information can be fully utilized at locations with high consistency between the thermal and optical structures, while the injection of untrusted optical textures can be suppressed at locations with low consistency, thereby improving the structural reliability and thermal radiation representation stability of thermal infrared image super-resolution reconstruction.
[0179] In one possible embodiment, the method steps shown in S1065 can be implemented by S1071 to S1074, which are described in detail below.
[0180] S1071. Based on the reconstructed thermal structure features, generate thermal infrared manifold preservation features.
[0181] Specifically, a manifold preservation branch can be set up, and the reconstructed thermal structure features can be input into the manifold preservation branch to generate thermal infrared manifold preservation features. The thermal infrared manifold preservation features are used to preserve the thermal structure, thermal radiation distribution, and reconstructed spatial manifold relationship of the thermal infrared mode itself, avoiding the destruction of the physical expression of the thermal infrared image itself due to over-reliance on optical images in subsequent fusion processes.
[0182] In one alternative implementation, the output of the manifold-preserving branch can be represented as .
[0183] in, This represents the output of the manifold preservation branch, i.e., the thermal infrared manifold preservation feature. This thermal infrared manifold preservation feature can serve as a basic branch in subsequent gated fusion, used to preferentially preserve the structure of the thermal infrared mode itself in locations where the optical structure is unreliable or has low consistency.
[0184] S1072. Generate optical compensation features based on optical structural features.
[0185] Specifically, a compensation retrieval branch can be set up, and optical compensation features can be generated based on optical structural features. The optical compensation features are used to extract compensable information from the optical modes, such as structural information related to edge structures, geometric contours, or local spatial topology, to help the thermal infrared reconstruction results recover clearer boundaries and local details.
[0186] In one alternative implementation, the output of the compensation retrieval branch can be represented as .
[0187] in, This indicates the output of the compensation retrieval branch, i.e., the optical compensation feature.
[0188] It should be noted that the compensation retrieval branch does not directly inject all optical textures into the thermal infrared features, but rather provides selectable optical structure compensation information under the control of the trust gate.
[0189] S1073. Determine the injection intensity of the optical compensation feature relative to the thermal infrared manifold preservation feature based on the trust gate.
[0190] Specifically, the trust gate This is used to control the injection intensity of the optical compensation feature relative to the thermal infrared manifold preservation feature. When the thermal and optical structures have high consistency, the trust gate... A higher value indicates that the optical structure has high reference value for thermal infrared reconstruction, and the injection of optical compensation features can be appropriately enhanced; when the consistency between the thermal structure and the optical structure is low, the trust gate... A lower value indicates that the optical texture may not match the thermal radiation structure, which can reduce the injection of optical compensation features and prioritize the preservation of thermal infrared manifold features.
[0191] Thus, the trust gate It can be used as a fusion weight between thermal infrared manifold preservation features and optical compensation features, so that the fusion process can be dynamically adjusted according to structural consistency, rather than using the same optical compensation intensity for all spatial locations.
[0192] S1074. The thermal infrared manifold preservation feature and optical compensation feature are fused according to the injection intensity to obtain the refined thermal feature.
[0193] Specifically, it can be done according to the trust gate. A defined injection intensity preserves the characteristics of the thermal infrared manifold. and optical compensation features Element-by-element fusion is performed to obtain refined thermal characteristics. This process can be represented as: ;in, This indicates the fusion result, i.e., the refined thermal characteristics. This indicates element-wise multiplication.
[0194] From the above formula, it can be seen that when When the value is small, the fusion result More closely resembles manifold-preserving branch output This indicates that the system prioritizes maintaining its own manifold structure in the thermal infrared mode; when When the value is large, the fusion result More compensation search branch outputs This demonstrates that the system enhances optical compensation information. In this way, the intensity of optical information injection can be dynamically controlled based on the consistency between the thermal and optical structures.
[0195] This application embodiment sets up a manifold preservation branch and a compensation retrieval branch, and utilizes a trust gate. Element-wise gating fusion of the two technologies allows for the selective introduction of reliable optical structure compensation information while preserving the inherent thermal structure of the thermal infrared mode. When the thermal and optical features have high consistency, the system appropriately enhances optical compensation; when the consistency is low, the system reduces the injection of optical information, prioritizing the preservation of the thermal infrared mode's inherent manifold structure. This mechanism effectively suppresses optical texture overflow and false hotspot generation, and improves the structural clarity and thermal structure realism of high-resolution thermal infrared images.
[0196] In one possible embodiment, the method steps shown in S107 can be implemented by S1071 to S1073, which are described in detail below.
[0197] S1071. Perform feature aggregation on the refined thermal features to obtain aggregated thermal features.
[0198] Specifically, the refined thermal features are those refined through uniform gating, which already include the manifold structure information of the thermal infrared mode itself and the reliable optical compensation information introduced under trust gate control. In order to integrate the above multi-source refined information into a feature representation suitable for subsequent image reconstruction, the refined thermal features can first be subjected to feature aggregation processing to obtain aggregated thermal features.
[0199] In one alternative implementation, the refined thermal features can be aggregated using 1×1 convolution. 1×1 convolution aggregation is used to fuse and compress the channel information of the refined thermal features without changing the spatial dimensions, unifying the thermal structure information, edge compensation information, and local detail information in different channels into aggregated thermal features. This process reduces channel redundancy in subsequent reconstruction output and transforms the refined thermal features into feature representations more suitable for reconstruction head processing.
[0200] In this embodiment, feature aggregation does not reintroduce new optical image texture information, but rather integrates the refined thermal features obtained after the aforementioned consistency-gated refinement at the channel level. Therefore, the aggregated thermal features still primarily consist of thermal infrared reconstruction information, while retaining the reliable optical structure compensation information filtered by the trust gate.
[0201] S1072. Perform nonlinear feature mapping on the polymerization thermal characteristics to obtain the reconstructed thermal characteristics.
[0202] Specifically, after obtaining the aggregated thermal features, these features can be input into a lightweight reconstruction head. The lightweight reconstruction head then performs nonlinear feature mapping to obtain the reconstructed thermal features. The lightweight reconstruction head is used to further map the aggregated thermal features into reconstructed features suitable for generating high-resolution thermal infrared images.
[0203] In one alternative implementation, the lightweight reconstruction head may include two 3×3 convolutional layers and a non-linear activation layer disposed between the two 3×3 convolutional layers.
[0204] The first 3×3 convolution layer is used to perform feature mapping within the local spatial receptive field of the aggregated thermal features, so that the boundary, local details and thermal structure information in the aggregated thermal features can be further expressed; the intermediate nonlinear activation layer is used to enhance the nonlinear expression capability of the reconstruction head, so that it can fit the complex mapping relationship between low-resolution thermal infrared images and high-resolution thermal infrared images; the second 3×3 convolution layer is used to further convert the nonlinearly activated features into reconstructed thermal features.
[0205] In this embodiment, the thermal structure representation and reliable optical structure compensation information in the aggregated thermal features can be further mapped into reconstructed thermal features suitable for upsampling output, thereby providing a feature basis for subsequent super-resolution reconstruction at the target scale.
[0206] S1073. Upsample the reconstructed thermal features to obtain a high-resolution thermal infrared image.
[0207] Specifically, after obtaining the reconstructed thermal features, upsampling can be performed on these features to restore them to the target scale, resulting in a high-resolution thermal infrared image. In one optional implementation, a PixelShuffle upsampling operation can be used to output a super-resolution thermal infrared image at the target scale. The PixelShuffle upsampling operation rearranges the channel dimension information in the reconstructed thermal features to the spatial dimension, thereby improving spatial resolution.
[0208] The target scale can be a preset super-resolution magnification scale, such as 2x, 4x, or other magnification scales set according to application requirements. Through PixelShuffle upsampling processing, the spatial size of thermal infrared images can be magnified with low computational complexity, while preserving the boundary information, local detail information, and thermal structure information already formed in the reconstructed thermal features.
[0209] During the inference phase, a low-resolution thermal infrared image and a high-resolution optical image spatially aligned with it can be input into a trained thermal infrared image super-resolution network. The network then sequentially performs thermal feature extraction, optical prior generation, physical constraint state space encoding, sparse manifold reconstruction, consistency gating refinement, and reconstruction output processing to obtain the corresponding high-resolution thermal infrared reconstruction result.
[0210] This embodiment first performs 1×1 convolution aggregation on the refined thermal features, then uses a lightweight reconstruction head containing two 3×3 convolution layers and an intermediate nonlinear activation layer for nonlinear feature mapping, and finally employs a PixelShuffle upsampling operation to output a high-resolution thermal infrared image at the target scale. This allows the thermal structure information and reliable optical compensation information in the refined thermal features to be effectively converted into the final reconstructed image. Therefore, while maintaining low reconstruction complexity, it improves the boundary sharpness, local detail representation, and thermal structure continuity of the high-resolution thermal infrared image.
[0211] In one possible embodiment, the method further includes training a thermal infrared image super-resolution network. The training process can be implemented through S1091 to S1094, which are described in detail below.
[0212] S1091. Obtain the training sample set.
[0213] The training sample set includes low-resolution thermal infrared sample images, high-resolution optical sample images spatially aligned with the low-resolution thermal infrared sample images, and high-resolution thermal infrared ground truth images.
[0214] Specifically, the training sample set is used for end-to-end training of the thermal infrared image super-resolution network to be trained. Each training sample set may include low-resolution thermal infrared sample images, high-resolution optical sample images, and high-resolution thermal infrared ground truth images.
[0215] Among them, low-resolution thermal infrared sample images are used as input images to be reconstructed, high-resolution optical sample images are used as structure-guided input images, and high-resolution thermal infrared ground truth images are used as supervision labels.
[0216] It should be noted that the high-resolution optical sample image and the low-resolution thermal infrared sample image have a spatial alignment relationship, allowing the edges, contours, and local spatial topology in the high-resolution optical sample image to correspond to the thermal radiation regions in the low-resolution thermal infrared sample image. The high-resolution thermal infrared ground truth image is used to characterize the desired output thermal infrared image result and to constrain the pixel reconstruction accuracy, perceptual structure consistency, and edge gradient consistency of the network output image during training.
[0217] In this embodiment, the training sample set can be derived from a dual-light camera on a drone, a vehicle-mounted multimodal imaging device, a fixed monitoring device, or other optical-thermal infrared synchronous acquisition systems. By acquiring a training sample set containing low-resolution thermal infrared sample images, high-resolution optical sample images, and high-resolution thermal infrared ground truth images, a data foundation can be provided for subsequent composite loss calculation and network parameter optimization.
[0218] S1092. Input the low-resolution thermal infrared sample image and the high-resolution optical sample image into the thermal infrared image super-resolution network to be trained to obtain the training output image.
[0219] Specifically, low-resolution thermal infrared sample images and spatially aligned high-resolution optical sample images can be input into the thermal infrared image super-resolution network to be trained. The thermal infrared image super-resolution network to be trained may include a thermal feature extraction part, an optical prior generation part, a physically constrained state space encoding part, a sparse manifold reconstruction part, a consistency gating thinning part, and a reconstruction output part.
[0220] During the training process, the thermal infrared image super-resolution network to be trained first extracts thermal features from low-resolution thermal infrared sample images to obtain shallow thermal infrared sample features; then it generates optical structure priors and diffusion coefficient maps based on high-resolution optical sample images; subsequently, it obtains refined thermal features through physical constraint state space encoding, sparse manifold reconstruction, and consistency gating refinement; finally, it obtains the training output image through reconstruction output processing.
[0221] The training output image can be denoted as This represents the super-resolution thermal infrared image output by the thermal infrared image super-resolution network to be trained, based on low-resolution thermal infrared sample images and high-resolution optical sample images. This training output image is used to compare with the high-resolution thermal infrared ground truth image to determine the subsequent composite loss.
[0222] S1093. Determine the composite loss based on the training output image and the high-resolution thermal infrared ground truth image.
[0223] The composite loss includes pixel reconstruction loss, perceptual structure loss, and edge supervision loss.
[0224] Specifically, a composite loss can be determined based on the training output image and the high-resolution thermal infrared ground truth image. The composite loss is used to jointly constrain the thermal infrared image super-resolution network to be trained from three aspects: pixel fidelity, perceptual structure consistency, and edge gradient consistency.
[0225] In one alternative implementation, the pixel reconstruction loss can be a multi-scale Charbonnier loss. The multi-scale Charbonnier loss constrains the pixel-level reconstruction accuracy between the training output image and the high-resolution thermal infrared ground truth image, and can be expressed as: ;in, This represents the multi-scale Charbonnier loss. Let {1,2,4} represent the scaling factor, and {1,2,4} represent multiple supervision scales. Representing scale The corresponding set of spatial locations Representing scale The corresponding number of spatial locations Represents the training output image In scale Lower spatial position Pixel value at that location, Represents high-resolution thermal infrared true image In scale Lower spatial position Pixel value at that location, This represents a constant used to improve numerical stability. By employing multi-scale Charbonnier loss, the pixel differences between the training output image and the high-resolution thermal infrared ground truth image can be constrained at different scales, thereby improving pixel-level reconstruction accuracy.
[0226] In one alternative implementation, perceptual structural loss can be used. Perceptual loss is used to constrain the high-level structural and semantic consistency between the training output image and the high-resolution thermal infrared ground truth image, and can be expressed as: ;in, Indicates perceived loss. This represents a perceptual feature extractor. By using perceptual loss, the thermal infrared image super-resolution network to be trained can focus not only on pixel-level errors, but also on the consistency between the reconstructed image and the ground truth image in terms of high-level structural representation.
[0227] In one alternative implementation, the edge-supervised loss can be used to constrain the gradient field consistency between the training output image and the high-resolution thermal infrared ground truth image. The edge-supervised loss can be expressed as: ;in, Indicates edge monitoring loss, Let ||·||1| denote the Sobel gradient magnitude calculation function, and ||·||1| denote the L1 norm. By using edge-supervised loss, the consistency between the trained output image and the high-resolution thermal infrared ground truth image in terms of edge gradients and boundary sharpness can be enhanced, thereby reducing edge blurring and structural distortion.
[0228] Furthermore, the final total loss can be determined based on the pixel reconstruction loss, perceptual structure loss, and edge supervision loss. The final total loss can be expressed as: ;in, The total loss represents the final loss, and α, β, and γ represent the loss weight parameters. By setting the loss weight parameters, the constraint strength of pixel-level reconstruction accuracy, perceptual structure consistency, and edge sharpness during the training and optimization process can be adjusted.
[0229] S1094. The thermal infrared image super-resolution network to be trained is trained based on the composite loss to obtain the trained thermal infrared image super-resolution network.
[0230] Specifically, in determining the final total loss Then, the thermal infrared image super-resolution network to be trained can be performed end-to-end based on the final total loss. During training, the final total loss can be used to determine the optimal training method. Backpropagation and iterative updates are performed on the network parameters, enabling the thermal infrared image super-resolution network to be trained to gradually learn the mapping relationship from low-resolution thermal infrared sample images and high-resolution optical sample images to high-resolution thermal infrared ground truth images.
[0231] During training, multi-scale Charbonnier loss is used to improve pixel-level reconstruction accuracy, making the training output image close to the high-resolution thermal infrared ground truth image at different scales; perceptual loss is used to enhance the high-level structural and semantic consistency between the training output image and the high-resolution thermal infrared ground truth image; edge supervision loss is used to constrain the gradient field and boundary sharpness of the training output image. Through the joint optimization of the above three types of losses, the trained thermal infrared image super-resolution network can simultaneously achieve pixel fidelity, structural rationality, and edge sharpness.
[0232] In this embodiment, the trained thermal infrared image super-resolution network can receive low-resolution thermal infrared images and spatially aligned high-resolution optical images during the inference phase, and output the corresponding high-resolution thermal infrared image. Because the training phase simultaneously constrains pixel-level reconstruction, perceptual structure, and edge gradients through a composite loss, the trained thermal infrared image super-resolution network can output super-resolution thermal infrared reconstruction results with clearer edges, more complete details, and a more reasonable thermal distribution.
[0233] This application embodiment acquires a training sample set including low-resolution thermal infrared sample images, high-resolution optical sample images, and high-resolution thermal infrared ground truth images. It then utilizes a composite loss consisting of multi-scale Charbonnier loss, perceptual loss, and edge supervision loss to perform end-to-end training on the thermal infrared image super-resolution network to be trained. This allows the network to simultaneously focus on pixel reconstruction accuracy, high-level structure consistency, and edge gradient fidelity during training. Consequently, the trained thermal infrared image super-resolution network can improve its overall reconstruction capabilities in terms of pixel fidelity, perceptual quality, and edge sharpness, thereby enhancing the objective metrics and subjective visual quality of high-resolution thermal infrared images.
[0234] The optically guided thermal infrared image super-resolution method provided in this application introduces the physical laws of thermal diffusion into the feature propagation process of a physically constrained state-space encoder, enabling the optical boundary determined by the high-resolution optical image to act as a boundary constraint during thermal feature propagation. Specifically, by determining the diffusion coefficient map through the optical edge map and modulating the state-space propagation parameters using the diffusion coefficient map, the cross-boundary propagation of thermal features in strong boundary regions can be suppressed, while continuous diffusion propagation is maintained in relatively smooth regions. This reduces the problems of excessive smoothing of thermal boundaries and cross-boundary diffusion distortion of thermal structures, improving the physical consistency and interpretability of the thermal infrared image super-resolution reconstruction process.
[0235] Furthermore, in this embodiment, the discrete time step and / or input injection parameters in the physical constraint state-space encoder are modulated using a diffusion coefficient map. This enables the state-space propagation process to no longer rely solely on ordinary data-driven methods, but instead combine thermal feature inputs and physical diffusion constraints to jointly complete feature propagation. Thus, the physical constraint state-space encoder can form a boundary-aware thermal feature propagation mode, maintaining long-range thermal dependency modeling capabilities while improving the preservation of target edges, thermal structure distribution, and the continuity of local thermal radiation.
[0236] Furthermore, this embodiment selects optical structural pivots from optical structural features based on optical edge maps, and uses these optical structural pivots to reconstruct sparse manifolds of physically constrained thermal prior features. This allows for spatial structural repair of thermal features using a small number of reliable optical structural locations. Instead of directly injecting all optical textures into the thermal infrared features, this approach uses optical structural pivots as geometric anchors to assist in restoring the two-dimensional spatial topological relationships within the thermal infrared features. This alleviates the problems of two-dimensional spatial manifold distortion, weakened local adjacency relationships, or topological breaks that may occur during one-dimensional sequence scanning of the state-space model, thereby improving the structural recovery capability for complex boundaries, slender targets, and local high-frequency regions.
[0237] Furthermore, this embodiment employs a consistency-gated refinement mechanism to generate a trust gate based on the structural consistency between the reconstructed thermal and optical structural features. This trust gate controls the injection intensity of optical compensation features relative to the thermal infrared manifold preservation features. When the consistency between the thermal and optical structures is high, the compensation effect of optical structural information on thermal infrared reconstruction can be appropriately enhanced; conversely, when the consistency is low, the injection of optical information can be weakened, and the manifold structure of the thermal infrared mode itself can be preferentially preserved. Thus, while utilizing optical images to provide structural references, it reduces the entry of optical textures mismatched with thermal radiation distribution into the thermal infrared reconstruction process, effectively suppressing optical texture overflow, false edges, and localized false hotspots.
[0238] Furthermore, during the training phase, embodiments of this application can employ a composite loss, including pixel reconstruction loss, perceptual structure loss, and edge supervision loss, to train the thermal infrared image super-resolution network. Specifically, the pixel reconstruction loss constrains the pixel-level differences between the training output image and the high-resolution thermal infrared ground truth image; the perceptual structure loss constrains the consistency of their high-level structural representations; and the edge supervision loss constrains the consistency of their edge gradients and boundary sharpness. Through joint optimization using these composite losses, the trained thermal infrared image super-resolution network can simultaneously focus on pixel reconstruction accuracy, high-level structural consistency, and edge gradient fidelity, thereby improving the objective metrics and subjective visual quality of the reconstructed image.
[0239] In summary, the embodiments of this application can improve the spatial resolution of thermal infrared images while taking into account the physical consistency of thermal structures, reliable utilization of optical structures, preservation of local boundaries, and efficiency of engineering applications. They are suitable for optically guided thermal infrared image enhancement scenarios such as UAV thermal infrared image enhancement, nighttime inspection, security monitoring, target recognition, disaster search and rescue, and environmental monitoring.
[0240] In one possible embodiment, the optically guided thermal infrared image super-resolution method provided in this application can be implemented using a hierarchical processing structure. This hierarchical processing structure may include a data input layer, a priori generation layer, a core model layer, a training optimization layer, and an application output layer. Through this hierarchical processing structure, low-resolution thermal infrared images and high-resolution optical images can sequentially complete data input, optical prior generation, physical constraint feature propagation, cross-modal reliable fusion, network training optimization, and super-resolution result output.
[0241] Specifically, at the data input layer, a low-resolution thermal infrared image to be processed and a high-resolution optical image spatially aligned with the low-resolution thermal infrared image can be acquired. The low-resolution thermal infrared image serves as the thermal imaging image to be reconstructed using super-resolution technology, while the high-resolution optical image serves as a structure-guided image, providing structural reference information such as edge structures, geometric contours, and local spatial topology. Since the high-resolution optical image and the low-resolution thermal infrared image have a spatial correspondence, the structural information in the high-resolution optical image can provide a reference basis for subsequent super-resolution reconstruction of the thermal infrared image.
[0242] In the prior generation layer, preliminary feature processing can be performed on the input image. On one hand, shallow feature extraction is performed on the low-resolution thermal infrared image to obtain shallow thermal infrared features, enabling subsequent processing to be based on the thermal radiation expression of the thermal infrared image itself. On the other hand, optical structural features and optical edge maps are generated based on the high-resolution optical image, and the diffusion coefficient map is determined based on the optical edge map. Among them, the optical structural features are used to characterize the edge structure, geometric contour, and local spatial topology in the high-resolution optical image; the optical edge map is used to characterize the strength of optical structural boundaries at different spatial locations; and the diffusion coefficient map is used to characterize the thermal feature diffusion capability at different spatial locations, enabling strong edge regions to suppress thermal feature diffusion across boundaries and relatively smooth regions to allow thermal features to propagate continuously.
[0243] At the core model layer, the shallow thermal infrared features, optical structure features, optical edge maps, and diffusion coefficient maps obtained from the prior generation layer can be input into the core model for processing. The core model can include a physically constrained state-space encoder, a sparse manifold reconstruction module, and a consistency-gated refinement module. The physically constrained state-space encoder modulates the state-space propagation parameters based on the diffusion coefficient map, subjecting the shallow thermal infrared features to diffusion constraints corresponding to the optical boundaries during propagation in the state space, thus forming physically constrained thermal prior features. The sparse manifold reconstruction module selects optical structure pivots based on the optical edge maps and uses these pivots to perform sparse manifold reconstruction of the physically constrained thermal prior features, repairing the two-dimensional spatial structure that may be weakened during sequential propagation. The consistency-gated refinement module controls the injection intensity of optical auxiliary information based on the structural consistency between the reconstructed thermal and optical structure features, thereby reducing the interference of mismatched optical textures on the thermal infrared reconstruction results while utilizing reliable optical structures to compensate for thermal infrared details.
[0244] In the training optimization layer, a composite loss can be used to train the thermal infrared image super-resolution network. This composite loss can include pixel reconstruction loss, perceptual structure loss, and edge supervision loss. Specifically, the pixel reconstruction loss constrains the pixel-level reconstruction error between the training output image and the high-resolution thermal infrared ground truth image; the perceptual structure loss constrains the high-level structural consistency between the training output image and the high-resolution thermal infrared ground truth image; and the edge supervision loss constrains the edge gradient consistency between the training output image and the high-resolution thermal infrared ground truth image. Through the joint optimization of these losses, the trained thermal infrared image super-resolution network can simultaneously achieve pixel fidelity, structural rationality, and boundary sharpness.
[0245] In the output layer, the low-resolution thermal infrared image to be processed and its spatially aligned high-resolution optical image can be input into the trained thermal infrared image super-resolution network to obtain the corresponding high-resolution thermal infrared image. The high-resolution thermal infrared image has higher spatial resolution than the low-resolution thermal infrared image and exhibits better representation in terms of edge sharpness, local detail integrity, and thermal structure continuity. This method can be applied to scenarios such as UAV thermal infrared image enhancement, nighttime patrol, security monitoring, target recognition, disaster search and rescue, and environmental monitoring to improve the reliability of subsequent intelligent perception and target analysis tasks.
[0246] Figure 5 This is a schematic diagram illustrating the layered workflow of an optically guided thermal infrared image super-resolution method provided in an embodiment of this application. Figure 5 As shown, this method uses a data input layer, a prior generation layer, a core model layer, a training optimization layer, and an application output layer to process data in sequence. This enables structural information in optical images to participate in thermal infrared image super-resolution reconstruction in a physically constrained and consistent gating manner. This improves the resolution of thermal infrared images while reducing the risks of optical texture overflow, false edges, and thermal structure distortion.
[0247] This application also provides an optically guided thermal infrared image super-resolution system, which is used to execute the optically guided thermal infrared image super-resolution method of any of the above embodiments. This system can be deployed in servers, edge computing devices, UAV-borne processing devices, vehicle-mounted multimodal imaging devices, or other electronic devices with image processing capabilities.
[0248] Figure 6 This is a schematic diagram of the structure of an optically guided thermal infrared imaging super-resolution system provided in an embodiment of this application. Figure 6 As shown, the thermal infrared image super-resolution system 600 includes an image acquisition module 601, a thermal feature extraction module 602, an optical prior generation module 603, a physical constraint encoding module 604, a manifold reconstruction module 605, a gating refinement module 606, and a reconstruction output module 607.
[0249] The image acquisition module 601 is used to acquire a low-resolution thermal infrared image to be processed and a high-resolution optical image spatially aligned with the low-resolution thermal infrared image. The low-resolution thermal infrared image serves as the thermal infrared input image to be reconstructed, while the high-resolution optical image provides structural reference information such as edge structure, geometric contours, and local spatial topology.
[0250] The thermal feature extraction module 602 is used to extract thermal features from low-resolution thermal infrared images to obtain shallow thermal infrared features. These shallow thermal infrared features are used to characterize the initial thermal radiation distribution, local grayscale changes, and basic thermal structure information in the low-resolution thermal infrared images, and serve as input for subsequent physical constraint coding.
[0251] The optical prior generation module 603 is used to generate optical structural features and optical edge maps based on high-resolution optical images, and to determine the diffusion coefficient map based on the optical edge map. The diffusion coefficient map is used to characterize the thermal feature diffusion capability at different spatial locations, enabling the optical boundary to act as a boundary constraint during subsequent thermal feature propagation.
[0252] The physical constraint encoding module 604 is used to input the shallow thermal infrared features and diffusion coefficient map into the physical constraint state-space encoder, and modulate the state-space propagation parameters in the physical constraint state-space encoder based on the diffusion coefficient map to obtain the physical constraint thermal prior features. This allows the thermal feature propagation process to be physically constrained by the diffusion coefficient map, reducing unreasonable propagation of thermal features across structural boundaries.
[0253] The manifold reconstruction module 605 is used to select optical structure pivots from optical structure features based on optical edge maps, and to perform sparse manifold reconstruction on physically constrained thermal prior features using optical structure pivots to obtain reconstructed thermal structure features. This module allows for the repair of spatial structures in thermal infrared features using a small number of reliable optical structure locations.
[0254] The gated refinement module 606 is used to perform optically guided gated refinement of the reconstructed thermal structural features based on the structural consistency between the reconstructed thermal structural features and the optical structural features, thereby obtaining refined thermal features. This module can control the injection intensity of optical structural information according to the level of trust. When the structural consistency is high, optical compensation information is introduced, and when the structural consistency is low, the structure of the thermal infrared mode itself is preserved first.
[0255] The reconstruction output module 607 is used for super-resolution reconstruction based on refined thermal features to obtain a high-resolution thermal infrared image. Specifically, the reconstruction output module can perform feature aggregation, nonlinear feature mapping, and upsampling processing on the refined thermal features to output a high-resolution thermal infrared image at the target scale.
[0256] The optically guided thermal infrared image super-resolution system provided in this application, through the collaborative work of the image acquisition module, thermal feature extraction module, optical prior generation module, physical constraint encoding module, manifold reconstruction module, gated thinning module and reconstruction output module, can suppress the interference of unreliable optical textures on the thermal infrared reconstruction results while using high-resolution optical images to provide structural references, thereby improving the boundary sharpness, thermal structural continuity and physical consistency of the thermal infrared image super-resolution reconstruction results.
[0257] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 7 As shown, the electronic device 700 provided in this embodiment includes a memory 701 and a processor 702.
[0258] The memory 701 can be a separate physical unit, connected to the processor 702 via a bus 703. Alternatively, the memory 701 and processor 702 can be integrated and implemented in hardware. The memory 701 stores program instructions, which the processor 702 calls to execute the operations performed by the thermal infrared image super-resolution system in any of the above method embodiments.
[0259] Optionally, when some or all of the methods in the above embodiments are implemented by software, the electronic device 700 may also include only the processor 702. A memory 701 for storing programs is located outside the electronic device 700, and the processor 702 is connected to the memory via circuits / wires to read and execute the programs stored in the memory. The processor 702 may be a central processing unit (CPU), a network processor (NP), or a combination of a CPU and an NP. The processor 702 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0260] The memory 701 may include volatile memory, such as random-access memory (RAM); the memory may also include non-volatile memory, such as flash memory, hard disk drive (HDD) or solid-state drive (SSD); the memory may also include a combination of the above types of memory.
[0261] For example, this application provides a chip including: an interface circuit and a logic circuit. The interface circuit is used to receive signals from other chips outside the chip and transmit them to the logic circuit, or to send signals from the logic circuit to other chips outside the chip. The logic circuit is used to perform the operations performed by the thermal infrared image super-resolution system in the above method embodiments.
[0262] For example, this application provides a computer-readable storage medium having computer program instructions stored thereon. The computer program instructions are executed by a processor of an electronic device to cause the electronic device to perform the operations performed by the thermal infrared image super-resolution system in the above method embodiments.
[0263] For example, this application provides a computer program product that, when run on an electronic device, causes the electronic device to perform the operations performed by the thermal infrared image super-resolution system in the above method embodiments.
[0264] The above are merely specific embodiments of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to these embodiments, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An optically guided super-resolution method for thermal infrared images, characterized in that, The method includes: Acquire the low-resolution thermal infrared image to be processed and the high-resolution optical image spatially aligned with the low-resolution thermal infrared image; Thermal features are extracted from low-resolution thermal infrared images to obtain shallow thermal infrared features; Optical structural features and optical edge maps are generated based on high-resolution optical images, and diffusion coefficient maps are determined based on the optical edge maps; wherein, the diffusion coefficient maps are used to characterize the thermal feature diffusion capability at different spatial locations; The shallow thermal infrared features and diffusion coefficient map are input into the physical constraint state space encoder, and the state space propagation parameters in the physical constraint state space encoder are modulated based on the diffusion coefficient map to obtain the physical constraint thermal prior features. Based on the optical edge map, optical structure pivots are selected from the optical structure features, and the physical constraint thermal prior features are reconstructed using the optical structure pivots to obtain the reconstructed thermal structure features. Based on the structural consistency between the reconstructed thermal structural features and the optical structural features, optically guided gating refinement is performed on the reconstructed thermal structural features to obtain refined thermal features. High-resolution thermal infrared images are obtained by super-resolution reconstruction based on refined thermal features.
2. The method according to claim 1, characterized in that, The process of generating optical structural features and optical edge maps based on the high-resolution optical image, and determining a diffusion coefficient map based on the optical edge maps, includes: The high-resolution optical image is scale-aligned to obtain a scale-aligned optical image corresponding to the thermal infrared shallow layer features. Structural features are extracted from the scale-aligned optical image to obtain the optical structural features; Gradient feature extraction and edge fusion processing are performed on the scale-aligned optical image to obtain the optical edge map; The diffusion coefficient map is determined based on the edge response intensity at different spatial locations in the optical edge map.
3. The method according to claim 1, characterized in that, The modulation of the state-space propagation parameters in the physical constraint state-space encoder based on the diffusion coefficient map to obtain physical constraint thermal prior features includes: The thermal infrared shallow layer features and the diffusion coefficient map are respectively serialized to obtain thermal feature sequences and diffusion coefficient sequences. Based on the thermal feature sequence, determine the data-driven modulation information; Based on the diffusion coefficient sequence, determine the physical diffusion modulation information; The state space propagation parameters are determined based on the data-driven modulation information and the physical diffusion modulation information; wherein the state space propagation parameters include discrete time step and / or input injection parameters; Based on the state-space propagation parameters, the thermal feature sequence is propagated in the state space to obtain the physical constraint thermal prior features.
4. The method according to claim 1, characterized in that, The physical constraint state space encoder includes multiple cascaded residual physical diffusion coding groups; Each of the residual physical diffusion coding groups is used to propagate the state space features of the input thermal features based on the diffusion coefficient map, and to perform local convolution correction on the thermal features after propagation of the state space features, so as to obtain the output thermal features of the corresponding residual physical diffusion coding group. Among them, the output thermal features of the last residual physical diffusion coding group are used to determine the physical constraint thermal prior features.
5. The method according to claim 1, characterized in that, The step of selecting an optical structure pivot from the optical structure features based on the optical edge map, and using the optical structure pivot to reconstruct the sparse manifold of the physically constrained thermal prior features to obtain the reconstructed thermal structure features includes: Based on the optical edge map, edge saliency information is determined; Based on the edge saliency information, the position of the optical structure feature that satisfies the preset saliency condition is selected from the optical structure features as the optical structure pivot; Based on the edge saliency information, the location of the thermal feature to be reconstructed in the physical constraint thermal prior features is determined; The thermal features corresponding to the positions of the thermal features to be reconstructed are masked, and sparse cross-attention processing is performed based on the masked thermal features and the optical structure pivot to obtain the reconstructed thermal structure features.
6. The method according to claim 1, characterized in that, The step of optically guiding gating refinement of the reconstructed thermal structure features based on the structural consistency between the reconstructed thermal structure features and the optical structure features to obtain refined thermal features includes: The reconstructed thermal structure features are mapped to a shared comparison space to obtain thermal structure comparison features; The optical structure features are mapped to the shared comparison space to obtain the optical structure comparison features; Based on the thermal structure comparison features and the optical structure comparison features, the structural consistency between the reconstructed thermal structure features and the optical structure features is determined; A trust gate is generated based on the aforementioned structural consistency; The reconstructed thermal structure features are optically guided and refined based on the trust gate to obtain the refined thermal features.
7. The method according to claim 6, characterized in that, The optically guided gating refinement of the reconstructed thermal structure features based on the trust gate, resulting in the refined thermal features, includes: Based on the reconstructed thermal structure features, thermal infrared manifold preservation features are generated; Based on the aforementioned optical structural features, optical compensation features are generated; The injection intensity of the optical compensation feature relative to the thermal infrared manifold preservation feature is determined based on the trust gate; The refined thermal feature is obtained by fusing the thermal infrared manifold preservation feature and the optical compensation feature according to the injection intensity.
8. The method according to claim 1, characterized in that, The process of super-resolution reconstruction based on the refined thermal features to obtain a high-resolution thermal infrared image includes: The refined thermal features are then subjected to feature aggregation to obtain aggregated thermal features; The polymerization thermal features are subjected to nonlinear feature mapping to obtain the reconstructed thermal features; The reconstructed thermal features are upsampled to obtain the high-resolution thermal infrared image.
9. The method according to claim 1, characterized in that, Before performing super-resolution reconstruction based on the refined thermal features to obtain a high-resolution thermal infrared image, the method further includes: Obtain a training sample set; wherein the training sample set includes low-resolution thermal infrared sample images, high-resolution optical sample images spatially aligned with the low-resolution thermal infrared sample images, and high-resolution thermal infrared ground truth images; The low-resolution thermal infrared sample image and the high-resolution optical sample image are input into the thermal infrared image super-resolution network to be trained to obtain the training output image. Based on the training output image and the high-resolution thermal infrared ground truth image, a composite loss is determined; wherein, the composite loss includes pixel reconstruction loss, perceptual structure loss and edge supervision loss; The thermal infrared image super-resolution network to be trained is trained based on the composite loss to obtain the trained thermal infrared image super-resolution network.
10. An optically guided thermal infrared imaging super-resolution system, characterized in that, The system includes: An image acquisition module is used to acquire a low-resolution thermal infrared image to be processed and a high-resolution optical image spatially aligned with the low-resolution thermal infrared image. A thermal feature extraction module is used to extract thermal features from the low-resolution thermal infrared image to obtain shallow thermal infrared features. An optical prior generation module is used to generate optical structural features and an optical edge map based on the high-resolution optical image, and to determine a diffusion coefficient map based on the optical edge map; wherein the diffusion coefficient map is used to characterize the thermal feature diffusion capability at different spatial locations; The physical constraint coding module is used to input the thermal infrared shallow layer features and the diffusion coefficient map into the physical constraint state space encoder, and modulate the state space propagation parameters in the physical constraint state space encoder based on the diffusion coefficient map to obtain the physical constraint thermal prior features. The manifold reconstruction module is used to select optical structure pivots from the optical structure features based on the optical edge map, and to perform sparse manifold reconstruction on the physical constraint thermal prior features using the optical structure pivots to obtain reconstructed thermal structure features. The gated refinement module is used to perform optically guided gated refinement on the reconstructed thermal structure features based on the structural consistency between the reconstructed thermal structure features and the optical structure features, so as to obtain refined thermal features; The reconstruction output module is used to perform super-resolution reconstruction based on the refined thermal features to obtain a high-resolution thermal infrared image.