Dual light fusion super-resolution integrated method and device based on semantic guidance
By adopting a semantically guided dual-light fusion super-resolution integrated method, the semantic gap and resolution mismatch problems in low-resolution infrared-high-resolution visible light image fusion are solved, and high-quality, highly semantically consistent image fusion and super-resolution reconstruction are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TSINGHUA UNIVERSITY
- Filing Date
- 2026-03-10
- Publication Date
- 2026-07-24
AI Technical Summary
In scenarios such as unmanned systems, border patrol and urban security, existing technologies are unable to achieve high-quality super-resolution reconstruction in low-resolution infrared-high-resolution visible light image fusion. There are degradation phenomena such as blurred edges of thermal targets, color bleeding and jagged edges. Furthermore, there is a lack of unified modeling of the symbiotic relationship between thermal targets and textures, resulting in semantic inconsistencies and low resolution.
A semantically guided dual-light fusion super-resolution integrated method is adopted. Multimodal features of infrared and visible light images are obtained through heteroscale shallow coding. The CLIP model is used for semantic recoding and cross-modal attention mechanism alignment. Combined with noise prediction diffusion model and sharpness perception vision-language model, high-level semantic embedding is generated, and finally high-quality fused image is generated.
It achieves high-quality, highly semantically consistent image fusion and super-resolution reconstruction, avoiding problems such as blurred details, insufficient saliency of hot targets, and semantic inconsistency, thereby improving image resolution and semantic consistency.
Smart Images

Figure CN122453604A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a semantically guided dual-light fusion super-resolution integrated method and apparatus. Background Technology
[0002] In practical scenarios such as unmanned systems, border patrol, and urban security, visible light sensors typically possess higher physical resolution (1920×1080 or even 4K), while infrared sensors, limited by detector pixel size and cost, often only output low-resolution signals (640×512 or lower). Therefore, dual-light joint observation presents a cross-scale and cross-modal contradiction: visible light is clear but lacks thermal features, while infrared sensors contain thermal targets but lack spatial detail. Directly performing traditional interpolation amplification and fusion on the infrared image not only blurs the edges of thermal targets but also introduces amplification artifacts into subsequent detection and segmentation processes. Conversely, performing single-modal super-resolution on the visible light alone fails to simultaneously obtain thermal radiation information, making it difficult to meet target recognition requirements under conditions such as nighttime, smoke, and obstruction. Therefore, how to achieve integrated super-resolution and fusion under the typical heteroscale input of low-resolution infrared and high-resolution visible light, so that the fusion result retains both the high-frequency texture of visible light and the clear edges and accurate positions of infrared thermal targets, is a core problem that urgently needs to be solved in the current field of dual-light imaging.
[0003] In related technologies, cascaded schemes of super-resolution followed by fusion or fusion followed by super-resolution are often adopted, which have fundamental drawbacks: First, step-by-step processing amplifies errors at each stage, leading to degradation phenomena such as misalignment of thermal target contours, color bleeding, and jagged edges. Second, independent optimization of network parameters lacks a unified modeling of the symbiotic relationship between thermal targets and textures, making it difficult to achieve optimal fusion. Crucially, when low-resolution infrared images are severely blurred, traditional methods, unable to perceive modal reliability, often adopt an averaging strategy, resulting in noise contamination of texture areas and weakened details in thermal target areas. Although diffusion models possess strong generative priors, their same-scale design struggles to adapt to heterogeneous modal-scale inputs, resulting in high redundancy in iterative computations. Furthermore, they are prone to misinterpreting weak thermal signals as noise filtering or generating incorrect textures that trigger false alarms. Summary of the Invention
[0004] This application provides a semantically guided dual-light fusion super-resolution integrated method and apparatus to solve the semantic gap and resolution mismatch problems in heterogeneous image fusion in related technologies, avoid the problems of blurred details, insufficient saliency of thermal targets, semantic inconsistency and low resolution in fused images, and achieve high-quality, highly semantically consistent integrated fusion and super-resolution reconstruction.
[0005] To achieve the above objectives, the first aspect of this application proposes a semantically guided dual-light fusion super-resolution integration method, comprising the following steps: Infrared and visible light images are acquired, and heteroscale shallow coding is performed on the infrared and visible light images respectively to obtain a first initial multimodal feature and a second initial multimodal feature; Multimodal semantic fusion features are obtained based on the first initial multimodal features and the second initial multimodal features; The multimodal semantic fusion features are input into a noise prediction diffusion model to output clarity-aware semantics, wherein the clarity-aware semantics fuses a clear latent representation of dual-modal information. The clarity-aware semantics are input into the clarity-aware visual-language model to output a high-level semantic embedding through the clarity-aware visual-language model; An initial fused image is generated based on the high-level semantic embedding, and a final fused image is generated based on the initial fused image.
[0006] According to one embodiment of this application, the step of performing heteroscale shallow coding on the infrared image and the visible light image respectively to obtain a first initial multimodal feature and a second initial multimodal feature includes: The infrared image is subjected to heteroscale shallow coding to obtain the first spatial-spectral initial features, and the first initial multimodal features are obtained based on the first spatial-spectral initial features; The visible light image is subjected to heteroscale shallow coding to obtain the second spatial-spectral initial features, and the second initial multimodal features are obtained based on the second spatial-spectral initial features.
[0007] According to one embodiment of this application, obtaining multimodal semantic fusion features based on the first initial multimodal features and the second initial multimodal features includes: Based on the preset CLIP model, the first initial multimodal features and the second initial multimodal features are respectively subjected to multi-scale semantic recoding to obtain the first semantic features and the second semantic features; The multimodal semantic fusion feature is obtained based on the first semantic feature and the second semantic feature.
[0008] According to one embodiment of this application, obtaining the multimodal semantic fusion feature based on the first semantic feature and the second semantic feature includes: The first semantic feature and the second semantic feature are weighted and concatenated to obtain the concatenated text features, and the concatenated text features are mapped to the text embedding space of CLIP to obtain the final text features; Based on a preset cross-modal attention mechanism, the final text features are used as a guide to perform semantic enhancement and alignment on visual features to obtain aligned multi-scale semantic features. The multimodal semantic fusion feature is obtained based on the aligned multi-scale semantic features.
[0009] According to one embodiment of this application, generating a final fused image from the initial fused image includes: Extract temporal-spatial consistency features from the initial fused image; The temporal-spatial consistency feature and its spatial coordinates are input into an implicit neural network to obtain the pixel value corresponding to the spatial coordinates. The final fused image is generated based on the pixel values corresponding to the spatial coordinates.
[0010] According to one embodiment of this application, the resolution of the infrared image is lower than that of the visible light image.
[0011] According to the semantically guided dual-light fusion super-resolution integrated method proposed in this application, infrared and visible light images are subjected to heteroscale shallow coding to obtain first and second initial multimodal features, and multimodal semantic fusion features are obtained. These multimodal semantic fusion features are input into a noise prediction diffusion model, which outputs sharpness-aware semantics. This sharpness-aware vision-language model is then input, outputting high-level semantic embeddings to generate an initial fused image, which is further used to generate a final fused image. This solves the semantic gap and resolution mismatch problems in heterogeneous image fusion in related technologies, avoiding issues such as blurred details, insufficient saliency of hot targets, semantic inconsistency, and low resolution in the fused image, thus achieving high-quality, highly semantically consistent integrated fusion and super-resolution reconstruction.
[0012] To achieve the above objectives, a second aspect of this application proposes a semantically guided dual-light fusion super-resolution integrated device, comprising: The acquisition module acquires infrared images and visible light images, and performs heteroscale shallow coding on the infrared images and the visible light images respectively to obtain first initial multimodal features and second initial multimodal features; The fusion module obtains multimodal semantic fusion features based on the first initial multimodal features and the second initial multimodal features; The first input module inputs the multimodal semantic fusion features into the noise prediction diffusion model to output clarity-aware semantics through the noise prediction diffusion model, wherein the clarity-aware semantics fuses a clear latent representation of dual-modal information; The second input module inputs the clarity-aware semantics into the clarity-aware visual-language model, so as to output high-level semantic embeddings through the clarity-aware visual-language model; The generation module generates an initial fused image based on the high-level semantic embedding, and generates a final fused image based on the initial fused image.
[0013] According to one embodiment of this application, the acquisition module is specifically used for: The infrared image is subjected to heteroscale shallow coding to obtain the first spatial-spectral initial features, and the first initial multimodal features are obtained based on the first spatial-spectral initial features; The visible light image is subjected to heteroscale shallow coding to obtain the second spatial-spectral initial features, and the second initial multimodal features are obtained based on the second spatial-spectral initial features.
[0014] According to one embodiment of this application, the fusion module is specifically used for: Based on the preset CLIP model, the first initial multimodal features and the second initial multimodal features are respectively subjected to multi-scale semantic recoding to obtain the first semantic features and the second semantic features; The multimodal semantic fusion feature is obtained based on the first semantic feature and the second semantic feature.
[0015] According to one embodiment of this application, the fusion module is specifically used for: The first semantic feature and the second semantic feature are weighted and concatenated to obtain the concatenated text features, and the concatenated text features are mapped to the text embedding space of CLIP to obtain the final text features; Based on a preset cross-modal attention mechanism, the final text features are used as a guide to perform semantic enhancement and alignment on visual features to obtain aligned multi-scale semantic features. The multimodal semantic fusion feature is obtained based on the aligned multi-scale semantic features.
[0016] According to one embodiment of this application, the generation module is specifically used for: Extract temporal-spatial consistency features from the initial fused image; The temporal-spatial consistency feature and its spatial coordinates are input into an implicit neural network to obtain the pixel value corresponding to the spatial coordinates. The final fused image is generated based on the pixel values corresponding to the spatial coordinates.
[0017] According to one embodiment of this application, the resolution of the infrared image is lower than that of the visible light image.
[0018] According to the semantically guided dual-light fusion super-resolution integrated device proposed in this application, infrared and visible light images are subjected to heteroscale shallow encoding to obtain first and second initial multimodal features, and multimodal semantic fusion features are obtained. These multimodal semantic fusion features are input into a noise prediction diffusion model, which outputs sharpness-aware semantics. This sharpness-aware vision-language model is then input, outputting high-level semantic embeddings to generate an initial fused image, and further, a final fused image is generated. This solves the semantic gap and resolution mismatch problems in heterogeneous image fusion in related technologies, avoiding issues such as blurred details, insufficient saliency of hot targets, semantic inconsistency, and low resolution in the fused image, achieving high-quality, highly semantically consistent integrated fusion and super-resolution reconstruction.
[0019] To achieve the above objectives, a third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the semantically guided dual-light fusion super-resolution integrated method as described in the above embodiments.
[0020] To achieve the above objectives, a fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement the semantically guided dual-light fusion super-resolution integrated method as described in the above embodiments.
[0021] To achieve the above objectives, a fifth aspect of this application provides a computer program product, which, when executed by a processor, implements the semantically guided dual-light fusion super-resolution integrated method as described in the above embodiments.
[0022] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0023] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a flowchart of a semantically guided dual-light fusion super-resolution integrated method according to an embodiment of this application; Figure 2 This is a flowchart of a semantically guided dual-light fusion super-resolution integrated method according to an embodiment of this application; Figure 3 This is a block diagram of a semantically guided dual-light fusion super-resolution integrated device according to an embodiment of this application; Figure 4This is a schematic diagram of the structure of an electronic device provided according to an embodiment of this application. Detailed Implementation
[0024] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0025] The semantically guided dual-light fusion super-resolution integrated method and apparatus proposed according to the embodiments of this application will be described below with reference to the accompanying drawings. First, the semantically guided dual-light fusion super-resolution integrated method proposed according to the embodiments of this application will be described with reference to the accompanying drawings.
[0026] Figure 1 This is a flowchart of a semantically guided dual-light fusion super-resolution integration method according to an embodiment of this application.
[0027] like Figure 1 As shown, this semantically guided dual-light fusion super-resolution integrated method includes the following steps: In step S101, infrared images and visible light images are acquired, and heteroscale shallow coding is performed on the infrared images and visible light images respectively to obtain the first initial multimodal features and the second initial multimodal features.
[0028] Optionally, in some embodiments, the resolution of the infrared image is lower than that of the visible light image.
[0029] Optionally, in some embodiments, different-scale shallow coding is performed on infrared images and visible light images respectively to obtain first initial multimodal features and second initial multimodal features, including: performing different-scale shallow coding on infrared images to obtain first spatial-spectral initial features, and obtaining first initial multimodal features based on the first spatial-spectral initial features; performing different-scale shallow coding on visible light images to obtain second spatial-spectral initial features, and obtaining second initial multimodal features based on the second spatial-spectral initial features.
[0030] Heterogeneous-scale shallow coding refers to the process of independently extracting and encoding features at different scales or adaptive scales in the shallow stages of a neural network for images of different modalities. The first initial multimodal feature refers to the initial feature representation extracted from the infrared image through heterogeneous-scale shallow coding. The second initial multimodal feature refers to the initial feature representation extracted from the visible light image through heterogeneous-scale shallow coding.
[0031] Specifically, in this embodiment, a low-resolution infrared image $I_{ir}^{lr} \in \mathbb{R}^{H \times W \times 1}$ and a high-resolution visible light image $I_{vis}^{hr} \in \mathbb{R}^{sH \times sW \times 3}$ (where $s$ is a super-resolution scaling factor) are acquired. Two independent shallow encoders $E_{ir}$ and $E_{vis}$ are used to perform lightweight shallow coding on the infrared and visible light images, respectively. The first spatial-spectral initial features corresponding to the infrared image and the second spatial-spectral initial features corresponding to the visible light image are extracted, thereby obtaining the first and second initial multimodal features. This completes cross-modal channel alignment and resolution matching, providing a feature base of a unified scale for subsequent joint processing. The shallow encoder consists of several convolutional layers used to capture low-level structural information such as edges and gradients, expressed as follows: ; ; in, This is the first initial multimodal feature. It is an infrared shallow encoder. For low-resolution infrared images, For a three-dimensional real number feature space, This is the second initial multimodal feature. Encoding for the shallow layer of visible light. For high-resolution visible light images, It is a three-dimensional real number feature space.
[0032] In step S102, multimodal semantic fusion features are obtained based on the first initial multimodal features and the second initial multimodal features.
[0033] Optionally, in some embodiments, obtaining multimodal semantic fusion features based on the first initial multimodal features and the second initial multimodal features includes: performing multi-scale semantic recoding on the first initial multimodal features and the second initial multimodal features based on a preset CLIP (Contrastive Language-Image Pre-training) model to obtain the first semantic features and the second semantic features respectively; and obtaining multimodal semantic fusion features based on the first semantic features and the second semantic features.
[0034] Optionally, in some embodiments, obtaining multimodal semantic fusion features based on the first semantic feature and the second semantic feature includes: weighting and concatenating the first semantic feature and the second semantic feature to obtain concatenated text features, and mapping the concatenated text features to the text embedding space of CLIP to obtain the final text features; based on a preset cross-modal attention mechanism, using the final text features as a guide, performing semantic enhancement and alignment on visual features to obtain aligned multi-scale semantic features; and obtaining multimodal semantic fusion features based on the aligned multi-scale semantic features.
[0035] The first semantic feature refers to the high-order semantic feature representation of the infrared modality obtained after multi-scale semantic recoding based on a preset CLIP model, using the first initial multimodal feature of the infrared image as input. The second semantic feature refers to the high-order semantic feature representation specific to the visible light modality obtained after multi-scale semantic recoding based on a preset CLIP model, using the second initial multimodal feature of the visible light image as input.
[0036] Specifically, this embodiment introduces a pre-defined CLIP model as a semantic encoder to perform global average pooling ($\text{GAP}(\cdot)$) on the first and second initial multimodal features to extract global visual features of the image. This embodiment employs two small, learnable multilayer perceptrons, $\text{MLP}{\theta{ir}}$ and $\text{MLP}{\theta{vis}}$ (MLP, Multi-Layer Perceptron), to map the global visual features into a learnable prompt vector, $P$. This embodiment concatenates the learnable prompt vector $P$ with a predefined category word (such as "infrared target"), a process represented as $\text{Concat}(P, \text{Category})$, thereby forming a complete text description, i.e., the concatenated text features. $E_{\text{text}}(\cdot)$ represents a pre-trained CLIP text encoder with frozen parameters. The CLIP text encoder maps the concatenated complete text description to the CLIP text embedding space to obtain the final text features $T$.
[0037] Furthermore, through a cross-modal attention mechanism, textual features are used as guidance to semantically enhance and align global visual features, resulting in multi-scale semantic features $T_{ir}$ and $T_{vis}$. The expression formula is as follows: ; ; ; ; in, These are infrared learnable cue word vectors. It is an infrared branch multilayer sensor. These are the learnable parameters for an infrared branched multilayer sensor. This is a global average pooling operation. For CLIP pre-trained text encoder, (·) indicates a splicing operation. For learnable cue word vectors in visible light, For visible light branched multilayer perceptron, These are the learnable parameters for the visible light branched multilayer perceptron.
[0038] In this embodiment, the aligned multi-scale semantic features are concatenated along the channel dimension to generate a multimodal semantic fusion feature, i.e., $F_{fus}^{sem}=\text{Concat}(F_{ir}^{clip}, F_{vis}^{clip})$.
[0039] Therefore, by using the pre-trained CLIP visual encoder as a semantic pillar, multi-scale semantic recoding is performed on the first and second initial multimodal features to obtain semantic tokens (minimum semantic units) rich in contextual information. Cross-modal attention is used to align and weightedly concatenate the semantic features of infrared and visible light, and output fused features that are semantically consistent and spatially complementary in both modes.
[0040] In step S103, the multimodal semantic fusion features are input into the noise prediction diffusion model to output the clarity-aware semantics, wherein the clarity-aware semantics fuses the clarity potential representation of the dual-modal information.
[0041] Specifically, in this embodiment, the multimodal semantic fusion feature $F_{fus}^{sem}$ obtained in step S102 is input into the noise prediction diffusion model. This model adopts a latent diffusion architecture and performs denoising in the latent space. To fully integrate bimodal information, the model incorporates a cross-modal attention layer to simultaneously accept conditional information from both infrared and visible light modes. The infrared conditional vector $C_{ir}$ and the visible light conditional vector $C_{vis}$ are obtained by projection operations on the infrared semantic feature $F_{ir}^{clip}$ and the visible light semantic feature $F_{vis}^{clip}$, respectively.
[0042] The estimation process for the noise latent representation $z_t$ in denoising step $t$ is as follows: ; in, For noise prediction, Let be the latent space representation at time t. For denoising UNet (U-Net Convolutional Neural Network). This is the infrared modal condition vector. This is the visible light mode condition vector. (·) represents cross-modal attention layer operations.
[0043] In the noise prediction diffusion model, $\mathcal{U}$ represents the denoising UNet, providing the basic computational architecture for noise estimation of the latent representation of noise in the latent space. $\theta$ is the set of learnable parameters for all convolutional layers, attention layers, and other modules in the denoising UNet network, and $\epsilon_\theta$ is the noise prediction value output by the denoising UNet with parameters $\theta$. The core learning objective of this denoising module is to predict noise under the joint constraints of dual-modal conditions composed of infrared and visible light conditional vectors, and to progressively remove noise to generate a clear latent representation that integrates dual-modal information, i.e., clarity-aware semantics.
[0044] Therefore, in this embodiment, the aligned fused features are fed into the latent diffusion noise prediction network, and a cross-modal cross attention layer is embedded in the U-shaped decoder. This supports multi-condition parallel input with fused features as spatial conditions and sharpness labels as modal conditions, so that each step of denoising is subject to both semantic and structural constraints.
[0045] In step S104, the clarity-aware semantics are input into the clarity-aware visual-language model to output high-level semantic embeddings through the clarity-aware visual-language model.
[0046] Specifically, to enhance the semantic consistency and discriminative details of the generated images, embodiments of this application introduce a sharpness-aware vision-language model (such as a CLIP-based image-text contrastive learning model) to extract discriminative semantic embeddings $E_{disc}$ from a high-resolution visible light image $I_{vis}^{hr}$. These semantic embeddings are then injected into multiple levels of the diffusion decoder through another cross-attention layer. ; in, This represents the latent space representation of the l-th layer of UNet at time t-1 in the diffusion model. Let t be the latent space representation of the l-th layer of UNet in the diffusion model at time t. This is for discriminative semantic embedding.
[0047] Here, $l$ represents the $l$th layer of UNet, which makes the denoising process not only guided by structural features, but also finely tuned by high-level semantic concepts.
[0048] Therefore, by constructing a clarity-aware visual-language model, the network adaptively judges the credibility of the current infrared and visible light modalities and generates corresponding high-level semantic embeddings. These embeddings are dynamically injected into the diffusion decoding path through a cross-attention layer, enabling the network to rely more on the semantic cues of the clear modal in ambiguous areas and suppressing error generation.
[0049] In step S105, an initial fused image is generated based on the high-level semantic embedding, and a final fused image is generated based on the initial fused image.
[0050] Optionally, in some embodiments, generating a final fused image based on an initial fused image includes: extracting temporal-spatial consistency features from the initial fused image; inputting the temporal-spatial consistency features and their spatial coordinates into an implicit neural network to obtain pixel values corresponding to the spatial coordinates; and generating the final fused image based on the pixel values corresponding to the spatial coordinates.
[0051] Among them, implicit neural networks are end-to-end neural network architectures that take continuous spatial coordinates as input and output corresponding pixel / feature values in the target space.
[0052] Specifically, guided by both semantic and structural information, a preliminary low-resolution fused image is generated through a $T$-step iterative denoising process, represented in the latent space as $z_0$: ; in, This represents the latent space representation of the initial low-resolution fused image. (·) represents the iterative denoising process. Let T be the latent space representation at time T. These are the fused features after semantic alignment.
[0053] Furthermore, in this embodiment of the application, the decoder $D$ in the diffusion model is used to decode the initial low-resolution fused image into an initial fused image, i.e., $I_{fus}^{lr} = D(z_0)$.
[0054] Thus, guided by both semantics (target category, edge location) and structure (gradient, texture details, etc.), the diffusion model gradually recovers the low-resolution fused image from pure noise. During the iteration process, semantic embedding adjusts the attention weights in real time to ensure that the edges of hot targets and visible light details are enhanced simultaneously, avoiding artifact accumulation.
[0055] Furthermore, in this embodiment, an Implicit Neural Representation (INR) network $M_\phi$ is constructed. Its input consists of temporal-spatial consistency features extracted from the initial fused image $I_{fus}^{lr}$ and their spatial coordinates $(x, y)$. The output is the RGB (Red, Green, Blue) pixel value at those coordinates. ; in, Let (x, y) be the high-resolution fused image pixel value at coordinates (x, y). For learnable parameters The implicit neural expression network, This is a temporal-spatial consistency feature.
[0056] The implicit neural representation network $M_\phi$ is typically composed of multilayer perceptrons. By learning a mapping function from continuous coordinates to pixel values, it can reconstruct images at arbitrary resolutions. Finally, by querying all high-resolution coordinate points, it generates the final high-resolution fused image $I_{fus}^{hr}$.
[0057] Therefore, by continuously upsampling the implicit neural representation network and using the diffused output features as input to the network, the feature vector is continuously decoded into high-density pixel values through coordinate mapping, achieving end-to-end upsampling at any multiplier of ×2, ×4, or ×8. This process does not require explicit interpolation, maintains consistency between edge sharpness and thermal radiation intensity, and ultimately outputs a high-resolution, high-fidelity dual-light fusion image.
[0058] Furthermore, the overall model in this application adopts a strategy combining phased training and end-to-end fine-tuning to ensure collaborative optimization and stable convergence of each module. Specifically, dataset construction and preprocessing are performed, using publicly available dual-light image datasets (such as LLVIP (A Visible-infrared Paired Dataset for Low-light Vision), FLIR (FLIR Thermal Dataset), etc.) and self-acquired registered infrared-visible image pairs as training samples. Each data pair includes: a low-resolution infrared image (obtained by downsampling real high-resolution infrared images) and a high-resolution visible light image, with a corresponding high-resolution real fused image (generated as a supervision signal through weighted averaging or existing fusion algorithms). All images undergo data augmentation operations such as normalization, random cropping, and horizontal flipping to improve the model's generalization ability.
[0059] Furthermore, the training is divided into three stages: pre-training of the shallow encoder and semantic alignment module, using mean squared error (MSE) and perceptual loss to jointly optimize the shallow encoder and CLIP projection layer, with the goal of minimizing the difference between reconstructed features and true features; joint training of the diffusion model and the implicit neural expression (INR) network, fixing the CLIP semantic encoder parameters, and training the denoising UNet and implicit neural expression network in the latent diffusion model using a multi-task loss function: ; in, For pixel-level L1 loss; For feature matching loss in VGG-19 (Visual Geometry Group-19, a 19-layer deep convolutional neural network proposed by the Visual Geometry Group at Oxford University); To combat the loss, a PatchGAN (Patch-based Generative Adversarial Network Discriminator) discriminator is used. For semantic consistency loss, the image and text embeddings are aligned using CLIP space pre-trained in the first stage.
[0060] Furthermore, end-to-end fine-tuning is performed by unfreezing all module parameters and fine-tuning them at a small learning rate to optimize the overall quality metrics between the output and the real fused image (such as PSNR (Peak Signal-to-Noise Ratio, core localization: image quality assessment) and SSIM (Structural Similarity Index)).
[0061] Therefore, introducing semantic guidance to provide the necessary high-level semantic priors and perception capabilities can enhance semantic targets and protect target thermal response contours while suppressing background redundancy. This can effectively solve the semantic gap and resolution mismatch problems in heterogeneous image fusion, avoid problems such as blurred details, insufficient saliency of thermal targets, semantic inconsistency and low resolution in fused images, and achieve high-quality, highly semantically consistent integrated fusion and super-resolution reconstruction.
[0062] To facilitate a deeper understanding of the semantically guided dual-light fusion super-resolution integrated method proposed in this application, the following section combines... Figure 2 Further explanation is needed.
[0063] like Figure 2 As shown, Figure 2This is a flowchart of a semantically guided dual-light fusion super-resolution integration method according to an embodiment of this application. The semantically guided dual-light fusion super-resolution integration method includes the following steps: S201 performs heteroscale shallow feature encoding on the input low-resolution infrared image and high-resolution visible light image to extract initial multimodal features.
[0064] S202 introduces a pre-trained CLIP model as a semantic encoder to perform multi-scale semantic enhancement and cross-modal feature alignment on the initial features, and then generates dual-modal semantic fusion features after concatenation.
[0065] S203 inputs the semantically aligned fused features into the noise prediction diffusion model and achieves noise estimation under multiple constraints by introducing a cross-modal cross-attention layer in the latent diffusion architecture.
[0066] S204 proposes a clarity-aware visual-language model to extract discriminative semantic embeddings and injects them into the diffusion decoding process through a cross-attention mechanism to enhance semantic consistency.
[0067] S205, guided by both semantic and structural information, generates a preliminary low-resolution fused image through iterative denoising and encodes it into a latent space representation.
[0068] S206 constructs an implicit neural representation network, extracts consistency features, maps feature coordinates to pixel values, and achieves continuous upsampling fusion reconstruction.
[0069] According to the semantically guided dual-light fusion super-resolution integrated method proposed in this application, infrared and visible light images are subjected to heteroscale shallow coding to obtain first and second initial multimodal features, and multimodal semantic fusion features are obtained. These multimodal semantic fusion features are input into a noise prediction diffusion model, which outputs sharpness-aware semantics. This sharpness-aware vision-language model is then input, outputting high-level semantic embeddings to generate an initial fused image, which is further used to generate a final fused image. This solves the semantic gap and resolution mismatch problems in heterogeneous image fusion in related technologies, avoiding issues such as blurred details, insufficient saliency of hot targets, semantic inconsistency, and low resolution in the fused image, thus achieving high-quality, highly semantically consistent integrated fusion and super-resolution reconstruction.
[0070] Next, referring to the accompanying drawings, a semantically guided dual-light fusion super-resolution integrated device based on an embodiment of this application is described.
[0071] Figure 3 This is a block diagram of a semantically guided dual-light fusion super-resolution integrated device according to an embodiment of this application.
[0072] like Figure 3 As shown, the semantically guided dual-light fusion super-resolution integrated device 10 includes: an acquisition module 100, a fusion module 200, a first input module 300, a second input module 400, and a generation module 500.
[0073] The acquisition module 100 acquires infrared images and visible light images, and performs heteroscale shallow coding on the infrared images and visible light images respectively to obtain the first initial multimodal features and the second initial multimodal features; The fusion module 200 obtains multimodal semantic fusion features based on the first initial multimodal features and the second initial multimodal features; The first input module 300 inputs the multimodal semantic fusion features into the noise prediction diffusion model to output the clarity-aware semantics through the noise prediction diffusion model, wherein the clarity-aware semantics fuses the clarity latent representation of dual-modal information; The second input module 400 inputs the clarity-aware semantics into the clarity-aware visual-language model so as to output the high-level semantic embedding through the clarity-aware visual-language model; The generation module 500 generates an initial fused image based on high-level semantic embeddings and generates a final fused image based on the initial fused image.
[0074] According to one embodiment of this application, the acquisition module 100 is specifically used for: The infrared image is encoded using a different scale shallow layer to obtain the first spatial-spectral initial features, and the first initial multimodal features are obtained based on the first spatial-spectral initial features; The visible light image is encoded using a different scale shallow layer to obtain the second spatial-spectral initial features, and the second initial multimodal features are obtained based on the second spatial-spectral initial features.
[0075] According to one embodiment of this application, the fusion module 200 is specifically used for: Based on the preset CLIP model, the first initial multimodal features and the second initial multimodal features are respectively subjected to multi-scale semantic recoding to obtain the first semantic features and the second semantic features; Multimodal semantic fusion features are obtained based on the first and second semantic features.
[0076] According to one embodiment of this application, the fusion module 200 is specifically used for: The first semantic feature and the second semantic feature are weighted and concatenated to obtain the concatenated text feature, and the concatenated text feature is mapped to the text embedding space of CLIP to obtain the final text feature; Based on a pre-defined cross-modal attention mechanism, the final text features are used as a guide to perform semantic enhancement and alignment on visual features to obtain aligned multi-scale semantic features. Multimodal semantic fusion features are obtained based on the aligned multi-scale semantic features.
[0077] According to one embodiment of this application, the generation module 500 is specifically used for: Extract temporal-spatial consistency features from the initial fused image; The temporal-spatial consistency features and their spatial coordinates are input into an implicit neural network to obtain the pixel values corresponding to the spatial coordinates. The final fused image is generated based on the pixel values corresponding to the spatial coordinates.
[0078] According to one embodiment of this application, the resolution of an infrared image is lower than that of a visible light image.
[0079] It should be noted that the foregoing explanation of the embodiment of the semantically guided dual-light fusion super-resolution integrated method also applies to the semantically guided dual-light fusion super-resolution integrated device of this embodiment, and will not be repeated here.
[0080] According to the semantically guided dual-light fusion super-resolution integrated device proposed in this application, infrared and visible light images are subjected to heteroscale shallow encoding to obtain first and second initial multimodal features, and multimodal semantic fusion features are obtained. These multimodal semantic fusion features are input into a noise prediction diffusion model, which outputs sharpness-aware semantics. This sharpness-aware vision-language model is then input, outputting high-level semantic embeddings to generate an initial fused image, and further, a final fused image is generated. This solves the semantic gap and resolution mismatch problems in heterogeneous image fusion in related technologies, avoiding issues such as blurred details, insufficient saliency of hot targets, semantic inconsistency, and low resolution in the fused image, achieving high-quality, highly semantically consistent integrated fusion and super-resolution reconstruction.
[0081] Figure 4 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. The electronic device may include: The memory 401, the processor 402, and the computer program stored on the memory 401 and capable of running on the processor 402.
[0082] When the processor 402 executes the program, it implements the semantically guided dual-light fusion super-resolution integrated method provided in the above embodiments.
[0083] Furthermore, electronic devices also include: Communication interface 403 is used for communication between memory 401 and processor 402.
[0084] The memory 401 is used to store computer programs that can run on the processor 402.
[0085] The memory 401 may include high-speed RAM (Random Access Memory) memory, and may also include non-volatile memory, such as at least one disk storage.
[0086] If the memory 401, processor 402, and communication interface 403 are implemented independently, then the communication interface 403, memory 401, and processor 402 can be interconnected via a bus to complete communication between them. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0087] Optionally, in a specific implementation, if the memory 401, processor 402, and communication interface 403 are integrated on a single chip, then the memory 401, processor 402, and communication interface 403 can communicate with each other through an internal interface.
[0088] Processor 402 may be a CPU (Central Processing Unit), an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement embodiments of the present invention.
[0089] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the semantically guided dual-light fusion super-resolution integrated method described above.
[0090] This application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described semantically guided dual-light fusion super-resolution integrated method embodiments.
[0091] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0092] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0093] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. A semantically guided dual-light fusion super-resolution integrated method, characterized in that, include: Infrared and visible light images are acquired, and heteroscale shallow coding is performed on the infrared and visible light images respectively to obtain a first initial multimodal feature and a second initial multimodal feature; Multimodal semantic fusion features are obtained based on the first initial multimodal features and the second initial multimodal features; The multimodal semantic fusion features are input into a noise prediction diffusion model to output clarity-aware semantics, wherein the clarity-aware semantics fuses a clear latent representation of dual-modal information. The clarity-aware semantics are input into the clarity-aware visual-language model to output a high-level semantic embedding through the clarity-aware visual-language model; An initial fused image is generated based on the high-level semantic embedding, and a final fused image is generated based on the initial fused image.
2. The method according to claim 1, characterized in that, The step of performing heteroscale shallow coding on the infrared image and the visible light image respectively to obtain the first initial multimodal feature and the second initial multimodal feature includes: The infrared image is subjected to heteroscale shallow coding to obtain the first spatial-spectral initial features, and the first initial multimodal features are obtained based on the first spatial-spectral initial features; The visible light image is subjected to heteroscale shallow coding to obtain the second spatial-spectral initial features, and the second initial multimodal features are obtained based on the second spatial-spectral initial features.
3. The method according to claim 1, characterized in that, The process of obtaining multimodal semantic fusion features based on the first initial multimodal features and the second initial multimodal features includes: Based on the preset CLIP model, the first initial multimodal features and the second initial multimodal features are respectively subjected to multi-scale semantic recoding to obtain the first semantic features and the second semantic features; The multimodal semantic fusion feature is obtained based on the first semantic feature and the second semantic feature.
4. The method according to claim 3, characterized in that, The process of obtaining the multimodal semantic fusion feature based on the first semantic feature and the second semantic feature includes: The first semantic feature and the second semantic feature are weighted and concatenated to obtain the concatenated text feature, and the concatenated text feature is mapped to the text embedding space of CLIP to obtain the final text feature; Based on a preset cross-modal attention mechanism, the final text features are used as a guide to perform semantic enhancement and alignment on visual features to obtain aligned multi-scale semantic features. The multimodal semantic fusion feature is obtained based on the aligned multi-scale semantic features.
5. The method according to claim 1, characterized in that, The step of generating the final fused image based on the initial fused image includes: Extract temporal-spatial consistency features from the initial fused image; The temporal-spatial consistency feature and its spatial coordinates are input into an implicit neural network to obtain the pixel value corresponding to the spatial coordinates. The final fused image is generated based on the pixel values corresponding to the spatial coordinates.
6. The method according to any one of claims 1-5, characterized in that, The resolution of the infrared image is lower than that of the visible light image.
7. A semantically guided dual-light fusion super-resolution integrated device, characterized in that, include: The acquisition module acquires infrared images and visible light images, and performs heteroscale shallow coding on the infrared images and the visible light images respectively to obtain first initial multimodal features and second initial multimodal features; The fusion module obtains multimodal semantic fusion features based on the first initial multimodal features and the second initial multimodal features; The first input module inputs the multimodal semantic fusion features into the noise prediction diffusion model to output clarity-aware semantics through the noise prediction diffusion model, wherein the clarity-aware semantics fuses a clear latent representation of dual-modal information; The second input module inputs the clarity-aware semantics into the clarity-aware visual-language model, so as to output high-level semantic embeddings through the clarity-aware visual-language model; The generation module generates an initial fused image based on the high-level semantic embedding, and generates a final fused image based on the initial fused image.
8. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the semantically guided dual-light fusion super-resolution integrated method as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the semantically guided dual-light fusion super-resolution integrated method as described in any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the semantically guided dual-light fusion super-resolution integrated method as described in any one of claims 1-6.