Degradation and semantic prior dual guided infrared and visible image fusion method and device
By employing a dual-guided approach of degradation and semantic priors for infrared and visible light image fusion, this method utilizes degradation and semantic priors embedded in a network model and combines a diffusion model to recover high-quality semantic priors in a compact latent space. This addresses the issues of high computational cost and generalization difficulties in existing methods under complex degradation scenarios, achieving efficient information aggregation and degradation suppression, and improving the robustness and efficiency of the fusion algorithm.
Patent Information
- Application Number
- CN202510122263.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-01-26
AI Technical Summary
Existing infrared and visible light image fusion methods are computationally intensive and time-consuming when dealing with complex degradation scenes, are difficult to generalize to other degradation types, and are sensitive to additional semantic information, making it difficult to cope with situations where infrared and visible light images degrade simultaneously.
We adopt a dual-guided approach of degradation and semantic prior. We extract degradation type and global scene context through degradation prior embedding network and semantic prior embedding network, and use enhancement and fusion network model for feature extraction and fusion. Combined with diffusion model, we recover high-quality semantic prior in compact latent space to achieve degradation suppression and information aggregation.
It improves the applicability and robustness of the fusion algorithm in complex real-world scenarios, increases computational efficiency by more than 200 times, generates high-quality fused images, and adapts to various degradation scenarios without the need for additional auxiliary information.
Smart Images

Figure CN120070201B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision image processing technology, and in particular to a method, apparatus, storage medium and electronic device for infrared and visible light image fusion guided by both degradation and semantic prior. Background Technology
[0002] Image fusion is an important enhancement technique that aims to integrate complementary information from multiple images to overcome the limitations of single-modality or single-type sensors. Infrared and visible light image fusion is a popular research area in image fusion, effectively aggregating the significant thermal radiation information in infrared images with the rich texture in visible light images to achieve a comprehensive representation of the imaging scene. The complete information integration and pleasing visual results make infrared and visible light image fusion widely applicable in fields such as military detection, security monitoring, driver assistance, target detection, and semantic segmentation.
[0003] Image fusion has gained widespread attention in recent years, and related algorithms have developed rapidly. These algorithms can be classified according to network architecture, including methods based on convolutional neural networks, autoencoders, generative adversarial networks, Transformers, and diffusion models. From a functional perspective, these algorithms can also be divided into vision-oriented fusion methods, degradation-aware fusion methods, semantically driven fusion methods, and joint registration and fusion methods. Although these methods have achieved satisfactory fusion performance, some challenges remain. On the one hand, while diffusion models with strong generative capabilities can bring performance gains, diffusion-based fusion methods are often computationally intensive and time-consuming. On the other hand, although some degradation-aware methods have been proposed to address interference problems in the imaging process, they still perform poorly in complex fusion scenarios. For example, DIVFusion and PAIF, while specifically designed for certain degradations (such as low light or noise), are difficult to generalize to other degradation types. In addition, there are some general degradation-aware methods that, with the assistance of additional semantic information (such as text cues), can handle multiple degradations within a unified framework. However, such methods are highly sensitive to text prompts and struggle to handle situations where both infrared and visible light images degrade simultaneously. Furthermore, developing tailored text descriptions for each fusion scenario is extremely time-consuming and labor-intensive. Summary of the Invention
[0004] This application provides a method, apparatus, storage medium, and electronic device for infrared and visible light image fusion guided by both degradation and semantic priors. It achieves effective degradation suppression and efficient information aggregation, greatly improving the applicability and robustness of the fusion algorithm in complex real-world scenarios.
[0005] This application provides a method for fusing infrared and visible light images guided by both degradation and semantic prior knowledge, including:
[0006] Acquire degraded infrared images, degraded visible light images, high-quality infrared images, and high-quality visible light images;
[0007] The degraded infrared image and the degraded visible light image are input into the degradation prior embedding network model to obtain the degradation prior;
[0008] The high-quality infrared image and the high-quality visible light image are input into the semantic prior embedding network model to obtain the high-quality semantic prior;
[0009] The degraded infrared image, the degraded visible light image, the degradation prior, and the high-quality semantic prior are input into the degradation and semantic prior-guided enhancement and fusion network model to obtain the fused image.
[0010] Furthermore, in the aforementioned degradation and semantic prior-guided infrared and visible light image fusion method, the degradation and semantic prior-guided enhancement and fusion network model includes a prior-adjusted coding layer, a prior-guided fusion module, a prior-modulated decoding layer, and a fusion reconstruction head connected in sequence.
[0011] The process involves inputting the degraded infrared image, the degraded visible light image, the degradation prior, and the high-quality semantic prior into a degradation and semantic prior-guided enhancement and fusion network model to obtain a fused image, including:
[0012] The degraded infrared image, the degraded visible light image, the degraded prior, and the high-quality semantic prior are input into the coding layer of the prior modulation for feature extraction and feature enhancement, resulting in enhanced infrared feature maps and enhanced visible light feature maps.
[0013] The enhanced infrared feature map and the enhanced visible light feature map are input into the prior-guided fusion module for fusion to obtain a fused feature map;
[0014] The fused feature map is input into the decoding layer of the prior modulation for feature enhancement to obtain an enhanced fused map;
[0015] The enhanced fusion map is input into the fusion reconstruction head to obtain the fused image.
[0016] Furthermore, in the aforementioned infrared and visible light image fusion method guided by both degradation and semantic priors, the step of inputting the degraded infrared image, the degraded visible light image, the degradation prior, and the high-quality semantic prior into the coding layer of the prior modulation for feature extraction and feature enhancement to obtain enhanced infrared feature maps and enhanced visible light feature maps includes:
[0017] The degraded prior is input into the coding layer of the prior modulation and integrated with the degraded infrared image and the degraded visible light image respectively to obtain the first infrared feature map and the first visible light feature map;
[0018] The high-quality semantic prior is input into the coding layer of the prior modulation and integrated with the degraded infrared image and the degraded visible light image respectively to obtain the second infrared feature map and the second visible light feature map;
[0019] The first infrared feature map and the second infrared feature map are combined to obtain an enhanced infrared feature map;
[0020] The first visible light feature map and the second visible light feature map are aggregated to obtain an enhanced visible light feature map.
[0021] Furthermore, in the aforementioned infrared and visible light image fusion method guided by both degradation and semantic prior, the step of inputting the degradation prior into the coding layer of the prior modulation and integrating it with the degraded infrared image and the degraded visible light image respectively to obtain a first infrared feature map and a first visible light feature map includes:
[0022] The degraded infrared image is compressed into a high-dimensional vector with the same shape as the degraded prior.
[0023] The high-dimensional vector is multiplied by the degenerate prior, and the result of the multiplication is passed through a linear layer to obtain the first modulation parameter and the second modulation parameter.
[0024] A first infrared feature map is obtained based on the first modulation parameter, the second modulation parameter, and the degraded infrared image;
[0025] The degraded visible light image is compressed into a high-dimensional vector with the same shape as the degraded prior.
[0026] The high-dimensional vector is multiplied by the degenerate prior, and the result of the multiplication is passed through a linear layer to obtain the first modulation parameter and the second modulation parameter.
[0027] A first visible light feature map is obtained based on the first modulation parameter, the second modulation parameter, and the degraded visible light image.
[0028] Furthermore, in the aforementioned infrared and visible light image fusion method guided by both degradation and semantic prior, the step of inputting the high-quality semantic prior into the coding layer of the prior modulation and integrating it with the degraded infrared image and the degraded visible light image respectively to obtain a second infrared feature map and a second visible light feature map includes:
[0029] The degraded infrared image is mapped to a first query vector, the degraded infrared image is mapped to a second query vector, and the high-quality semantic prior is mapped to a key and a value.
[0030] A second infrared feature map is obtained based on the degraded infrared image, the key, and the value;
[0031] A second visible light feature map is obtained based on the degraded visible light image, the key, and the value.
[0032] Furthermore, in the aforementioned degradation and semantic prior-guided infrared and visible light image fusion method, the step of inputting the enhanced infrared feature map and the enhanced visible light feature map into the prior-guided fusion module for fusion to obtain a fused feature map includes:
[0033] Channel-level fusion weights are obtained based on the aforementioned semantic priors;
[0034] Spatial weights are generated based on the enhanced infrared feature map and the enhanced visible light feature map, respectively.
[0035] The enhanced infrared feature map and the enhanced visible light feature map are weighted and fused based on the channel-level fusion weight and the spatial weight to obtain a fused feature map.
[0036] Furthermore, in the aforementioned infrared and visible light image fusion method guided by both degradation and semantic prior, the method further includes:
[0037] The high-quality infrared image and the high-quality visible light image are embedded into the high-quality semantic prior. This embedding process is denoted as x0, where x0 serves as the starting point of a forward Markov chain, and is iterated through T iterations.
[0038]
[0039] Where, x t β is the noise variable at step t. t Controlling the variance of noise, α t =1-β t ;
[0040] Through reparameterization and iterative derivation, the forward Markov process can be re-expressed as:
[0041]
[0042] in, When t approaches a large value T q(x) approaches 0 T |x0) approximates a normal distribution The forward diffusion process has ended.
[0043] Furthermore, in the aforementioned infrared and visible light image fusion method guided by both degradation and semantic prior, the method further includes:
[0044] A high-quality semantic prior is generated by progressively denoising the low-quality semantic prior using a T-step Markov chain. The T-step Markov chain is defined as follows:
[0045]
[0046] in, ∈ represents noise;
[0047] Utilizing reparameterization techniques, and replacing ∈ with We can obtain:
[0048]
[0049] in, As a low-quality semantic prior, This represents high-quality semantic priors extracted from high-quality visible light images. This represents high-quality semantic priors extracted from high-quality infrared images.
[0050] This application also provides an infrared and visible light image fusion device guided by both degradation and semantic prior knowledge, including:
[0051] The acquisition module is used to acquire degraded infrared images, degraded visible light images, high-quality infrared images, and high-quality visible light images;
[0052] The degradation prior generation module is used to input the degradation infrared image and the degradation visible light image into the degradation prior embedding network model to obtain the degradation prior;
[0053] The semantic prior generation module is used to input the high-quality infrared image and the high-quality visible light image into the semantic prior embedding network model to obtain high-quality semantic priors;
[0054] The fusion module is used to input the degraded infrared image, the degraded visible light image, the degradation prior, and the high-quality semantic prior into the degradation and semantic prior-guided enhancement and fusion network model to obtain the fused image.
[0055] This application also provides a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute any of the above-described degradation and semantic prior guided infrared and visible light image fusion methods.
[0056] This application also provides an electronic device, including a processor and a memory, wherein the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used in the steps of the degradation and semantic prior dual-guided infrared and visible light image fusion method described in any of the above claims.
[0057] This application provides a method, apparatus, storage medium, and electronic device for infrared and visible light image fusion guided by both degradation and semantic priors. It proposes a fusion framework based on dual modulation of degradation and semantic priors, where the degradation prior represents the degradation type of the source image, and the semantic prior represents the global scene context. These two complement each other, achieving effective degradation suppression and efficient aggregation of complementary information through an enhancement and fusion network model guided by degradation and semantic priors. Secondly, the infrared and visible light image fusion method guided by both degradation and semantic priors can adaptively perceive the degradation type from the source image without additional auxiliary information, and recovers high-quality semantic priors in a compact high-dimensional latent space by leveraging the generative capabilities of a diffusion model. Guided by degradation type priors and high-quality semantic priors, the enhancement and fusion network generates a high-quality fused image, achieving effective degradation suppression and efficient information aggregation, greatly improving the applicability and robustness of the fusion algorithm in complex real-world scenarios. Finally, existing fusion methods based on diffusion models perform the diffusion process in the image domain, resulting in a heavy computational burden and difficulty in meeting the real-time requirements of image fusion tasks. This invention proposes an efficient semantic prior diffusion model that can recover high-quality semantic priors in a compact latent space to help enhance and fuse networks suppress degenerate interference and aggregate complementary information. Its operating efficiency is more than 200 times higher than that of mainstream diffusion model fusion methods. Attached Figure Description
[0058] The technical solution and other beneficial effects of this application will become apparent from the following detailed description of specific embodiments in conjunction with the accompanying drawings.
[0059] Figure 1 A flowchart of the infrared and visible light image fusion method guided by both degradation and semantic prior knowledge provided in the embodiments of this application.
[0060] Figure 2 Another flowchart of the infrared and visible light image fusion method guided by both degradation and semantic priors provided in the embodiments of this application.
[0061] Figure 3 This is a flowchart illustrating the processing of the degenerate prior embedding network and the semantic prior embedding network provided in the embodiments of this application.
[0062] Figure 4 This is a schematic diagram of the structure of the degradation and semantic prior guidance enhancement and fusion network model provided in the embodiments of this application.
[0063] Figure 5 The flowchart for obtaining the first infrared feature map and the first visible light feature map is provided for the embodiments of this application.
[0064] Figure 6 The flowchart for obtaining the second infrared feature map and the second visible light feature map is provided for the embodiments of this application.
[0065] Figure 7 This is a flowchart illustrating the process of obtaining the fused feature map, as provided in an embodiment of this application.
[0066] Figure 8 A schematic diagram of the semantic diffusion model provided in this application.
[0067] Figure 9 This is a schematic diagram comparing the method of the present invention with advanced image fusion methods in the same degradation scenario provided in the embodiments of this application.
[0068] Figure 10 This diagram illustrates a comparison between the method of the present invention and advanced image enhancement + image fusion methods under different degradation scenarios provided in the embodiments of this application.
[0069] Figure 11 This is a schematic diagram of the structure of the infrared and visible light image fusion device with dual guidance of degradation and semantic prior provided in the embodiments of this application.
[0070] Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0071] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0072] This application provides a method, apparatus, storage medium, and electronic device for infrared and visible light image fusion guided by both degradation and semantic priors. The infrared and visible light image fusion apparatus provided in this application can be integrated into an electronic device, such as a terminal or server. The terminal can include tablet computers, laptops, personal computers (PCs), microprocessor boxes, or other devices.
[0073] Please see Figure 1 and Figure 2 , Figure 1 The flowchart illustrates the infrared and visible light image fusion method guided by both degradation and semantic prior knowledge, as provided in the embodiments of this application. Figure 2 Another flowchart of the degradation and semantic prior guided infrared and visible light image fusion method provided in this application embodiment, which is applied in an electronic device, includes the following steps:
[0074] S1, acquire degraded infrared image, degraded visible light image, high-quality infrared image and high-quality visible light image.
[0075] The infrared image and the visible light image are images of the same scene.
[0076] S2, input the degraded infrared image and the degraded visible light image into the degradation prior embedding network model to obtain the degradation prior.
[0077] Figure 3 The flowcharts for the degenerate prior embedding network and semantic prior embedding network provided in the embodiments of this application are as follows: Figure 3 As shown, the degraded infrared image and degraded visible light images The input is fed into the degenerate prior embedding network model, and the degenerate prior embedding network is utilized. Extract modality-specific degenerate priors Degradation priors of visible light images This can be expressed by the following formula:
[0078]
[0079] in, and C d This indicates the number of embedded tokens and the channel dimension. and Extracted from degraded visible light and infrared images, respectively, it is possible to characterize different degradation types in different modes.
[0080] S3 inputs high-quality infrared and high-quality visible light images into the semantic prior embedding network model to obtain high-quality semantic priors.
[0081] High-quality infrared images and high-quality visible light images The input is fed into the semantic prior embedding network model, and the semantic prior embedding network is used. Extracting high-quality semantic priors p from cascaded high-quality infrared and visible light images s (In another embodiment, a semantic prior p recovered from a low-quality semantic prior can be used.) s ′ Replacing high-quality semantics (described in detail later) can be represented by the following formula:
[0082]
[0083] in, The scene context used to represent the global context. For semantic prior embedding network models, since N s Much smaller than H×W, thus enabling a higher compression ratio. This can effectively reduce the computational burden of subsequent semantic prior diffusion models.
[0084] S4 inputs the degraded infrared image, the degraded visible light image, the degradation prior, and the high-quality semantic prior into the enhancement and fusion network model guided by the degradation and semantic prior, and obtains the fused image.
[0085] Figure 4 This is a schematic diagram of the structure of the degradation and semantic prior guidance enhancement and fusion network model provided in the embodiments of this application, as shown below. Figure 4 As shown, the degradation and semantic prior-guided enhancement and fusion network model includes a prior-adjusted coding layer, a prior-guided fusion module, a prior-modulated decoding layer, and a fusion reconstruction head connected in sequence. In one embodiment, step S4 includes the following steps:
[0086] S41, the degraded infrared image, degraded visible light image, degraded prior, and high-quality semantic prior are input into the coding layer of prior modulation for feature extraction and feature enhancement, resulting in enhanced infrared feature map and enhanced visible light feature map.
[0087] Specifically, firstly, in the feature extraction stage, two parallel prior-modulated coding branches are used to extract multi-scale infrared and visible light features respectively. Each coding branch has k layers, and the feature extraction process of the k-th layer can be represented as:
[0088]
[0089] Among them, E k This represents the coding layer with prior modulation at layer k. This represents the infrared or visible light characteristics of the output of the k-th layer. This represents the infrared or visible light characteristics output by the (k-1)th layer. p represents a high-quality degradation prior. s This indicates semantic prior.
[0090] In one embodiment, step S41 includes the following steps:
[0091] S411, the degraded prior input is integrated with the prior modulation coding layer and the degraded infrared image and the degraded visible light image respectively to obtain the first infrared feature map and the first visible light feature map.
[0092] Specifically, step S411 includes:
[0093] S4111 compresses the degraded infrared image into a high-dimensional vector with the same shape as the degraded prior.
[0094] S4112, multiply the high-dimensional vector with the degenerate prior, and pass the multiplication result through a linear layer to obtain the first modulation parameter and the second modulation parameter;
[0095] S4113, a first infrared feature map is obtained based on the first modulation parameter, the second modulation parameter and the degraded infrared image;
[0096] S4114 compresses the degraded visible light image into a high-dimensional vector with the same shape as the degraded prior;
[0097] S4115, multiply the high-dimensional vector with the degenerate prior, and pass the multiplication result through a linear layer to obtain the first modulation parameter and the second modulation parameter;
[0098] S4116, a first visible light feature map is obtained based on the first modulation parameter, the second modulation parameter and the degraded visible light image.
[0099] Figure 5 The flowchart for obtaining the first infrared feature map and the first visible light feature map provided in the embodiments of this application is as follows: Figure 5 As shown, the input features Compressed into a degenerate prior High-dimensional vectors with the same shape, then with The product is multiplied, and the result is passed through a linear layer to output the modulation parameters. and This process can be represented as:
[0100]
[0101] in, This represents the first infrared feature map and the first visible light feature map.
[0102] It should be noted that in the first layer of the priori modulation coding layer, This refers to degraded infrared and visible light images, in other layers of the prior modulation coding layer. This refers to the features output from the previous layer (including features of visible light images and infrared images).
[0103] S412, high-quality semantic prior input is fed into the coding layer of prior modulation and integrated with the degraded infrared image and the degraded visible light image respectively to obtain the second infrared feature map and the second visible light feature map.
[0104] Specifically, S412 includes:
[0105] S4121, map the degraded infrared image to the first query vector, map the degraded infrared image to the second query vector, and map the high-quality semantic prior to the key and value;
[0106] S4122, a second infrared feature map is obtained based on the degraded infrared image, key and value;
[0107] S4123, a second visible light feature map is obtained based on the degraded visible light image, the key and the value.
[0108] Figure 6 The flowchart for obtaining the second infrared feature map and the second visible light feature map provided in the embodiments of this application is as follows: Figure 6 As shown, high-quality semantic priors are integrated into In order to enhance its global awareness of high-quality scene context. Mapped to query High-quality semantic prior p s Mapped to key Sum Then, a cross-attention mechanism is applied to implement semantic prior embedding. This process can be represented as:
[0109]
[0110] Where, d k It is a learnable scaling factor. This represents the second infrared feature map and the second visible light feature map.
[0111] It should be noted that in the first layer of the priori modulation coding layer, This refers to degraded infrared and visible light images, in other layers of the prior modulation coding layer. This refers to the features output from the previous layer (including features of visible light images and infrared images).
[0112] S413, the first infrared feature map and the second infrared feature map are aggregated to obtain an enhanced infrared feature map.
[0113] S414, the first visible light feature map and the second visible light feature map are aggregated to obtain an enhanced visible light feature map.
[0114] Specifically, it utilizes cross-attention mechanisms to aggregate... and This generates the final enhanced features (including a strong infrared feature map and an enhanced visible light feature map):
[0115]
[0116] Q spi from Obtained by mapping, and K dpm and V dpm Then by The mapping yields the result. Ultimately, p s Provides global semantic guidance, p d Clearly characterizing the degradation type can effectively reduce the overall training difficulty of enhancement and fusion tasks.
[0117] S42, the enhanced infrared feature map and the enhanced visible light feature map are input into the prior-guided fusion module for fusion to obtain the fused feature map.
[0118] Figure 7 The flowchart for obtaining the fused feature map provided in the embodiments of this application is as follows: Figure 7 As shown, step S42 includes the following steps:
[0119] S421, infrared fusion weights and visible light fusion weights are obtained based on semantic priors;
[0120] S422, based on the enhanced infrared feature map and the enhanced visible light feature map, respectively, spatial weights are generated;
[0121] S423, the enhanced infrared feature map and the enhanced visible light feature map are weighted and fused based on channel-level fusion weights and spatial weights to obtain a fused feature map.
[0122] Considering p s It is jointly extracted from multimodal input, implicitly aggregating comprehensive and high-quality scene information. This invention utilizes semantic channel attention to generate channel-level fusion weights. and On the other hand, this invention also employs spatial attention to perform spatial activity level measurement. Infrared or visible light features are compressed using global max pooling and global average pooling. The pooling results are then concatenated along the channel dimension and fed into a convolutional layer to generate spatial weights. and Finally, by integrating the channel and spatial weights, the fusion weights for the k-th layer are obtained:
[0123]
[0124] in, This represents element-wise multiplication with a broadcast mechanism, where σ is the Sigmoid function.
[0125] The final fusion process is defined as:
[0126]
[0127] in, This represents the enhanced infrared feature map output of the k-th layer of the coding layer with prior modulation. This represents the enhanced visible light feature map of the output of the k-th layer of the coded layer that has been modulated a priori.
[0128] S43, the fused feature map is input into the decoding layer of the prior modulation for feature enhancement, resulting in an enhanced fused map.
[0129] Specifically, the decoding layer of prior modulation utilizes high-quality semantic priors to enhance features.
[0130] S44. Input the enhanced fusion map into the fusion reconstruction head to obtain the fused image.
[0131] Specifically, the fusion reconstruction head will enhance the fusion characteristics. Mapping to the image domain to generate a fused image I f .
[0132] The above describes the application process of the degenerate prior embedding network model, the semantic prior embedding network model, and the degenerate and semantic prior guided enhancement and fusion network model. The above steps use the trained degenerate prior embedding network model, the semantic prior embedding network model, and the semantic prior guided enhancement and fusion network model. The following describes the calculation process of the loss function of the degenerate prior embedding network model, the semantic prior embedding network model, and the semantic prior guided enhancement and fusion network model:
[0133] Since semantic and degenerate priors cannot be constrained by ground truth, this invention uses fusion and contrastive losses to jointly optimize... (Semantic prior-guided augmented and fusion network model) (Semantic prior embedding network model) and (Degenerate prior embedding network model). The fusion loss includes content loss, structural similarity loss, and color consistency loss. To eliminate the interference of degradation factors, high-quality source images acquired manually are used to construct the above losses. Content loss is defined as:
[0134]
[0135] in, Let represent the Sobel operator, max(·) represent the maximum value selection strategy for preserving salient objects and textures, and |·|1 and γ represent the l1 norm and the hyperparameters used for balancing, respectively. The structural similarity loss is used to maintain the structural similarity between the fused image and the high-quality source image, and is specifically defined as:
[0136]
[0137] Here, SSIM(·,·) measures the structural similarity between two images. Furthermore, a color consistency loss is constructed to ensure that the fused image fully preserves the color information from the high-quality visible light image, and its definition is as follows:
[0138]
[0139] Where, Φ CbCr (·) is a color mapping function that converts the RGB color space to the CbCr color space. The degradation prior embedding module aims to adaptively identify various degradation types; therefore, for different degradation inputs, the corresponding p... d There should be a distinction, even if their image content is the same. To this end, this invention designs a contrastive loss, thereby bringing degradation priors representing the same degradation type closer together and pushing back degradation priors representing different degradation types. For degradation prior p... d , and Let represent the corresponding positive and negative samples, respectively. The definition of contrastive loss is as follows:
[0140]
[0141] Where K and M represent the number of positive and negative samples, respectively, and τ is the temperature parameter. Specifically, if p d Extract from images with specific degradation, then Extracted from other scenes with the same degradation, and This involves extracting images from the same scene but with different degradation or modalities. Finally, step 1 is used for constraint. and The total loss is defined as the weighted sum of content loss, structural similarity loss, color consistency loss, and contrast loss:
[0142]
[0143] Where, λ cont , λ ssim , λ color and λ cl It is a hyperparameter that controls the balance of various losses.
[0144] Furthermore, this method also includes:
[0145] Constructing a semantic prior diffusion model: High-quality semantic priors are recovered by establishing a diffusion model in a compact latent space, which effectively guides the enhancement and fusion process. Figure 8 A schematic diagram of the semantic diffusion model provided in this application, such as Figure 8 As shown, the semantic prior diffusion model includes a forward diffusion process and a backward denoising process, with the specific steps as follows:
[0146] S51, Forward Diffusion Process: First, cascaded high-quality infrared and visible light images are... and Embedded into high-quality semantic prior p s In this step, it is simplified and represented as x0. x0 serves as the starting point of the forward Markov chain, and Gaussian noise is gradually added through T iterations. The specific process of T is as follows:
[0147]
[0148] Where, x t β is the noise variable at step t. t Controlling the variance of noise, α t =1-β t Through reparameterization and iterative derivation, the forward Markov process can be re-expressed as:
[0149]
[0150] in, When t approaches a large value T q(x) approaches 0 T |x0) approximates a normal distribution The forward diffusion process has ended.
[0151] S52, Reverse Denoising Process: The reverse denoising process starts with a pure Gaussian distribution and uses a T-step Markov chain to progressively denoise the low-quality semantic prior to generate a high-quality semantic prior, as defined below:
[0152]
[0153]
[0154] in, ∈ represents noise.
[0155] This invention designs a denoising U-Net network, which leverages semantic priors extracted from low-quality images. and degenerate priors as well as To estimate the noise ∈. Utilizing a reparameterization technique, and replacing ∈ with We can obtain:
[0156]
[0157] Due to latent semantic space The distribution ratio of the image space (R) H×W×3 It's simpler, semantic prior (p) s ′ This can be generated with fewer iterations. Therefore, this invention runs a complete T (<<1000)-step iterative reverse process to infer p. s ′ Therefore, the present invention uses The semantic prior diffusion model is trained using a combination of loss functions: content loss, structural similarity loss, and color consistency loss. Furthermore, content loss, structural similarity loss, and color consistency loss are used to jointly constrain the training of the semantic prior diffusion model. Therefore, the total loss of the semantic prior diffusion model is defined as:
[0158]
[0159] Where, λ diff These are the hyperparameters used to balance various losses. After training, this step recovers the high-quality semantic prior p. ′ ′ Used to replace the high-quality semantic prior p in step S2 s To guide the enhancement and integration process.
[0160] Figure 9 This diagram illustrates a comparison between the method of the present invention and advanced image fusion methods in the same degradation scenario provided in an embodiment of this application. Figure 10 This diagram illustrates a comparison between the method of this invention and advanced image enhancement + image fusion methods under different degradation scenarios provided in embodiments of this application. Compared to existing fusion methods, this invention achieves adaptive perception of degradation types, efficient recovery of scene semantic priors, and an organic combination of image enhancement and fusion, without requiring additional auxiliary information. Even in harsh, highly interfering environments, it can generate high-quality fusion results. To objectively evaluate the performance of the method of this invention, various degradation scenarios were selected for assessment. For example... Figure 8 As shown, when existing fusion algorithms are directly applied to degraded source images, the fused image is significantly affected by interference factors. Furthermore, as... Figure 9 As shown, when a pre-enhancement algorithm is used to process the source image, it can only utilize the context within a single modality to enhance the source image. Although it can reduce the influence of interfering factors, the fusion algorithm still cannot achieve satisfactory visual results. In contrast, this invention does not rely on additional pre-enhancement algorithms and auxiliary information. Its multimodal fusion results are not affected by degradation and can provide satisfactory visual perception, greatly improving the applicability and robustness of image fusion tasks in real-world complex environments.
[0161] Based on the method described in the above embodiments, this embodiment will further describe the infrared and visible light image fusion device guided by both degradation and semantic prior. The infrared and visible light image fusion device guided by both degradation and semantic prior can be implemented as an independent entity or integrated into an electronic device. The electronic device can be a terminal, server, or other devices. The terminal can include a tablet computer, a laptop computer, a personal computer (PC), a microprocessor box, or other devices.
[0162] Please see Figure 11 , Figure 11 This application provides a specific description of a degradation- and semantic prior-guided infrared and visible light image fusion apparatus, which is applied in electronic devices. This degradation- and semantic prior-guided infrared and visible light image fusion apparatus may include:
[0163] The acquisition module is used to acquire degraded infrared images, degraded visible light images, high-quality infrared images, and high-quality visible light images;
[0164] The degradation prior generation module is used to input the degradation infrared image and the degradation visible light image into the degradation prior embedding network model to obtain the degradation prior;
[0165] The semantic prior generation module is used to input the high-quality infrared image and the high-quality visible light image into the semantic prior embedding network model to obtain high-quality semantic priors;
[0166] The fusion module is used to input the degraded infrared image, the degraded visible light image, the degradation prior, and the high-quality semantic prior into the degradation and semantic prior-guided enhancement and fusion network model to obtain the fused image.
[0167] In specific implementation, the above modules and / or units can be implemented as independent entities, or they can be arbitrarily combined and implemented as the same or several entities. For the specific implementation of the above modules and / or units, please refer to the previous method embodiments. For the specific beneficial effects that can be achieved, please also refer to the beneficial effects in the previous method embodiments, which will not be repeated here.
[0168] In addition, this application also provides an electronic device, which may be a computer, tablet computer, or other similar device. This electronic device can implement the steps of any embodiment of the degradation and semantic prior guided infrared and visible light image fusion method provided in this application. Therefore, it can achieve the beneficial effects that any degradation and semantic prior guided infrared and visible light image fusion method provided in this invention can achieve, as detailed in the preceding embodiments, and will not be repeated here.
[0169] Figure 12 A specific structural block diagram of an electronic device provided in an embodiment of the present invention is shown. This electronic device can be used to implement the infrared and visible light image fusion method guided by both degradation and semantic prior knowledge provided in the above embodiments. The electronic device 500 can be a terminal, server, or other device. The terminal can include a tablet computer, laptop computer, personal computer (PC), microprocessor box, or other devices.
[0170] RF circuit 510 is used to receive and transmit electromagnetic waves, converting electromagnetic waves into electrical signals and vice versa, thereby enabling communication with communication networks or other devices. RF circuit 510 may include various existing circuit elements used to perform these functions, such as antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, Subscriber Identity Module (SIM) cards, memory, etc. RF circuit 510 can communicate with various networks such as the Internet, corporate intranets, and wireless networks, or communicate with other devices via wireless networks. The aforementioned wireless networks may include cellular telephone networks, wireless local area networks (WLANs), or metropolitan area networks (MANs). The aforementioned wireless networks may use various communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging, and short messages, and any other suitable communication protocols, including those that have not yet been developed.
[0171] The memory 520 can be used to store software programs and modules, such as the program instructions / modules corresponding to those in the above embodiments. The processor 580 executes various functional applications and data processing by running the software programs and modules stored in the memory 520, such as taking pictures with the front-facing camera, processing the captured images, and switching the display colors of the content displayed on the screen. The memory 520 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 520 may further include memory remotely located relative to the processor 580, and these remote memories can be connected to the electronic device 500 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0172] The input unit 530 can be used to receive input numeric or character information, and to generate a keyboard and mouse related to user settings and function control.
[0173] Display unit 540 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces, which can be composed of graphics, text, icons, video, and any combination thereof. Display unit 540 may include display panel 541, which may optionally be configured in the form of LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), or other similar forms.
[0174] Audio circuitry 560, speaker 561, and microphone 562 provide an audio interface between the user and electronic device 500. Audio circuitry 560 converts received audio data into electrical signals and transmits them to speaker 561, where speaker 561 converts them into sound signals for output. Conversely, microphone 562 converts collected sound signals into electrical signals, which are then received by audio circuitry 560, converted back into audio data, and processed by processor 580. The audio data is then transmitted via RF circuitry 510 to, for example, another terminal, or output to memory 520 for further processing. Audio circuitry 560 may also include an earphone jack to facilitate communication between external headphones and electronic device 500.
[0175] Electronic device 500, through transmission module 570 (e.g., Wi-Fi module), can help users receive requests, send information, etc., providing users with wireless broadband internet access. Although transmission module 570 is shown in the figure, it is understood that it is not an essential component of electronic device 500 and can be omitted as needed without changing the essence of the invention.
[0176] The processor 580 is the control center of the electronic device 500. It connects to various parts of the phone via various interfaces and lines, and performs various functions and processes data of the electronic device 500 by running or executing software programs and / or modules stored in the memory 520, and by calling data stored in the memory 520, thereby providing overall monitoring of the electronic device. Optionally, the processor 580 may include one or more processing cores; in some embodiments, the processor 580 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 580.
[0177] Electronic device 500 also includes a power supply 590 (such as a battery) for supplying power to various components. In some embodiments, the power supply may be logically connected to processor 580 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. The power supply 590 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0178] Although not shown, the electronic device 500 also includes cameras (such as front-facing cameras and rear-facing cameras), Bluetooth modules, etc., which will not be described in detail here. Specifically, in this embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal also includes a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors. One or more programs contain instructions for performing the following operations:
[0179] Acquire degraded infrared images, degraded visible light images, high-quality infrared images, and high-quality visible light images;
[0180] The degraded infrared image and the degraded visible light image are input into the degradation prior embedding network model to obtain the degradation prior;
[0181] The high-quality infrared image and the high-quality visible light image are input into the semantic prior embedding network model to obtain the high-quality semantic prior;
[0182] The degraded infrared image, the degraded visible light image, the degradation prior, and the high-quality semantic prior are input into the degradation and semantic prior-guided enhancement and fusion network model to obtain the fused image.
[0183] In practice, the above modules can be implemented as independent entities or combined in any way to be implemented as the same or several entities. For the specific implementation of the above modules, please refer to the previous method implementation examples, which will not be repeated here.
[0184] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. Therefore, embodiments of the present invention provide a storage medium storing multiple instructions that can be loaded by a processor to execute the steps of any embodiment of the degradation and semantic prior guided infrared and visible light image fusion method provided by the present invention.
[0185] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0186] Since the instructions stored in the storage medium can execute the steps in any embodiment of the degradation and semantic prior dual-guided infrared and visible light image fusion method provided in the embodiments of the present invention, the beneficial effects that the degradation and semantic prior dual-guided infrared and visible light image fusion method provided in the embodiments of the present invention can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.
[0187] The foregoing has provided a detailed description of a degradation- and semantic prior-guided infrared and visible light image fusion method, apparatus, storage medium, and electronic device provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for fusing infrared and visible light images guided by both degradation and semantic priors, characterized in that, The method includes: Acquire degraded infrared images, degraded visible light images, high-quality infrared images, and high-quality visible light images; The degraded infrared image and the degraded visible light image are input into the degradation prior embedding network model to obtain the degradation prior; The high-quality infrared image and the high-quality visible light image are input into the semantic prior embedding network model to obtain the high-quality semantic prior; The degraded infrared image, the degraded visible light image, the degradation prior, and the high-quality semantic prior are input into the degradation and semantic prior-guided enhancement and fusion network model to obtain the fused image. The degradation and semantic prior-guided enhancement and fusion network model includes a prior-modulated coding layer, a prior-guided fusion module, a prior-modulated decoding layer, and a fusion reconstruction head connected in sequence. The processing steps of the degradation and semantic prior-guided enhancement and fusion network model include: inputting the degraded infrared image, the degraded visible light image, the degradation prior, and the high-quality semantic prior into the prior-modulated coding layer for feature extraction and enhancement, resulting in enhanced infrared feature maps and enhanced visible light feature maps; inputting the enhanced infrared feature maps and the enhanced visible light feature maps into the prior-guided fusion module for fusion, resulting in a fused feature map; inputting the fused feature map into the prior-modulated decoding layer for feature enhancement, resulting in an enhanced fused map; and inputting the enhanced fused map into the fusion reconstruction head to obtain a fused image. The steps of obtaining the enhanced infrared feature map and the enhanced visible light feature map include: inputting the degraded prior into the prior-modulated coding layer and integrating it with the degraded infrared image and the degraded visible light image respectively to obtain a first infrared feature map and a first visible light feature map; inputting the high-quality semantic prior into the prior-modulated coding layer and integrating it with the degraded infrared image and the degraded visible light image respectively to obtain a second infrared feature map and a second visible light feature map; aggregating the first infrared feature map and the second infrared feature map to obtain an enhanced infrared feature map; and aggregating the first visible light feature map and the second visible light feature map to obtain an enhanced visible light feature map.
2. The infrared and visible light image fusion method guided by both degradation and semantic prior as described in claim 1, characterized in that, The step of integrating the degraded prior input into the coding layer of the prior modulation with the degraded infrared image and the degraded visible light image respectively to obtain a first infrared feature map and a first visible light feature map includes: The degraded infrared image is compressed into a high-dimensional vector with the same shape as the degraded prior. The high-dimensional vector is multiplied by the degenerate prior, and the result of the multiplication is passed through a linear layer to obtain the first modulation parameter and the second modulation parameter. A first infrared feature map is obtained based on the first modulation parameter, the second modulation parameter, and the degraded infrared image; The degraded visible light image is compressed into a high-dimensional vector with the same shape as the degraded prior. The high-dimensional vector is multiplied by the degenerate prior, and the result of the multiplication is passed through a linear layer to obtain the first modulation parameter and the second modulation parameter. A first visible light feature map is obtained based on the first modulation parameter, the second modulation parameter, and the degraded visible light image.
3. The infrared and visible light image fusion method guided by both degradation and semantic prior as described in claim 1, characterized in that, The step of inputting the high-quality semantic prior into the coding layer of the prior modulation and integrating it with the degraded infrared image and the degraded visible light image respectively to obtain the second infrared feature map and the second visible light feature map includes: The degraded infrared image is mapped to a first query vector, the degraded visible light image is mapped to a second query vector, and the high-quality semantic prior is mapped to a key and a value. A second infrared feature map is obtained based on the degraded infrared image, the first query vector, the key, and the value; A second visible light feature map is obtained based on the degraded visible light image, the second query vector, the key, and the value.
4. The infrared and visible light image fusion method guided by both degradation and semantic prior as described in claim 1, characterized in that, The step of inputting the enhanced infrared feature map and the enhanced visible light feature map into the prior-guided fusion module for fusion to obtain a fused feature map includes: Channel-level fusion weights are obtained based on the aforementioned semantic priors; Spatial weights are generated based on the enhanced infrared feature map and the enhanced visible light feature map, respectively. The enhanced infrared feature map and the enhanced visible light feature map are weighted and fused based on the channel-level fusion weight and the spatial weight to obtain a fused feature map.
5. The infrared and visible light image fusion method guided by both degradation and semantic prior as described in claim 1, characterized in that, The method further includes: The high-quality infrared image and the high-quality visible light image are embedded into the high-quality semantic prior. This embedding process is represented as follows: , As the starting point of a positive Markov chain, through Gaussian noise is gradually added in each iteration, as expressed by the following formula: in, It is the first noise variables of the step, Controlling the variance of noise, ; Through reparameterization and iterative derivation, the forward Markov process can be re-expressed as: in, ,when Approaching a larger value hour, Approaching 0, Approximately normal distribution The forward diffusion process has ended.
6. A degradation- and semantic prior-guided infrared and visible light image fusion apparatus, wherein the degradation- and semantic prior-guided infrared and visible light image fusion apparatus is used to implement the degradation- and semantic prior-guided infrared and visible light image fusion method of claim 1, characterized in that, include: The acquisition module is used to acquire degraded infrared images, degraded visible light images, high-quality infrared images, and high-quality visible light images; The degradation prior generation module is used to input the degradation infrared image and the degradation visible light image into the degradation prior embedding network model to obtain the degradation prior; The semantic prior generation module is used to input the high-quality infrared image and the high-quality visible light image into the semantic prior embedding network model to obtain high-quality semantic priors; The fusion module is used to input the degraded infrared image, the degraded visible light image, the degradation prior, and the high-quality semantic prior into the degradation and semantic prior-guided enhancement and fusion network model to obtain the fused image.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted to be loaded by a processor to execute the degradation and semantic prior dual-guided infrared and visible light image fusion method according to any one of claims 1 to 5.