Degradation and semantic prior dual-guided infrared and visible light image fusion method and device
Through the dual guidance of degeneration and semantic priors, the degeneration priors and high-quality semantic priors enhancement and fusion network is used to solve the problem of poor processing of existing technology in complex scenarios, and efficient infrared and visible light images fusion and degradation suppression are achieved, which improves the applicability and robustness of the algorithm.
Patent Information
- Application Number
- CN202510122263.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-26
AI Technical Summary
Existing infrared and visible image fusion methods do not perform well when dealing with complex scenarios, especially in the case of multiple degradation types, making it difficult to achieve effective degradation suppression and information aggregation.
The dual guidance method of degeneration and semantic prior is adopted to enhance and fusion networks through degeneration prior and high-quality semantic prior guidance to achieve efficient fusion of infrared and visible light images. The method includes acquiring degraded and high-quality images, extracting degraded and semantic priors, and inputting them into an enhancement and fusion network for fusion.
It significantly improves the applicability and robustness of the image fusion algorithm in complex real scenarios, realizes effective degradation suppression and efficient information aggregation, and has an operating efficiency of 200 times higher than that of traditional methods.
Smart Images

Figure CN120070201A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of computer vision image processing, and in particular, to an infrared and visible light image fusion method, device, storage medium, and electronic device that are dual-guided by degradation and semantic priors. Background Art
[0002] Image fusion is an important enhancement technology aimed at integrating complementary information in multiple images to overcome the limitations of single-modal or single-type sensors. Infrared and visible light image fusion is a popular research field in image fusion, which can effectively aggregate the significant thermal radiation information in infrared images and the rich textures in visible light images to achieve a comprehensive characterization of the imaging scene. Complete information integration and pleasing visual results make infrared and visible light image fusion widely used in military detection, security monitoring, assisted driving, target detection, semantic segmentation, and other fields.
[0003] In recent years, image fusion has received extensive attention, and related algorithms have developed rapidly. These algorithms can be classified according to the network architecture, including methods based on convolutional neural networks, autoencoders, generative adversarial networks, Transformers, and diffusion models. From a functional perspective, these algorithms can also be divided into vision-oriented fusion methods, degradation-aware fusion methods, semantic-driven fusion methods, and methods for joint registration and fusion. Although these methods have achieved satisfactory fusion performance, there are still some challenges. On the one hand, although diffusion models with powerful generative capabilities can bring performance gains, fusion methods based on diffusion models are often computationally intensive and time-consuming. On the other hand, although some degradation-aware methods have been proposed to solve the interference problems in the imaging process, they still perform poorly in complex fusion scenarios. For example, DIVFusion and PAIF are designed specifically for specific degradations (such as low light or noise), but it is difficult to generalize them to other degradation types. In addition, there are some general degradation-aware methods that can handle multiple degradations in a unified framework with the assistance of additional semantic information (such as text prompts). However, such methods are very sensitive to text prompts and difficult to handle the situation where infrared and visible light images degrade simultaneously. Moreover, formulating targeted text descriptions for each fusion scenario is also very time-consuming and laborious. Summary of the Invention
[0004] Embodiments of this application provide an infrared and visible light image fusion method, device, storage medium, and electronic device that are dual-guided by degradation and semantic priors, which realizes effective degradation suppression and efficient information aggregation, and greatly improves the applicability and robustness of the fusion algorithm in complex real-world scenarios.
[0005] The embodiment of the present application provides an infrared and visible light image fusion method guided by degradation and semantic prior, including:
[0006] Obtain a degraded infrared image, a degraded visible light image, a high-quality infrared image, and a high-quality visible light image;
[0007] Input the degraded infrared image and the degraded visible light image into a degradation prior embedding network model to obtain a degradation prior;
[0008] Input the high-quality infrared image and the high-quality visible light image into a semantic prior embedding network model to obtain a high-quality semantic prior;
[0009] Input the degraded infrared image, the degraded visible light image, the degradation prior, and the high-quality semantic prior into a degradation and semantic prior guided enhancement and fusion network model to obtain a fused image.
[0010] Further, in the above infrared and visible light image fusion method guided by degradation and semantic prior, the degradation and semantic prior guided enhancement and fusion network model includes a coding layer with prior adjustment, a prior-guided fusion module, a decoding layer with prior modulation, and a fusion reconstruction head connected in sequence;
[0011] The step of inputting the degraded infrared image, the degraded visible light image, the degradation prior, and the high-quality semantic prior into a degradation and semantic prior guided enhancement and fusion network model to obtain a fused image includes:
[0012] Input the degraded infrared image, the degraded visible light image, the degradation prior, and the high-quality semantic prior into the prior-modulated coding layer for feature extraction and feature enhancement to obtain an enhanced infrared feature map and an enhanced visible light feature map;
[0013] Input the enhanced infrared feature map and the enhanced visible light feature map into the prior-guided fusion module for fusion to obtain a fused feature map;
[0014] Input the fused feature map into the prior-modulated decoding layer for feature enhancement to obtain an enhanced fused map;
[0015] Input the enhanced fused map into the fusion reconstruction head to obtain a fused image.
[0016] Further, in the above infrared and visible light image fusion method guided by both degradation and semantic prior, where the degraded infrared image, the degraded visible light image, the degradation prior, and the high-quality semantic prior are input into the prior-modulated encoding layer for feature extraction and feature enhancement to obtain an enhanced infrared feature map and an enhanced visible light feature map, including:
[0017] Input the degradation prior into the prior-modulated encoding layer and integrate it with the degraded infrared image and the degraded visible light image respectively to obtain a first infrared feature map and a first visible light feature map;
[0018] Input the high-quality semantic prior into the prior-modulated encoding layer and integrate it with the degraded infrared image and the degraded visible light image respectively to obtain a second infrared feature map and a second visible light feature map;
[0019] Aggregate the first infrared feature map and the second infrared feature map to obtain an enhanced infrared feature map;
[0020] Aggregate the first visible light feature map and the second visible light feature map to obtain an enhanced visible light feature map.
[0021] Further, in the above infrared and visible light image fusion method guided by both degradation and semantic prior, where the degradation prior is input into the prior-modulated encoding layer and integrated with the degraded infrared image and the degraded visible light image respectively to obtain a first infrared feature map and a first visible light feature map, including:
[0022] Compress the degraded infrared image into a high-dimensional vector with the same shape as the degradation prior;
[0023] Multiply the high-dimensional vector by the degradation prior, and pass the multiplication result through a linear layer to obtain a first modulation parameter and a second modulation parameter;
[0024] Obtain a first infrared feature map based on the first modulation parameter, the second modulation parameter, and the degraded infrared image;
[0025] Compress the degraded visible light image into a high-dimensional vector with the same shape as the degradation prior;
[0026] Multiply the high-dimensional vector by the degradation prior, and pass the multiplication result through a linear layer to obtain a first modulation parameter and a second modulation parameter;
[0027] Obtain a first visible light feature map based on the first modulation parameter, the second modulation parameter, and the degraded visible light image.
[0028] Further, in the above infrared and visible light image fusion method guided by both degradation and semantic prior, wherein integrating the high-quality semantic prior into the prior-modulated encoding layer and respectively integrating it with the degraded infrared image and the degraded visible light image to obtain a second infrared feature map and a second visible light feature map includes:
[0029] Mapping the degraded infrared image to a first query vector, mapping the degraded infrared image to a second query vector, and mapping the high-quality semantic prior to keys and values;
[0030] Obtaining a second infrared feature map based on the degraded infrared image, the keys, and the values;
[0031] Obtaining a second visible light feature map based on the degraded visible light image, the keys, and the values.
[0032] Further, in the above infrared and visible light image fusion method guided by both degradation and semantic prior, wherein inputting the enhanced infrared feature map and the enhanced visible light feature map into the prior-guided fusion module for fusion to obtain a fusion feature map includes:
[0033] Obtaining channel-level fusion weights based on the semantic prior;
[0034] Generating respective spatial weights based on the enhanced infrared feature map and the enhanced visible light feature map respectively;
[0035] Performing weighted fusion on the enhanced infrared feature map and the enhanced visible light feature map based on the channel-level fusion weights and the spatial weights to obtain a fusion feature map.
[0036] Further, in the above infrared and visible light image fusion method guided by both degradation and semantic prior, wherein the method further includes:
[0037] Embedding the high-quality infrared image and the high-quality visible light image into the high-quality semantic prior, and representing the above embedding process as x 0 , x 0 As the starting point of a forward Markov chain, through T iterations
[0038]
[0039] wherein, x t is the noise variable at the t-th step, β t controls the variance of the noise, α t = 1 - β t ;
[0040] Through reparameterization and iterative derivation, the forward Markov process can be re-expressed as:
[0041]
[0042] Among them, When t approaches a relatively large value T, tends to 0, and q(x T |x 0 ) approaches a normal distribution The forward diffusion process ends.
[0043] Furthermore, for the above-mentioned infrared and visible light image fusion method guided by degradation and semantic prior, among which, the method further includes:
[0044] Gradually denoise the low-quality semantic prior through a T-step Markov chain to generate a high-quality semantic prior. The definition of the T-step Markov chain is as follows:
[0045]
[0046] Among them, ∈ is noise;
[0047] Using the reparameterization trick and replacing ∈ with We can obtain:
[0048]
[0049] Among them, is the low-quality semantic prior, represents the high-quality semantic prior extracted from the high-quality visible light image, represents the high-quality semantic prior extracted from the high-quality infrared image.
[0050] The embodiment of the present application also provides a device for infrared and visible light image fusion guided by degradation and semantic prior, including:
[0051] An acquisition module, configured to acquire a degraded infrared image, a degraded visible light image, a high-quality infrared image, and a high-quality visible light image;
[0052] A degradation prior generation module, configured to input the degraded infrared image and the degraded visible light image into a degradation prior embedding network model to obtain a degradation prior;
[0053] A semantic prior generation module, configured to input the high-quality infrared image and the high-quality visible light image into a semantic prior embedding network model to obtain a high-quality semantic prior;
[0054] A fusion module for inputting the degraded infrared image, the degraded visible light image, the degradation prior, and the high-quality semantic prior into a degradation and semantic prior-guided enhancement and fusion network model to obtain a fused image.
[0055] An embodiment of the present application also provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions are suitable for being loaded by a processor to execute any one of the above-mentioned infrared and visible light image fusion methods with dual guidance of degradation and semantic prior.
[0056] An embodiment of the present application also provides an electronic device, including a processor and a memory, the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used for the steps in any one of the above-mentioned infrared and visible light image fusion methods with dual guidance of degradation and semantic prior.
[0057] The infrared and visible light image fusion method, device, storage medium and electronic device provided by the present application propose a fusion framework based on dual modulation of degradation and semantic prior. Among them, the degradation prior represents the degradation type of the source image, and the semantic prior represents the global scene context. The two complement each other, and through the degradation and semantic prior-guided enhancement and fusion network model, effective degradation suppression and efficient complementary information aggregation are achieved. Secondly, the infrared and visible light image fusion method with dual guidance of degradation and semantic prior of the present invention can adaptively perceive the degradation type from the source image without additional auxiliary information, and with the help of the generation ability of the diffusion model, recover high-quality semantic prior in a compact high-dimensional latent space. Guided by the degradation type prior and the high-quality semantic prior, the enhancement and fusion network generates high-quality fused images, realizing effective degradation suppression and efficient information aggregation, and greatly improving the applicability and robustness of the fusion algorithm in complex real scenes. Finally, the existing fusion methods based on the diffusion model perform the diffusion process in the image domain, with a heavy computational burden and difficult to meet the real-time requirements of the image fusion task. The present invention proposes an efficient semantic prior diffusion model, which can recover high-quality semantic prior in a compact latent space to assist the enhancement and fusion network in suppressing degradation interference and aggregating complementary information, and its operating efficiency is more than 200 times higher than that of the mainstream diffusion model fusion method. Description of the Drawings
[0058] The following combines the drawings and details the specific implementation manners of the present application, and the technical solutions and other beneficial effects of the present application will be obvious.
[0059] Figure 1 It is a flowchart of the infrared and visible light image fusion method with dual guidance of degradation and semantic prior provided by an embodiment of the present application.
[0060] Figure 2 Another flowchart of the infrared and visible light image fusion method guided by degradation and semantic prior provided by the embodiment of the present application.
[0061] Figure 3 A processing flowchart of the degradation prior embedding network and the semantic prior embedding network provided by the embodiment of the present application.
[0062] Figure 4 A structural schematic diagram of the enhancement and fusion network model guided by degradation and semantic prior provided by the embodiment of the present application.
[0063] Figure 5 A flowchart of obtaining the first infrared feature map and the first visible light feature map provided by the embodiment of the present application.
[0064] Figure 6 A flowchart of obtaining the second infrared feature map and the second visible light feature map provided by the embodiment of the present application.
[0065] Figure 7 A flowchart of obtaining the fused feature map provided by the embodiment of the present application.
[0066] Figure 8 A schematic diagram of the semantic diffusion model provided by the present application.
[0067] Figure 9 A schematic diagram for comparing the method of the present invention with advanced image fusion methods under the same degradation scenario provided by the embodiment of the present application.
[0068] Figure 10 A schematic diagram for comparing the method of the present invention with advanced image enhancement + image fusion methods under different degradation scenarios provided by the embodiment of the present application.
[0069] Figure 11 A structural schematic diagram of the infrared and visible light image fusion device guided by degradation and semantic prior provided by the embodiment of the present application.
[0070] Figure 12 A structural schematic diagram of the electronic device provided by the embodiment of the present application. Detailed implementation manners
[0071] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0072] The embodiments of the present application provide a method, apparatus, storage medium, and electronic device for infrared and visible light image fusion guided by degradation and semantic prior. An apparatus for infrared and visible light image fusion guided by degradation and semantic prior provided by the embodiments of the present application can be integrated into an electronic device, which can be a device such as a terminal or a server. Among them, the terminal can include a tablet computer, a laptop computer, a personal computer (PC), a microprocessing box, or other devices, etc.
[0073] Please refer to Figure 1 and Figure 2 , Figure 1 which is a flowchart of the method for infrared and visible light image fusion guided by degradation and semantic prior provided by the embodiments of the present application. Figure 2 Another flowchart of the method for infrared and visible light image fusion guided by degradation and semantic prior provided by the embodiments of the present application, which is applied to an electronic device. The method for infrared and visible light image fusion guided by degradation and semantic prior includes the following steps:
[0074] S1. Obtain a degraded infrared image, a degraded visible light image, a high-quality infrared image, and a high-quality visible light image.
[0075] Among them, the infrared image and the visible light image are images of the same scene.
[0076] S2. Input the degraded infrared image and the degraded visible light image into the degradation prior embedding network model to obtain a degradation prior.
[0077] Figure 3 The processing flowchart of the degradation prior embedding network and the semantic prior embedding network provided by the embodiments of the present application is as shown in Figure 3 Input the degraded infrared image and the degraded visible light image into the degradation prior embedding network model, and use the degradation prior embedding network to extract the modality-specific degradation prior and the degradation prior of the visible light image which is represented by the following formula:
[0078]
[0079] Among them, and C d represent the number of embedding tokens and the channel dimension, and are respectively extracted from the degraded visible light and infrared images, and can represent different degradation types of different modalities.
[0080] S3. Input the high-quality infrared image and the high-quality visible light image into the semantic prior embedding network model to obtain a high-quality semantic prior.
[0081] Input the high-quality infrared image and the high-quality visible light image into the semantic prior embedding network model, and use the semantic prior embedding network to extract the high-quality semantic prior p from the cascaded high-quality infrared and visible light images s (In another embodiment, the semantic prior p recovered from the low-quality semantic prior can be used s ′ to replace the high-quality semantics, which will be specifically described later). This process can be represented by the following formula:
[0082]
[0083] where is used to characterize the global scene context. is the semantic prior embedding network model. Since N s is much smaller than H×W, a higher compression ratio can be obtained which can effectively reduce the computational burden of the subsequent semantic prior diffusion model.
[0084] S4. Input the degraded infrared image, the degraded visible light image, the degraded prior, and the high-quality semantic prior into the degradation and semantic prior-guided enhancement and fusion network model to obtain a fused image.
[0085] Figure 4 is the structural schematic diagram of the degradation and semantic prior-guided enhancement and fusion network model provided by the embodiment of the present application. As Figure 4 shown, the degradation and semantic prior-guided enhancement and fusion network model includes a sequentially connected prior-adjusted encoding layer, a prior-guided fusion module, a prior-modulated decoding layer, and a fusion reconstruction head. In one embodiment, step S4 includes the following steps:
[0086] S41. Input the degraded infrared image, the degraded visible light image, the degraded prior, and the high-quality semantic prior into the prior-modulated encoding layer for feature extraction and feature enhancement to obtain an enhanced infrared feature map and an enhanced visible light feature map.
[0087] Specifically, first, two parallel prior-modulated encoding branches are used in the feature extraction stage to separately extract multi-scale infrared features and visible light features. In each encoding branch, there are k layers, and the feature extraction process of the k-th layer can be represented as:
[0088]
[0089] Among them, E k represents the encoding layer of the k-th layer of prior modulation, represents the infrared feature or visible light feature output by the k-th layer, represents the infrared feature or visible light feature output by the (k - 1)-th layer, represents the high-quality degradation prior, p s represents the semantic prior.
[0090] In one embodiment, step S41 includes the following steps:
[0091] S411, input the degradation prior into the encoding layer of prior modulation and integrate it with the degraded infrared image and the degraded visible light image respectively to obtain the first infrared feature map and the first visible light feature map.
[0092] Specifically, step S411 includes:
[0093] S4111, compress the degraded infrared image into a high-dimensional vector with the same shape as the degradation prior;
[0094] S4112, multiply the high-dimensional vector by the degradation prior, and pass the multiplication result through a linear layer to obtain the first modulation parameter and the second modulation parameter;
[0095] S4113, obtain the first infrared feature map based on the first modulation parameter, the second modulation parameter and the degraded infrared image;
[0096] S4114, compress the degraded visible light image into a high-dimensional vector with the same shape as the degradation prior;
[0097] S4115, multiply the high-dimensional vector by the degradation prior, and pass the multiplication result through a linear layer to obtain the first modulation parameter and the second modulation parameter;
[0098] S4116, obtain the first visible light feature map based on the first modulation parameter, the second modulation parameter and the degraded visible light image.
[0099] Figure 5 is the flowchart for obtaining the first infrared feature map and the first visible light feature map provided by the embodiment of the present application. As Figure 5 shown, the input feature is compressed into a high-dimensional vector with the same shape as the degradation prior , and then multiplied by , and the obtained result passes through a linear layer to output the modulation parameters and This process can be expressed as:
[0100]
[0101] Among them, represents the first infrared feature map and the first visible light feature map.
[0102] It should be noted that in the first layer of the encoding layer of the prior modulation, refers to the degraded infrared image and the degraded visible light image. In other layers of the encoding layer of the prior modulation, refers to the features output by the previous layer (including the features of the visible light image and the infrared image).
[0103] S412. Input the high-quality semantic prior into the encoding layer of the prior modulation and integrate it with the degraded infrared image and the degraded visible light image respectively to obtain the second infrared feature map and the second visible light feature map.
[0104] Specifically, S412 includes:
[0105] S4121. Map the degraded infrared image to a first query vector, map the degraded infrared image to a second query vector, and map the high-quality semantic prior to keys and values;
[0106] S4122. Obtain the second infrared feature map based on the degraded infrared image, keys, and values;
[0107] S4123. Obtain the second visible light feature map based on the degraded visible light image, keys, and values.
[0108] Figure 6 is the flowchart for obtaining the second infrared feature map and the second visible light feature map provided by the embodiments of the present application. As Figure 6 shown, integrate the high-quality semantic prior into to enhance its global perception of the high-quality scene context. is mapped to a query while the high-quality semantic prior p s is mapped to keys and values Then, apply the cross-attention mechanism to achieve semantic prior embedding. This process can be expressed as:
[0109]
[0110] Among them, d k is a learnable scaling factor, represents the second infrared feature map and the second visible light feature map.
[0111] It should be noted that in the first layer of the encoding layer of the prior modulation, refers to the degraded infrared image and the degraded visible light image. In other layers of the encoding layer of the prior modulation, Refers to the features output by the previous layer (including the features of the visible light image and the features of the infrared image).
[0112] S413. Aggregate the first infrared feature map and the second infrared feature map to obtain an enhanced infrared feature map.
[0113] S414. Aggregate the first visible light feature map and the second visible light feature map to obtain an enhanced visible light feature map.
[0114] Specifically, use the cross-attention mechanism to aggregate and to generate the final enhanced features (including the enhanced infrared feature map and the enhanced visible light feature map):
[0115]
[0116] Q spi is obtained from by mapping, while K dpm and V dpm are obtained by mapping. Finally, p s provides global semantic guidance, and p d clearly characterizes the degradation type, which can effectively reduce the overall training difficulty of the enhancement and fusion tasks.
[0117] S42. Input the enhanced infrared feature map and the enhanced visible light feature map into the prior-guided fusion module for fusion to obtain a fused feature map.
[0118] Figure 7 This is the flowchart for obtaining the fused feature map provided by the embodiments of the present application. As Figure 7 shown, step S42 includes the following steps:
[0119] S421. Obtain the infrared fusion weight and the visible light fusion weight based on the semantic prior;
[0120] S422. Generate respective spatial weights based on the enhanced infrared feature map and the enhanced visible light feature map respectively;
[0121] S423. Perform weighted fusion on the enhanced infrared feature map and the enhanced visible light feature map based on the channel-level fusion weight and the spatial weight to obtain a fused feature map.
[0122] Considering that p s is jointly extracted from the multi-modal input and implicitly aggregates comprehensive and high-quality scene information, the present invention uses semantic channel attention to generate the channel-level fusion weight ( and ) On the other hand, the present invention also uses spatial attention to perform spatial activity level measurement. Infrared features or visible light features are compressed in information through global max pooling and global average pooling. Then, the pooling results are concatenated in the channel dimension and input into a convolutional layer to generate spatial weights ( and ). Finally, by integrating the channel and spatial weights, the fusion weight of the k-th layer is obtained:
[0123]
[0124] wherein, represents element-wise multiplication with a broadcast mechanism, and σ is the Sigmoid function.
[0125] The final fusion process is defined as:
[0126]
[0127] wherein, represents the enhanced infrared feature map of the output of the k-th layer of the encoding layer with prior modulation, represents the enhanced visible light feature map of the output of the k-th layer of the encoding layer with prior modulation.
[0128] S43. Input the fusion feature map into the decoding layer with prior modulation for feature enhancement to obtain an enhanced fusion map.
[0129] Specifically, the decoding layer with prior modulation uses high-quality semantic priors to strengthen features.
[0130] S44. Input the enhanced fusion map into the fusion reconstruction head to obtain a fused image.
[0131] Specifically, the fusion reconstruction head maps the enhanced fusion features to the image domain to generate the fused image I f .
[0132] The above is the application process of the degradation prior embedding network model, the semantic prior embedding network model, and the degradation and semantic prior-guided enhancement and fusion network model. The above steps use the trained degradation prior embedding network model, the semantic prior embedding network model, and the semantic prior-guided enhancement and fusion network model. The following introduces the calculation process of the loss functions of the degradation prior embedding network model, the semantic prior embedding network model, and the semantic prior-guided enhancement and fusion network model:
[0133] Since semantic and degradation priors cannot be constrained by ground truth, the present invention uses fusion and contrast losses to jointly optimize (the semantic prior-guided enhancement and fusion network model), (Semantic Prior Embedding Network Model) and (Degradation Prior Embedding Network Model). The fusion loss includes content loss, structural similarity loss, and color consistency loss. To eliminate the interference of degradation factors, high-quality source images obtained manually are used to construct the above losses. The content loss is defined as:
[0134]
[0135] where, represents the Sobel operator, max(·) represents the maximum selection strategy for retaining significant objects and textures, and |·| 1 and γ represent the l 1 norm and the hyperparameter for balancing, respectively. The structural similarity loss is used to maintain the structural similarity between the fused image and the high-quality source image, and is specifically defined as:
[0136]
[0137] where, SSIM(·,·) measures the structural similarity between two images. In addition, a color consistency loss is constructed to ensure that the fused image fully retains the color information in the high-quality visible light image, and its definition is as follows:
[0138]
[0139] where, Φ CbCr (·) is a color mapping function that converts the RGB color space to the CbCr color space. The degradation prior embedding module aims to adaptively identify various degradation types. Therefore, for different degraded inputs, the corresponding p d should be distinguished, even if their image contents are the same. To this end, the present invention designs a contrastive loss to bring closer the degradation priors representing the same degradation type and push away the degradation priors representing different degradation types. For the degradation prior p d , and represent the corresponding positive sample and negative sample, respectively. The definition of the contrastive loss is as follows:
[0140]
[0141] where, K and M represent the number of positive samples and negative samples, respectively, and τ is the temperature parameter. Specifically, if p d is extracted from an image with a specific degradation, then is extracted from other scenes with the same degradation, while is extracted from images in the same scene but with different degradations or modalities. Finally, step 1 is used to constrain and The total loss is defined as the weighted sum of the content loss, structural similarity loss, color consistency loss, and contrast loss:
[0142]
[0143] where λ cont 、λ ssim 、λ color and λ cl are hyperparameters that control the balance of each loss.
[0144] Furthermore, this method also includes:
[0145] Constructing a semantic prior diffusion model: By establishing a diffusion model in a compact latent space to recover high-quality semantic priors, effectively guiding the enhancement and fusion process. Figure 8 For the schematic diagram of the semantic diffusion model provided by this application, as Figure 8 shown, the semantic prior diffusion model includes a forward diffusion process and a backward denoising process. The specific steps are as follows:
[0146] S51, Forward diffusion process: First, the cascaded high-quality infrared and visible light images and are embedded into the high-quality semantic prior p s . In this step, it is simplified as x 0 . x 0 serves as the starting point of the forward Markov chain and gradually adds Gaussian noise through T iterations. The specific process of T is as follows:
[0147]
[0148] where x t is the noise variable at the t-th step, β t controls the variance of the noise, and α t = 1 - β t . Through reparameterization and iterative derivation, the forward Markov process can be re-expressed as:
[0149]
[0150] where, When t approaches a relatively large value T, tends to 0, and q(x T |x 0 ) approaches the normal distribution The forward diffusion process ends.
[0151] S52, Backward denoising process: The backward denoising process starts from a pure Gaussian distribution and gradually denoises the low-quality semantic prior through a T-step Markov chain to generate a high-quality semantic prior, which is defined as follows:
[0152]
[0153]
[0154] Among them, ∈ is noise.
[0155] The present invention designs a denoising U-Net network, which estimates the noise ∈ by means of semantic priors extracted from low-quality images and degradation priors as well as to estimate the noise ∈. By using the reparameterization trick and replacing ∈ with
[0156]
[0157] Since the distribution of the latent semantic space is simpler than the image space (R H×W×3 ), the semantic prior (p s ′ ) can be generated with fewer iterations. Therefore, the present invention runs a complete iterative reverse process of T (<<1000) steps to infer p s ′ . Therefore, the present invention uses to train the semantic prior diffusion model. In addition, content loss, structural similarity loss, and color consistency loss are jointly used to constrain the training of the semantic prior diffusion model. Therefore, the total loss of the semantic prior diffusion model is defined as:
[0158]
[0159] Among them, λ diff is a hyperparameter used to balance various losses. After training is completed, the high-quality semantic prior p ′ ′ restored in this step is used to replace the high-quality semantic prior p s in step S2 to guide the enhancement and fusion process.
[0160] Figure 9 This is a schematic diagram for comparing the method of the present invention with an advanced image fusion method under the same degradation scenario provided by an embodiment of the present application. Figure 10Schematic diagram for comparing the method of the present invention with advanced image enhancement + image fusion methods under different degradation scenarios provided by the embodiments of the present application. Compared with the existing fusion methods, the present invention realizes the adaptive perception of degradation types, the efficient restoration of scene semantic priors, the organic combination of image enhancement and fusion, and does not require additional auxiliary information. Even in harsh and strongly interfering environments, high-quality fusion results can be generated. To objectively measure the performance of the method of the present invention, a variety of degradation scenarios are selected for evaluation. As Figure 8 shown, when the existing fusion algorithm is directly applied to the degraded source image, the fused image is significantly affected by interference factors. In addition, as Figure 9 shown, when the source image is processed using a pre-enhancement algorithm, since it can only utilize the context within a single modality to enhance the source image, although the influence of interference factors can be weakened, the fusion algorithm still cannot obtain a satisfactory visual effect. On the contrary, the present invention does not rely on additional pre-enhancement algorithms and auxiliary information. Its multi-modal fusion results are not affected by degradation, can provide satisfactory visual perception, and greatly improve the applicability and robustness of the image fusion task in actual complex environments.
[0161] According to the method described in the above embodiments, in this embodiment, the infrared and visible light image fusion device guided by degradation and semantic priors will be further described from the perspective of the device. The infrared and visible light image fusion device guided by degradation and semantic priors can be specifically implemented as an independent entity, or integrated in an electronic device, and the electronic device can be a device such as a terminal, a server, etc. Among them, the terminal can include a tablet computer, a notebook computer, a personal computer (PC), a microprocessing box, or other devices, etc.
[0162] Please refer to Figure 11 , Figure 11 which specifically describes the infrared and visible light image fusion device guided by degradation and semantic priors provided by the embodiments of the present application, applied in an electronic device. The infrared and visible light image fusion device guided by degradation and semantic priors may include:
[0163] An acquisition module, configured to acquire a degraded infrared image, a degraded visible light image, a high-quality infrared image, and a high-quality visible light image;
[0164] A degradation prior generation module, configured to input the degraded infrared image and the degraded visible light image into a degradation prior embedding network model to obtain a degradation prior;
[0165] A semantic prior generation module, configured to input the high-quality infrared image and the high-quality visible light image into a semantic prior embedding network model to obtain a high-quality semantic prior;
[0166] A fusion module, configured to input the degraded infrared image, the degraded visible light image, the degradation prior, and the high-quality semantic prior into a degradation and semantic prior-guided enhancement and fusion network model to obtain a fused image.
[0167] In specific implementation, each of the above modules and / or units may be implemented as an independent entity, or may be combined arbitrarily to be implemented as the same or several entities. For the specific implementation of each of the above modules and / or units, reference may be made to the foregoing method embodiments, and the beneficial effects that can be specifically achieved may also be referred to the beneficial effects in the foregoing method embodiments, which will not be elaborated herein.
[0168] In addition, an embodiment of the present application further provides an electronic device, which may be a device such as a computer or a tablet computer. The electronic device may implement the steps in any embodiment of the infrared and visible light image fusion method with dual guidance of degradation and semantic prior provided by the embodiments of the present application. Therefore, the beneficial effects that can be achieved by any of the infrared and visible light image fusion methods with dual guidance of degradation and semantic prior provided by the embodiments of the present invention can be achieved. For details, refer to the foregoing embodiments, which will not be elaborated herein.
[0169] Figure 12 The specific structural block diagram of the electronic device provided by the embodiment of the present invention is shown. The electronic device may be used to implement the infrared and visible light image fusion method with dual guidance of degradation and semantic prior provided in the foregoing embodiment. The electronic device 500 may be a device such as a terminal or a server. Among them, the terminal may include a tablet computer, a notebook computer, a personal computer (PC), a micro processing box, or other devices, etc.
[0170] The RF circuit 510 is used to receive and transmit electromagnetic waves, realizing the mutual conversion between electromagnetic waves and electrical signals, so as to communicate with a communication network or other devices. The RF circuit 510 may include various existing circuit elements for performing these functions, such as antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, subscriber identity module (SIM) cards, memories, and so on. The RF circuit 510 can communicate with various networks such as the Internet, enterprise intranets, wireless networks or communicate with other devices through wireless networks. The above-mentioned wireless networks may include cellular phone networks, wireless local area networks or metropolitan area networks. The above-mentioned wireless networks can use various communication standards, protocols and technologies, including but not limited to Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as Institute of Electrical and Electronics Engineers standards IEEE 802.11a, IEEE 802.11b, IEEE 802.11g and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging and short messages, and any other suitable communication protocols, and may even include those protocols that have not been developed yet.
[0171] The memory 520 can be used to store software programs and modules, such as the corresponding program instructions / modules in the above embodiments. The processor 580 executes various functional applications and data processing by running the software programs and modules stored in the memory 520, that is, to implement functions such as taking pictures with the front camera, processing the captured images, and switching the display colors of the display content on the display screen. The memory 520 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 520 may further include a memory remotely disposed relative to the processor 580, and these remote memories can be connected to the electronic device 500 through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.
[0172] The input unit 530 can be used to receive input digital or character information, as well as generate keyboards and mice related to user settings and function controls.
[0173] The display unit 540 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces, and these graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. The display unit 540 may include a display panel 541. Optionally, the display panel 541 can be configured in the form of an LCD (Liquid Crystal Display) or an OLED (Organic Light-Emitting Diode).
[0174] The audio circuit 560, speaker 561, and microphone 562 can provide an audio interface between the user and the electronic device 500. The audio circuit 560 can transmit the electrical signal converted from the received audio data to the speaker 561, and the speaker 561 converts it into a sound signal for output; on the other hand, the microphone 562 converts the collected sound signal into an electrical signal, which is received by the audio circuit 560 and then converted into audio data. After the audio data is output to the processor 580 for processing, it is sent to another terminal, for example, through the RF circuit 510, or the audio data is output to the memory 520 for further processing. The audio circuit 560 may also include an earphone jack to provide communication between the peripheral earphone and the electronic device 500.
[0175] The electronic device 500 can help the user receive requests, send information, etc. through the transmission module 570 (such as a Wi-Fi module), and it provides the user with wireless broadband Internet access. Although the transmission module 570 is shown in the figure, it can be understood that it does not belong to the essential components of the electronic device 500 and can be omitted completely within the scope of not changing the essence of the invention according to needs.
[0176] The processor 580 is the control center of the electronic device 500, connecting various parts of the entire mobile phone through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 520, and by calling the data stored in the memory 520, it executes various functions of the electronic device 500 and processes data, thereby monitoring the electronic device as a whole. Optionally, the processor 580 may include one or more processing cores; in some embodiments, the processor 580 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 580 either.
[0177] The electronic device 500 further includes a power supply 590 (such as a battery) for supplying power to each component. In some embodiments, the power supply can be logically connected to the processor 580 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 590 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0178] Although not shown, the electronic device 500 further includes a camera (such as a front camera and a rear camera), a Bluetooth module, etc., which will not be elaborated here. Specifically, in this embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal further includes a memory, and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors. One or more programs include instructions for performing the following operations:
[0179] Obtain a degraded infrared image, a degraded visible light image, a high-quality infrared image, and a high-quality visible light image;
[0180] Input the degraded infrared image and the degraded visible light image into a degradation prior embedding network model to obtain a degradation prior;
[0181] Input the high-quality infrared image and the high-quality visible light image into a semantic prior embedding network model to obtain a high-quality semantic prior;
[0182] Input the degraded infrared image, the degraded visible light image, the degradation prior, and the high-quality semantic prior into a degradation and semantic prior-guided enhancement and fusion network model to obtain a fused image.
[0183] In specific implementation, each of the above modules can be implemented as an independent entity, or can be arbitrarily combined and implemented as the same or several entities. For the specific implementation of each of the above modules, reference can be made to the foregoing method embodiments, which will not be elaborated herein.
[0184] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. For this purpose, an embodiment of the present invention provides a storage medium, in which multiple instructions are stored, and the instructions can be loaded by a processor to execute the steps of any one of the embodiments of the infrared and visible light image fusion method with dual guidance of degradation and semantic prior provided by the embodiments of the present invention.
[0185] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.
[0186] Since the instructions stored in the storage medium can execute the steps of any one of the embodiments of the infrared and visible light image fusion method with dual guidance of degradation and semantic prior provided by the embodiments of the present invention, the beneficial effects that can be achieved by any of the infrared and visible light image fusion methods with dual guidance of degradation and semantic prior provided by the embodiments of the present invention can be realized. For details, see the foregoing embodiments, which will not be elaborated herein.
[0187] The above has introduced in detail an infrared and visible light image fusion method, device, storage medium and electronic device with dual guidance of degradation and semantic prior provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for fusion of infrared and visible light images guided by dual factors of degradation and semantic prior, characterized in that: The method comprises: Acquire a degraded infrared image, a degraded visible light image, a high-quality infrared image, and a high-quality visible light image; Inputting the degraded infrared image and the degraded visible light image into a degradation prior embedding network model to obtain a degradation prior; Inputting the high-quality infrared image and the high-quality visible light image into a semantic prior embedding network model to obtain a high-quality semantic prior; The degraded infrared image, the degraded visible light image, the degradation prior and the high-quality semantic prior are input into an enhancement and fusion network model guided by degradation and semantic prior to obtain a fused image.
2. The infrared and visible light image fusion method based on dual guidance of degradation and semantic prior according to claim 1 is characterized in that: The degradation and semantic prior-guided enhancement and fusion network model includes a prior-adjusted encoding layer, a prior-guided fusion module, a prior-modulated decoding layer and a fusion reconstruction head connected in sequence; The step of inputting the degraded infrared image, the degraded visible light image, the degradation prior and the high-quality semantic prior into an enhancement and fusion network model guided by degradation and semantic prior to obtain a fused image comprises: Inputting the degraded infrared image, the degraded visible light image, the degradation prior and the high-quality semantic prior into the coding layer of the prior modulation to perform feature extraction and feature enhancement, so as to obtain an enhanced infrared feature map and an enhanced visible light feature map; Inputting the enhanced infrared feature map and the enhanced visible light feature map into the prior-guided fusion module for fusion to obtain a fused feature map; Inputting the fused feature map into the decoding layer of the prior modulation to perform feature enhancement to obtain an enhanced fused map; The enhanced fusion image is input into the fusion reconstruction head to obtain a fused image.
3. The infrared and visible light image fusion method based on dual guidance of degradation and semantic prior according to claim 2 is characterized in that: The step of inputting the degraded infrared image, the degraded visible light image, the degradation prior and the high-quality semantic prior into the coding layer of the prior modulation for feature extraction and feature enhancement to obtain an enhanced infrared feature map and an enhanced visible light feature map comprises: Inputting the degradation prior into the coding layer of the prior modulation and integrating it with the degraded infrared image and the degraded visible light image respectively to obtain a first infrared feature map and a first visible light feature map; Inputting the high-quality semantic prior into the coding layer of the prior modulation and integrating them with the degraded infrared image and the degraded visible light image respectively to obtain a second infrared feature map and a second visible light feature map; Aggregating the first infrared characteristic image and the second infrared characteristic image to obtain an enhanced infrared characteristic image; The first visible light characteristic map and the second visible light characteristic map are aggregated to obtain an enhanced visible light characteristic map.
4. The infrared and visible light image fusion method based on dual guidance of degradation and semantic prior according to claim 3 is characterized in that: The step of inputting the degradation prior into the coding layer of the prior modulation and integrating the degradation prior with the degraded infrared image and the degraded visible light image to obtain a first infrared feature map and a first visible light feature map comprises: compressing the degraded infrared image into a high-dimensional vector having the same shape as the degraded prior; Multiplying the high-dimensional vector by the degradation prior, and passing the multiplication result through a linear layer to obtain a first modulation parameter and a second modulation parameter; Obtaining a first infrared characteristic map based on the first modulation parameter, the second modulation parameter and the degraded infrared image; Compressing the degraded visible light image into a high-dimensional vector having the same shape as the degraded prior; Multiplying the high-dimensional vector by the degradation prior, and passing the multiplication result through a linear layer to obtain a first modulation parameter and a second modulation parameter; A first visible light characteristic map is obtained based on the first modulation parameter, the second modulation parameter and the degraded visible light image.
5. The infrared and visible light image fusion method based on dual guidance of degradation and semantic prior according to claim 3 is characterized in that: The step of inputting the high-quality semantic prior into the coding layer of the prior modulation and integrating them with the degraded infrared image and the degraded visible light image to obtain a second infrared feature map and a second visible light feature map comprises: Mapping the degraded infrared image to a first query vector, mapping the degraded infrared image to a second query vector, and mapping the high-quality semantic prior to a key and a value; obtaining a second infrared feature map based on the degraded infrared image, the key and the value; A second visible light characteristic map is obtained based on the degraded visible light image, the key and the value.
6. The infrared and visible light image fusion method with dual guidance of degradation and semantic prior according to claim 2 is characterized in that: The step of inputting the enhanced infrared feature map and the enhanced visible light feature map into the a priori guided fusion module for fusion to obtain a fused feature map comprises: Obtaining channel-level fusion weights based on the semantic prior; Generate respective spatial weights based on the enhanced infrared feature map and the enhanced visible light feature map; The enhanced infrared feature map and the enhanced visible light feature map are weightedly fused based on the channel-level fusion weight and the spatial weight to obtain a fused feature map.
7. The infrared and visible light image fusion method based on dual guidance of degradation and semantic prior according to claim 1, characterized in that: The method further comprises: The high-quality infrared image and the high-quality visible light image are embedded into the high-quality semantic prior, and the embedding process is represented as x0, where x0 is the starting point of the forward Markov chain. Among them, x t is the noise variable at step t, β t Controls the variance of the noise, α t =1-β t ; By reparameterization and iterative derivation, the forward Markov process can be reformulated as: in, When t approaches a larger value T, tends to 0, q(x T |x0) is close to normal distribution The forward diffusion process ends.
8. The infrared and visible light image fusion method with dual guidance of degradation and semantic prior according to claim 7 is characterized in that: The method further comprises: The low-quality semantic prior is gradually denoised by a T-step Markov chain to generate a high-quality semantic prior. The T-step Markov chain is defined as follows: in, ∈ is noise; Using the reparameterization technique, and replacing ∈ with You can get: in, is a low-quality semantic prior, represents the high-quality semantic prior extracted from high-quality visible light images, Represents high-quality semantic priors extracted from high-quality infrared images.
9. A device for fusion of infrared and visible light images with dual guidance of degradation and semantic prior, characterized in that: include: An acquisition module, used for acquiring a degraded infrared image, a degraded visible light image, a high-quality infrared image and a high-quality visible light image; A degradation prior generation module, used for inputting the degraded infrared image and the degraded visible light image into a degradation prior embedding network model to obtain a degradation prior; A semantic prior generation module, used for inputting the high-quality infrared image and the high-quality visible light image into a semantic prior embedding network model to obtain a high-quality semantic prior; The fusion module is used to input the degraded infrared image, the degraded visible light image, the degradation prior and the high-quality semantic prior into the enhancement and fusion network model guided by the degradation and semantic prior to obtain a fused image.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the infrared and visible light image fusion method with dual guidance of degradation and semantic prior as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Infrared and visible light visual information fusion method based on gradient transformation prior
CN117173063A
Image generation method and device, electronic equipment and storage medium
CN117474748A
Recovery method and device based on series model
CN118552447A
Degradation parameter assisted spatial adaptive multi-frame image restoration method
CN119048372A
Multimodal medical image fusion method based on darts network
US20230196528A1
Cited By
Interactive integrated image restoration fusion method and device
CN121353127A