A lightweight end-to-end infrared-visible light adaptive image fusion method
By employing a lightweight end-to-end convolutional neural network and a multi-dimensional loss-constrained infrared-visible image fusion method, the problems of insufficient feature extraction capability and high computational complexity in existing technologies are solved, achieving high-quality, real-time image fusion in different environments.
Patent Information
- Application Number
- CN202511194609.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-08-26
AI Technical Summary
Existing infrared and visible light image fusion methods suffer from insufficient feature extraction capabilities, high computational complexity, and a lack of multi-dimensional constraints, resulting in poor fused image quality and poor real-time performance.
A lightweight end-to-end convolutional neural network is used, combined with a brightness adaptive weight allocation mechanism and multi-dimensional loss constraints, to fuse infrared and visible light images through local brightness perception and strong light semantic detection modules.
It achieves high real-time performance and high fusion quality on embedded devices, maintaining stable image fusion effects under different lighting and scenes, and improving edge sharpness and texture details.
Smart Images

Figure CN120746864B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and more specifically to a lightweight end-to-end adaptive infrared-visible image fusion method. Background Technology
[0002] In recent years, with the rapid development of infrared technology and the decrease in its cost, infrared cameras have been widely used in many fields. Infrared technology captures information based on the thermal radiation of objects and is unaffected by lighting conditions, but it can cause a loss of color and texture details, limiting its application in some computer vision applications. Visible light cameras, on the other hand, can capture the color and texture details of objects very well, but are affected by the environment. Therefore, how to complement and integrate the advantages of infrared and visible light to highlight the thermal radiation information of a target while preserving its color and texture details is a current research hotspot and challenge.
[0003] Currently, infrared and visible light image fusion methods fall into two categories: traditional image fusion methods and deep learning image fusion methods. Traditional image fusion methods include spatial domain, transform domain, sparse representation, and multi-scale transform methods. They fuse infrared and visible light images by extracting base layers and detail layers separately. These methods are faster than deep learning-based image fusion methods and are well-suited for real-time deployment. However, traditional methods rely on manually designed fusion rules, which can lead to poor feature representation and unsatisfactory results with environmental changes, resulting in poor robustness. Deep learning-based image fusion methods include convolutional neural networks, generative adversarial networks, and autoencoders. These methods can adaptively train fusion network models using large amounts of data, employing different branches of the network for differentiated feature extraction, resulting in higher feature fit. Furthermore, well-designed loss functions can guide the model to learn better feature fusion strategies during training, enabling it to adapt to different environments and producing higher-quality fused images. Currently, deep learning-based image fusion methods often employ complex network structures and cumbersome fusion strategies to obtain high-quality fused images, leading to poor real-time performance and difficulty in deployment.
[0004] It is evident that existing infrared-visible image fusion methods generally suffer from the following problems:
[0005] Insufficient feature extraction capability; traditional methods rely on manually designed feature extraction strategies, making it difficult to cope with interference in different environments.
[0006] High computational complexity: Although existing deep learning methods can extract effective features well, their high computational cost makes them difficult to apply in resource-constrained scenarios and lacks real-time performance.
[0007] The lack of multi-dimensional constraints leads to deviations in target edges, textures, and brightness in the fused image.
[0008] Therefore, how to provide a lightweight, end-to-end adaptive infrared-visible image fusion method that can effectively solve the problems existing in the above-mentioned infrared-visible image fusion methods is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0009] In view of this, the present invention provides a lightweight end-to-end adaptive infrared-visible image fusion method. By constructing an end-to-end network based on a convolutional neural network, and using a brightness-adaptive weight allocation mechanism during network model training, the brightness differences in visible light images are analyzed through local brightness perception and strong light semantic detection modules. Overly bright and overexposed areas are suppressed to avoid interfering with infrared target information.
[0010] To achieve the above objectives, the present invention adopts the following technical solution:
[0011] A lightweight end-to-end adaptive infrared-visible image fusion method includes:
[0012] Step 1: Acquire the registered infrared and visible light images, as well as the temporary fused GT image of the infrared and visible light images, and perform image preprocessing;
[0013] Step 2: Construct a lightweight end-to-end infrared and visible light adaptive image fusion network model and train it by inputting the preprocessed image. Update and iterate the training model according to multi-dimensional loss constraints, including pixel-level loss, structure-level loss, region-level loss and weight-level loss.
[0014] Step 3: Input the infrared and visible light images to be fused into the updated training model to obtain the image fusion result.
[0015] Optionally, in step 1, a temporary fused GT image is generated from the registered infrared and visible light images based on the image fusion method TIF.
[0016] Optionally, in step 1, image preprocessing specifically involves converting the infrared and visible light images and the temporary fused GT image into grayscale images containing only brightness information, and scaling them to a fixed size [H,H]; where [H,H] indicates that the scaled image is a square with equal pixel height and pixel width, both being H.
[0017] Optionally, in step 2, a lightweight end-to-end infrared-visible adaptive image fusion network model includes: shallow feature extraction, deep feature extraction, and fused image output.
[0018] The input for shallow feature extraction is a visible light image, which includes a first convolutional layer and a brightness weight generation module connected in sequence; wherein, the first convolutional layer is used to extract the basic feature layer, and the brightness weight generation module is used to extract the brightness weight map;
[0019] Deep feature extraction includes: a convolutional layer connected in sequence, multiple residual blocks and a convolutional layer, used to extract semantic features of thermal targets from infrared images and semantic features of texture structure from visible light images;
[0020] The fused image output is used to fuse shallow and deep features, and then passes through a convolutional layer to obtain the final fused image.
[0021] Optionally, the brightness weight generation module includes local brightness extraction and strong light region semantic detection, and a comprehensive calculation of the brightness weight probability extracted from the visible light image based on the results of local brightness extraction and strong light region semantic detection, to obtain the final comprehensive weight map.
[0022] Optionally, in step 2, pixel-level loss is adopted. The losses are as follows:
[0023]
[0024]
[0025] in, This is a pixel-level loss; Visible light image; Infrared image; This is a fused image; the visible light image, infrared image, and fused image are all the same size. ;in, The number of images used during model training; Number of channels; , The image has the same height and width; For image exist Pixel value at; For image exist Pixel value at;
[0026] Structural losses include: global structural losses and local structural losses;
[0027] The global structural loss employs a multi-scale structural similarity index, as follows:
[0028]
[0029]
[0030] in, This represents the global structural loss. These are weighting coefficients at the global structural level. For image and images In the Brightness similarity at each scale; , Images and images In the Contrast and structural similarity at various scales; , , Set the parameters used to balance the components. ,and ;
[0031] The local structural losses are as follows:
[0032]
[0033]
[0034] in, This is a local structural loss; These are weighting coefficients at the local structural level; To make the image A convolution operation that shifts 1 pixel to the right; To make the image A convolution operation that shifts down by 1 pixel; For local summation convolution;
[0035] The structural level loss is as follows:
[0036]
[0037] in, For structural level loss;
[0038] Regional-level losses include: infrared thermal target retention loss and edge retention loss;
[0039] The loss in infrared thermal target retention is as follows:
[0040]
[0041]
[0042] in, Loss is preserved for infrared thermal targets; The loss weighting coefficient is retained for infrared thermal targets;
[0043] The edge preservation loss is calculated using the Sobel edge detection operator, as follows:
[0044]
[0045] in, Preserve loss at the edges; Preserve the loss weight coefficients at the edges; This represents the gradient magnitude output by the Sobel edge detection operator.
[0046] Regional losses are as follows:
[0047]
[0048] in, This represents a regional-level loss.
[0049] The weighted loss is as follows:
[0050]
[0051]
[0052] in, For weighted loss; Weights are used for weight regularization loss weights; To Extracted brightness weights.
[0053] Optionally, it also includes adding a guided loss to the multi-dimensional loss constraint, as follows:
[0054]
[0055] in, To guide losses; To guide the loss weight; For temporary fusion of GT images;
[0056] The multi-dimensional loss constraints are as follows:
[0057]
[0058] in, For multi-dimensional loss constraints.
[0059] Optionally, step 2, during model training, also includes: processing the fused images. The best fused image is selected as the new ground truth (GT) by comparing it with the temporary fused GT image. Specifically:
[0060] The fused images are calculated using the following formulas respectively. Temporarily fused GT images and visible light images and infrared images The correlation coefficients are as follows:
[0061]
[0062]
[0063] in, For image With images and images The correlation coefficient; , Images and images The average pixel value; , Images and images In Pixel value at; , For image and images Height and width;
[0064] Compare and ,when The resulting image will be merged. As the new GT.
[0065] As can be seen from the above technical solution, compared with the prior art, this invention discloses a lightweight end-to-end adaptive infrared and visible light image fusion method. It achieves the following beneficial effects: Lightweight and Real-time Performance: Through end-to-end network structure design, the network is lightweight, ensuring real-time fusion and enabling deployment on embedded devices; Improved Fusion Quality: Through joint constraints of multi-dimensional losses, the edge sharpness and texture details of the fused image are significantly improved, and the thermal target enhancement loss enhances the contrast of the thermal target region in the fused image; Enhanced Adaptability: The dynamic weight allocation mechanism enables the model to maintain stable performance under different lighting conditions (such as strong light and weak light) and scene types (such as urban and outdoor), and its generalization ability is significantly better than fixed-weight methods. Attached Figure Description
[0066] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0067] Figure 1 This is a schematic diagram of the method flow of the present invention.
[0068] Figure 2 This is a schematic diagram of the lightweight end-to-end infrared-visible adaptive image fusion network model structure of the present invention.
[0069] Figure 3 This is a schematic diagram of the brightness weight generation module of the present invention. Detailed Implementation
[0070] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0071] Example 1:
[0072] Embodiment 1 of this invention discloses a lightweight end-to-end adaptive infrared-visible image fusion method, such as... Figure 1 As shown, it includes:
[0073] Step 1: Acquire the registered infrared and visible light images, as well as the temporary fused GT image of the infrared and visible light images, and perform image preprocessing.
[0074] Based on the image fusion method TIF, a temporary fused GT image is generated from the registered infrared and visible light images.
[0075] Image preprocessing specifically involves converting infrared and visible light images, as well as the temporary fused GT image, into grayscale images containing only brightness information, and scaling them to a fixed size; where [H,H] indicates that the scaled image is a square with equal pixel height and pixel width, both being H.
[0076] Step 2: Construct a lightweight end-to-end infrared-visible adaptive image fusion network model and train it using preprocessed images as input. Update and iterate the training model based on multi-dimensional loss constraints, including pixel-level loss, structural loss, region-level loss, and weight-level loss.
[0077] A lightweight, end-to-end adaptive infrared-visible image fusion network model, designed based on convolutional neural networks, achieves efficient fusion through three core technologies: lightweight architecture design, brightness-adaptive weight allocation mechanism, and multi-dimensional loss constraints. Figure 2 As shown, it includes: shallow feature extraction, deep feature extraction, and fused image output;
[0078] The input for shallow feature extraction is a visible light image, which includes a first convolutional layer and a brightness weight generation module connected in sequence; wherein, the first convolutional layer is used to extract the basic feature layer, and the brightness weight generation module is used to extract the brightness weight map;
[0079] Deep feature extraction includes: a convolutional layer connected in sequence, multiple residual blocks and a convolutional layer, used to extract semantic features of thermal targets from infrared images and semantic features of texture structure from visible light images. It can resist illumination interference and improve the quality of fused images.
[0080] The fused image output is used to fuse shallow and deep features, and then passes through a convolutional layer to obtain the final fused image.
[0081] Brightness weight generation module, such as Figure 3 As shown, the process includes local brightness extraction and strong light region semantic detection, as well as the comprehensive calculation of the brightness weight probability of the visible light image based on the results of local brightness extraction and strong light region semantic detection, to obtain the final comprehensive weight map.
[0082] The fusion of infrared and visible light images is a multimodal heterogeneous image fusion process. It requires consideration at multiple levels to prevent the loss of useful information from both images during model training, ensuring the fused image closely approximates the ideal result. Therefore, this invention designs a multi-dimensional loss constraint from four levels: pixel-level, structure-level, region-level, and weight-level.
[0083] Pixel-level loss adopts To minimize losses, ensure the overall brightness of the fused image lies between visible light and infrared, avoiding it being too bright or too dark, as follows:
[0084]
[0085]
[0086] in, This is a pixel-level loss; Visible light image; Infrared image; This is a fused image; the visible light image, infrared image, and fused image are all the same size. ;in, The number of images used during model training; Number of channels; , The image has the same height and width; For image exist Pixel value at; For image exist Pixel value at;
[0087] Structural losses include: global structural losses and local structural losses;
[0088] The global structural loss employs the Multi-Scale Structural Similarity Index (MSSSIM), which comprehensively constrains the brightness, contrast, and structural information of the image. It is one of the core metrics for measuring fusion quality, as follows:
[0089]
[0090]
[0091] in, This represents the global structural loss. These are weighting coefficients at the global structural level. For image and images In the Brightness similarity at each scale; , Images and images In the Contrast and structural similarity at various scales; , , Set the parameters used to balance the components. ,and ;
[0092] Local structure loss is used to improve the detail preservation ability of the fused image, that is, to preserve texture and edge details, as follows:
[0093]
[0094]
[0095] in, This is a local structural loss; These are weighting coefficients at the local structural level; To make the image A convolution operation that shifts 1 pixel to the right; To make the image A convolution operation that shifts down by 1 pixel; For local summation convolution;
[0096] The structural level loss is as follows:
[0097]
[0098] in, For structural level loss;
[0099] Regional-level losses include: infrared thermal target preservation loss and edge preservation loss, to avoid interference from non-thermal target areas and improve the edge sharpness and integrity of the fused image;
[0100] The loss in infrared thermal target retention is as follows:
[0101]
[0102]
[0103] in, Loss is preserved for infrared thermal targets; The loss weighting coefficient is retained for infrared thermal targets;
[0104] The edge preservation loss is calculated using the Sobel edge detection operator, as follows:
[0105]
[0106] in, Preserve loss at the edges; Preserve the loss weight coefficients at the edges; This is the gradient magnitude output by the Sobel edge detection operator (the sum of the absolute values of the gradients in the x and y directions).
[0107] Regional losses are as follows:
[0108]
[0109] in, This represents a regional-level loss.
[0110] The weighted loss constraint constrains the brightness-adaptive weights generated by the model, ensuring negative correlation between brightness and suppression of strong light regions, as follows:
[0111]
[0112]
[0113] in, For weighted loss; Weights are used for weight regularization loss weights; To The extracted brightness weights are used by the brightness adaptive weight generation module of the network model for visible light images. Extracted.
[0114] This also includes adding guided loss to the multi-dimensional loss constraints for customized optimization of the fusion results, as follows:
[0115]
[0116] in, To guide losses; To guide the loss weight; For temporary fusion of GT images;
[0117] The multi-dimensional loss constraints are as follows:
[0118]
[0119] in, For multi-dimensional loss constraints.
[0120] The model training process also includes: processing the fused images. The best fused image is selected as the new ground truth (GT) by comparing it with the temporary fused GT image. Specifically:
[0121] The fused images are calculated using the following formulas respectively. Temporarily fused GT images and visible light images and infrared images The correlation coefficients are as follows:
[0122]
[0123]
[0124] in, For image With images and images The correlation coefficient; , Images and images The average pixel value; , Images and images In Pixel value at; , For image and images Height and width;
[0125] Compare and ,when The resulting image will be merged. As the new GT.
[0126] The model is updated and trained iteratively based on the overall loss. When the overall loss decreases, converges, and stabilizes, the model with the best performance is selected as the final model based on the model's output.
[0127] Step 3: Input the infrared and visible light images to be fused into the updated training model to obtain the image fusion result.
[0128] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0129] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A lightweight end-to-end adaptive infrared-visible image fusion method, characterized in that, include: Step 1: Acquire the registered infrared and visible light images and the temporary fused GT image of the infrared and visible light images, and perform image preprocessing; Step 2: Construct a lightweight end-to-end infrared and visible light adaptive image fusion network model and train it by inputting the preprocessed image. Update and iterate the training model according to multi-dimensional loss constraints, including pixel-level loss, structure-level loss, region-level loss and weight-level loss. Step 3: Input the infrared and visible light images to be fused into the updated and iterated training model to obtain the image fusion result; In step 2, the lightweight end-to-end infrared-visible adaptive image fusion network model includes: shallow feature extraction, deep feature extraction, and fused image output. The input for the shallow feature extraction is a visible light image, including: a first convolutional layer and a brightness weight generation module connected in sequence; wherein, the first convolutional layer is used to extract the basic feature layer, and the brightness weight generation module is used to extract the brightness weight map; The deep feature extraction includes: a convolutional layer connected in sequence, multiple residual blocks and a convolutional layer, used to extract semantic features of thermal targets from infrared images and semantic features of texture structure from visible light images; The fused image output is used to fuse shallow and deep features, and then passes through a convolutional layer to obtain the final fused image. The brightness weight generation module includes local brightness extraction and strong light region semantic detection, and a comprehensive calculation of the brightness weight probability extracted from the visible light image based on the results of the local brightness extraction and strong light region semantic detection, to obtain the final comprehensive weight map.
2. The lightweight end-to-end adaptive infrared-visible image fusion method according to claim 1, characterized in that, In step 1, a temporary fused GT image is generated from the registered infrared and visible light images based on the image fusion method TIF.
3. The lightweight end-to-end adaptive infrared-visible image fusion method according to claim 1, characterized in that, In step 1, the image preprocessing specifically involves converting the infrared and visible light images and the temporary fused GT image into grayscale images containing only brightness information, and scaling them to a fixed size [H,H]; where [H,H] indicates that the scaled image is a square with equal pixel height and pixel width, both being H in size.
4. The lightweight end-to-end adaptive infrared-visible image fusion method according to claim 1, characterized in that, In step 2, the pixel-level loss is adopted The losses are as follows: in, For the pixel-level loss; Visible light image; Infrared image; The image is a fused image; the visible light image, infrared image, and fused image are all the same size. ;in, The number of images used during model training; Number of channels; , The image has the same height and width; For image exist Pixel value at; For image exist Pixel value at; The structural-level loss includes: global structural loss and local structural loss; The global structural loss employs a multi-scale structural similarity index, as follows: in, This is the global structural loss; These are weighting coefficients at the global structural level. For image and images In the Brightness similarity at each scale; , Images and images In the Contrast and structural similarity at various scales; , , Set the parameters used to balance the components. ,and ; The local structural loss is as follows: in, This refers to the local structural loss; These are weighting coefficients at the local structural level; To make the image A convolution operation that shifts 1 pixel to the right; To make the image A convolution operation that shifts down by 1 pixel; For local summation convolution; The structural-level loss is as follows: in, This refers to the structural level loss; The regional-level loss includes: infrared thermal target retention loss and edge retention loss; The loss in infrared thermal target retention is as follows: in, Loss is retained for the infrared thermal target; The loss weighting coefficient is retained for infrared thermal targets; The edge preservation loss is calculated using the Sobel edge detection operator, as follows: in, Preserve the loss for the edges; Preserve the loss weight coefficients at the edges; This represents the gradient magnitude output by the Sobel edge detection operator. The regional-level loss is as follows: in, This refers to the regional level loss; The weighted loss is as follows: in, The weight level loss; Weighted regularization loss weights; To Extracted brightness weights.
5. A lightweight end-to-end adaptive infrared-visible image fusion method according to claim 4, characterized in that, Also includes: Add a guided loss to the multi-dimensional loss constraint as follows: in, The guiding loss; To guide the loss weight; For temporary fusion of GT images; The multi-dimensional loss constraint is as follows: in, The multi-dimensional loss constraint is defined as follows.
6. A lightweight end-to-end adaptive infrared-visible image fusion method according to claim 1, characterized in that, Step 2, during model training, also includes: processing the fused images. The best fused image is selected as the new ground truth (GT) by comparing it with the temporary fused GT image. Specifically: The fused images are calculated using the following formulas respectively. Temporarily fused GT images and visible light images and infrared images The correlation coefficients are as follows: in, For image With images and images The correlation coefficient; , Images and images The average pixel value; , Images and images In Pixel value at; , For image and images Height and width; Compare and ,when The resulting image will be merged. As the new GT.
Citation Information
Patent Citations
Low-illumination target detection method based on MSFAF-Net
CN117456330A
Dual-energy X-ray image fusion method and device, computer equipment and storage medium
CN119991472A