Lightweight real-time infrared and visible light image fusion method based on deep learning
By constructing a lightweight deep learning infrared and visible light image fusion network, the problem of imbalance between model size and inference efficiency is solved, achieving efficient real-time image fusion, which is suitable for application scenarios with limited computing resources.
Patent Information
- Application Number
- CN202511090998.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-11-21
AI Technical Summary
Existing deep learning infrared and visible light image fusion methods suffer from an imbalance between model size and inference efficiency in practical applications, leading to computational delays and making it difficult to meet real-time requirements. This is particularly true in military UAV perception systems where real-time image fusion is challenging to achieve.
A lightweight real-time infrared and visible light image fusion method based on deep learning is adopted. By converting the image color gamut, generating masks with U2Net, and constructing a lightweight fusion network structure, the model is trained using the PyTorch framework and a specific loss function. Combined with the Adam optimization function and learning rate decay strategy, a lightweight feature extraction and fusion module is designed, including a spatial encoder-decoder type module and a fly-scale fusion layer, to optimize the feature extraction and fusion process.
It achieves improved image quality while maintaining lightweight design, with fewer parameters and faster computation speed, enabling image fusion to be completed in milliseconds, meeting real-time requirements and suitable for application scenarios with limited computing resources.
Smart Images

Figure CN120997056A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image fusion, and particularly relates to a lightweight real-time infrared and visible light image fusion method based on deep learning. BACKGROUND
[0002] For many years, infrared and visible light image fusion (IVIF) has been a very popular research topic. It is the complementary characteristics of visible light and infrared images that make IVIF applicable to many tasks, such as target tracking, target detection and biometric recognition, and have a wide range of applications in military and medical fields.
[0003] Specifically, visible light images contain rich texture details, have high spatial resolution and rich appearance and gradient information, but are easily affected by obstacles and light reflection, and are sensitive to changes in light. In contrast, infrared images capture thermal radiation information that is not sensitive to changes in light, but have less texture detail and lower resolution. Therefore, the fusion image of infrared and visible light images retains both the thermal radiation information of infrared images and the texture information of visible light images, which will be beneficial to target recognition and tracking, thus making it easier to carry out downstream applications.
[0004] IVIF methods can be mainly divided into traditional methods and deep learning-based methods. Traditional IVIF methods can be divided into several types according to their corresponding theories, i.e. multi-scale transform-based methods, saliency-based methods, sparse representation-based methods, subspace-based methods, optimization-based methods, etc. However, in traditional methods, the feature extraction rules and fusion rules are dependent on artificial formulation, so this method becomes more and more difficult to develop and more and more difficult to deploy in practice. In recent years, GPU computing power has been continuously developed and broken through, and deep learning technology, which is powerful but requires high-end GPU computing power, has been favored by researchers. The powerful feature extraction capability of the convolution technology in deep learning has solved the problem of the need for artificial formulation of feature extraction rules and fusion rules.
[0005] The IVIF method based on deep learning is classified into a neural network model type, namely a CNN-based IVIF method and a Transformer-based IVIF method. There is a branch of the CNN-based method, namely a GAN-based IVIF method. The above methods usually enhance the feature fusion capability of the model by a large number of stacked multi-layer feature extraction modules and image enhancement units. However, the performance improvement at the cost of increasing the parameter quantity and the computational complexity of the model is often accompanied by a significant decrease in inference speed. It is worth noting that in actual application scenarios, the balance between model size and inference efficiency is increasingly prominent. Taking a military unmanned aerial vehicle perception system as an example, the IVIF algorithm carried by the unmanned aerial vehicle needs to complete multi-source image fusion processing within milliseconds, and a model with excessive parameter quantity is easy to cause calculation delay, which is difficult to meet the real-time requirement, and thus affects the combat. In such a specific scenario, the lightweight design of the model has become a key factor determining the usability of the system.
[0006] The present subject matter focuses on the lightweight research of the infrared and visible image fusion model driven by deep learning, improves the quality of the fused image while ensuring the lightweight, and the research of the present subject matter is of great significance for the actual deployment of the lightweight IVIF model. SUMMARY
[0007] In view of the deficiencies in the prior art, the present application provides a lightweight real-time infrared and visible image fusion method based on deep learning.
[0008] The present application discloses a lightweight real-time infrared and visible image fusion method based on deep learning, comprising the following steps:
[0009] Step 1, input the aligned infrared and visible images, perform image color gamut conversion processing on the visible image, and obtain a single-channel infrared image I ir and a Y channel I vi of the visible image.
[0010] Step 2, call U2Net to generate a mask I mask of the single-channel infrared image, which is used for calculation of a loss function in a model training process.
[0011] Step 3, select a deep learning framework, establish a suitable infrared and visible image fusion network structure, and design a loss function.
[0012] Step 4, select an optimization function and a learning rate decay strategy, set hyperparameters including the number of iterations and the batch size, input I ir , I vi , and I mask into an image fusion network for training, and save the model for testing.
[0013] As a further improvement of the present application, in the step 1, the image gamut conversion processing comprises:
[0014] The visible light image is converted from the RGB gamut to the YCbCr gamut; wherein the conversion formula is:
[0015]
[0016] After the conversion is completed, the Y channel of the YCbCr gamut is separated to obtain the Y channel of the visible light image.
[0017] As a further improvement of the present application, the step 3 specifically comprises:
[0018] Step 3.1, PyTorch is selected as a deep learning framework to construct an infrared and visible light image fusion network; wherein the network is composed of a feature extraction module-a "spatial encoder-decoder type module" (SEDM) and a fusion module-a "flyweight fusion layer" (FFL), the feature extraction module is composed of an encoder, an intermediate layer and a decoder, the encoder and the decoder are both single convolution operators with a convolution kernel of 3, which are used to integrate the feature maps and adjust the number of channels; the intermediate layer is composed of an attention sub-module (ATB) and a feature enhancement sub-module (FEB), the attention sub-module is composed of a single convolution operator with a convolution kernel of 3 and a HardSigmoid activation function, which is used to enhance the salient objects of the image; the feature enhancement sub-module is composed of a bias-free convolution with a convolution kernel size of 3 and a ReLU activation function in parallel, which is used to enhance the image feature details; the fusion module is composed of three groups of "convolution-activation function" blocks, which are used to fuse the feature details and output a fused image;
[0019] Step 3.2, I ir , I vi is input to the feature extraction module to perform feature extraction to obtain I ir-embed , I vi-embed , I ir-re , I vi-re for subsequent tasks. Wherein, I ir-embed , I vi-embed are the infrared image and the visible light image respectively, and the feature map groups obtained by encoding through the encoder have the characteristics of a large number of channels and retaining the original features. I ir-re , I vi-r are the feature maps obtained by decoding through the decoder, and the number of channels is only one, which contains the core features extracted from the original features by the feature extraction module;
[0020] Step 3.3, I ir-embed , vi-embed , ir-re , vi-re is input to the fusion module for feature fusion to obtain a fusion result Fuse; wherein, in the inference stage, the fusion result is additionally subjected to normalization and data format conversion operations as formula (2), which is used for image saving;
[0021]
[0022] wherein, UINT8 represents converting the image data format to unsigned int8 format;
[0023] Step 3.4, designing a loss function L for model training total , the loss function L total is composed of a content loss L content , a similarity loss L SSIM , a saliency loss L saliency and an image reconstruction loss L recons , the content loss function is used to enhance the contrast information and texture details in the fusion image; the similarity loss guides the model to retain the pixel intensity and detail texture information of the input image during fusion; the saliency loss guides the network model to better extract the salient information in the infrared image and the texture information in the visible light image; the image reconstruction loss function guides the decoder in the feature extraction module to retain the pixel intensity and detail texture information of the source image.
[0024] As a further improvement of the present application, in the step 4, the Adam optimization function is selected, and the LambdaLR learning rate decay strategy is selected.
[0025] Compared with the prior art, the present application has the following beneficial effects:
[0026] The infrared and visible light image fusion method of the present application can improve the quality of the fusion image while ensuring lightweight, and has certain research significance and application value in the field of computer vision, infrared and visible light image fusion. BRIEF DESCRIPTION OF DRAWINGS
[0027] Figure 1 is a flow chart of the lightweight real-time infrared and visible light image fusion method based on deep learning disclosed by the present application;
[0028] Figure 2 is a structure diagram of the infrared and visible light image fusion network disclosed by the present application;
[0029] Figure 3 is a training set infrared image disclosed by the present application;
[0030] Figure 4 The training set visible light image disclosed by the present application;
[0031] Figure 5 The training set mask image disclosed by the present application;
[0032] Figure 6 The test set infrared image disclosed by the present application;
[0033] Figure 7 The test set visible light image disclosed by the present application;
[0034] Figure 8 The test set fusion image disclosed by the present application. DETAILED DESCRIPTION
[0035] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0036] The present application will be described in further detail below with reference to the drawings:
[0037] As shown in Figure 1 , the present application provides a lightweight real-time infrared and visible light image fusion method based on deep learning, which selects public open source mainstream infrared and visible light image fusion data sets RoadScene, MSRS and M3FD, including the following steps:
[0038] Step 1, read the infrared and visible light images with python-opencv library, as shown in Figures 3-4 ; the read infrared image is single channel, denoted as I ir . The read visible light image is three channels of RGB color domain, which is converted to YCbCr color domain, and its Y channel (single channel) is taken, denoted as I vi ; wherein,
[0039] The conversion formula is:
[0040]
[0041] Step 2, call U2Net to generate the mask I mask of the single channel infrared image I ir , as shown in Figure 5 , which is used for the calculation of loss function in the model training process;
[0042] Step 3, select a deep learning framework, build a suitable infrared and visible light image fusion network structure, and design a loss function;
[0043] Specifically, it includes:
[0044] PyTorch is selected as the deep learning framework, and an infrared and visible light image fusion network is constructed, as shown in Figure 2 ; wherein the network is composed of a feature extraction module-"spatial encoder-decoder type module" (SEDM) and a fusion module-"flyweight fusion layer" (FFL), the feature extraction module is composed of an encoder, an intermediate layer and a decoder, the encoder and the decoder are both single convolution operators with a convolution kernel of 3, and the function is to integrate the feature map features and adjust the channel number; the intermediate layer is composed of an attention sub-module (ATB) and a feature enhancement sub-module (FEB), the attention sub-module is composed of a single convolution operator with a convolution kernel of 3 and a HardSigmoid activation function, which is used to enhance the image salient target; the feature enhancement sub-module is composed of a parallel bias-free convolution with a convolution kernel size of 3 and a ReLU activation function, which is used to enhance the image feature details; the fusion module is composed of three groups of "convolution-activation function" blocks, which are used to fuse the feature details and output the fusion image;
[0045] Step 3.2, I ir , I vi is input to the feature extraction module for feature extraction to obtain I ir-embed , I vi-embed , I ir-re , I vi-re for subsequent tasks. Among them, I ir-embed , I vi-embed are the feature maps obtained by encoding the infrared image and the visible light image by the encoder, respectively, which are characterized by a large number of channels and retain the original features. I ir-re , I vi-re are the feature maps obtained by decoding the infrared image and the visible light image by the decoder, respectively, which have only one channel and contain the core features extracted from the original features by the feature extraction module;
[0046] Step 3.3, I ir-embed , I vi-embed , I ir-re , I vi-Input to the fusion module for feature fusion, get the fusion result Fuse;Wherein, in the inference stage, additionally to the fusion result as formula (2) Normalization and data format conversion operation, for image saving;
[0047]
[0048] Wherein, UINT8 represents the image data format conversion unsigned int8 format;
[0049] Step 3.4, design the loss function L for model training total , the loss function L total By content loss L content , similarity loss L SSIM , saliency loss L saliency And image reconstruction loss L recons Composed of, the content loss function is used to enhance the contrast information and texture details in the fusion image;Similarity loss guides the model to retain the pixel intensity and detail texture information of the input image during fusion;Saliency loss guides the network model to better extract the salient information in the infrared image and the texture information in the visible light image;Image reconstruction loss function guides the infrared and visible light images reconstructed by the decoder in the feature extraction module to retain the pixel intensity and detail texture information of the source image.
[0050] Step 4, select Adam optimization function, select LambdaLR learning rate decay strategy, set iteration number, batch size and other hyperparameters, input I ir , I vi , I mask Into the network for training, save the model for testing;Wherein, the test set infrared image as shown in Figure 6 , the test set visible light image as shown in Figure 7 , the test set fusion image as shown in Figure 8 .
[0051] The trained model of the application has the following characteristics:
[0052] 1. Lightweight: the parameter amount of the method is about 4.5k, which is the least parameter amount in the existing work;
[0053] 2. Strong fusion ability: a large number of experiments are carried out on three mainstream data sets RoadScene, MSRS and M3FD, six evaluation indexes are used for evaluation, compared with the lightweight infrared and visible light image fusion work in recent years, the results show that the fusion ability of the application is better than the existing work;
[0054] 3. Strong real-time performance: On three mainstream datasets, RoadScene, MSRS, and M3FD, the inference speed of the GPU can reach up to 3 ms, and the inference speed of the CPU can reach up to 57 ms, which is better than existing work.
[0055] The above merely illustrates the preferred embodiments of the present application, and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A lightweight real-time infrared and visible image fusion method based on deep learning, characterized in that, The method comprises the following steps: Step 1, input the aligned infrared and visible light images, perform image color gamut conversion processing on the visible light image to obtain a single-channel infrared image I ir and the Y channel of the visible light image I vi ; Step 2, call U2Net to generate a mask I of the single-channel infrared image mask ; Step 3, selecting a deep learning framework, establishing an infrared and visible light image fusion network structure, and designing a loss function; Step 4, select the optimization function and learning rate decay strategy, set the hyperparameters including the number of iterations and batch size, train the I ir , vi , mask Input image fusion network training, save the model for testing.
2. The image fusion method of claim 1, wherein, In the step 1, the image gamut conversion processing comprises: Converting the visible light image from the RGB color gamut to the YCbCr color gamut; wherein the conversion formula is: After the conversion is completed, the Y channel of the YCbCr color gamut is separated to obtain the Y channel of the visible light image.
3. The image fusion method of claim 1, wherein, The step 3 specifically comprises: Step 3.1, selecting PyTorch as the deep learning framework to construct an infrared and visible light image fusion network; wherein the network is composed of a feature extraction module, a "spatial domain encoder-decoder type module" and a fusion module, a "Flyweight Fusion Layer" (FFL); the feature extraction module is composed of an encoder, an intermediate layer and a decoder; the encoder and the decoder are both single convolution operators with a convolution kernel of 3, which are used to integrate the feature maps and adjust the channel number; the intermediate layer is composed of an attention submodule and a feature enhancement submodule; the attention submodule is composed of a single convolution operator with a convolution kernel of 3 and a HardSigmoid activation function, which is used to enhance the significant targets of the image; the feature enhancement submodule is composed of a parallel unbiased convolution with a convolution kernel size of 3 and a ReLU activation function, which is used to enhance the image feature details; the fusion module is composed of three groups of "convolution-activation function" blocks, which are used to fuse the feature details and output the fused image; Step 3.2, I ir , vi is input to the feature extraction module for feature extraction to obtain I ir-embed , vi-embed , ir-re , vi-r for subsequent tasks; wherein I ir-embed , vi-embed are feature map groups obtained by encoding the infrared image and the visible light image respectively; I ir-re , vi- are feature maps obtained by decoding the decoder, and the channel number is only one, which contains the core features extracted from the original features by the feature extraction module. Step 3.3, I ir-embed , I vi-embed , I ir-r , I vi-r is input to the fusion module for feature fusion to obtain a fusion result Fuse; wherein in the inference stage, the fusion result is additionally subjected to normalization and data format conversion operations as in equation (2) for image saving; Wherein, UINT8 represents converting the image data format to unsigned int8 format; Step 3.4, design a loss function L for model training total , loss function L total consists of content loss L content , similarity loss L SSIM , saliency loss L saliency and image reconstruction loss L recons , the content loss function is used to enhance the contrast information and texture details in the fusion image; the similarity loss guides the model to retain the pixel intensity and detail texture information of the input image during fusion; the saliency loss guides the network model to better extract the salient information in the infrared image and the texture information in the visible light image; the image reconstruction loss function guides the decoder in the feature extraction module to retain the pixel intensity and detail texture information of the source image when reconstructing the infrared and visible light images.
4. The image fusion method of claim 1, wherein, In the step 4, an Adam optimization function is selected, and a LambdaLR learning rate decay strategy is selected.