Infrared and visible light image fusion method and system based on deep learning
Through the infrared and visible image fusion method based on deep learning, combined with CNN and adaptive fusion algorithm, the problems of poor fusion effect and information loss in the prior art are solved, and the fusion effect of high accuracy and low computing overhead in complex environments is achieved.
Patent Information
- Application Number
- CN202510263320.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-17
AI Technical Summary
The existing infrared and visible image fusion methods have problems such as poor fusion effect, loss of information and high computing resource consumption, especially in complex environments.
Using deep learning-based methods, image preprocessing, feature extraction, feature fusion and image reconstruction are combined with CNN and adaptive fusion algorithms to improve image detail retention and information integration.
It effectively improves the detail retention, information integration and visibility of the fusion image, especially in complex environments such as low light and haze, ensuring high accuracy of target detection, identification and classification, and optimizing the algorithm structure to reduce computing overhead.
Smart Images

Figure CN120164068A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and image processing, and more specifically, to an infrared and visible light image fusion method and system based on deep learning. Background Art
[0002] Currently, the methods for infrared and visible light image fusion mainly include those based on wavelet transform, principal component analysis, traditional image fusion algorithms (such as weighted average method, Laplacian pyramid, etc.), and those based on deep learning. Traditional algorithms usually have problems in maintaining image details and color consistency and are difficult to cope with challenges in complex environments. The methods based on deep learning, especially the application of convolutional neural network (CNN) and generative adversarial network (GAN), have gradually become a new trend in the field of image fusion, but the existing methods still have problems such as uneven processing of image details and information loss.
[0003] Traditional methods usually have disadvantages such as poor fusion effect, information loss, and manual design of feature extraction rules; while deep learning methods require a large amount of data and computing resources during processing, and have limited processing effects on low-quality images. Although the existing deep learning technologies perform well in many applications, there is still room for optimization in the balance between fusion effect and computing efficiency. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies of poor fusion effect and information loss in the prior art, and provide an infrared and visible light image fusion method and system based on deep learning, which can effectively improve the detail retention, information integration degree, and visibility of the fused image.
[0005] To solve the above technical problems, the technical solution adopted by the present invention is:
[0006] Provide an infrared and visible light image fusion method based on deep learning, including the following steps:
[0007] S1. Image preprocessing: Preprocess the input infrared and visible light images, including noise removal and contrast enhancement;
[0008] S2. Feature extraction: Use CNN to extract the features of infrared and visible light images;
[0009] S3. Feature fusion: Fuse the obtained features respectively;
[0010] S4. Image reconstruction: Reconstruct the fused feature image into a fused image;
[0011] S5. Object detection and evaluation: Perform object detection on the fused image and evaluate the accuracy rate.
[0012] Further, the step S1 includes:
[0013] S11. Input a pair of infrared and visible light images, and after registration, ensure that the image sizes are the same;
[0014] S12. Use the MSRCR algorithm to enhance the visible light image therein to improve the contrast and color fidelity of the night image;
[0015] S13. Use the DDE algorithm to process the infrared image to enhance the image contrast and optimize the image noise.
[0016] Further, the step S12 includes: First, decompose the brightness and reflection information of the image; then apply multi-scale filtering for image processing; finally, enhance the visual effect of the image through color restoration technology.
[0017] Further, the step S2 includes:
[0018] S21. Extract basic image features using the first convolutional layer of ResNet101 pre-trained on ImageNet;
[0019] S22. Use an adaptive fusion algorithm to synthesize the features of the two images;
[0020] S23. Input the extracted features into two parallel processing lines, namely global feature extraction and local feature extraction; in the local branch, there are two serial convolutional layers, using a 3*3 convolutional kernel, and the activation function is ReLU; in the global branch, there are two serial convolutional layers, using a 5*5 convolutional kernel, and the activation function is ReLU.
[0021] Further, in the step S3, the obtained features are respectively fused, where the maximum value fusion strategy is adopted in the local feature extraction branch, and the average value fusion strategy is adopted in the global feature extraction branch; finally, the addition fusion strategy is used to fuse the features extracted in the global feature extraction and the features extracted in the local feature extraction.
[0022] Further, the step S4 includes: performing one-layer convolutional reconstruction on the fused features to output a fused image, where the reconstruction uses a 1*1 convolutional kernel, the activation function is ReLU, and the loss function uses the mean square error.
[0023] The present invention also provides an infrared and visible light image fusion system based on deep learning, including:
[0024] Image preprocessing module: used to preprocess the input infrared and visible light images, including noise removal and contrast enhancement;
[0025] Feature extraction module: used to extract features of infrared and visible light images by using CNN;
[0026] Feature fusion module: used to fuse the obtained features respectively;
[0027] Image reconstruction module: used to reconstruct the fused feature image into a fused image;
[0028] Target detection and evaluation module: used to perform target detection on the fused image and evaluate the accuracy rate.
[0029] Furthermore, when the image preprocessing module performs image preprocessing, it first obtains a pair of infrared and visible light images, and after registration, ensures that the image sizes are the same; then uses the MSRCR algorithm to enhance the visible light image among them to improve the contrast and color fidelity of the night image; finally uses the DDE algorithm to process the infrared image to enhance the image contrast and optimize the image noise; the MSRCR algorithm includes first decomposing the brightness and reflection information of the image; then applying multi-scale filtering for image processing; finally enhancing the visual effect of the image through color restoration technology.
[0030] Furthermore, when the feature extraction module performs feature extraction, it includes extracting basic image features by using the first convolutional layer of ResNet101 pre-trained on ImageNet; using an adaptive fusion algorithm to synthesize the features of the two images; inputting the extracted features into two parallel processing lines, namely global feature extraction and local feature extraction; in the local branch, it contains two serial convolutional layers, using a 3*3 convolutional kernel, and the activation function is ReLU; in the global branch, it contains two serial convolutional layers, using a 5*5 convolutional kernel, and the activation function is ReLU.
[0031] Furthermore, when the feature fusion module performs image fusion, it fuses the obtained features respectively, where the maximum fusion strategy is adopted in the local feature extraction branch, and the average fusion strategy is adopted in the global feature extraction branch; finally, the addition fusion strategy is used to fuse the features extracted in the global feature extraction and the features extracted in the local feature extraction.
[0032] Compared with the prior art, the beneficial effects of the present invention are:
[0033] A method and system for infrared and visible light image fusion based on deep learning of the present invention can effectively improve the detail retention, information integration degree and visibility of the fused image, especially under complex environments such as low light and haze conditions, ensure high accuracy of the fused image in target detection, recognition and classification; at the same time, optimize the algorithm structure, reduce the computational overhead, and improve the feasibility in practical applications. Description of the Drawings
[0034] Figure 1 This is a schematic flow diagram of the method of the present invention;
[0035] Figure 2 This is a schematic structural diagram of the system of the present invention. Detailed implementation manners
[0036] The present invention will be further described below in conjunction with the detailed implementation manners. Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and should not be construed as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, which do not represent the dimensions of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0037] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0038] Embodiment 1
[0039] This embodiment is the first embodiment of an infrared and visible light image fusion method based on deep learning, including the following steps:
[0040] Step S1. Image preprocessing: Preprocess the input infrared and visible light images, including noise removal and contrast enhancement.
[0041] First, a pair of infrared-visible light images are manually input, which have been registered and have the same image size.
[0042] Then, the Multi-Scale Retinex with Color Restoration (MSRCR) algorithm is used to enhance the visible light image, aiming to improve the contrast and color fidelity of the night image. Its core idea is to simultaneously adjust the brightness and color information of the image through multi-scale processing, so as to better simulate the human eye's perception of light and color. The algorithm includes the following steps: (1) Decompose the brightness and reflection information of the image; (2) Apply multi-scale filtering for image processing; (3) Enhance the visual effect of the image through color restoration technology.
[0043] Finally, the DDE algorithm is used to process the infrared image to enhance the image contrast and optimize the image noise.
[0044] Step S2. Feature extraction: CNN is used to extract the features of infrared and visible light images.
[0045] After enhancement, the first convolutional layer of ResNet101 pre-trained on ImageNet is used to extract basic image features. Feature fusion: An adaptive fusion algorithm is used to synthesize the features of the two images. Thereafter, the extracted features are input into two parallel processing lines, namely global feature extraction and local feature extraction. In the local branch, there are two serial convolutional layers with a 3*3 convolutional kernel and a ReLU activation function. In the global branch, there are also two serial convolutional layers with a 5*5 convolutional kernel and a ReLU activation function.
[0046] Step S3. Feature fusion: The obtained features are fused respectively. The obtained features are fused respectively, where the maximum fusion strategy is adopted in the local feature extraction branch and the average fusion strategy is adopted in the global feature extraction branch. Finally, the addition fusion strategy is used to fuse the above two groups of features.
[0047] Step S4. Image reconstruction: The fused feature image is reconstructed into a fused image.
[0048] The fused features are reconstructed by one layer of convolution, and the fused image is output. A 1*1 convolutional kernel is used, the activation function is ReLU, and the loss function is the mean square error (MSE).
[0049] Step S5. Object detection and evaluation: Object detection is performed on the fused image, and the accuracy is evaluated.
[0050] The position of pedestrians has been marked for each image in the dataset. Some images are selected from the fused image for object detection experiments, hoping to improve the performance of the image under different lighting conditions, so as to help the object detection model better identify pedestrians. The YOLO model is selected for the experiment because of its high speed and good detection accuracy, which is suitable for real-time detection tasks.
[0051] In the subjective evaluation, the detection effects of the original image and the fused image are compared. By observing the detection results manually, it can be found that in scenes with low light or large light changes, the fused image shows stronger detail retention and object detection ability. Especially in the case of a large number of people, the fused image can identify pedestrians more clearly and the false alarm rate is significantly reduced.
[0052] Common object detection performance metrics are used for objective evaluation, including:
[0053] Precision: It represents the proportion of real targets in the detection results.
[0054] Recall rate: It represents the proportion of detected targets among real targets.
[0055] F1 score: The harmonic mean of precision and recall rate, which is used to comprehensively evaluate the performance of the model.
[0056] Mean Average Precision (mAP): It represents the average of the object detection precision at multiple IoU thresholds.
[0057] A deep learning-based infrared and visible light image fusion method proposed in this embodiment can effectively improve the detail retention, information integration degree, and visibility of the fused image. Especially in complex environments such as low light and haze conditions, it can ensure high accuracy of the fused image in object detection, recognition, and classification. At the same time, it optimizes the algorithm structure, reduces the computational overhead, and improves the feasibility in practical applications.
[0058] Embodiment 2
[0059] This embodiment is an embodiment of a deep learning-based infrared and visible light image fusion system. This embodiment is similar to Embodiment 1 and includes:
[0060] Image preprocessing module: It is used to preprocess the input infrared and visible light images, including noise removal and contrast enhancement.
[0061] Feature extraction module: It is used to extract the features of infrared and visible light images by using CNN.
[0062] Feature fusion module: It is used to fuse the obtained features respectively.
[0063] Image reconstruction module: It is used to reconstruct the fused feature image into a fused image.
[0064] Object detection and evaluation module: It is used to perform object detection on the fused image and evaluate the accuracy rate.
[0065] In this embodiment, when the image preprocessing module performs image preprocessing, it first obtains a pair of infrared and visible light images and performs registration to ensure that the image sizes are the same. Then, it uses the MSRCR algorithm to enhance the visible light image to improve the contrast and color fidelity of the night image. Finally, it uses the DDE algorithm to process the infrared image to enhance the image contrast and optimize the image noise. The MSRCR algorithm includes first decomposing the brightness and reflection information of the image, then applying multi-scale filtering for image processing, and finally enhancing the visual effect of the image through color restoration technology.
[0066] In this embodiment, when the feature extraction module performs feature extraction, it includes extracting basic image features using the first convolutional layer of ResNet101 pre-trained on ImageNet; synthesizing the features of two images using an adaptive fusion algorithm; inputting the extracted features into two parallel processing lines, namely global feature extraction and local feature extraction; in the local branch, there are two consecutive convolutional layers with a 3*3 convolutional kernel and a ReLU activation function; in the global branch, there are two consecutive convolutional layers with a 5*5 convolutional kernel and a ReLU activation function.
[0067] In this embodiment, when the feature fusion module performs image fusion, it fuses the obtained features respectively. Among them, the maximum fusion strategy is adopted in the local feature extraction branch, and the average fusion strategy is adopted in the global feature extraction branch; finally, the addition fusion strategy is used to fuse the features extracted in the global feature extraction and the features extracted in the local feature extraction.
[0068] In this embodiment, there is also a resource layer, whose task is to efficiently store, read, and process raw image data, fusion results, model parameters, etc. It is necessary to ensure the unobstructed flow of data among multiple computing nodes, disks, and memories, and minimize the IO bottleneck.
[0069] Storage of raw images: Infrared and visible light images usually need to be stored in large-sized file formats. It is planned to use efficient storage formats such as HDF5, TIFF, PNG, etc., and compress them as needed.
[0070] Storage of fusion results: The fused image can be stored in the same format, and a database can be used to manage the generated images. Especially for images that need to be stored for a long time and accessed frequently, a database can be used to store metadata (such as image generation time, fusion parameters, etc.) and provide retrieval capabilities.
[0071] Cache management: During high-performance computing, consider using in-memory caches or distributed cache systems to reduce the time of frequent hard disk reads and optimize the data reading speed.
[0072] GPU / CPU resource management performs GPU scheduling through tools such as CUDA and NVIDIA Docker to support multi-GPU parallel computing, while CPU management optimizes resource utilization through multi-threading and asynchronous computing. The system uses task queues and load balancing strategies to ensure efficient allocation of computing resources, and dynamically adjusts resource usage through monitoring tools to ensure the stability and efficiency of task execution.
[0073] Embodiment III
[0074] This embodiment is an embodiment of a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method described in Embodiment 1 are implemented.
[0075] Embodiment 4
[0076] This embodiment is an embodiment of a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in Embodiment 1 are implemented.
[0077] In the specific content of the above specific implementation manners, each technical feature can be combined arbitrarily without contradiction. For the sake of concise description, not all possible combinations of the above technical features are described. However, as long as the combinations of these technical features do not exist in contradiction, they should be considered as the scope described in this specification.
[0078] Obviously, the above embodiments of the present invention are merely examples for clearly explaining the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made on the basis of the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A method for fusion of infrared and visible light images based on deep learning, characterized in that: The following steps are involved: S1. Image preprocessing: preprocess the input infrared and visible light images, including noise removal and contrast enhancement; S2. Feature extraction: CNN is used to extract features of infrared and visible light images; S3. Feature fusion: fuse the acquired features separately; S4. Image reconstruction: reconstructing the fused feature image into a fused image; S5. Object detection and evaluation: Perform object detection on the fused image and evaluate the accuracy.
2. The infrared and visible light image fusion method based on deep learning according to claim 1, characterized in that: The step S1 comprises: S11. Input a pair of infrared and visible light images and align them to ensure that the image sizes are consistent; S12. The visible light image is enhanced by using the MSRCR algorithm to improve the contrast and color fidelity of the night image; S13. Use DDE algorithm to process infrared images to improve image contrast and optimize image noise.
3. The infrared and visible light image fusion method based on deep learning according to claim 2, characterized in that: The step S12 includes: firstly decomposing the brightness and reflection information of the image; then applying multi-scale filtering to perform image processing; and finally enhancing the visual effect of the image by using color restoration technology.
4. The infrared and visible light image fusion method based on deep learning according to claim 1, characterized in that: The step S2 comprises: S21. Use the first convolutional layer of ResNet101 pre-trained on ImageNet to extract basic image features; S22. Using an adaptive fusion algorithm to synthesize the features of the two images; S23. The extracted features are input into two parallel processing lines, namely global feature extraction and local feature extraction; in the local branch, two convolution layers are included in two series, a 3*3 convolution kernel is used, and the activation function is ReLU; in the global branch, two convolution layers are included in two series, a 5*5 convolution kernel is used, and the activation function is ReLU.
5. The infrared and visible light image fusion method based on deep learning according to claim 4 is characterized in that: In step S3, the acquired features are fused separately, wherein the maximum fusion strategy is adopted in the local feature extraction branch, and the average fusion strategy is adopted in the global feature extraction branch; finally, the features extracted in the global feature extraction and the features extracted in the local feature extraction are fused using the addition fusion strategy.
6. The infrared and visible light image fusion method based on deep learning according to claim 5, characterized in that: Step S4 includes: performing a layer of convolution reconstruction on the fused features and outputting a fused image, wherein the reconstruction uses a 1*1 convolution kernel, the activation function is ReLU, and the loss function uses mean square error.
7. A deep learning-based infrared and visible light image fusion system, characterized in that: include: Image preprocessing module: used to preprocess the input infrared and visible light images, including noise removal and contrast enhancement; Feature extraction module: used to extract features of infrared and visible light images using CNN; Feature fusion module: used to fuse the acquired features separately; Image reconstruction module: used to reconstruct the fused feature image into a fused image; Target detection and evaluation module: used to perform target detection on the fused image and evaluate the accuracy.
8. The infrared and visible light image fusion system based on deep learning according to claim 7, characterized in that: When performing image preprocessing, the image preprocessing module first obtains a pair of infrared and visible light images, and performs registration to ensure that the image sizes are consistent; then the MSRCR algorithm is used to enhance the visible light image to improve the contrast and color fidelity of the night image; finally, the DDE algorithm is used to process the infrared image to improve the image ratio and optimize the image noise; The MSRCR algorithm includes first decomposing the brightness and reflectance information of the image; then applying multi-scale filtering to process the image; and finally enhancing the visual effect of the image through color restoration technology.
9. The infrared and visible light image fusion system based on deep learning according to claim 7, characterized in that: The feature extraction module, when performing feature extraction, includes using the first convolution layer of ResNet101 pre-trained on ImageNet to extract basic image features; using an adaptive fusion algorithm to synthesize the features of two images; inputting the extracted features into two parallel processing lines, namely global feature extraction and local feature extraction; in the local branch, including two serial convolution layers, using a 3*3 convolution kernel, and an activation function of ReLU; in the global branch, including two serial convolution layers, using a 5*5 convolution kernel, and an activation function of ReLU.
10. The infrared and visible light image fusion system based on deep learning according to claim 9, characterized in that: When performing image fusion, the feature fusion module fuses the acquired features separately, wherein the maximum fusion strategy is adopted in the local feature extraction branch, and the average fusion strategy is adopted in the global feature extraction branch; finally, the features extracted in the global feature extraction and the features extracted in the local feature extraction are fused using the addition fusion strategy.
Citation Information
Cited By
Multispectral fusion imaging camera module and image processing method
CN121391636A
A multispectral fusion imaging camera module and image processing method
CN121391636B