Rapid infrared and visible light image fusion method based on adaptive saliency map
Through the adaptive significant graph structure and mean filtering decomposition method, combined with infrared and visible light information for image fusion, the problem of complex information loss and calculation in existing algorithms is solved, and efficient and real-time image fusion effect is achieved.
Patent Information
- Application Number
- CN202510447202.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing infrared and visible light image fusion algorithm ignores visible light information during significance detection, resulting in the loss of a large amount of visible light information in the fused image. The algorithm framework is complex, the calculation amount is large, and the real-time performance is poor.
Adaptive significant graph construction method is adopted, infrared images are dominated by combining visible light information, and the image is decomposed into the basic layer and the detail layer through mean filtering, and the maximum value and weighted fusion rules are used to fuse, simplifying the algorithm framework and improving computing efficiency.
It has achieved outstanding infrared targets and rich visible background information. It also has high real-time performance, conforms to human visual characteristics and excellent fusion image quality.
Smart Images

Figure CN120374416A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of infrared and visible light image processing, and particularly relates to a fast infrared and visible light image fusion method based on an adaptive saliency map. Background Technique
[0002] Image fusion is an important branch of information fusion, that is, information fusion processing is performed on the data of the same scene collected by multiple image sensors to provide richer information than a single image source. Among them, the fusion of infrared images and visible light images is one of the most studied fusion technologies. Due to the characteristics of radiation imaging, infrared images have extremely strong anti-interference ability in imaging, and it is very easy to distinguish the thermal radiation information of the target position. Visible light images are rich in color and details, but are easily affected by complex environments such as smoke and exposure, resulting in poor imaging effects. At this time, it is necessary to rely on infrared image information. Studying the infrared and visible light image fusion technology helps to combine their complementary advantages to obtain more effective imaging information.
[0003] In recent years, various infrared and visible light fusion algorithms have emerged in an endless stream. Among them, the fusion method of multi-scale decomposition has always been the research focus in the fusion field. Multi-scale decomposition is to decompose the source image into different feature layers and determine the fusion method according to the feature layers, which is a very important method in the fusion technology. Gonzalo et al. successfully applied wavelet transformation (WT) to image fusion, overcoming the problems of correlation and data redundancy of multi-scale sub-images, but having the problems of lack of translational invariance and limited extraction of directional information. Later, fusion methods based on contourlet transform, curvelet transform, etc. emerged one after another, improving some problems of wavelet transform fusion. At the same time, many scholars proposed methods of using image filters for multi-scale decomposition, such as cross bilateral filtering, guided filtering, hybrid decomposition of bilateral filtering and Gaussian filtering, etc., all of which achieved good fusion effects. Image saliency detection has always been one of the research hotspots in the field of image processing, and can extract the salient regions of images. In 2017, Ma combined the saliency algorithm with multi-scale decomposition to obtain a fusion image with clearer infrared targets. An proposed an improved saliency algorithm and applied it to the multi-scale decomposition framework, and the obtained fusion image introduced fewer artifacts, had prominent targets and high contrast. Lin proposed a three-scale decomposition framework based on the rolling guidance filter and saliency detection method, achieving good results and expanding new ideas for the research of image fusion.
[0004] However, most fusion algorithms only consider the salient information of infrared images when using the saliency detection method, while ignoring visible light images, resulting in a large amount of visible light information being lost in the fusion image. And although most algorithms can achieve good fusion effects, the algorithm framework is too complex and the computational amount is huge, resulting in poor real-time performance of the algorithm. Summary of the Invention
[0005] To solve the above-mentioned technical problems existing in the prior art, the present invention provides a fast infrared and visible light image fusion method based on an adaptive saliency map, including the following steps: adaptive saliency map construction, source image decomposition, image fusion strategy, and image reconstruction. The image fusion strategy includes base layer fusion and detail layer fusion.
[0006] Furthermore, the construction of the adaptive saliency map is dominated by the infrared image and combines the visible light image information to guide the fusion of the base layer in the fusion framework; the calculation process is as follows:
[0007] First, use Equation (1) and Equation (2) to calculate the saliency map S1 and the saliency map S2
[0008]
[0009] where Iir and Ivis represent the source infrared image and the source visible light image respectively, and S1 and S2 represent the calculated saliency maps; then calculate the maximum value of S1 and S2, and then multiply it by the source infrared image Iir to obtain the primary saliency map SM, which is expressed as:
[0010] SM = max(S1, S2) * Iir#(3)
[0011] where max(·) represents the maximum value comparison operator, and Iir represents the source infrared image;
[0012] Finally, normalize the result to obtain the normalized saliency map S, which is expressed as:
[0013]
[0014] Furthermore, the source image decomposition uses mean filtering as the filter for image decomposition, and decomposes the image into a base layer and a detail layer. The process of obtaining different feature layers of the image from the source image is as follows:
[0015] Bir = MF(Iir), Bvis = MF(Ivis) (5)
[0016] Dir = Iir - Bir, Dvis = Ivis - Bvis (6)
[0017] where Bir and Bvis represent the obtained infrared base layer and visible light base layer respectively, Dir and Dvis represent the obtained infrared detail layer and visible light detail layer respectively, Iir and Ivis represent the source infrared image and the source visible light image respectively, and MF(·) represents mean filtering.
[0018] Furthermore, the basic layer fusion selects an adaptive saliency map to guide the fusion of the basic layer, and the fusion strategy is as follows:
[0019] Bf = S * Bir + (1 - S) * Bvis #(7)
[0020] where S is the normalized saliency map, Bir and Bvis are the basic layer images of infrared and visible light respectively, and Bf is the fused basic layer image.
[0021] Furthermore, the detail layer fusion uses a fusion rule that combines the maximum value and the weighted method, and the fusion strategy is as follows:
[0022] First, the maximum value strategy is used to obtain the detail maximum value, which is expressed as:
[0023] DM = max(Dvis, Dir) #(8)
[0024] where DM represents the obtained detail maximum value, max(·) represents the maximum value comparison operator, and Dir and Dvis represent the infrared detail layer and the visible light detail layer respectively.
[0025] Then, the weight coefficient of the detail layer is calculated, which is expressed as:
[0026]
[0027] where WDvis represents the weight coefficient of the visible light detail layer, and u is a constant factor set to 0.00001;
[0028] Finally, the fused detail layer is calculated, which is expressed as:
[0029] Df = WDvis * Dvis + (1 - WDvis) * DM #(10)
[0030] where Df is the obtained fused detail layer image.
[0031] Furthermore, the fused image obtained by the image reconstruction is:
[0032] F = Bf + Df #(11)
[0033] where F is the final fused image, and Bf and Df are the fused basic layer image and the detail layer image respectively.
[0034] The present invention proposes an adaptive saliency calculation that takes the infrared image as the dominant and combines the visible light information for use in the fusion of the feature layer. The framework of the algorithm is very simple and the calculation efficiency is high. According to the comparative analysis of the experimental results, the fusion effect of this method has prominent infrared targets, rich visible light background information, and high real-time performance at the same time. Description of the Drawings
[0035] Figure 1 It is a framework diagram of a fast infrared and visible light image fusion method based on an adaptive saliency map;
[0036] Figure 2 They are the source infrared, visible light images, and the adaptive saliency map;
[0037] Figure 3 They are the source image and the basic layer and detail layer diagrams decomposed from it;
[0038] Figure 4 It is a comparison diagram of the first group of experimental results;
[0039] Figure 5 It is a comparison diagram of the second group of experimental results;
[0040] Figure 6 It is a comparison diagram of the third group of experimental results. Specific implementation manner
[0041] The present invention will be further described below in conjunction with the accompanying drawings.
[0042] As Figure 1 shown, the fast infrared and visible light image fusion method based on an adaptive saliency map of the present invention uses mean filtering to divide each source image into a basic layer and a detail layer, and then uses different fusion strategies for fusion. An adaptive saliency map designed with the infrared image as the dominant and combined with the visible light image information is used to guide the fusion of the basic layer, and the fusion of the detail layer is guided by a method combining the maximum value fusion and the weighted fusion method. Finally, the fused basic layer and detail layer are reconstructed to obtain an infrared and visible light fusion image with better visual effects.
[0043] The fast infrared and visible light image fusion method based on an adaptive saliency map includes adaptive saliency map construction, source image decomposition, image fusion strategy, and image reconstruction. The specific implementation technical solutions are as follows:
[0044] Adaptive saliency map construction
[0045] Most algorithms directly use existing saliency detection algorithms for fusion, ignoring the different feature differences between infrared and visible light images themselves, resulting in the loss of a large amount of source image information. The present invention proposes an adaptive saliency map construction method with the infrared image as the dominant and combined with the visible light image information to guide the fusion of the basic layer in the fusion framework. The obtaining process is as follows:
[0046] First, use equations (1) and (2) to obtain the saliency map S1 and the saliency map S2
[0047]
[0048] Where Iir and Ivis represent the source infrared image and the source visible light image respectively, and S1 and S2 represent the obtained saliency maps. Then calculate the maximum value of S1 and S2, and then multiply it by the source infrared image Iir to obtain the primary saliency map SM, which is expressed as:
[0049] SM = max(S1, S2) * Iir #(3)
[0050] Where max(·) represents the maximum value comparison operator, and Iir represents the source infrared image.
[0051] Finally, normalize the result to obtain the normalized saliency map S, which is expressed as:
[0052]
[0053] As Figure 2 shown, the obtained adaptive saliency map has obvious infrared targets and can simultaneously reflect the rich environmental information of infrared and visible light. Using this map for the fusion of the basic layer can highlight the target while integrating more detailed information.
[0054] Source image decomposition
[0055] Since the efficiency of decomposition using mean filtering is high and the effect of feature decomposition is good, the present invention uses mean filtering as the filter for image decomposition, and decomposes the image into a basic layer and a detail layer. The process of obtaining different feature layers of the image from the source image is as follows:
[0056] Bir = MF(Iir), Bvis = MF(Ivis) (5)
[0057] Dir = Iir - Bir, Dvis = Ivis - Bvis (6)
[0058] Where Bir and Bvis represent the obtained infrared basic layer and visible light basic layer respectively, Dir and Dvis represent the obtained infrared detail layer and visible light detail layer respectively, Iir and Ivis represent the source infrared image and the source visible light image respectively, and MF(·) represents mean filtering. The source images (a, d), the basic layer images (b, e), and the detail layer images (c, f) are as Figure 3 shown.
[0059] Image fusion strategy. The image fusion strategy includes the fusion of the basic layer and the detail layer:
[0060] a) Basic layer fusion
[0061] The base layer contains most of the energy information of the image, representing the basic brightness of the image. The fusion of the base layer is a crucial step in the fusion algorithm, directly affecting the quality of the fusion result. In the present invention, the adaptive saliency map obtained in the previous step is selected to guide the fusion of the base layer, and the fusion strategy is as follows:
[0062] Bf = S * Bir+(1 - S) * Bvis #(7)
[0063] where S is the normalized saliency map, Bir and Bvis are the base layer images of the infrared and visible light respectively, and Bf is the fused base layer image.
[0064] b) Detail layer fusion
[0065] The detail layer contains rich detail information of the source images. The common maximum rule is prone to cause noise problems. The present invention adopts a fusion rule combining the maximum value and the weighted method, effectively retaining the detail information of the source images. The fusion strategy is as follows:
[0066] First, the maximum value strategy is used to obtain the detail maximum value, which is expressed as:
[0067] DM = max(Dvis, Dir) #(8)
[0068] where DM represents the obtained detail maximum value, max(·) represents the maximum value comparison operator, and Dir and Dvis represent the infrared detail layer and the visible light detail layer respectively.
[0069] Then, the weight coefficient of the detail layer is calculated, which is expressed as:
[0070]
[0071] where WDvis represents the weight coefficient of the visible light detail layer, and u is a constant factor, set to 0.00001.
[0072] Finally, the fused detail layer is calculated, which is expressed as:
[0073] Df = WDvis * Dvis+(1 - WDvis) * DM #(10)
[0074] where Df is the obtained fused detail layer image.
[0075] Image reconstruction
[0076] After the above calculation process, the finally obtained fused image is:
[0077] F = Bf + Df #(11)
[0078] where F is the final fused image, and Bf and Df are the fused base layer image and detail layer image respectively.
[0079] To verify the effectiveness of the method proposed in the present invention, three sets of registered image data in the TNO dataset were selected as experimental images and compared with a variety of classical fusion algorithms, including GFF (guided filtering fusion), TIF (Two-scale Image Fusion), HMSD (hybrid multi-scale decomposition), VSMWLS (visual saliency map weighted least square), FPDE (Fourth Order Partial Differential Equation), IFEVIP (Infrared Feature Extraction and Visual Information Preservation), and MGFF (multi-scale guided filtering fusion). The experiment was run in Matlab R2021a with a hardware configuration of 11th Gen Intel(R) Core(TM) i5-11320H@3.20GHz, 16GB to obtain the experimental results.
[0080] As Figures 4 - 6 shown, from the subjective comparison of the experimental results of these three groups, it can be seen that the fusion results of the method of the present invention can retain relatively prominent infrared target information, and the contrast between the infrared target and the background is very high, which is more obvious in the experimental results of the first and second groups. In addition, the method of the present invention can retain richer visible light background and detail information, and the obtained fusion image is more in line with human vision. For example, in the second and third group experiments, the restoration degree of the sky is higher, and the background information is more biased towards visible light information. In summary, the method of the present invention has good performance in enhancing the target contrast and restoring the visible light background information, and is more in line with human vision.
[0081] To qualitatively evaluate the effectiveness of this method, five objective evaluation indicators of image fusion were used to quantitatively compare the experimental results. These include mutual information (MI), structural similarity index measure (SSIM), standard deviation (SD), and a human-inspired perception-based metric (Chen-Varshney metric, Qcv). Among them, the larger the MI, SSIM, and SD indicators, the better the fusion quality, and the smaller the Qcv indicator, the better the fusion quality.
[0082] Comparison of objective evaluation indicators for the experimental results of the first group in Table 1
[0083]
[0084] Comparison of objective evaluation indicators for the experimental results of the second group in Table 2
[0085]
[0086] Comparison of objective evaluation indicators for the experimental results of the third group in Table 3
[0087]
[0088] The comparison of objective evaluation indicators for the experimental results of each group is shown in Tables 1, 2, and 3. It can be seen from the tables that the performance of this method is very high in terms of the MI, SSIM, SD, and Qcv indicators, and the results are all in the optimal or sub-optimal positions. Especially in the Qcv indicator, the indicators of the experimental results of the three groups have reached the optimal. This shows that the fused images obtained by this method have high contrast, richer source image information, are more in line with the human visual characteristics, and can adapt to the applications of various fusion scenarios.
[0089] Real-time analysis
[0090] Comparison of average running times of different algorithms in Table 4
[0091]
[0092] The average running times of each algorithm in the three groups of experiments are calculated as shown in Table 4, and the optimal value of the running time is in bold. It can be seen from the average running time that the algorithm proposed by this method is significantly superior to the other seven compared algorithms in terms of real-time performance, and the real-time performance is optimal. At the same time, it can be seen from the above that the fusion performance of the algorithm is also relatively superior, and it can be applied in occasions with high real-time performance, which is convenient for algorithm transplantation and application. The experimental results show that the fusion algorithm proposed by this method can better retain the effective information of infrared and visible light source images, and has relatively superior real-time performance, which is a relatively reliable fusion algorithm.
Claims
1. A fast infrared and visible image fusion method based on an adaptive saliency map, characterized in that It includes the following steps: adaptive saliency map construction, source image decomposition, image fusion strategy, and image reconstruction. The image fusion strategy includes base layer fusion and detail layer fusion.
2. The fast infrared and visible light image fusion method based on an adaptive saliency map according to claim 1, characterized in that: The adaptive saliency map construction is dominated by the infrared image and combines visible light image information to guide the fusion of the base layer in the fusion framework. The calculation process is as follows: First, use Equation (1) and Equation (2) to calculate the saliency map S1 and the saliency map S2 where Iir and Ivis respectively represent the source infrared image and the source visible light image, and S1 and S2 respectively represent the calculated saliency maps. Then calculate the maximum value of S1 and S2, and then multiply it by the source infrared image Iir to obtain the primary saliency map SM, which is expressed as: SM = max(S1, S2) * Iir #(3) where max(·) represents the maximum value comparison operator, and Iir represents the source infrared image; Finally, normalize the result to obtain the normalized saliency map S, which is expressed as:
3. The fast infrared and visible light image fusion method based on an adaptive saliency map according to claim 2, characterized in that: The source image decomposition uses mean filtering as the filter for image decomposition, and decomposes the image into a base layer and a detail layer. The process of obtaining different feature layers of the image from the source image is as follows: Bir = MF(Iir), Bvis = MF(Ivis) (5) Dir = Iir - Bir, Dvis = Ivis - Bvis (6) where Bir and Bvis respectively represent the obtained infrared base layer and visible light base layer, Dir and Dvis respectively represent the obtained infrared detail layer and visible light detail layer, Iir and Ivis respectively represent the source infrared image and the source visible light image, and MF(·) represents mean filtering.
4. The fast infrared and visible light image fusion method based on an adaptive saliency map according to claim 3, characterized in that: The base layer fusion selects the adaptive saliency map to guide the fusion of the base layer, and the fusion strategy is: Bf = S * Bir + (1 - S) * Bvis #(7) where S is the normalized saliency map, Bir and Bvis are the base layer images of infrared and visible light respectively, and Bf is the fused base layer image.
5. The fast infrared and visible light image fusion method based on an adaptive saliency map according to claim 4, wherein: The detail layer fusion uses a fusion rule that combines the maximum value and the weighted method. The fusion strategy is as follows: First, use the maximum value strategy to calculate the detail maximum value, which is expressed as: DM = max(Dvis, Dir) #(8) where DM represents the calculated detail maximum value, max(·) represents the maximum value comparison operator, and Dir and Dvis represent the infrared detail layer and the visible light detail layer. Then calculate the weight coefficient of the detail layer, which is expressed as: where WDvis represents the weight coefficient of the visible light detail layer, and u is a constant factor, set to 0.00001; Finally, calculate the fused detail layer, which is expressed as: Df = WDvis * Dvis + (1 - WDvis) * DM #(10) where Df is the obtained fused detail layer image.
6. The fast infrared and visible light image fusion method based on an adaptive saliency map according to claim 5, characterized in that: The image reconstruction obtains the fused image as: F = Bf + Df #(11) where F is the final fused image, and Bf and Df are the fused base layer image and detail layer image respectively.