Infrared and Visible Light Image Fusion Method and Device

Through the three-scale decomposition and sparse representation, the problem of noise processing in infrared and visible light images is solved, and efficient fusion and denoising of noise-free and noise-containing images is achieved, improving the fusion performance and efficiency.

CN114862710BActive Publication Date: 2025-07-11ARMY ENG UNIV OF PLA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210454565.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-26
Publication Date
2025-07-11
Estimated Expiration
2042-04-26

AI Technical Summary

Technical Problem

The prior art is difficult to effectively handle noise-containing infrared and visible image fusion, and the traditional methods do not perform well when noise-free and noisy images fusion.

Method used

The three-scale decomposition and sparse representation method is used to decompose the image into a base layer and a detail layer using a rolling guide filter, and the basic layer is further decomposed into the basic structure layer and the basic texture layer through the structure-texture decomposition model, and the layers are pre-fused using different fusion rules to finally reconstruct the fusion image.

Benefits of technology

It realizes effective fusion of noise-free and noise-containing images, maintains image details and effectively denoise, and improves fusion performance and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114862710B_ABST
    Figure CN114862710B_ABST
Patent Text Reader

Abstract

An embodiment of this specification provides an infrared and visible light image fusion method and apparatus. Among them, the method includes: decomposing a source image into a base layer and a detail layer by using a rolling guidance filter, where the detail layer includes most details and external noise, and the base layer includes: residual details and energy; decomposing the base layer again based on a constructed structure-texture decomposition model to decompose the base layer into a base structure layer and a base texture layer; using different fusion rules corresponding to each layer to perform pre-fusion on the detail layer, the base structure layer, and the base texture layer; and obtaining a fused image by reconstructing the three pre-fusion layers. It can not only effectively handle the fusion problem of noisy images, but also has good fusion performance for noiseless images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to the field of image fusion technology, and particularly to an infrared and visible light image fusion method and device. Background Art

[0002] In recent years, unmanned aerial vehicles (UAVs) have played an increasingly important role in many fields due to their high flexibility, low cost, and easy operation. They are often used to perform tasks such as reconnaissance and situation monitoring. However, with the diversification and complexity of actual needs, a single imaging sensor is limited by its own physical imaging principle and can only perceive some targets or information, making it difficult to complete detection and recognition tasks in various target backgrounds. Therefore, by fusing the image data of multiple sensors to achieve complementarity between different sensor data, a more intuitive, reliable, and comprehensive target or scene can be obtained, which can then provide strong data support for subsequent tasks such as feature extraction, target recognition, and detection, facilitating more reasonable decision-making.

[0003] Most image fusion research is carried out at the pixel level. Among them, according to the different image representations and fusion processes, image fusion can be roughly divided into four categories: methods based on the spatial domain, methods based on the transform domain, methods based on neural networks, and methods based on dictionary learning. Currently, many image fusion methods mostly assume that the source images are noise-free and rarely study the situation of noise perturbation. However, due to the influence of many factors such as imaging equipment and shooting environment, the images obtained in actual tasks inevitably contain noise. When directly performing image fusion, the noise and detail information in the source images may be processed equally, resulting in poor fusion effects. For the fusion of noisy images, a step-by-step method of first fusing and then denoising or first denoising and then fusing is usually adopted, that is, combining the denoising algorithm and the fusion algorithm to achieve fusion denoising. However, from the perspective of efficiency and fusion performance, the step-by-step approach may not be the best choice.

[0004] To solve this problem, some synchronous fusion denoising methods have emerged. Among them, the method based on Sparse Representation (SR) can solve both image fusion and denoising problems simultaneously. An effective SR-based image denoising algorithm has been proposed in the prior art. This method establishes the connection between the noise standard deviation and the sparse reconstruction error, and can effectively achieve image denoising in a parameter-adaptive manner. Subsequently, some scholars proposed a joint image denoising and fusion algorithm based on SR, realizing the synchronization of fusion and denoising. To further improve the fusion performance and solve the problem of long fusion time, an adaptive SR method is proposed for image fusion and denoising. This method pre-trains multiple feature dictionaries according to the gradient features of training samples, and then adaptively selects a suitable dictionary according to the image gradient features, which can effectively achieve image fusion and denoising. In the prior art, feature clustering is also carried out by introducing kernel local regression weights, and a multi-modal image fusion method based on dictionary learning is designed. This method can effectively suppress noise generation and has good fusion and denoising performance. In addition, to reduce the damage to the edge information of the image that may be caused when directly denoising the noise source image processed by SR, a medical image fusion method based on sparse low-rank dictionary learning is also proposed. In this method, the source image is regarded as the superposition of coarse-scale and fine-scale components, effectively solving the above problems. However, the above research cannot take into account both noiseless and noisy image fusion at the same time, which is an urgent problem to be solved at present. Summary of the Invention

[0005] The purpose of the present invention is to provide an infrared and visible light image fusion method and device, aiming to solve the above problems in the prior art.

[0006] The present invention provides an infrared and visible light image fusion method, including:

[0007] Using a rolling guidance filter to decompose the source image into a base layer and a detail layer, where the detail layer includes most details and external noise, and the base layer includes residual details and energy;

[0008] Based on the constructed structure-texture decomposition model, decomposing the base layer again, and decomposing the base layer into a base structure layer and a base texture layer;

[0009] Using different fusion rules corresponding to each layer to pre-fuse the detail layer, the base structure layer, and the base texture layer;

[0010] Obtaining a fused image by reconstructing the three-layer pre-fusion layer.

[0011] The present invention provides an infrared and visible light image fusion device, including:

[0012] The first decomposition module is used to decompose the source image into a base layer and a detail layer by using a rolling guidance filter. Among them, the detail layer includes most details and external noise, and the base layer includes residual details and energy;

[0013] The second decomposition module is used to decompose the base layer again based on the constructed structure-texture decomposition model, and decompose the base layer into a base structure layer and a base texture layer;

[0014] The pre-fusion module is used to perform pre-fusion on the detail layer, the base structure layer and the base texture layer by using different fusion rules corresponding to each layer;

[0015] The fusion module is used to obtain a fused image by reconstructing three pre-fused layers.

[0016] Adopting the embodiment of the present invention can not only effectively handle the fusion problem of noisy images, but also has good fusion performance for noiseless images. Description of the Drawings

[0017] In order to more clearly illustrate the technical solutions in one or more embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0018] Figure 1 is a flowchart of the infrared and visible light image fusion method according to the embodiment of the present invention;

[0019] Figure 2 is a schematic diagram of the image structure-texture decomposition according to the embodiment of the present invention;

[0020] Figure 3 is a structural diagram of the fusion denoising method according to the embodiment of the present invention;

[0021] Figure 4 is a schematic diagram of the three-scale decomposition image with Gaussian white noise of σ = 30 added to the input images in the second and fourth rows according to the embodiment of the present invention;

[0022] Figure 5 is a schematic diagram of the detail layer fusion process according to the embodiment of the present invention;

[0023] Figure 6 is a schematic diagram of five pairs of source images according to the embodiment of the present invention;

[0024] Figure 7 is a schematic diagram of the results of detail layer fusion and denoising with different C values (noise level 20) according to the embodiment of the present invention;

[0025] Figure 8 It is a schematic diagram of the fusion result of the noise-free infrared and visible light grayscale images according to an embodiment of the present invention;

[0026] Figure 9 It is a schematic diagram of the fusion result of the noise-free infrared and visible light color images according to an embodiment of the present invention;

[0027] Figure 10 It is a schematic diagram of the fusion result of the noisy infrared and visible light grayscale images according to an embodiment of the present invention;

[0028] Figure 11 It is a schematic diagram of the fusion result of the noisy infrared and visible light color images according to an embodiment of the present invention;

[0029] Figure 12 It is a schematic diagram of the infrared and visible light image fusion device according to an embodiment of the present invention. Detailed implementation manners

[0030] To improve the processing effect of the noise source image, an embodiment of the present invention proposes an infrared and visible light image fusion method and device based on three-scale decomposition and sparse representation. The source image is decomposed into a base layer and a detail layer by using a rolling guidance filter, and the maximum sparse reconstruction error parameter is adaptively determined according to the image features, so as to simultaneously realize the fusion and denoising of the detail components; a structure-texture decomposition model is constructed to decompose the base layer again to make full use of the details and energy in the base components, and different fusion rules are used to fuse the structure layer and the texture layer. Finally, the fused image is obtained by reconstructing the detail, base structure, and base texture layers. Experimental results show that the embodiment of the present invention can not only effectively process the fusion problem of noisy images, but also has good fusion performance for noise-free images.

[0031] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the following will clearly and completely describe the technical solutions in one or more embodiments of this specification with reference to the accompanying drawings in one or more embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this document.

[0032] Method embodiment

[0033] According to an embodiment of the present invention, an infrared and visible light image fusion method is provided. Figure 1 It is a flowchart of the infrared and visible light image fusion method according to an embodiment of the present invention. As Figure 1 shown, the infrared and visible light image fusion method according to an embodiment of the present invention specifically includes:

[0034] Step 101: Decompose the source image into a base layer and a detail layer using a rolling guidance filter. Among them, the detail layer includes most details and external noise, and the base layer includes residual details and energy;

[0035] Step 102: Decompose the base layer again based on the constructed structure-texture decomposition model, and decompose the base layer into a base structure layer and a base texture layer;

[0036] Step 103: Use different fusion rules corresponding to each layer to pre-fuse the detail layer, the base structure layer, and the base texture layer;

[0037] Step 104: Obtain the fused image by reconstructing the three-layer pre-fusion layer.

[0038] In step 101, the specific process of decomposing the source image into a base layer and a detail layer using a rolling guidance filter includes:

[0039] Perform base and detail decomposition on the source image according to Formula 1 and Formula 2, and obtain the detail layer of image I by solving n where I

[0040]

[0041] is the nth source image, n ∈ {1, 2,..., N}, n and represents the base layer of I n .

[0042] In step 102, the specific process of decomposing the base layer into a base structure layer and a base texture layer again based on the constructed structure-texture decomposition model includes:

[0043] According to Formula 3 and Formula 4, decompose based on the structure-texture decomposition model to obtain the base structure layer and the base texture layer

[0044]

[0045] where and λ are the scale parameter and the smoothing parameter respectively.

[0046] In step 103, the specific process of pre-fusing the detail layer, the base structure layer, and the base texture layer using different fusion rules corresponding to each layer includes:

[0047] Based on the SR method, by establishing the connection between the sparse reconstruction error and the noise standard deviation, fusion denoising is achieved, and pre-fusion of the detail layer is carried out;

[0048] The weighted average technology based on the Visual Saliency Map (VSM) is used to pre-fuse the basic structure layer;

[0049] The principal component analysis method is used to pre-fuse the basic texture layer.

[0050] Among them, based on the SR method, by establishing the connection between the sparse reconstruction error and the noise standard deviation, fusion denoising is achieved, and the specific steps of pre-fusing the detail layer include:

[0051] The detail layer of the training data is generated by a rolling guidance filter. Blocks of size 8×8 are collected from the detail images to construct the final training set, and the dictionary D is obtained using the KSVD algorithm;

[0052] For each source image, a block of size 8×8 is taken and normalized. By solving the following objective function, the orthogonal matching pursuit algorithm (OMP) is used to generate the SR coefficients of the detail layer:

[0053]

[0054] Where, is the k-th small block of the source image I n and is the corresponding sparse vector. is the maximum sparse reconstruction error, σ is the Gaussian standard deviation, and C > 0 is a parameter that controls when σ > 0;

[0055] The "absolute value - maximum" scheme is used to generate the fused sparse coefficients:

[0056]

[0057] The fused detail vector is linearly represented as follows:

[0058]

[0059] Each is reshaped into an 8×8 small block and then arranged according to the initial position to obtain the pre-fused detail layer;

[0060] Among them, the specific steps of using the weighted average technology based on the Visual Saliency Map (VSM) to pre-fuse the basic structure layer include:

[0061] Construct the VSM. Set I p to represent the intensity value of a pixel p in the image I. The saliency value V(p) of the pixel p is defined as

[0062]

[0063] Among them, N represents the total number of pixels in I, j represents the pixel intensity, M j represents the number of pixels with intensity equal to j, L represents the number of gray levels. If two pixels have the same intensity value, their saliency values are equal;

[0064] Then, V(p) is normalized to [0, 1];

[0065] Let V1 and V2 represent the VSMs of different source images respectively, and represent the basic structure layer images of different source images. The final pre-fused image F of the basic structure layer is obtained through weighted averaging b,s :

[0066]

[0067] Among them, the weight W b is defined as:

[0068]

[0069] Among them, the pre-fusion of the basic texture layer using the principal component analysis method specifically includes:

[0070] Taking the basic texture images of the visible light and infrared images and as the column vectors of matrix γ, then taking each row as a reference and each column as a variable, the covariance matrix C of γ is obtained;

[0071] Calculating the eigenvalues λ1, λ2 of C and the corresponding eigenvectors and

[0072] Finding the largest eigenvalue from the two eigenvalues, that is, λ max = max(λ1, λ2), and taking the eigenvector corresponding to λ max as the maximum eigenvector φ max , calculating the principal components P1 and P2 corresponding to φ max , and normalizing their values:

[0073]

[0074] Taking the principal components P1 and P2 as weights, they are fused into the final pre-fused image F of the basic texture layer b,t :

[0075]

[0076] In step 104, obtaining the fused image by reconstructing the three-layer pre-fusion layer specifically includes:

[0077] The final fused image F according to formula 15 is:

[0078] F = F d + F b,s + F b,t Formula 15.

[0079] The above technical solutions of the embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings.

[0080] The embodiment of the present invention proposes an infrared and visible light image fusion method based on three-scale decomposition and sparse representation. SR can make full use of the advantages of the method based on the spatial domain in terms of fusion performance and computational efficiency, handle the noise problem in image fusion, and construct a three-scale decomposition model. The image is decomposed into a basic component and a detail component through rolling guidance filtering, and then the basic layer is processed through structure-texture decomposition to effectively extract the detailed texture information in the basic component, so as to improve the ability of the fused image to express detailed information.

[0081] The key theories involved above will be described in detail below.

[0082] 1. Rolling guidance filter (RGF):

[0083] RGF has the characteristics of scale perception and edge preservation. Therefore, it not only has good noise removal ability, but also can maintain the structure and edge characteristics of the source image. RGF includes two main steps: small structure removal and edge restoration.

[0084] The first step is to use a Gaussian filter to remove small structures. The filtered image G from the input image I can be expressed as:

[0085] G = Gaussian(I, σ s ) (1)

[0086] where Gaussian(I, σ s ) represents Gaussian filtering with the standard deviation σ s as the scale parameter. This filter can remove structures smaller than σ s in the scale space theory.

[0087] The second step is to use a guided filter for iterative edge restoration because it has high computational efficiency and good edge preservation characteristics. This step is a process of iterative update of a restored image J t and the initial image J 1is the Gaussian smoothed image G. The t-th iteration can be expressed as

[0088]

[0089] where is the guided filter, where I, σ s i.e., the parameters in Eq. (1), J t is the guidance image, σ r controls the distance weights. In our method, we set σ r = 0.05. RGF is completed by combining Eq. (1) and Eq. (2), and can be simply expressed as

[0090] u = RGF(I, σ s , σ r , T) (3)

[0091] where T is the number of iterations and u is the filtered output.

[0092] 2. Structure-texture decomposition

[0093] In structure-texture decomposition, the image I can be decomposed into I = S + T, i.e., regarded as the superposition of the structure component S and the texture component T. In the structure-texture decomposition model, the structure component is mainly composed of some non-repetitive edges and smooth energy regions, while the texture component is repetitive oscillatory information and noise. For this reason, the local features of the structure texture are defined, and the interval gradient operator is proposed. The interval gradient filter (IGF) of the pixel p in the discrete signal I is defined as which can be expressed as

[0094]

[0095] where Ω is the region near the central pixel p, and are the left and right shear one-dimensional Gaussian functions respectively, where

[0096]

[0097] where is the shear exponential weighting function, is the scale parameter.

[0098]

[0099] Define k r and k l as the normalization coefficients, expressed as

[0100]

[0101] Different from the traditional forward differentiation, the interval gradient measures the weighted average difference between the left and right sides of a pixel. To obtain structural information, there should be no texture information in the structural area. The local window Ω(p) in the structural area can only contain increasing or decreasing information and cannot contain repeated textures.

[0102] Refine the gradient of the input signal I into the following interval gradient

[0103]

[0104] where is the scaled gradient, ω p is the scaling weight, expressed as:

[0105]

[0106] where ε is a small constant, usually set to 10 -4 .

[0107] To eliminate the residual oscillation signal in the signal and obtain the structural component of the input signal, based on the guided filter, the temporary filtering result is obtained from the regression gradient

[0108]

[0109] where N p represents the number of pixels in I, and I0 is the minimum value (the leftmost) in I. The optimal coefficients a p and b p of the one-dimensional signal guided filter can be obtained by solving Equation (13).

[0110]

[0111] where ω n represents the Gaussian weight, is the scaling parameter defined in Equation (5), λ represents the smoothing parameter, and the coefficients a p and b p are defined as:

[0112]

[0113] After obtaining the coefficients a p and b p , the structural component of the signal I can be obtained through .

[0114] For a two-dimensional image signal, use the interval gradient filtering of the one-dimensional signal alternately in the x and y directions, and converge in an iterative manner to obtain the structural layer of the final image. From Figure 2As can be seen from the enlarged area of the middle image, almost all vibration and repetitive information is retained in the texture layer, while brightness and weak edge information is retained in the structure layer.

[0115] 3. Three-scale decomposition and sparse representation model:

[0116] To solve the problem of effectively retaining details while denoising, a new image fusion model is constructed as Figure 3 shown. Different from the traditional two-scale decomposition scheme, in order to better denoise and utilize the useful information of the base layer, first use RGF to decompose the source image into base and detail components. At this time, most details and external noise can be effectively retained in the detail layer, and the base layer contains residual details and energy; then perform structure-texture decomposition on the base layer, and finally the source image is decomposed into three components: details, base structure, and base texture. According to the characteristics of each layer, three established fusion rules are used to generate the pre-fusion of each layer. Among them, for the fusion of the detail layer, by establishing the connection between the sparse reconstruction error and the noise standard deviation, fusion denoising is effectively realized; for the base structure layer, a weighted average technique based on the Visual Saliency Map (VSM) is used for pre-fusion; for the base texture layer, the principal component analysis method is used for pre-fusion. Finally, the fusion result is obtained by reconstructing the three pre-fusion layers.

[0117] Decomposition model:

[0118] To specifically remove the noise attached to the detail layer of the image, first perform base and detail decomposition on the source image:

[0119]

[0120] where I n is the nth source image, n ∈ {1, 2,..., N}, represents the base layer of I n , and the detail layer of image I is obtained by solving n

[0121]

[0122] For the fusion of the base layer, the absolute value-maximum or average method is usually used, but the fusion results generated by these methods may deteriorate due to reduced contrast and edge degradation. However, multiple decompositions of the image still cannot well separate the base and detail information of the image, and multiple decompositions will inevitably increase the complexity of the reconstruction process, resulting in poor results. To solve this problem, a structure-texture decomposition model is introduced to perform decomposition to obtain its structure layer

[0123]

[0124] Among them, and λ are the scale parameter and the smoothing parameter respectively. The texture layer of

[0125]

[0126] can be generated in the following way:

[0127] Figure 4 Each group of images in Figure 4 shows that:

[0128] (1) Through the decomposition by RGF, most of the noise and details are retained in the detail layer. At the same time, it can be seen that the base layer still contains certain detail information.

[0129] (2) After the structure-texture decomposition process, the structure layer hardly contains local oscillation information, and the local area usually only contains intensity information or a few obvious edge structures.

[0130] (3) The structure and texture layers generated by the noise-free and noisy images are very similar, that is, the noise information is almost completely present in the detail layer.

[0131] Fusion rule:

[0132] According to the characteristics of the three parts, the present invention provides three different fusion rules in real time.

[0133] 1. Detail layer fusion:

[0134] The method based on SR can well achieve the fusion and denoising of the detail layer. It includes two steps: dictionary learning and sparse coefficient generation. In the first stage, the detail layer of the training data is generated by the rolling guidance filter of Equation (14). Blocks of size 8×8 are collected from the detail images to construct the final training set, and an over-complete dictionary D can be obtained by using the KSVD algorithm. In the second stage, blocks of size 8×8 are taken from each source image and normalized. By solving the following objective function, the SR coefficients of the detail layer are generated by using the Orthogonal Matching Pursuit (OMP) algorithm:

[0135]

[0136] In the formula is the source image In The k-th small block of is the corresponding sparse vector. is the maximum sparse reconstruction error, defined as

[0137]

[0138] where σ is the Gaussian standard deviation, and C > 0 is a parameter that controls when σ > 0. Then, the "absolute value - maximum" scheme is adopted to generate the fused sparse coefficients:

[0139]

[0140] The fused detail vector can be obtained from the following linear representation:

[0141]

[0142] Reshape each into an 8×8 small block, and then arrange them according to the initial position to obtain the fused detail layer, Figure 4 which is the fusion process of the sparse representation of the detail layer.

[0143] 2. Fusion of the basic structure layer:

[0144] Since the basic structure layer comes from the basic component of the source image, this layer contains less detail, as shown in Figure 5 the image in the fourth column. Therefore, the weighted average technique based on (Visual saliency map) VSM is used to fuse the basic structure layer F b,s .

[0145] The embodiments of the present invention adopt a method to construct the VSM, and let I P represent the intensity value of a pixel p in the image I. The saliency value V(p) of the pixel p is defined as

[0146]

[0147] where N represents the total number of pixels in I, j represents the pixel intensity, M j represents the number of pixels with intensity equal to j, and L represents the number of gray levels (preferably 256 in the embodiments of the present invention). If two pixels have the same intensity value, their saliency values are equal. Then, V(p) is normalized to [0, 1].

[0148] Let V1 and V2 represent the VSMs of different source images respectively, and represent the basic structure layer images of different source images. The final fused image of the basic structure layer is obtained by weighted average

[0149]

[0150] Where the weight W b is defined as

[0151]

[0152] 3. Fusion of the basic texture layer

[0153] Compared with the basic structure layer, the basic texture layer contains visually important information or image features, such as active information like edges, straight lines, and contours, which can reflect the main details of the original basic image. Therefore, the principal component analysis method is used to effectively detect these features.

[0154] Take the basic texture images of the visible light and infrared images and as the column vectors of matrix γ. Then, taking each row as a reference and each column as a variable, calculate the covariance matrix C of γ.

[0155] Calculate the eigenvalues λ1, λ2 of C and the corresponding eigenvectors and

[0156] Find the largest eigenvalue from the two eigenvalues, that is, λ max = max(λ1, λ2). Take the eigenvector corresponding to λ max as the maximum eigenvector φ max . Calculate the principal components P1 and P2 corresponding to φ max and normalize their values:

[0157]

[0158]

[0159] These principal components P1 and P2 are used as weights to fuse into the final basic texture image F b,t

[0160]

[0161] After obtaining these three pre-fusion components, the final fusion image F is:

[0162] F = F d + F b,s + F b,t (31)

[0163] The following is a detailed description of the experimental analysis and results.

[0164] 1. Experimental setup

[0165] The five pairs of source images used in the experiment can be obtained from the public website http: / / imagefusion.org / . As Figure 6 shown. And five recent methods, including ADF, FPDE, GTF, IFEVIP, and TIF, were selected and compared and verified in the same experimental environment. The entropy EN, edge information retention degree Q AB / F , the index Q proposed by Chen-Blum CB , mutual information MI, the index Q proposed by Wang et al. W , and the index Q proposed by Yang et al. Y were used to quantitatively evaluate the fusion results with 6 indexes.

[0166] 2. Parameter Settings

[0167] Here, the free parameter C in Equation (22) is mainly analyzed. Since the denoising process in the model only targets the detail component, for an intuitive analysis of the fusion denoising performance under different Cs, only the detail fusion results are analyzed, as Figure 7 shown. Taking the two pairs of images (a1, a2) and (b1, b2) in Figure 7 as examples, Gaussian noise with σ = 20 was added respectively to generate fusion images with different Cs.

[0168] For the fusion of images (a1) and (a2), it can be seen that: ① When C < 0.0035, the noise in the fused detail layer is relatively obvious Figure 7 (a3 - a7)], especially when C is far lower than 0.0035, the denoising effect is limited; ② When C = 0.0035, most of the noise has been eliminated; ③ When C > 0.0035, the noise is hardly noticeable, but at this time, the fused detail layer encounters an over-smoothing effect. Therefore, C = 0.0035 is an obvious demarcation value. From the perspective of detail protection, C = 0.0035 can achieve the best visual effect for the fusion of this pair of images. Therefore, for the fusion of grayscale images, the optimal value of C is 0.0035.

[0169] For the fusion of images (b1) and (b2), it can be seen that: ① When C < 0.002, some noise can be easily seen in the fused detail layer Figure 7 (b3 - b4)]; ② When C = 0.002, the noise is greatly reduced and suppressed, and at the same time, the details are well preserved Figure 7 (b5)]; ③ When C > 0.002, the denoising performance is good, but Figure 7 the details in (b6 - b10) are more or less damaged, and some fine edges are smoothed. As C increases, more details are missed [Compare Figure 7(b6 - b10)]. Therefore, considering the performance of both fusion and denoising comprehensively, for color image fusion, the optimal value of C is 0.002.

[0170] In addition, the parameter P in Equation (22) is set to 0.001, and the interval gradient filtering parameter in Equation (19) is set to λ = 0.03 2 , and the parameter T of RGF in Equation (17) is set to 4, and σ s is set to 3.

[0171] 3. Fusion and Evaluation of Noiseless Images

[0172] Figure 8 These are three pairs of examples of fusing noiseless infrared and visible gray - scale images. Figure 9 These are two pairs of examples of fusing infrared - visible color images, where Figure 8 (a1, b1, c1) and Figure 9 (a1, b1) are infrared images, Figure 8 (a2, b2, c2) and Figure 9 (a2, b2) are visible - light images; Figure 8 (a3 - a8, b3 - b8, c3 - c8) and Figure 9 (a3 - a8, b3 - b8) are the fusion results obtained by different methods.

[0173] From Figure 8 it can be seen that the fused images obtained by the ADF, FPDE, and GTF methods have lower contrast compared with the result images obtained by the proposed method; the IFEVIP method maintains good contrast, but the visual effect is overly enhanced, resulting in Figure 8 (a6) having obvious errors; the TIF method has the phenomenon of blurred internal features. Therefore, in the fusion results, the embodiments of the present invention can effectively separate the component information of different images, and combine their respective fusion rules to transfer the useful information of the source images into the fused images, obtaining the best visual performance in terms of contrast and detail preservation.

[0174] Figure 9 Two groups of infrared / visible - light color image fusions are shown. It can be seen that the brightness of the ADF, FPDE, and GTF methods is significantly lower than that of the IFEVIP and TIF methods, but the TIF method has a noise effect, and the IFEVIP method introduces artifacts. Although the FPDE and GTF methods have better preservation of the structure, the details are relatively weakened and lost. Generally speaking, Figure 9(a8) and (b8) show better performance in terms of brightness, structure, and details compared to other methods, which means that the proposed method can produce better visual effects. In addition to subjective visual analysis, the above fusion results are quantitatively evaluated, and the results are shown in Table 1. According to the data in the table, it can be seen that the objective evaluation of the embodiments of the present invention is significantly higher than that of other methods, especially for the indicators EN, Q CB and Q Y always perform better. In all quantitative evaluations, only a few places are not the best, but this does not affect the advantages of the method in this paper.

[0175] In summary, for noiseless image fusion, the method in this paper has good performance both subjectively and objectively.

[0176] Table 1 Quantitative indicators of noiseless image fusion results

[0177]

[0178] Figure 10 is an example of the fusion of a pair of noisy infrared and visible light grayscale images. Figure 11 is an example of the fusion of a pair of noisy infrared and visible light color images, where (a1 - a2, b1 - b2, c1 - c2) are the perturbed source images with 10-level, 20-level, and 30-level Gaussian noise added respectively, and (a3 - a8, b3 - b8, c3 - c8) are the fusion results obtained by different methods.

[0179] From Figure 10 , Figure 11 the noisy image fusion results, it can be seen that:

[0180] ① When the Gaussian noise level is 10, the denoising capabilities of the ADF, FPDE, and IFEVIP methods are limited, and their fusion results cannot well retain useful information; the GTF and TIF methods can effectively denoise to a certain extent, but they cannot protect the brightness of the source images. These two methods can perform fusion in a noisy environment, but they will introduce some irrelevant information, resulting in an unrealistic visual effect. Compared with other methods, the method in this paper has the best fusion ability in terms of detail preservation, and at the same time, the noise in the fusion result is significantly reduced, having good denoising performance.

[0181] ② When the noise level reaches 20, the structures of the fusion results generated by the ADF, FPDE, IFEVIP, and TIF methods will be severely damaged, and a large amount of obvious error information will be introduced into the fusion results. The GTF method suppresses the noise to a certain extent, but there are phenomena of reduced contrast and over-smoothing. In contrast, the method proposed in this paper can not only retain the details, brightness, and structure of the source images in the fused image, but also effectively eliminate the noise, having good denoising performance.

[0182] ③ When the noise level is 30, the details and small structures in the ADF, FPDE, GTF, IFEVIP, and TIF methods are damaged, and there is noise in all of them. By comparison, the method in this paper not only retains the contrast and structure, but also effectively and appropriately denoises, thus obtaining better fusion performance.

[0183] Based on the above subjective analysis, the method in this paper can effectively synchronize image fusion and denoising, and produce better visual effects compared with some of the latest methods.

[0184] The objective evaluations of the noise fusion results produced by different methods are shown in Tables 2, 3, and 4. Compared with five advanced image fusion methods, this method can obtain better quantitative evaluations, which are basically consistent with the objective evaluation results of the noise-free image fusion, verifying the effectiveness and superiority of the proposed method.

[0185] In summary, for the fusion of noisy images, the method in this paper has good performance both subjectively and objectively.

[0186] Table 2 Quantitative indicators of image fusion results when σ = 10

[0187]

[0188]

[0189] Table 3 Quantitative indicators of image fusion results when σ = 20

[0190]

[0191] Table 4 Quantitative indicators of image fusion results when σ = 30

[0192]

[0193] In summary, the present invention proposes a method for infrared and visible light image fusion and denoising based on three-scale decomposition and sparse representation in real time, making full use of the advantages of the RGF filter and the SR method. The sparse representation is used to fuse the detail layer images, and the fused detail layer images are adaptively obtained. The potential details in the basic components are effectively utilized through structure-texture decomposition. This method is easy to implement and can take into account both noise-free image fusion and noisy image fusion. It should be noted that the embodiments of the present invention only discuss the case of two source images, and in practical applications, the proposed method can be extended to the fusion problem of more than two source images.

[0194] Device Embodiment

[0195] According to the embodiments of the present invention, an infrared and visible light image fusion device is provided. Figure 12It is a schematic diagram of the infrared and visible light image fusion device according to an embodiment of the present invention. As Figure 12 shown, the infrared and visible light image fusion device according to an embodiment of the present invention specifically includes:

[0196] A first decomposition module 120, configured to decompose the source image into a base layer and a detail layer by using a rolling guidance filter. Among them, the detail layer includes most details and external noise, and the base layer includes: residual details and energy;

[0197] A second decomposition module 122, configured to decompose the base layer again based on a constructed structure-texture decomposition model, and decompose the base layer into a base structure layer and a base texture layer;

[0198] A pre-fusion module 124, configured to perform pre-fusion on the detail layer, the base structure layer, and the base texture layer by using different fusion rules corresponding to each layer;

[0199] A fusion module 126, configured to obtain a fused image by reconstructing three pre-fusion layers.

[0200] The first decomposition module 120 is specifically configured to:

[0201] Perform base and detail decomposition on the source image according to Formula 1 and Formula 2, and obtain the image I by solving to obtain the detail layer of the image I n where I

[0202]

[0203] where I n is the nth source image, n ∈ (1, 2,..., N), represents the base layer of I n ;

[0204] The second decomposition module 122 is specifically configured to:

[0205] Based on Formula 3 and Formula 4, decompose based on the structure-texture decomposition model to obtain the base structure layer and the base texture layer

[0206]

[0207]

[0208] where, and λ are respectively a scaling parameter and a smoothing parameter;

[0209] The pre-fusion module 124 is specifically configured to:

[0210] Based on the SR method, by establishing the connection between the sparse reconstruction error and the noise standard deviation, fusion denoising is achieved, and pre-fusion of the detail layer is carried out;

[0211] The weighted average technique based on the visual saliency map (VSM) is used to pre-fuse the basic structure layer;

[0212] The principal component analysis method is used to pre-fuse the basic texture layer.

[0213] The pre-fusion module 124 is specifically used for:

[0214] Generate the detail layer of the training data through a rolling guidance filter, collect 8×8-sized blocks from the detail image, construct the final training set, and obtain the dictionary D using the KSVD algorithm;

[0215] Take 8×8-sized blocks from each source image, normalize them, and generate the SR coefficients of the detail layer using the orthogonal matching pursuit algorithm (OMP) by solving the following objective function:

[0216]

[0217] where, is the k-th small block of the source image I n and is the corresponding sparse vector. is the maximum sparse reconstruction error, σ is the Gaussian standard deviation, and C > 0 is a parameter that controls when σ > 0;

[0218] Generate the fusion sparse coefficients using the "absolute value - maximum" scheme:

[0219]

[0220] The fusion detail vector is linearly represented as follows:

[0221]

[0222] Reshape each into an 8×8 small block, and then arrange them according to the initial position to obtain the pre-fused detail layer;

[0223] Construct the VSM, set I P to represent the intensity value of a pixel p in the image I, and the saliency value V(p) of the pixel p is defined as

[0224]

[0225] where, N represents the total number of pixels in I, j represents the pixel intensity, and M jLet $n_j$ represent the number of pixels with intensity equal to $j$, and $L$ denote the number of gray levels. If two pixels have the same intensity value, their saliency values are equal.

[0226] Then normalize $V(p)$ to $[0, 1]$.

[0227] Let $V_1$ and $V_2$ represent the VSMs of different source images respectively. and Let $I_1$ and $I_2$ represent the basic structure layer images of different source images, and obtain the pre - fused image $F$ of the final basic structure layer through weighted average: b,s :

[0228]

[0229] where the weight $W$ b is defined as:

[0230]

[0231] Take the basic texture images of visible light and infrared images and as the column vectors of the matrix $\gamma$. Then, taking each row as a reference and each column as a variable, calculate the covariance matrix $C$ of $\gamma$.

[0232] Calculate the eigenvalues $\lambda_1$, $\lambda_2$ of $C$ and the corresponding eigenvectors and

[0233] Find the largest eigenvalue from the two eigenvalues, that is, $\lambda$ max $=\max(\lambda_1,\lambda_2)$, and take the eigenvector corresponding to $\lambda$ max as the maximum eigenvector $\varphi$ max , calculate the principal components $P_1$ and $P_2$ corresponding to $\varphi$ max , and normalize their values:

[0234]

[0235]

[0236] Take the principal components $P_1$ and $P_2$ as weights to fuse into the pre - fused image $F$ of the final basic texture layer: b,t :

[0237]

[0238] The fusion module is specifically used for:

[0239] The final fused image $F$ according to formula (15) is:

[0240] $F = F$ d $+ F$b,s +F b,t Formula 15.

[0241] The embodiments of the present invention are system embodiments corresponding to the above method embodiments. The specific operations of each module can be understood with reference to the description of the method embodiments, and will not be elaborated herein.

[0242] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An infrared and visible light image fusion method, characterized in that, Including: Decompose the source image into a base layer and a detail layer using a rolling guidance filter. Among them, the detail layer includes most details and external noise, and the base layer includes: residual details and energy. Specifically, decomposing the source image into a base layer and a detail layer using a rolling guidance filter includes: Perform basic and detail decomposition on the source image according to Formula 1 and Formula 2, and obtain the detail layer of image I by solving to get image I n ​ where I n is the n-th source image, b ∈ {1, 2,..., N}, denotes the n base layer of I; Based on the constructed structure-texture decomposition model, decompose the base layer again, and decompose the base layer into a base structure layer and a base texture layer. Specifically including: According to Formula 3 and Formula 4, based on the structure-texture decomposition model, is decomposed to obtain the basic structure layer and the basic texture layer wherein, and λ are a scale parameter and a smoothing parameter, respectively; Use different fusion rules corresponding to each layer to pre-fuse the detail layer, the base structure layer, and the base texture layer. Specifically including: Based on the super-resolution SR method, by establishing the connection between the sparse reconstruction error and the noise standard deviation, realize fusion denoising and perform pre-fusion of the detail layer; Adopt the weighted average technology based on the visual saliency map VSM to pre-fuse the base structure layer; Adopt the principal component analysis method to pre-fuse the base texture layer; Obtain the fused image by reconstructing the three pre-fusion layers.

2. The method according to claim 1, characterized in that Based on the SR method, by establishing the connection between the sparse reconstruction error and the noise standard deviation, realizing fusion denoising, and the specific process of pre-fusing the detail layer includes: Generate the detail layer of the training data through a rolling guidance filter, collect 8×8-sized blocks from the detail image, construct the final training set, and obtain the dictionary D using the KSVD algorithm; Take 8×8-sized blocks for each source image and normalize them. By solving the following objective function, use the orthogonal matching pursuit algorithm OMP to generate the detail layer SR coefficients: Among them, is the k-th patch of the source image I n , is the corresponding sparse vector. is the maximum sparse reconstruction error, σ is the Gaussian standard deviation, and C > 0 is a parameter that controls when σ > 0; Adopt the "absolute value - maximum" scheme to generate the fused sparse coefficients: Where N represents the total number of pixels in I; Fusion Detail Vector Obtained by the following linear representation: Reshape each into 8×8 small blocks, and then arrange them according to the initial position to obtain the pre-fused detail layer; The specific process of pre-fusing the base structure layer using the weighted average technology based on the visual saliency map VSM includes: Construct VSM, set I p Representing the intensity value of a pixel p in image I, the saliency value V(p) of pixel p is defined as where j represents the pixel intensity, M j represents the number of pixels with intensity equal to j, L represents the number of gray levels, and if two pixels have the same intensity value, their saliency values are equal; Then normalize V(p) to [0,1]; Let V1 and V2 represent the VSMs of different source images respectively, and represent the basic structure layer images of different source images, and the final pre-fusion image F of the basic structure layer is obtained by weighted average b,s : Among them, the weight W b is defined as: The specific process of pre-fusing the base texture layer using the principal component analysis method includes: The base texture images of visible light and infrared images and are used as the column vectors of matrix γ. Then, taking each row as a reference and each column as a variable, the covariance matrix C of γ is calculated; Calculate the eigenvalues λ1, λ2 of C and the corresponding eigenvectors and Find the largest eigenvalue from the two eigenvalues, i.e., λ max = max(λ1, λ2), and take the eigenvector corresponding to λ max as the largest eigenvector φ max , calculate the principal components P1 and P2 corresponding to φ max , and normalize their values: Using the main components P1 and P2 as weights, fuse them into the pre-fused image F of the final basic texture layer b,t :

3. The method according to claim 2, wherein The specific process of obtaining the fused image by reconstructing the three pre-fusion layers includes: The final fused image F according to formula 15 is: F = F d + F b,s + F b,t Formula 15.

4. An infrared and visible light image fusion device, characterized in that, Including: The first decomposition module is used to decompose the source image into a base layer and a detail layer using a rolling guidance filter. Among them, the detail layer includes most details and external noise, and the base layer includes: residual details and energy. The first decomposition module is specifically used for: Perform basic and detail decomposition on the source image according to Formula 1 and Formula 2, and solve to obtain image I n of the detail layer Where I n is the n-th source image, n ∈ {1, 2, ..., N}, denotes the base layer of I n ; The second decomposition module is used to decompose the base layer again based on the constructed structure-texture decomposition model, and decompose the base layer into a base structure layer and a base texture layer. Specifically used for: According to Formula 3 and Formula 4, decompose based on the structure-texture decomposition model to obtain the basic structure layer and the basic texture layer wherein, and λ are a scale parameter and a smoothing parameter, respectively; The pre-fusion module is used to pre-fuse the detail layer, the base structure layer, and the base texture layer using different fusion rules corresponding to each layer. The pre-fusion module is specifically used for: Based on the super-resolution SR method, by establishing the connection between the sparse reconstruction error and the noise standard deviation, realize fusion denoising and perform pre-fusion of the detail layer; Adopt the weighted average technology based on the visual saliency map VSM to pre-fuse the base structure layer; Adopt the principal component analysis method to pre-fuse the base texture layer; The fusion module is used to obtain the fused image by reconstructing the three pre-fusion layers.

5. The device according to claim 4, characterized in that, The pre-fusion module is specifically used for: Generate the detail layer of the training data through a rolling guidance filter, collect 8×8-sized blocks from the detail image, construct the final training set, and obtain the dictionary D using the KSVD algorithm; Take 8×8-sized blocks from each source image, normalize them, and generate the detail layer SR coefficients using the orthogonal matching pursuit algorithm OMP by solving the following objective function: Among them, is the k-th patch of the source image I n , is the corresponding sparse vector. is the maximum sparse reconstruction error, σ is the Gaussian standard deviation, and C > 0 is a parameter that controls when σ > 0; Generate the fused sparse coefficients using the "absolute value - maximum" scheme: where N represents the total number of pixels in I; Fusion detail vector Obtained by the following linear representation: Reshape each into 8×8 small blocks, and then arrange them according to the initial positions to obtain the pre-fused detail layer; Construct VSM, set I P Representing the intensity value of a pixel p in image I, the saliency value V(p) of pixel p is defined as where j represents the pixel intensity, M j represents the number of pixels with intensity equal to j, and L represents the number of gray levels. If two pixels have the same intensity value, their significance values are equal; Then normalize V(p) to [0, 1]; Let V1 and V2 represent the VSMs of different source images, and represent the basic structure layer images of different source images, and the final pre-fusion image F of the basic structure layer is obtained by weighted average b,s : Among them, the weight W b is defined as: The base texture images of visible light and infrared images and are used as the column vectors of matrix γ. Then, taking each row as a reference and each column as a variable, the covariance matrix C of γ is calculated; Calculate the eigenvalues λ1, λ2 of C and the corresponding eigenvectors and Find the largest eigenvalue from the two eigenvalues, i.e., λ max = max(λ1, λ2), and take the eigenvector corresponding to λ max as the largest eigenvector φ max , calculate the principal components P1 and P2 corresponding to φ max and normalize their values: Using the principal components P1 and P2 as weights, fuse them into the pre-fused image F of the final base texture layer b,t :

6. The device according to claim 5, characterized in that The fusion module is specifically used for: The final fused image F according to Equation 15 is: F = F d + F b,s + F b,t Formula 15.

Citation Information

Patent Citations

  • Artificial Intelligence Based Image Fusion Apparatus and Method for Fusing Infrared and Visible Image

    KR102257752B1