An infrared and visible light image fusion method and application
Through a fusion method based on the principle of light propagation and retinal imaging theory, combined with the fusion of medium-frequency and high-frequency layers, as well as histogram enhancement and Mean Shift clustering algorithm, the problems of information loss and artifacts during the fusion of infrared and visible light images in the prior art are solved, and high-quality image fusion is achieved.
Patent Information
- Application Number
- CN202110640871.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-18
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-07-18
AI Technical Summary
The prior art is difficult to effectively retain all information of the source map during the fusion process of infrared rays and visible light images, and it is easy to generate information that is independent of the source map, resulting in low quality of the fused image.
Using a fusion method based on the principle of light propagation and retinal imaging theory, light hybrid images are constructed through the fusion of medium frequency layers and the linear fusion of adaptive weights of high frequency layers, combined with histogram enhancement algorithm and Mean Shift image clustering algorithm.
The information of infrared and visible light images is fully retained during the fusion process, avoiding the generation of information independent of the source image, and improving the quality of the fusion image and the retention of target and environmental textures.
Smart Images

Figure CN113379659B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to an infrared and visible light image fusion method and application. Background Art
[0002] Combining images obtained from multi-source sensors into a new fused image, this fused image can provide richer texture information than any single sensor's image currently collected. For example, infrared images are sensitive to heat sources and can collect information that cannot be perceived by the human eye. Therefore, infrared images are less affected by environmental changes. However, in the actual environment, there are few heat-emitting objects, resulting in the inability of infrared sensors to collect the distribution of the environment. The images collected by visible light sensors have the characteristics that conform to human visual habits and can effectively reflect the scene distribution. However, in the actual environment, the intensity of light is often uncontrollable, making it difficult to identify target information in low-light environments. Therefore, fusing infrared images and visible light images, the obtained fused image contains the ability of infrared images to be sensitive to heat source targets and the ability of visible light images to reflect the environment distribution. This technology has been widely applied in the military field, medical field, and remote sensing field.
[0003] Currently, the fusion strategies of most image fusion algorithms can be divided into pixel level, feature level, symbol level, and hybrid level. The complexity of these levels increases in turn. The fusion theory technologies established on this basis can be summarized into eight categories: the theory technology using multi-scale geometric decomposition of images, the theory technology using sparse representation, the theory technology using deep learning, the theory technology using subspaces, the theory technology using significant features, the technology using hybrid theory, the theory technology using total variation models, and the theory technology using neural networks. The proposal of each theory technology can promote the development of the image fusion field. A good image fusion theory technology usually has the following three characteristics: First, the fused image should contain all the information of the source images; second, no information unrelated to the source images should be generated in the fused image; third, the fused image should be robust to the adverse factors of the source images (poor source image quality, registration error). To address the above three problems of image fusion, most image fusion theory technologies have reached a bottleneck period. To break through the bottleneck of the image fusion field again, a new fusion theory needs to be proposed. Summary of the Invention
[0004] To solve the above problems, the present invention provides an infrared and visible light image fusion method and application. This method is an infrared and visible light image fusion method based on the principle of light propagation and the theory of retinal imaging. It can not only completely retain the information of infrared images and visible light images, but also effectively avoid generating information unrelated to the source images during the fusion process by traditional image fusion theories.
[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] An infrared and visible light image fusion method, comprising
[0007] S1. Realize the fusion of the intermediate frequency layers of the visible light image and the infrared image;
[0008] First, compare the intermediate frequency layer of the visible light image and the intermediate frequency layer of the infrared image, and respectively select the pixels with the largest pixel values at the corresponding coordinates of the two images to form a maximum layer, and then select the pixels with the smallest pixel values at the corresponding coordinates of the two images to form a minimum layer. Then, the intermediate frequency layer fusion model of the image satisfies:
[0009]
[0010] In the formula, I M represents the fusion result of the intermediate frequency image; I MSLS1 represents the intermediate frequency layer of the visible light image, I MSLS2 represents the intermediate frequency layer of the visible light image, max(I MSLS1 , I MSLS2 ) represents the pixel standard deviation corresponding to the maximum layer of the intermediate frequency image, denoted by σ3, and min(I MSLS1 , I MSLS2 ) represents the pixel standard deviation corresponding to the minimum layer of the intermediate frequency image, denoted by σ4;
[0011] Extract the maximum mixed image max(I E , I I ) and the minimum mixed image min(I E , I I ) of the source images respectively, and linearly fuse max(I E , I I ) and min(I E , I I ) with the intermediate frequency layer I M by a weight of 0.1 to obtain the fusion result I EF of the new intermediate frequency image after texture enhancement, and the expression satisfies:
[0012] I MF = I M + 0.1×(max(I E , I I ) + min(I E , I I ));
[0013] S2. On the basis of the fusion result of the intermediate frequency layer, linearly fuse the high-frequency layer and the low-frequency layer of each source image to I MFIn this way, the final image fusion result is obtained, satisfying:
[0014] I F = I MF + ω1 × I H1 + ω2 × I H2 - I L1 - I L2
[0015] In the formula, I F represents the result of the light mixing diagram; I H1 represents the high-frequency layer of the visible light image, and I H2 represents the high-frequency layer of the infrared image; I L1 represents the low-frequency layer of the visible light image; I L2 represents the low-frequency layer of the infrared image; let I L1 , I L2 in the low-frequency layer, the pixel values less than 0 are set to 0; where ω1 and ω2 respectively represent the fusion weights of the visible light high-frequency layer and the infrared high-frequency image layer, and their expressions satisfy:
[0016]
[0017]
[0018]
[0019] In the formula, W represents the total weight of the fusion of the high-frequency detail layers of each source image; σ5 represents the standard deviation of the pixel values of the visible light enhanced image I E , and σ6 represents the standard deviation of the pixel values of the infrared image I I ; σ7 represents the standard deviation of the pixel values of the visible light intermediate-frequency layer I MSLS1 ; σ8 represents the standard deviation of the pixel values of the visible light intermediate-frequency layer I MSLS2 .
[0020] Furthermore, in the step S1, the histogram enhancement algorithm and the source image adaptive weighted linear fusion (HSAWE) are used, and the expression is as follows:
[0021]
[0022] In the formula, I E represents the enhancement of the visible light image, I V represents the visible light image, H isteq (·) represents the histogram enhancement process, σ1 represents the standard deviation of the pixels of the I V image, and σ2 represents the standard deviation of the pixels of the H isteq (·) image.
[0023] Further, in the step S1, the layer extraction in the source image is implemented by the Mean Shift based Layer Smoothing clustering algorithm (MSLS); specifically:
[0024] The mathematical model of MSLS: Given n pixel points in the d-dimensional space R d and let i = 1,..., n. Arbitrarily select a point x in the space R d , then the basic form of the MSLS vector is defined as:
[0025]
[0026] At this point, S k is a high-dimensional circular region with a radius of h, which is the set of y points satisfying the following relationship:
[0027] S k (x) = {y: (y - x i ) T (y - x i ) < h 2}
[0028] k represents that among these n sample points x i , there are k points falling into the S k region;
[0029] Let K(x) be a d-dimensional kernel function, then it needs to satisfy:
[0030] K(x) = c k,d k(||x|| 2 )
[0031] At this point, k(x) is called the kernel prototype function, where x ≥ 0; c k,d represents a normalization constant that makes the integral of K(x) equal to 1, satisfying:
[0032]
[0033] Here we use the Gaussian kernel function as the kernel prototype function k(x), then the multivariate normalized Gaussian kernel function satisfies:
[0034]
[0035] To effectively estimate the kernel density, the Parzen window is used as the multivariate density estimator; therefore, in the R d space, the multivariate kernel density estimator at x is:
[0036]
[0037] Among them,
[0038] KH f(x) = |H| -1 / 2 K(H -1 / 2 x)
[0039] Wherein, H is a diagonal matrix with equal diagonal elements, satisfying:
[0040]
[0041] Wherein, h represents a broadband parameter, satisfying h > 0; Therefore, the kernel density function can be simplified to:
[0042]
[0043] With the kernel function, the density estimation formula of MSLS can be rewritten as:
[0044] Substituting Equation (5) into Equation (11), the density estimation formula expressed by the kernel prototype function is obtained:
[0045]
[0046] Wherein, K(x) represents the kernel function, h represents the radius of the high-dimensional circular region, c k,d / nh d represents the unit density. To make the result in the above formula the largest, the derivative of this function should be 0. Therefore, the result of its derivative satisfies:
[0047]
[0048] Let:
[0049] g(x) = -k'(x)
[0050] Suppose that the derivative of the kernel function k(x) does not exist at finite points within x ∈ [0, ∞), and the derivative exists at other points; Then using g(x) as the primitive function, the kernel function G(x) is defined as:
[0051] G(x) = c g,d g(||x|| 2 )
[0052] Wherein, c g,d is the relevant normalization constant; Since the sum function K(x) and G(x) are not very different, K(x) is called the projection of G(x);
[0053] Substituting Equation (14) into Equation (15), we get
[0054]
[0055] Observing Equation (16), where the first term is the time value and the vector direction of the second term is consistent with the gradient direction. The separate expressions satisfy:
[0056]
[0057] where m h,G (x) is the mean shift amount;
[0058] (a) MSLS adjusts the window according to the calculated mean shift vector. When the movement amount is less than the threshold, similar pixels are merged and represented by I MS (x);
[0059] (b) The I MS (x) image is hierarchically processed in descending order according to the pixel values, satisfying:
[0060] I MS (x) = [I 255 , …, I0]
[0061] (c) Assume an n×n - dimensional window W n×n , and use this window to move in the highest - pixel - layer image I 255 . When all pixels within the window are 255, mark the coordinates of all pixels within the window. When one or more non - 255 values appear within the window, all pixels within the window are not marked, and continue to move the window to traverse the entire I 255 layer;
[0062] (d) The pixel values at the marked coordinates remain unchanged, and the pixel values at the unmarked coordinates are decreased by 1, and are used as the pixel values of the I 254 layer to perform the operations in step (c);
[0063] (e) Repeat steps (c) and (d) until after traversing the I1 layer and then stop; The I0 layer consists of its own pixel points and the pixels with pixel values decreased by 1 at the unmarked coordinates of I1;
[0064] (f) Re - compose all the processed layers into the image I MSLS , which is the clustering result of MSLS.
[0065] Furthermore, in step S2, let I MSLS represent the smoothed image of the source image after being smoothed by MSLS, as the intermediate - frequency layer of the source image. It should be noted that in the fusion algorithm process, the dimension n of the smoothing window W n×n is 150, then the high - frequency layer image I H is obtained by subtracting the intermediate - frequency layer I E from the source image I MSLS :
[0066] IH = I E -I MSLS
[0067] Low-frequency layer image I L From the intermediate-frequency layer I MSLS Subtract the source image I E To obtain:
[0068] I L = I MSLS -I E .
[0069] The present invention has the following beneficial effects:
[0070] Provided is an infrared and visible light image fusion method based on the principle of light propagation and the theory of retinal imaging, which can not only completely retain the information of infrared images and visible light images, but also effectively avoid the information unrelated to the source images generated in the fusion process by traditional image fusion theories;
[0071] Also proposed is a simple and effective image enhancement algorithm, which improves the Mean Shift image clustering algorithm to make the structure of the clustered image more perfect;
[0072] In order to construct the light mixing image required in the image fusion theory technology, a new image fusion algorithm is proposed to simulate the mixed light image. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Figure 1 It is a schematic diagram of image fusion;
[0074] In the figure: (a) Infrared image, (b) Visible light image,
[0075] (c) GTF, (d) ADF, (e) LP-SR, (f) Ou r.
[0076] Figure 2 It is a schematic diagram of the principle of the embodiment of the present invention.
[0077] Figure 3 It is a schematic diagram of the imaging of a color image.
[0078] Figure 4 It is a schematic diagram of the experiment on the coexistence of information after fusing the source images by the traditional fusion theory and the fusion theory of the present invention;
[0079] In the figure: (a) Source Figure 1 ; (b) Source Figure 2; (c) CP; (d) CVT; (e) DTCWT; (f) GFF; (g) GT (h) GTF; (i) LP; (j) LP-SR; (k) MSVD; (I) RP; (m) SAL; (n) Wavelet; (o) PCNN; (p) ADF; (q) Linear fusion; (r) Our.
[0080] Figure 5 It is a schematic diagram of trichromatic fusion.
[0081] Figure 6 It is the experimental result diagram without considering the mixing of mixed light and considering the mixing of light;
[0082] In the figure: (a) Two primary colors; (b) Three primary colors.
[0083] Figure 7 It is a comparison diagram of the fusion results of different algorithms of the present invention as the final fusion result of the light mixing image of the fusion theory of the present invention;
[0084] In the figure: (a) ADF; (b) CP; (c) CVT; (d) DTCWT; (e) GFF; (f) GT; (g) GTF; (h) LP; (i) LP-SR; (j) MSVD; (k) PCNN; (I) RP.
[0085] Figure 8 It is a schematic diagram of the influence of whether the low-light image is enhanced on the fusion result;
[0086] In the figure: (a) Visible light source diagram, (b) Infrared light source diagram, (c) ADF fusion result, (d) Visible light enhancement diagram, (e) Infrared light source diagram, (f) ADF fusion result after the visible light image is enhanced.
[0087] Figure 9 It is an enhancement comparison diagram of HE and HSAWE;
[0088] In the figure: (a) Low-light image; (b) HE; (c) HSAWE.
[0089] Figure 10 It is a comparison diagram of the extraction results of the intermediate frequency images of the source images;
[0090] In the figure: (a) Source image; (b) Mean Shift; (c) MSLS.
[0091] Figure 11 It is the pixel surf diagram of image clustering;
[0092] In the figure: (a) Represents the pixel surf diagram of the source image, (b) Represents the pixel surf diagram of the image after mean shift clustering, (c) Represents the pixel surf diagram of the image after MSLS clustering.
[0093] Figure 12 It is a schematic diagram of the layers after the source image is decomposed by MSLS at multiple scales;
[0094] In the figure: (a) Visible light image; (b) Visible light intermediate frequency image; (c) Visible light high frequency image; (d) Visible light low frequency image; (e) Infrared image; (f) Infrared intermediate frequency image; (g) Infrared high frequency image; (h) Infrared low frequency image.
[0095] Figure 13 Schematic diagram of the separation of the maximum layer and the minimum layer of the pixels in the intermediate frequency layer of the source image;
[0096] In the figure: (a) Maximum layer (b) Minimum layer.
[0097] Figure 14 It is the fusion result of the intermediate frequency layer I M of.
[0098] Figure 15 Schematic diagram of the separation of the maximum layer and the minimum layer of the pixels of the source image;
[0099] In the figure: (a) Maximum mixed image; (b) Minimum mixed image.
[0100] Figure 16 It is the MF effect picture;
[0101] Figure 17 It is the MSLSWF fusion result.
[0102] Figure 18 It is the subjective comparison diagram of the fusion results of each fusion algorithm and the present invention
[0103] Figure 19 It is the objective comparison diagram of the fusion results of each fusion algorithm and the present invention Detailed implementation manners
[0104] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several deformations and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0105] Inspired by the imaging of the transparent glass on the window at night, that is, the layout of the indoor environment is reflected on the glass, and the layout of the outdoor environment is refracted on the glass. Finally, the light converges on the retina to form a fused image containing the indoor and outdoor scenes. To illustrate the main idea of the method we proposed, we are in Figure 1A legend is shown in which (a) represents an infrared image that can clearly display the target position information in the image, but the environmental layout in the infrared image is very blurred; (b) represents a visible light image that can reflect the layout of the environment, but the image clarity is easily affected by the light intensity, resulting in the target being covered by the environment. (c) represents the fusion result of GTF. We can see that the target is easily recognizable, but the distribution of the environment is not clear. This disadvantage becomes more obvious when the light weakens. (d) represents the fusion result of ADF. We can see that compared with GTF, the heat source target of ADF is darker, but it is better than GTF in reflecting the target texture and environmental layout. (e) represents the fusion result of LP-SR. We can see that the fusion result is excellent in both the brightness of the target display and the brightness of the environment display, but there is interference from halos on the right side of the image. (f) represents our fusion result. Obviously, our result maximally retains and reproduces the information of the source image both in reflecting the target brightness and the environmental texture information, and the result of the present invention is more in line with the visual observation habits of humans.
[0106] Fusion principle of the present invention
[0107] Figure 2 It is a schematic diagram showing the fusion theory of the present invention. The gray part represents the outdoor area at night, the white part represents the indoor area with lights on at night, and the middle dark gray area represents the translucent glass. We assume that the 'visible light image' is the layout of the indoor scenery, the infrared ray is the layout of the outdoor scenery, and the eyeball position represents the angle and position of the observer observing the transparent glass. The imaging principle of this image fusion theory is as follows: (1) The light of the 'visible light image' projects the information into our eyes through the reflection of the transparent glass, and this process is represented by an arrow; (2) The light of the 'infrared image' projects the information into our eyes through the refraction of the transparent glass, and this process is represented by an arrow; (3) The phenomenon of fusion occurs during the propagation of light. For example, sunlight is white but the white light contains different color beams. Our fusion theory also takes this phenomenon into account and is represented by an arrow symbol. The final image presented in the observer's brain is as shown in Figure 2 'Brain imaging' behind the eyeball.
[0108] The basic expression of this principle satisfies:
[0109] B = [I B , I V , I I (1)
[0110] In the formula, B represents the imaging result of the fused image in the brain, IB Indicates the information represented by the red arrow, I V Indicates the information represented by the green arrow, I I Indicates the information represented by the blue arrow.
[0111] The benefits of applying this discovery to image fusion are as follows: It can simultaneously display the information of infrared images and visible light images, effectively solving the problems of source image information loss and artifact generation during the image fusion process. However, to apply this theory, the following two problems need to be solved. First, since the eyes can present images in the brain simultaneously through zooming, while in reality, once a display shows an image, it cannot be zoomed to display the remaining images. Therefore, a hardware technology is required, similar to the retina, that can present multiple images in the same area simultaneously. Second, although in Equation (1), / V 、I I are represented by 'visible light image' and 'infrared image' respectively, but the I B representing the light mixing phenomenon during the light propagation process needs to be reasonably constructed.
[0112] 1. Solution to Problem One
[0113] The human brain can automatically integrate the information seen by the left and right eyes into a single image. Its feature is that the two images complement each other to expand the field of view, and such an ability does not cause discomfort to humans. One of the reasons is that the brain can receive the input of two images simultaneously. Based on this principle, we hope that the display can, like the retina, also display the information of multiple images in the same area simultaneously, so as to maximize the restoration of the process of the retina receiving images. Fortunately, in actual engineering, people have successfully introduced the display from the black-and-white screen era to the color screen era through the research on the three primary colors. The principle is as Figure 3 shown. A color image is composed of images showing the red wavelength band, green wavelength band, and blue wavelength band displayed simultaneously on the display. The reason why the display has the ability to display images of different wavelength bands simultaneously is that each brightness point on the display is composed of three small brightness points of red, green, and blue. Therefore, the color display can help us solve the problem of Problem One.
[0114] Combined with Figure 3 、 Figure 4 and Figure 5 analysis, the advantages of using a color display to fit the retina imaging principle are as follows: (i) Compared with other fusion theory technologies that only display the grayscale image of a single fused image, our fusion theory gives each source image the opportunity to be displayed, effectively avoiding the problem that traditional image fusion theories often sacrifice the information of each source image to achieve coexistence. To further verify the advantages of the fusion theory of the present invention, the present invention designed a fusion experiment, as Figure 4As shown, (a) represents the source Figure 1 , which is a combined image where the pixel value of the white area is 200 and the pixel value of the five-pointed star is 100. (b) represents the source Figure 2 , which is a combined image where the pixel value of the white area is 100 and the pixel value of the five-pointed star is 200. The fusion results of each algorithm are as Figure 4 shown.
[0115] Observation Figure 4 shows that the MSVD, Wavelet, and linear fusion algorithms cannot identify the information of the five-pointed star. The CVT and DTCWT can only see the artifacts in the shape of a five-pointed star. The GFF, PCNN, and ADF can only see the edge contour of the five-pointed star. CP, GT, LP, LP-SR, RP, and SAL can see the overall information of the five-pointed star, but these images all have the phenomena of artifacts and information loss. The GTF can clearly see the contour and detail information of the fused image, but the sharp corners of the five-pointed star have the phenomenon of color variation. Observing the fusion result of the present invention, it can be clearly observed that the five-pointed star retains all the information of the source image and no artifacts are introduced.
[0116] (ii) Observation Figure 3 shows that the pistil of the colored picture is yellow. We can then Figure 5 judge that only the red-band R image and the green-band G image have information output. Observing Figure 3 the band pictures, the fact is consistent with the inference result, and the blue-band B image has no information about the pistil. Similar examples include the G image representing the green band. Only this image has a large amount of information about the lotus leaves. According to the color speculation of the color image, the red-band R image and the blue-band B image do not contain or contain only a small amount of information about the lotus leaves. Therefore, the superimposed part and the implicit information of the image redundancy part can be judged according to the color change; observing again Figure 4 shows that in the fusion result of the present invention, the color of the five-pointed star is purple, rather than a single primary color. Combining Figure 5 it can be judged that this fused picture is obtained by fusing two different source pictures. Observing the fusion results of the remaining theories, only knowing that the fused image contains a five-pointed star, it is impossible to judge whether the five-pointed star comes from the source Figure 1 or the source Figure 2 or both source pictures contain the information of the five-pointed star.
[0117] 3. Solution to Problem 2
[0118] In actual engineering, there will be a phenomenon of mixed light in the light propagation. The present invention uses one of the primary colors to represent the image in the mixing process. In this way, we need to answer two questions. First, if we do not consider the mixing phenomenon of light propagation, that is, only two primary colors are used to represent the infrared image and the visible light image respectively among the three primary colors, can it meet the fusion requirements; second, if we want to consider the mixing phenomenon, then what mathematical model is used to represent the mixed image;
[0119] 3.1 Experiments on not considering light mixing and considering light mixing
[0120] In order to verify the impact of not considering light mixing (two primary colors) and considering light mixing (three primary colors) in our fusion theory, relevant experiments were conducted as follows Figure 6 shown
[0121] Observation Figure 6 , in Figure (a), the infrared image is displayed in red among the three primary colors, the visible light image is displayed in green among the three primary colors, and blue among the three primary colors is not displayed. Figure (b) shows that the mixed light image is displayed in red of the three primary colors, the visible light image is displayed in green of the three primary colors, and the infrared image is displayed in blue of the three primary colors. First of all, both Figure (a) and (b) can fuse the information containing the target and the information containing the environment in each source image into one image, and the target can be easily identified in the environment. Secondly, it can be observed that the color of Figure (a) is not as rich as that of Figure (b), and the texture details of the whole image of Figure (b) are significantly better than those of Figure (a). Therefore, it can be proved that the fusion effect considering mixed light is significantly better than that without considering mixed light. Then, determining the mathematical model of the light mixing image becomes the focus of our research
[0122] 3.2 Mathematical model of light mixing image
[0123] For the convenience of research, we consider taking the results of the existing fusion theory as the part of the light mixing image. The purposes are as follows: First, it can continue the research theory and research results of predecessors; Second, it makes the fusion theory of the present invention have the value of further research and optimization
[0124] A. Fusion effect of using existing fusion algorithms
[0125] We take the fusion result of the existing fusion algorithm as the red component in the three primary colors, the visible light image as the green component in the three primary colors, and the infrared image as the blue component in the three primary colors. The experimental results are as follows Figure 7 shown
[0126] Observation Figure 7 , our fusion theory fully gives the opportunity for each source image to present information to users at the same time. Therefore, the target in each experimental image can be clearly identified, and the environmental layout can be quickly grasped, which proves that it is feasible to use the fusion image instead of the light mixing image. However, observing the experimental results Figure 7 , there is still a large room for optimization in the image brightness and clarity of each fusion result
[0127] B. Mathematical model of the light mixing image of the present invention
[0128] In order to ensure that the information in the infrared image and the visible light image in our fusion theory does not change at all, the quality of the light-mixed image largely determines the quality of the final result of the fusion of the present invention. Therefore, improving the quality (brightness, clarity) of the light-mixed image has become the focus of our research on the light-mixed image, and we use the multi-scale fusion theory to construct a mathematical model of the light-mixed image.
[0129] i) During the fusion process of the infrared image and the visible light image, compared with the infrared sensor, the performance of the visible light sensor is more likely to be affected by the environment, resulting in a poor quality of the captured image. And it will cause the fusion result not to reach the ideal result. Taking ADF as an example, we explore the influence of whether the visible light image is enhanced on the fusion result, and the result is as Figure 8 shown.
[0130] Observation Figure 8 , comparing the fusion image results of figures (c) and (f) which are both of ADF, as shown within the red frame for the display of the internal environment of the street store, no seat information can be obtained from figure c compared with figure f. Therefore, enhancing the low-light visible light image is a necessary step.
[0131] The present invention proposes an image enhancement method, the principle of which is to use the histogram enhancement algorithm and the source image adaptive weighted linear fusion (HSAWE), and the enhancement result is as Figure 8 shown in figure (d) in the middle. The purposes are as follows: First, to avoid the problem of unnatural image brightness after the source image is enhanced by the histogram algorithm; Second, this enhancement algorithm has a simple model and a fast operation speed.
[0132] The expression is as follows:
[0133]
[0134] In the formula, I E represents the enhancement of the visible light image, I V represents the visible light image, H isteq (·) represents the histogram enhancement process, σ1 represents the standard deviation of the pixels of the I V image, and σ2 represents the standard deviation of the pixels of the H isteq (·) image. The enhancement result is as Figure 9 shown.
[0135] ii) Image multi-scale decomposition algorithm
[0136] The present invention proposes a clustering algorithm based on Mean Shift for layer smoothing (MSLS) as the image multi-scale decomposition algorithm. The extraction results of the intermediate frequency images of the source image by Mean Shift and MSLS are as Figure 10 shown.
[0137] Observation Figure 10 , although the intermediate frequency map extracted by Mean Shift can remove some small texture details, compared with the result of MSLC, most of the texture lines are not erased. Therefore, it is more reasonable to adopt MSLC for the multi-scale algorithm of the source image.
[0138] Mathematical model of MSLS: Given n pixel points in the d-dimensional space R d , let i = 1,..., n. Arbitrarily select a point x in the space R d . Then the basic form of the MSLS vector is defined as:
[0139]
[0140] At this point, S k is a high-dimensional circular region with a radius of h, which is the set of y points satisfying the following relationship:
[0141] S k (x) = {y: (y - x i ) T (y - x i ) < h 2}} (4)
[0142] k represents that among these n sample points x i , k points fall into the S k region.
[0143] Let K(x) be a d-dimensional kernel function, then it needs to satisfy:
[0144] K(x) = c k,d k(||x|| 2 ) (5)
[0145] At this point, k(x) is called the kernel prototype function, where x ≥ 0. c k,d represents a normalization constant that makes the integral of K(x) equal to 1 and satisfies:
[0146]
[0147] Here we adopt the Gaussian kernel function as the kernel prototype function k(x), then the multi-dimensional normalized Gaussian kernel function satisfies:
[0148]
[0149] To effectively estimate the kernel density, the Parzen window is adopted as the multi-dimensional density estimator. Therefore, in the R d space, the multi-dimensional kernel density estimator at x is:
[0150]
[0151] Among them,
[0152] K H (x) = |H| -1 / 2 K(H -1 / 2 x) (9)
[0153] In the formula, H is a diagonal matrix with equal diagonal elements, satisfying:
[0154]
[0155] In the formula, h represents the broadband parameter, satisfying h > 0. Therefore, the kernel density function can be simplified to:
[0156]
[0157] With the kernel function, the density estimation formula of MSLS can be rewritten as:
[0158] Substitute Equation (5) into Equation (11) to obtain the density estimation formula expressed by the kernel prototype function:
[0159]
[0160] In the formula, K(x) represents the kernel function, h represents the radius of the high-dimensional circular region, and c k,d / nh d represents the unit density. To make the result in the above formula the largest, the derivative of this function should be 0. Therefore, the result of taking its derivative satisfies:
[0161]
[0162] Let:
[0163] g(x) = -k′(x) (14)
[0164] Suppose the derivative of the kernel function k(x) does not exist at finite points within x ∈ [0, ∞), and the derivative exists at other points. Then, using g(x) as the primitive function, the kernel function G(x) is defined as:
[0165] G(x) = c g,d g(||x|| 2 ) (15)
[0166] In the formula, c g,d is the relevant normalization constant. Since the sum function K(x) and G(x) do not differ much, K(x) is called the projection of G(x).
[0167] Substitute Equation (14) into Equation (15) to obtain
[0168]
[0169] Observing Equation (16), where the first term is the time value and the vector direction of the second term is consistent with the gradient direction, the separate expressions satisfy:
[0170]
[0171] where m h,G (x) is the mean shift amount.
[0172] (a) MSLS adjusts the window according to the calculated mean shift vector. When the shift amount is less than the threshold, similar pixels are merged and represented by I MS (x).
[0173] (b) The I MS (x) image is processed in layers from largest to smallest according to the pixel values, satisfying:
[0174] I MS (x) = [I 255 , …, I0] (18)
[0175] (c) Let there be an n×n - dimensional window W n×n , and use this window to move in the highest - pixel - layer image I 255 . When all pixels within the window are 255, mark the coordinates of all pixels within the window. When there is one or more non - 255 values within the window, all pixels within the window are not marked, and continue to move the window to traverse the entire I 255 layer.
[0176] (d) The pixel values at the marked coordinates remain unchanged, and the pixel values at the unmarked coordinates are decreased by 1, and used as the pixel values of the I 254 layer to perform the operations in step (c).
[0177] (e) Repeat steps (c) and (d) until after traversing the I1 layer and then stop. The I0 layer consists of its own pixel points and the pixels with pixel values decreased by 1 at the unmarked coordinates of I1.
[0178] (f) Re - compose all the processed layers into the image I MSLS , which is the clustering result of MSLS.
[0179] To clearly present the clustering effects of MSLS and the mean - shift algorithm, we will Figure 10 in the following, draw some pixels of the surf of each effect diagram as follows:
[0180] Observing Figure 11, both the mean shift algorithm and MSLS can smooth and cluster the image. However, by observing the structure of the source image, it is obvious that the clustering algorithm of MSLS preserves the structure of the region more completely.
[0181] Let I MSLS represent the smoothed image after the source image is smoothed by MSLS. As the intermediate frequency layer of the source image, it should be noted that in the fusion algorithm process of the present invention, the smoothing window W n×n has a dimension n of 150. Then the high-frequency layer image I H is obtained by subtracting the intermediate frequency layer image I E from the source image I MSLS :
[0182] I H = I E - I MSLS (19)
[0183] The low-frequency layer image I L is obtained by subtracting the source image I MSLS from the intermediate frequency layer image I E :
[0184] I L = I MSLS - I E (20)
[0185] The multi-scale decomposition result of the source image is as Figure 12 shown:
[0186] iii) Fusion strategy
[0187] The present invention divides the fusion into two parts. The first part: is to fuse the intermediate frequency layers of the visible light image and the infrared image. First, compare the visible light intermediate frequency layer and the infrared intermediate frequency layer, and respectively select the pixels with the largest pixel values at the corresponding coordinates of the two images to form a maximum layer, and then select the pixels with the smallest pixel values at the corresponding coordinates of the two images to form a minimum layer. The obtained results are as Figure 13 shown:
[0188] Then, the image intermediate frequency layer fusion model satisfies:
[0189]
[0190] In the formula, I M represents the fusion result of the intermediate frequency image. I MSLS1 represents the intermediate frequency layer of the visible light image, I MSLS2 represents the intermediate frequency layer of the visible light image, max(I MSLS1 , I MSLS2 ) represents the pixel standard deviation corresponding to the maximum layer of the intermediate frequency image, denoted by σ3, min(I MSLS1 , IMSLS2 ) It is indicated that the pixel standard deviation corresponding to the minimum layer of the intermediate-frequency image is represented by σ4.
[0191] Intermediate-frequency layer I M The fusion result of Figure 14 is shown as
[0192] Observe Figure 14 , I M Although the basic texture of the source image is retained, the outline of the target is not highlighted. To solve this problem, we separately extract the maximum mixed image max(I E , I I ) and the minimum mixed image min(I E , I I ), and the results are shown as Figure 15 shown.
[0193] Separate max(I E , I I ) and min(I E , I I ) are linearly fused with the intermediate-frequency layer I M by a weight of 0.1 to obtain the fusion result I EF of the new intermediate-frequency image after texture enhancement, and the expression satisfies:
[0194] I MF = I M + 0.1×(max(I E , I I ) + min(I E , I I )) (22)
[0195] The obtained I MF result is shown as Figure 16 shown.
[0196] The second part: Based on the fusion result of the intermediate-frequency layer, the high-frequency layer and the low-frequency layer of each source image are linearly fused into I MF according to the principle of expanding the texture gradient to obtain the final image fusion result. It satisfies:
[0197] I F = I MF + ω1×I H1 + ω2×I H2 - I L1 - I L2 (23)
[0198] In the formula, I F represents the result of the light mixing graph proposed by the present invention. I H1 represents the high-frequency layer of the visible light image, IH2 Represents the high-frequency layer of the infrared image. I L1 Represents the low-frequency layer of the visible light image. I L2 Represents the low-frequency layer of the infrared image. Let I L1 , I L2 For the low-frequency layer, the pixel values less than 0 are set to 0. Among them, ω1 and ω2 respectively represent the fusion weights of the high-frequency layer of the visible light image and the layer of the high-frequency infrared image, and their expressions satisfy:
[0199]
[0200]
[0201]
[0202] In the formula, W represents the total weight of the fusion of the high-frequency detail layers of each source image (for example, when water and ink are fused, the obtained volume is not simply the sum of the volumes, but less than the sum of the two volumes). σ5 represents the standard deviation of the pixel values of the image I after visible light enhancement E of, σ6 represents the standard deviation of the pixel values of the infrared image I I of. σ7 represents the standard deviation of the pixel values of the intermediate-frequency layer I of the visible light MSLS1 of. σ8 represents the standard deviation of the pixel values of the intermediate-frequency layer I of the visible light MSLS2 of.
[0203] For the convenience of subsequent description, the present invention proposes a clustering algorithm based on Mean Shift and layer smoothing (MSLS), and names the fusion result of the mixed light I F as a weighted image fusion algorithm MSLSWF based on Mean Shift and layer smoothing clustering. The result of the mixed light I F , that is, the fusion result of MSLSWF is as Figure 17 shown.
[0204] 4. Experiments and Analysis
[0205] All the experiments of the present invention were run on the Windows 10 system, with a CPU of 2.60 GHz and a running memory of 8 GB. The software used was MATLAB 2016a version. The present invention mainly verified the advantages of the infrared and visible light image fusion method proposed based on the principle of light propagation and the theory of retinal imaging from subjective evaluation, and verified the quality of the light mixed image MSLSWF proposed by the present invention by subjective and objective evaluations. The comparison algorithms included ADF, CP, CVT, DTCWT, GFF, GT, GTF, LP, LP-SR, MSVD, PCNN, SAL, and Wavelet. The objective evaluation parameters were average gradient AG, information entropy H, standard deviation SD, spatial frequency SF, edge intensity EI, fusion quantity function Q ab / f , amount of artifacts N ab / f , and fusion loss function L ab / f . Among them, the larger the evaluation values of AG, H, SD, SF, and EI, the better the image quality. The larger the Q AB / F value, the better. The larger the value, the more information of the source images the fused image contains. The smaller the N AB / F value, the better. The smaller the value, the fewer artifacts generated by the fused image. The smaller the L AB / F value, the better. The smaller the value, the smaller the loss of source image information during the fusion process.
[0206] The present invention designed a total of 10 groups of experiments. The experimental results are as Figure 18 shown, and the corresponding objective evaluation values are as Figure 19As shown. Observe the first group of experiments, which represent military fortresses hidden in the mountains and forests. It is very difficult to discover the fortresses just by looking at the visible light images. Although the outline of the fortresses can be seen clearly in the visible light, the environmental layout around the fortresses cannot be seen clearly. The purpose of image fusion is to be able to see the fortresses clearly while also seeing the surrounding environment. GTF can highlight the outline of the fortresses, but its performance in highlighting the environmental texture is relatively poor compared to other algorithms. The GFF algorithm has the opposite effect to the GTF algorithm, with a better effect in highlighting the environmental texture but basically unable to see the outline of the fortresses clearly. In order to fuse the information of the two source images into one image, ADF, CVT, DTCWT, MSVD, and Wavelet weaken the texture of the visible light image and the infrared image in the fused image. Although LP-SR, PCNN, SAL, and MSLSWF can fuse the texture of each source image into one image, it is relatively difficult to distinguish the texture of the fortresses compared to CP, GT, and LP. The reason is that the fusion theory of the purpose highlights the target by the brightness of the grayscale image, which is easily interfered with by the brightness in the environment. When observing the image of '0ur', the present invention uses purple to highlight the target, effectively solving the defect of the traditional fusion theory. Observe the second group of experiments, which represent an airplane hidden in the forest. Artifacts appear in the fused images of CP, GFF, and SAL, and a halo appears in the fused image of PCNN. The fusion effects of the other algorithms are not very different. Observing the fusion result of MSLSWF, not only can the texture of the tree branches and leaves be observed, but also the outline of the airplane can be seen in the image. The fusion result of 'our' can bring good visual effects both in highlighting the environmental texture and the outline of the image target. Observe the third group of experiments, which represent a picture of a person squatting on the river bank. Observe the fused image. The fusion result of CP does not incorporate the information of the infrared image into the image. The fused result of GFF has a relatively incomplete infrared target. The fused result of GTF not only does not incorporate the information of the infrared image into the image, but also the texture of the visible light image is relatively blurred in the fused image. The effects of the other algorithms are not very different, but they all have a common feature that due to the relatively complex environmental brightness, it is difficult for users to distinguish the layout of the environment in the fused image. Observing the fusion effect of '0ur', users can effectively distinguish the texture levels of the target and the environment clearly.Experiment 4 shows the scene of the street at night. Just looking at the visible light image, it is impossible to distinguish the distribution of targets such as people and cars. Observing the infrared image can clearly see the textures of heat-emitting targets such as people and cars, but it is impossible to distinguish the information of non-heat source targets such as street-side texts. The fusion result of CP only enhances the visible light image and does not integrate the information of the infrared image into the image. Artifacts such as spots appear in the fusion results of GFF and SAL. Jagged cracks appear on the ground in the fusion result of PCCN. The fusion effects of the remaining algorithms except MSLSWF are not very different, but the fusion results of these algorithms fail to reflect the information of the tables and chairs in the street-side shops. MSLSWF can not only effectively integrate the information contained in the visible light image and the infrared image into the image, but also effectively display the information of the tables and chairs in the street-side stores in the fused image. The fusion result of '0ur' can effectively and reasonably integrate the information of the visible light image and the infrared image into one image, which is beneficial to target detection. Experiment 5 shows the fusion of the visible light image and the infrared image of a military vehicle. The visible light image can only judge the parking position of the vehicle and cannot know the driving situation of the vehicle. The infrared image can judge whether the vehicle has just parked here after driving or has been parked here without driving according to the heat radiation phenomena of the tires and the engine. The fusion result of CP does not reflect the heat source information in the infrared image. Image distortion appears in the fusion result of GFF. The fusion result of GTF does not reflect the texture information of the scene in the visible light. The fusion result of PCNN fails to integrate the information of the clouds in the infrared into the image. Artifacts appear in the fusion result of SAL. MSLSWF and the remaining algorithms can all effectively integrate the information of the visible light image and the infrared image into one image, but MSLSWF can effectively judge that the vehicle is a military vehicle by depicting the body texture. The fusion result of '0ur' not only has the advantages of the above algorithms but also has higher clarity. Experiment 6 shows the discovery of people hidden in the forest. The fusion result of ADF weakens the texture of the heat source targets in the infrared image, resulting in incomplete information of the people. Texture blurring appears in some areas in the fusion results of CP, GFF, and GTF. The overall brightness of the fusion result of MSLSWF is relatively high, especially the brightness of the road, which is very likely to attract users to observe the target information. The fusion result of '0ur' uses the result of MSLSWF as the light mixing image. Under the combined action of the visible light image and the infrared image, it uses color to distinguish the brightness, making it easy for users to discover the target information.The seventh group of experiments represents the scenario of detecting people in a low-light environment. For the fusion results of CP, CVT, and SAL, artifacts appear. The fusion result of GFF is affected by a large number of patches. The fusion result of GTF weakens the texture information of the visible light image. For the fusion result of PCNN, tearing occurs on the ground. The fusion result of MSLSWF can not only clearly show the texture information of the scene but also observe the texture information on the person. It can be judged that the person is walking forward based on the person's texture information. However, the fusion results of the other algorithms only reflect the contour information of the person and cannot determine whether the person is walking forward or to the right. The fusion result of '0ur' not only has all the advantages of MSLSWF but also can distinguish the brighter regions and target information according to the color. Experiment 8 represents the scenario where people are hidden in the forest. All algorithms can meet the requirement of fusing the information of the visible light image and the infrared image into one image. However, the brightness of the leaves easily distracts the user's attention. The fusion result of '0ur' can effectively avoid the situation where the brightness interferes with the target. Experiment 9 represents the scenario where people are in the mountains. For the fusion result of GFF, patches appear. For the fusion results of GTF and SAL, the scene texture becomes blurred. For the fusion result of PCNN, serrations appear in the texture. The other algorithms can integrate the information of the infrared image and the visible light image into one image. The fusion result of '0ur' has obvious advantages in distinguishing the scene texture levels. Experiment 10 represents the scenario of detecting people or weapons in the smoke, which is very common on the battlefield. For the fusion result of ADF, it is difficult to find the position information of the person. For the fusion result of GFF, not only can it not find the position information of the person, but also a large number of patches affect the quality of the fused image. For the fusion result of GTF, only the information of the infrared image is observed, and it is impossible to judge whether there is smoke in this area. For the fusion result of PCNN, texture tearing occurs and the position information of the person cannot be observed. The fusion result of SAL is relatively poor. The fusion result of MSLSWF can not only clearly reflect the position information of the person and the weapon but also present the layout of the scene in the visible light image completely to the user. The fusion result of '0ur' uses colors to highlight the target information and easily attracts the user's attention.
[0207] Observation Figure 19 , which is a line chart of the objective evaluation values of each algorithm, mainly to judge the performance of the objective evaluation value of the MSLSWF algorithm proposed in the present invention. When the performance of the MSLSWF algorithm is excellent, it proves that the quality of '0ur' composed of the fusion results of MSLSWF is also excellent. Because 'our' fully gives the opportunity for the visible light image and the infrared image to be presented simultaneously in the fused image. Therefore, the biggest influencing factor for the quality of the '0ur' fused image is the light mixed image MSLSWF. Observation Figure 19From the line charts of the evaluation values of AG, H, SD, SF, and EI, it can be clearly observed that the results of MSLSWF are all at a high level. For the evaluation value L of the loss function ab\f , the evaluation value of MSLSWF is at a low level, indicating that the fusion result of MSLSWF causes less loss of information in the visible light image and the infrared image during the fusion process. The evaluation value Q for judging the amount of information fusion ab\f , the evaluation value of MSLSWF is at a medium level, and MSLSWF performs mediocre in this evaluation value. The N for judging the amount of artifacts ab\f , the sub-evaluation value of MSLSWF is at a high level, indicating that there are more artifacts in the MSLSWF image. However, no artifacts are observed in the fusion result of MSLSWF in the subjective evaluation. Therefore, the reason why MSLSWF performs unsatisfactorily in the objective evaluation values Q ab\f and N ab\f is that in the present invention, the visible light image is enhanced, resulting in too large a difference in the evaluation value when comparing the differences between the fused image and the source image.
[0208] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined arbitrarily with each other.
Claims
1. An infrared and visible light image fusion method, characterized in that: It includes the following steps: S1. Achieve the fusion of the intermediate frequency layers of visible light images and infrared images; First, compare the visible light intermediate frequency layer and the infrared intermediate frequency layer. Respectively select the pixels with the largest pixel values at the corresponding coordinates of the two images to form a maximum layer, and then select the pixels with the smallest pixel values at the corresponding coordinates of the two images to form a minimum layer. Then, the intermediate frequency layer fusion model of the image satisfies: Wherein, I M represents the fusion result of the intermediate-frequency image; I MSLS1 represents the intermediate-frequency layer of the visible-light image, I MSLS2 represents the intermediate-frequency layer of the visible-light image, max(I MSLS1 , I MSLS2 ) represents the pixel standard deviation corresponding to the maximum layer of the intermediate-frequency image, denoted by σ3, and min(I MSLS1 , I MSLS2 ) represents the pixel standard deviation corresponding to the minimum layer of the intermediate-frequency image, denoted by σ4; Extract the maximum mixed image max(I E , I I ) and the minimum mixed image min(I E , I I ) of the source image respectively. Linearly fuse max(I E , I I ) and min(I E , I I ) with the intermediate frequency layer I M according to the weight of 0.1 to obtain the fusion result I EF of the new intermediate frequency image after texture enhancement. The expression satisfies: I MF = I M + 0.1 × (max(I E , I I ) + min(I E , I I )); S2. On the basis of the intermediate frequency layer fusion result, linearly fuse the high-frequency layer and the low-frequency layer of each source image into I according to the principle of expanding the texture gradient with an adaptive weight, MF to obtain the final image fusion result, satisfying: I F = I MF + ω1 × I H1 + ω2 × I H2 - I L1 - I L2 Where, I F represents the result of the light mixing diagram; I H1 represents the high-frequency layer of the visible light image, and I H2 represents the high-frequency layer of the infrared image; I L1 represents the low-frequency layer of the visible light image; I L2 represents the low-frequency layer of the infrared image; let I L1 , I L2 the pixel values less than 0 in the low-frequency layer are 0; where, ω1 and ω2 respectively represent the fusion weights of the visible light high-frequency layer and the infrared high-frequency image layer, and their expressions satisfy: Wherein, W represents the total weight of the fusion of the high-frequency detail layers of each source image; σ5 represents the standard deviation of the pixel values of the image I after visible light enhancement E , σ6 represents the standard deviation of the pixel values of the infrared image I I ; σ7 represents the standard deviation of the pixel values of the intermediate-frequency layer I of visible light MSLS1 ; σ8 represents the standard deviation of the pixel values of the intermediate-frequency layer I of visible light MSLS2 .
2. The infrared and visible light image fusion method according to claim 1, wherein: In the step S1, use the histogram enhancement algorithm to perform adaptive weighted linear fusion with the source image, and the expression is as follows: Wherein, I E represents visible light image enhancement, I V represents a visible light image, H isteq (·) represents histogram enhancement processing, and σ1 represents the standard deviation of the pixels of the I V image, and σ2 represents the standard deviation of the pixels of the H isteq (·) image.
3. The infrared and visible light image fusion method according to claim 1, characterized in that: In the said step S2, let I MSLS represent the smoothed image after the source image is smoothed by MSLS, which is used as the intermediate frequency layer of the source image. It should be noted that in the fusion algorithm process, the dimension n of the smoothing window W n×n is 150, then the high-frequency layer image I H is obtained by subtracting the intermediate frequency layer I E from the source image I MSLS : I H = I E - I MSLS Low-frequency layer image I L Subtracted from the intermediate-frequency layer I MSLS Subtract the source image I E Obtain: I L = I MSLS - I E 。 4. Application of an infrared and visible light image fusion method as described in claim 1, characterized in that: It is used to simulate a mixed light image.
5. The application according to claim 4, characterized in that: It is used to form a pseudo-color image with infrared and visible light.
Citation Information
Patent Citations
Image fusion algorithm based on texture features
CN111507913A
Image fusion scheme for differential phase contrast imaging
WO2014194995A1