A Multimodal Remote Sensing Image Fusion Method Based on Multi-Directional Gradient Filtering
The multi-directional gradient filtering technique addresses low contrast and edge blurring in multi-modal image fusion by optimizing low-frequency layers and enhancing edge preservation, resulting in higher quality fused images.
Patent Information
- Application Number
- CN202210563729.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-23
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-05-23
AI Technical Summary
In the existing multimodal image fusion method, the image quality has problems such as low contrast and blurred edges, and the image quality needs to be further improved.
Using a multi-directional gradient filtering method, multiple Laplace pyramid sequences are obtained through Laplace decomposition, and the fusion factor of the significant layer and the residual layer is obtained by filtering optimization problems. Image reconstruction is carried out in combination with Laplace inverse transformation to preserve the brightness and edge information of the image.
Improves contrast and edge clarity of multimode fusion images, enhances the overall quality of the image, and preserves significant information in infrared or SAR images and texture details in visible images.
Smart Images

Figure CN115147691B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing, and particularly relates to a multi-modal image fusion method based on multi-directional gradient filtering. Background Art
[0002] In recent years, the multi-sensor technology and computer vision technology have been developing vigorously. Under this background, the multi-sensor image fusion technology has been widely applied. For example, in the field of remote sensing, the fusion of a large number of remote sensing images combines source images with complementarity to obtain a fused image with high clarity, large amount of information, and rich texture details, realizing an accurate description of the detection scene and providing a technical approach for more convenient and comprehensive understanding of the environment and natural resources. In the medical field, the multi-modal image fusion technology has greatly improved the ability of computer-aided diagnosis and treatment, saving a great deal of medical examination time to a large extent. Driven by the demands in various fields, scholars in various countries have paid more and more attention to the research on multi-sensor image fusion technology.
[0003] Multi-modal image fusion refers to fusing the scene images obtained by different-modal sensors or the scene images obtained by the same sensor at different times under the same geographical coordinates after filtering and denoising, time registration, spatial alignment, and resampling, through image fusion technology; the fused image breaks through the limitations of single-source images in terms of spatial resolution, physical properties, information amount, clarity, etc., and has high image quality, thus being conducive to the positioning, recognition, and accurate interpretation of physical phenomena and events.
[0004] According to different information representation levels, multi-modal image fusion can be divided into three levels, namely pixel level, feature level, and decision level. Compared with feature-level and decision-level image fusion, pixel-level image fusion is more universal. Its processing object is image pixels, using a certain image fusion strategy to extract useful information from the source image and eliminate redundant information, retaining the features and background information of the source image to a greater extent, and the obtained fusion result shows a high degree of comprehensiveness, accuracy, and reliability.
[0005] There are various existing pixel-level multi-modal image fusion methods. For example, there are NSST (Nonsubsampled Shearlet Transform) methods and LatLRR (latent low-rank representation) methods, etc. The general process of such existing methods includes: decomposing the multi-modal image into layers to obtain the representation information of different properties of the image, performing layer fusion by constructing a corresponding fusion strategy, and finally obtaining the corresponding multi-modal fusion image through inverse reconstruction operations.
[0006] However, the multi-modal fusion images obtained by existing methods have problems such as low contrast and blurred edges, and the image quality needs to be further improved. Summary of the Invention
[0007] To solve the above problems existing in the prior art, the present invention provides a multi-modal remote sensing image fusion method based on multi-directional gradient filtering.
[0008] The technical problems to be solved by the present invention are realized through the following technical solutions:
[0009] A multi-modal image fusion method based on multi-directional gradient filtering, comprising:
[0010] Obtain multiple source images of different modalities to be fused, and perform Laplace decomposition on each of the source images to obtain a plurality of Laplacian pyramid sequences; wherein, each Laplacian pyramid sequence includes a low-frequency layer and a plurality of high-frequency layers;
[0011] Substitute each of the low-frequency layers into a filtering optimization problem for iterative solution to obtain a filtering result corresponding to the low-frequency layer, and use the filtering result as the significant layer corresponding to the low-frequency layer to calculate the residual layer corresponding to the low-frequency layer; wherein, the filtering optimization problem is an optimization problem with the filtering effect as the optimization target, and the optimization target at least includes the similarity of the brightness distribution of the image before and after filtering and the total eight-direction gradient after filtering.
[0012] Calculate a significant layer fusion factor based on the significant layer of the target source image, and calculate a residual layer fusion factor based on each of the residual layers; wherein, the target source image is an infrared image or a SAR image among the multiple source images of different modalities.
[0013] Perform significant layer fusion and residual layer fusion on the source images of different modalities respectively based on the significant layer fusion factor and the residual layer fusion factor, and use the significant layer fusion result and the residual layer fusion result to obtain a low-frequency layer fusion result.
[0014] Perform high-frequency layer fusion on the source images of different modalities, and perform image reconstruction using the inverse Laplace transform method based on the high-frequency layer fusion result and the low-frequency layer fusion result to obtain a multi-modal fusion image.
[0015] Optionally, the filtering optimization problem is:
[0016]
[0017] Wherein, use O to represent the filtered image during the iteration process, O c represents the pixel value of any pixel point c of the filtered image, the coordinates of the pixel point c are (x, y), I crepresents the pixel value of the pixel points with the same coordinates (x, y) in the low-frequency layer I; ||·||2 represents the L2 norm; respectively represent the gradient operators in eight directions: left, right, up, down, upper right, lower right, upper left, and lower left; P c,l , P c,r , P c,u , P c,d , P c,ru , P c,rd , P c,lu , P c,ld are respectively corresponding penalty coefficients; each of the penalty coefficients is calculated according to the corresponding gradient operator; O * is the optimal solution obtained by iterative solution, that is, the filtering result.
[0018] Optionally, the calculation method of the penalty coefficient is as follows:
[0019]
[0020] where direct represents the direction, which are left, right, up, down, upper right, lower right, upper left, and lower left; α is a preset gradient threshold, β is a preset slope control factor, and P max is a preset maximum penalty coefficient.
[0021] Optionally, the method for calculating the saliency layer fusion factor based on the saliency layer of the target source image includes:
[0022] m = mean2(A1),
[0023] λ = 10 * m,
[0024] s = std2(A1),
[0025] R = |A1|,
[0026]
[0027]
[0028] where the bold A1 represents the saliency layer of the target source image, the non-bold A1 represents any pixel point of the saliency layer, R is the modulus value of A1, and R max is the maximum modulus value among all pixel points of the saliency layer; mean2(·) represents calculating the average value, std2(·) represents calculating the standard deviation; w is a saliency layer fusion factor corresponding to the target source image, and the saliency layer fusion factor corresponding to the non-target source image is determined according to w.
[0029] Optionally, calculating the residual layer fusion factor based on each residual layer includes:
[0030] Calculating the Sobel gradient map of each residual layer using the Sobel operator;
[0031] Diffusing each of the Sobel gradient maps using a Gaussian filter to obtain the corresponding diffused gradient map;
[0032] Calculating the residual layer fusion factor corresponding to each residual layer based on each diffused gradient map.
[0033] Optionally, calculating the residual layer fusion factor corresponding to each residual layer based on each diffused gradient map includes:
[0034]
[0035] Wherein, represents the diffused gradient map corresponding to the k-th ∈ [1, K] residual layer, and S R represents the sum of all diffused gradient maps, and K is the total number of the multiple source images of different modalities, where K ≥ 2.
[0036] Optionally, both the Gaussian filter template and the standard deviation value of the Gaussian filter are taken as 5.
[0037] Optionally, performing high-frequency layer fusion on the source images of different modalities includes:
[0038] Performing high-frequency layer fusion on the source images of different modalities using the maximum absolute value strategy.
[0039] Optionally, the multiple source images of different modalities have non-negligible parallax;
[0040] Before performing Laplace decomposition on each of the source images, the multi-modal image fusion method further includes: registering and aligning the multiple source images of different modalities obtained.
[0041] Optionally, the multiple source images of different modalities include visible light images;
[0042] The multi-modal image fusion method further includes:
[0043] Before registering and aligning the multiple source images of different modalities obtained, performing IHS transformation on the visible light image and extracting the I component in the transformation result to form a new image, and replacing the visible light image with the new image as the source image;
[0044] And, after obtaining the multi-modal fusion image, performing inverse IHS transformation on the multi-modal fusion image.
[0045] In the multi-modal remote sensing image fusion method based on multi-directional gradient filtering provided by the present invention, the multi-scale property of the Laplacian pyramid and the multi-directional property of multi-directional gradient filtering are used to extract the basic information at different scales and the texture information in different directions in the source images. Thereby, the noise edges and fine texture information in the source images can be removed, and the overall brightness and edge information of the images are better retained, so that the multi-modal fusion image has a higher contrast and clear edges, improving the image quality of the multi-modal fusion image.
[0046] In the present invention, the saliency layer fusion factor is dynamically calculated based on the saliency layer of the target source image, that is, an adaptive fuzzy coupling mechanism is adopted to achieve saliency layer fusion; moreover, in the process of calculating the residual layer fusion factor based on each residual layer in the present invention, a gradient saliency mapping strategy is adopted to achieve the fusion of the residual layers. Specifically, a Gaussian filter is used to diffuse the Sobel gradient map of the residual layer, and the residual layer fusion factors corresponding to each residual layer are respectively calculated based on each diffused gradient map. Thus, after obtaining the low-frequency layer fusion result by using the saliency layer fusion result and the residual layer fusion result, the low-frequency layer fusion result can retain as much as possible the significant information in the infrared image or SAR image and the texture detail information in the visible light image, further improving the quality of the multi-modal fusion image.
[0047] In addition, in the present invention, the maximum absolute value strategy is also used to perform high-frequency layer fusion on the source images of different modalities, so that the fused high-frequency layer can retain the edge information of the source images to the greatest extent, thereby further improving the quality of the multi-modal fusion image.
[0048] The following will further elaborate on the present invention in conjunction with the accompanying drawings. Description of the Drawings
[0049] Figure 1 is a flowchart of a multi-modal image fusion method based on multi-directional gradient filtering provided by an embodiment of the present invention;
[0050] Figure 2 is a schematic diagram of the process of Laplacian pyramid decomposition;
[0051] Figure 3 is according to Figure 1 is a schematic diagram of the process when fusing two multi-modal images containing visible light images according to the method shown;
[0052] Fig. 4(a) shows a set of visible light image and SAR image to be fused;
[0053] Fig. 4(b) respectively shows the multi-modal fusion images obtained by fusing the two images in Fig. 4(a) using the existing NSST method, the existing LatLRR method, and the method provided by the embodiment of the present invention;
[0054] Figure 5(a) shows another set of visible light images and SAR images to be fused;
[0055] Figure 5(b) respectively shows the multi-modal fusion images obtained by fusing the two images in Figure 5(a) using the existing NSST method, the existing LatLRR method, and the method provided by the embodiment of the present invention;
[0056] Figure 6(a) is a set of visible light images and infrared images to be fused;
[0057] Figure 6(b) respectively shows the multi-modal fusion images obtained by fusing the two images in Figure 6(a) using the existing NSST method, the existing LatLRR method, and the method provided by the embodiment of the present invention;
[0058] Figure 7(a) is a set of visible light images and infrared images to be fused;
[0059] Figure 7(b) respectively shows the multi-modal fusion images obtained by fusing the two images in Figure 7(a) using the existing NSST method, the existing LatLRR method, and the method provided by the embodiment of the present invention. Detailed implementation manners
[0060] The following further describes the present invention in detail with reference to specific embodiments, but the implementation manners of the present invention are not limited thereto.
[0061] The multi-modal image fusion technology is also called the heterogeneous image fusion processing technology, which can perform multi-faceted and in-depth combination and screening on the multi-sensor image information corresponding to the same ground object target, while integrating and optimizing the complementary information and eliminating the redundant information.
[0062] The imaging mechanisms and image representation forms of heterogeneous images are significantly different, and the information of the target in different electromagnetic spectrum bands can be obtained respectively, which has obvious complementarity and fusion value. However, it cannot be ignored that the gray difference between heterogeneous images and the geometric distortion caused by the non-negligible parallax make the general image fusion results show phenomena such as spectral distortion and low contrast, and the image quality is poor.
[0063] Therefore, in view of the above problems, in order to improve the multi-modal image fusion quality, the embodiment of the present invention provides a multi-modal image fusion method based on multi-directional gradient filtering to achieve high-quality fusion of heterogeneous images, so as to obtain a fusion result with rich information. The multi-modal image fusion method provided by the embodiment of the present invention is a pixel-level multi-modal image fusion method.
[0064] See Figure 1 As shown, the multi-modal image fusion method provided by the embodiment of the present invention includes the following steps:
[0065] S1: Obtain multiple source images of different modalities to be fused, and perform Laplacian decomposition on each source image to obtain multiple Laplacian pyramid sequences; where each Laplacian pyramid sequence includes a low-frequency layer and multiple high-frequency layers.
[0066] Here, the multiple source images of different modalities to be fused can include any two or all three of visible light images, infrared images, and SAR (Synthetic Aperture Radar) images. Among them, when the source image includes a visible light image, the visible light image can be subjected to IHS (also known as HSI) transformation, and the I component in the transformation result is extracted to form a new image to replace the visible light image as the source image. IHS is a digital image model that perceives colors with three basic characteristic components: hue (abbreviated as H), saturation (abbreviated as S), and intensity (abbreviated as I).
[0067] Among them, the process of IHS transformation is as follows:
[0068]
[0069]
[0070]
[0071] Among them, R, G, and B are the R component, G component, and B component of the visible light image respectively; V1 and V2 are intermediate variables; H, S, and I are the H component, S component, and I component in the HIS transformation result respectively.
[0072] Pyramid decomposition is a main implementation means in the field of multi-scale fusion and is also the most widely studied image decomposition algorithm. The Laplacian pyramid is an improvement of the Gaussian pyramid and is also called the prediction residual pyramid. Its advantage is that it can restore the image to the greatest extent and has a relatively fast decomposition speed. The process of Laplacian pyramid decomposition can be referred to Figure 2 as shown Figure 2 which represents the decomposition process from the low layer to the high layer in one-dimensional space. The actual decomposition process is carried out in two-dimensional space.
[0073] Specifically, the original source image is used as the bottom layer image of the Gaussian pyramid. This image can be regarded as an array I0 with R rows and C columns, and each pixel value represents the intensity at that position. The image of the first layer of the pyramid is the result of low-pass filtering and downsampling of the bottom layer image, and then the image of the second layer of the pyramid is obtained by performing low-pass filtering and downsampling on the first layer image, and so on to construct the Gaussian pyramid. The decomposition process can be expressed by the formula as follows:
[0074]
[0075] 1 ≤ l ≤ N, 0 ≤ j ≤ R l , 0 ≤ j ≤ C l ;
[0076] Among them, I l (i, j) represents the Gaussian pyramid image of the l-th layer, and N represents the total number of pyramid layers. R l and C l represent the number of rows and columns of the Gaussian pyramid image of the l-th layer, ω(m, n) is a two-dimensional window function, and its definition is: ω(m, n) = h(m) × h(n); among them, h(·) is a Gaussian density distribution function that satisfies the following constraint conditions:
[0077]
[0078] h(i) = h(-i),
[0079] h(0) + h(-2) + h(2) = h(-1) + h(1).
[0080] According to the definition of the window function, the window function ω(m, n) used in the actual decomposition process can be obtained by calculation as:
[0081]
[0082] Through the above formula, Laplacian decomposition is performed on each source image, and a sequence of Gaussian pyramid sub-images I0, I1,..., I N .
[0083] After constructing all the sequences of Gaussian pyramid sub-images, subtracting each pair of adjacent layers can obtain the corresponding Laplacian pyramid sequence. In this process, since the sizes of the Gaussian pyramid sub-images of adjacent layers are different and do not meet the requirements of matrix difference operations; therefore, before each matrix difference operation, upsampling is performed on the adjacent Gaussian pyramid sub-image of the upper layer to make the sizes of the Gaussian pyramid sub-images of the upper and lower layers consistent, so as to meet the requirements of matrix difference operations.
[0084] Specifically, after interpolating and dilating I l it becomes At this time the size of is the same as that of I l-1 and is expressed by the formula as follows:
[0085]
[0086] 1 ≤ l ≤ N, 0 ≤ j ≤ R l , 0 ≤ j ≤ C l ;
[0087] Among them,
[0088]
[0089] After interpolation expansion, a series of interpolation sequences can be obtained Then, pairwise subtraction is performed between adjacent layers using this interpolation sequence, and the Laplacian pyramid sequences LP1, LP2, …, LP can be obtained N . Among them, the definition of the Laplacian pyramid image of the l-th layer is as follows:
[0090]
[0091] Among them, LP with the most upsampling times and the most blurred image N is the low-frequency layer, and the remaining are all high-frequency layers.
[0092] S2: Substitute each low-frequency layer into a filtering optimization problem for iterative solution to obtain the filtering result corresponding to this low-frequency layer, and use this filtering result as the significant layer corresponding to this low-frequency layer, and calculate the residual layer corresponding to this low-frequency layer; among them, the filtering optimization problem is an optimization problem with the filtering effect as the optimization goal, and the optimization goal includes at least the similarity of the brightness distribution of the image before and after filtering and the total gradient in eight directions after filtering.
[0093] In one implementation, the above filtering optimization problem can be:
[0094]
[0095]
[0096]
[0097] Among them, in order to retain the brightness distribution of the source image, the brightness distribution of the filtered image is restricted to make its pixel brightness distribution as similar as possible to that of the source image, so the error e1 is defined; and, in order to remove as much as possible the small useless gradient changes, such as the gradient changes at the noise edges, the total gradient e2 of the filtered image in eight directions is restricted.
[0098] In the expression of this filtering optimization problem, O represents the filtered image during the iterative process, and O c represents the pixel value of any pixel point c of the filtered image, and the coordinates of this pixel point c are (x, y), and I c represents the pixel value of the pixel point with the same coordinates (x, y) in the low-frequency layer I; ||·||2 represents the L2 norm; Respectively represent the gradient operators in eight directions: left, right, up, down, upper right, lower right, upper left, and lower left; O * is the optimal solution obtained by iterative solution, that is, the filtering result obtained by iterative solution.
[0099] It can be understood that the filtering result obtained based on the filtering optimization problem can not only keep the brightness distribution similar to the source image, but also remove useless small gradient values such as fine textures and noise, and retain more useful gradients.
[0100] In another implementation, in order to retain more edges with large gradients while removing fine textures or noise, a gradient threshold can be introduced to distinguish edges, fine textures, and noise. At the same time, a penalty coefficient is designed based on the gradient threshold, and a smaller penalty coefficient is given to edges with larger gradients, while a larger penalty coefficient should be given to textures and noise with smaller gradients; thus, the above filtering optimization problem can be further optimized as follows:
[0101]
[0102] Among them, P c,l , P c,r , P c,u , P c,d , P c,ru , P c,rd , P c,lu , P c,ld They are The corresponding penalty coefficient; each penalty coefficient is calculated according to the corresponding gradient operator, and the calculation method is as follows:
[0103]
[0104] Wherein, direct represents the direction, which are left, right, up, down, upper right, lower right, upper left, and lower left; α is the preset gradient threshold, β is the preset slope control factor, and P max is the preset maximum penalty coefficient.
[0105] After obtaining the filtering result corresponding to the low-frequency layer by using the filtering optimization problem, the filtering result is the significant layer corresponding to the low-frequency layer. Then, the residual layer corresponding to the low-frequency layer is obtained by subtracting the significant layer from the low-frequency layer.
[0106] S3: calculating a significant layer fusion factor based on a significant layer of a target source image, and calculating a residual layer fusion factor based on each residual layer; wherein the target source image is an infrared image or a SAR image among multiple source images of different modalities.
[0107] Specifically, when the multiple source images of different modalities to be fused include visible light images and infrared images, the target source image is the infrared image among them. When the multiple source images of different modalities to be fused include visible light images and SAR images, the target source image is the SAR image among them. When the multiple source images of different modalities to be fused include infrared images and SAR images, the target source image can be selected between the two. The specific selection principle can be determined according to the imaging quality of the source images. The image in the modality with higher imaging quality is preferentially selected as the target source image. Similarly, when the multiple source images of different modalities to be fused include three images: visible light images, infrared images, and SAR images, the principle for selecting the target source image remains unchanged.
[0108] Among them, the method for calculating the saliency layer fusion factor based on the saliency layer of the target source image can be calculated based on an adaptive fuzzy coupling function, and its definition is as follows:
[0109] m = mean2(A1),
[0110] λ = 10 * m,
[0111] s = std2(A1),
[0112] R = |A1|,
[0113]
[0114]
[0115] Among them, the bold A1 represents the saliency layer of the target source image, which is a matrix image; the non-bold A1 represents any pixel point of this matrix image, R is the modulus value of A1, and R max is the maximum modulus value among all pixel points of this matrix image; mean2(·) represents calculating the average value, std2(·) represents calculating the standard deviation; P ∈ [0, 1]; w is a saliency layer fusion factor corresponding to the saliency layer of the target source image, specifically the factor corresponding to the pixel point A1 in the saliency layer of the target source image; the saliency layer fusion factor corresponding to the non-target source image is determined according to w. Specifically, if the non-target source image only includes visible light images, then the saliency layer fusion factor corresponding to the visible light image is 1 - w; if the non-target source image includes, in addition to visible light images, an infrared image or a SAR image, then the saliency layer fusion factor corresponding to this infrared image or SAR image is w', 0 < w' < w, and the saliency layer fusion factor corresponding to the visible light image is 1 - w - w'.
[0116] In the embodiments of the present invention, w reflects the distribution of the light and dark characteristics of the significant layer of the target source image. The larger w is, the more obvious the light and dark characteristics are. In the actual fusion process, to avoid the loss of the contrast of the fused image, when constructing the adaptive fuzzy coupling function, w should be made as large as possible. Therefore, in other methods of calculating the fusion factor of the significant layer, λ can also take other values other than 10*m, such as 8*m or 12*m, etc., to try to obtain a larger w. Alternatively, other methods can also be used to normalize R to obtain P, which are all reasonable and achievable; any method of dynamically and adaptively calculating the fusion factor of the significant layer based on the significant layer of the target source image can be applied to the embodiments of the present invention.
[0117] The residual layer mainly reflects the small gradient changes in the low-frequency image layer, including fine textures and noises. Therefore, the gradient saliency mapping matrix of the residual layer can be constructed according to the gradient of the image, and the fusion of the residual layer can be realized based on this gradient saliency mapping matrix, so that the fused residual layer can retain the texture information in the visible light image as much as possible.
[0118] Specifically, there are multiple ways to calculate the fusion factor of the residual layer based on each residual layer in step S3.
[0119] Exemplarily, in one calculation method, calculating the fusion factor of the residual layer based on each residual layer may include:
[0120] (1) Using the Sobel operator to calculate the Sobel gradient map of each residual layer;
[0121] Here, the Sobel gradient map is the constructed gradient saliency mapping matrix; the Sobel operator used is expressed by the formula:
[0122] Sobel(S(x,y)) = |(S (x+1,y-1) +2S (x+1,y) +S (x+1,y+1) -S (x-1,y-1) +2S (x-1,y) +S (x-1,y+1) )| + |(S (x-1,y+1) +2S (x,y+1) +S (x+1,y+1) -S (x-1,y-1) +2S (x,y-1) +S (x+1,y-1) )|;
[0123] Among them, Sobel(·) represents the Sobel operator, S(x,y) represents the pixel point with coordinates (x,y) in the residual layer, S (x,y-1) 、S (x,y+1) 、S (x-1,y) 、S (x+1,y) 、S (x-1,y-1) 、S (x+1,y-1) 、S(x-1,y+1) , S (x+1,y+1) are respectively the pixel points above, below, to the left, to the right, upper left, upper right, lower left, and lower right of this pixel point.
[0124] Assume that S k represents the residual layer image, and G k represents the Sobel gradient map corresponding to the residual layer image. Then the process of calculating the Sobel gradient map of the residual layer can be expressed by the following formula:
[0125] G k = Sobel(S k );
[0126] where k ∈ [1, K], K is the total number of multiple different modality source images to be fused, and K ≥ 2.
[0127] (2) Calculate the residual layer fusion factor corresponding to each residual layer based on each Sobel gradient map.
[0128] The specific calculation method is as follows:
[0129]
[0130] where S k represents the Sobel gradient map corresponding to the k-th residual layer, and S represents the sum of all Sobel gradient maps.
[0131] In another calculation method, in order to increase the influence of the gradient on the residual layer fusion factor, calculating the residual layer fusion factor based on each residual layer may include:
[0132] (1) Use the Sobel operator to calculate the Sobel gradient map of each residual layer.
[0133] The calculation method of this step (1) is the same as that in the previous calculation method.
[0134] (2) Use the Gaussian filter to diffuse each Sobel gradient map to obtain the corresponding diffused gradient map.
[0135] Specifically, this step can be expressed by the formula as follows:
[0136]
[0137] where Gaussian(·) represents the Gaussian filtering operator, G k represents the Sobel gradient map corresponding to the residual layer image, δ represents its standard deviation, r represents its Gaussian filtering template, and the specific values of δ and r can be tried out in advance; for example, both δ and r can be taken as 5, which is not limited to this.
[0138] It can be understood that at this time, the diffusion gradient map is the constructed gradient saliency mapping matrix.
[0139] (3) Calculate the residual layer fusion factors corresponding to each residual layer based on each diffusion gradient map.
[0140] Specifically, the calculation method is as follows:
[0141]
[0142] Among them, represents the diffusion gradient map corresponding to the k-th (k ∈ [1, K]) residual layer, and S R represents the sum of each diffusion gradient map.
[0143] S4: Perform significant layer fusion and residual layer fusion on the source images of different modalities based on the significant layer fusion factor and the residual layer fusion factor, and obtain the low-frequency layer fusion result by using the significant layer fusion result and the residual layer fusion result.
[0144] Among them, the process of significant layer fusion is to calculate the weighted average of the significant layers corresponding to the source images of each modality.
[0145] For example, if the source images of each modality include visible light images and infrared images, the process of significant layer fusion can be expressed as FA = w * A1 + (1 - w) * A2; where A1 represents the significant layer of the infrared image at this time, w is the corresponding significant layer fusion factor, A2 represents the significant layer of the visible light image at this time; FA represents the significant layer fusion result.
[0146] Similarly, the process of residual layer fusion is also to calculate the weighted average of the residual layers corresponding to the source images of each modality.
[0147] For example, if the source images of each modality include visible light images and SAR images, the method of significant layer fusion can be expressed as where S1 represents the residual layer corresponding to the SAR image, is its corresponding residual layer fusion factor; S2 represents the residual layer corresponding to the visible light image, is its corresponding residual layer fusion factor, and FS represents the residual layer fusion result.
[0148] After obtaining the significant layer fusion result and the residual layer fusion result, the low-frequency layer fusion result FL = FA + FS can be obtained.
[0149] S5: Perform high-frequency layer fusion on the source images of different modalities, and perform image reconstruction using the inverse Laplace transform method based on the high-frequency layer fusion result and the low-frequency layer fusion result to obtain the multi-modal fusion image.
[0150] Preferably, the maximum absolute value strategy can be adopted to perform high-frequency layer fusion on source images of different modalities. It can be understood that the high-frequency part of an image usually contains the edge contour information of the image, and this part of the information can represent the information richness of the corresponding position of the image. During the image decomposition process, the high-frequency coefficients with larger absolute values correspond to the points where the brightness changes suddenly, that is, the edge features with larger contrast changes in the image, such as boundaries, bright lines, and regional contours. The larger the absolute value of the high-frequency part, the more detailed the information at the corresponding position. Therefore, using the maximum absolute value strategy to achieve high-frequency layer fusion can obtain a high-frequency fusion image with rich edge information.
[0151] For example, assuming that the source images of different modalities include two, the process of performing high-frequency layer fusion using the maximum absolute value strategy can be expressed by the following formula:
[0152]
[0153] where \(0\leq l < N\), and are two pixel points at the same position in the high-frequency layers of two source images of different modalities, \(F_H\) l is the high-frequency layer fusion value of these two pixel points, which is located in the target source image.
[0154] After obtaining the high-frequency layer fusion result, combined with the low-frequency layer fusion result, inverse Laplace transform is used for image reconstruction, and the specific process is as follows:
[0155]
[0156] where the low-frequency layer fusion result \(F_L\) is used as the high-low frequency fusion result \(F\) of the \(l = N\) layer, N and layer by layer, the high-frequency layer fusion result \(F_H\) l of the \(l\in[0, N - 1]\) layer and the high-low frequency fusion result \(F\) l+1 of the \(l + 1\) layer are fused to finally obtain the multi-modal fusion image \(F_0\). During the fusion process, in order to facilitate matrix summation, \(F\) l+1 is interpolated and dilated to obtain
[0157] In one embodiment, as shown in Figure 3 , if the source images of multiple different modalities to be registered include visible light images, then after obtaining the multi-modal fusion image, IHS inverse transform is performed on the multi-modal fusion image to restore the color information and obtain a colored multi-modal fusion image.
[0158] The process of IHS inverse transform is as follows:
[0159] \(V_1 = S\cos(H)\),
[0160] V2 = S sin(H),
[0161]
[0162] Among them, component I is the multi-modal fusion image to be subjected to IHS transformation. For the interpretations of the remaining parameters, refer to the above description of IHS transformation and will not be elaborated here.
[0163] In the multi-modal remote sensing image fusion method based on multi-directional gradient filtering provided by the embodiments of the present invention, the multi-scale property of the Laplacian pyramid and the multi-directionality of multi-directional gradient filtering are used to extract the basic information at different scales and the texture information in different directions in the source images. Thereby, the noise edges and fine texture information in the source images can be removed, and the overall brightness and edge information of the images can be better retained, so that the multi-modal fusion image has a high contrast and clear edges, improving the image quality of the fusion image.
[0164] In the embodiments of the present invention, the significant layer fusion factor is dynamically calculated based on the significant layer of the target source image, that is, an adaptive fuzzy coupling mechanism is adopted to achieve significant layer fusion; and, in the present invention, a gradient saliency mapping strategy is adopted to achieve the fusion of the residual layers. Specifically, in the process of calculating the residual layer fusion factor based on each residual layer, the Sobel gradient map of the residual layer is diffused by using a Gaussian filter, and the residual layer fusion factor corresponding to each residual layer is calculated based on each diffused gradient map. Thus, after obtaining the low-frequency layer fusion result by using the significant layer fusion result and the residual layer fusion result, the low-frequency layer fusion result can retain as much as possible the significant information in the infrared image or SAR image and the texture detail information in the visible light image, further improving the quality of the multi-modal fusion image.
[0165] In addition, in the embodiments of the present invention, the high-frequency layer fusion of source images of different modalities can also be performed by using the strategy of taking the larger absolute value, so that the fused high-frequency layer can retain the edge information of the source images to the greatest extent, thereby further improving the quality of the multi-modal fusion image.
[0166] Optionally, in one implementation, if there is a non-negligible parallax among multiple source images of different modalities, before performing Laplacian decomposition on each source image, the multiple source images of different modalities obtained can also be registered and aligned first.
[0167] Here, the registration and alignment of the images can be implemented by using the existing SIFT (Scale-invariant Feature Transform) method, and the embodiments of the present invention will not elaborate on it.
[0168] The beneficial effects of the embodiments of the present invention are further verified and described below by using experimental results.
[0169] Figure 4(a) shows a set of visible light image and SAR image to be fused; among them, the visible light image is on the left and the SAR image is on the right; Figure 4(b) respectively shows the fused images obtained by using the existing NSST method, the existing LatLRR method, and the method provided by the embodiment of the present invention for fusing the two images in Figure 4(a).
[0170] Figure 5(a) shows another two sets of visible light image and SAR image to be fused, where the visible light image is on the left and the SAR image is on the right; Figure 5(b) respectively shows the multi-modal fused images obtained by using the existing NSST method, the existing LatLRR method, and the method provided by the embodiment of the present invention for fusing the two images in Figure 5(a).
[0171] Both Figure 4(b) and Figure 5(b) are color images in practice. Among them, the multi-modal fused image obtained by using the method provided by the embodiment of the present invention has a better visual effect, small color distortion, and better preservation of color information.
[0172] Based on the above experimental results, Table 1 gives the performance analysis and comparison results of the multi-modal fused images obtained by using three different methods.
[0173] Table 1
[0174]
[0175] In this Table 1, AG represents the average gradient, E represents the information entropy, PSNR represents the peak signal-to-noise ratio, QABF represents the fusion quality, RMSE represents the root mean square error, SF represents the spatial frequency, SSIM represents the structural similarity, and VIF is an index based on visual information fidelity.
[0176] In Table 1, the bolded values are the best results in the same index. It can be seen from Table 1 that the multi-modal fused image obtained by using the method provided by the embodiment of the present invention has more high-performance indicators, which shows that the method provided by the embodiment of the present invention has better advantages.
[0177] Figure 6(a) shows a set of visible light image and infrared image to be fused, where the visible light image is on the left and the infrared image is on the right; Figure 6(b) respectively shows the multi-modal fused images obtained by using the existing NSST method, the existing LatLRR method, and the method provided by the embodiment of the present invention for fusing the two images in Figure 6(a).
[0178] Figure 7(a) shows another group of visible light image and infrared image to be fused, where the left side is the visible light image and the right side is the infrared image; Figure 7(b) respectively shows the multi-modal fusion images obtained by fusing the two images in Figure 7(a) using the existing NSST method, the existing LatLRR method, and the method provided by the embodiment of the present invention.
[0179] Both Figure 6(b) and Figure 7(b) are actually color images in practice. Among them, the multi-modal fusion image obtained by using the method provided by the embodiment of the present invention has a better visual effect and better preserves color information; moreover, in the multi-modal fusion image obtained by using the method provided by the embodiment of the present invention, the significant regions of the infrared image are better preserved; for example, at the lower left corner of each subfigure in Figure 6(b), the present invention retains more significant features of the roof and the tree crown, while the significance in the same region in the other two comparison methods is not strong enough. And in each subfigure of Figure 7(b), the present invention retains more significant features of the pedestrians in the distance, while the significance in the same region in the other two comparison methods is not strong enough.
[0180] Based on the above test results, Table 2 gives the performance analysis and comparison results of the multi-modal fusion images obtained by using three different methods.
[0181] Table 2
[0182]
[0183]
[0184] Similarly, it can be seen from Table 2 that the method provided by the embodiment of the present invention has better advantages in comparison.
[0185] In the embodiment of the present invention, first, the Laplacian pyramid decomposition is combined with multi-directional gradient filtering to extract the basic information of different scales and the detailed information of different directions of the source image, so as to eliminate the noise edges and fine texture information in the source image and better retain the overall brightness and edge information of the image. Secondly, the embodiment of the present invention designs an adaptive fuzzy coupling mechanism and a gradient saliency mapping strategy to realize the fusion of the significant layer and the residual layer, and then the fused significant layer and residual layer are added and reconstructed as the fused low-frequency layer, so that the fusion result of the low-frequency layer can retain as much as possible the significant information in the infrared image or SAR image and the texture detail information in the visible light image. Finally, an absolute value taking the larger strategy is designed to realize the fusion of the high-frequency layer, so that the fused high-frequency layer can retain the edge information of the source image to the greatest extent, and the fused low-frequency layer and high-frequency layer are reconstructed to obtain a multi-modal fusion image with high image quality.
[0186] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. means that the specific features or characteristics described in connection with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0187] Although the present application has been described herein in connection with various embodiments, however, in the process of implementing the claimed present application, those skilled in the art can understand and achieve other variations of the disclosed embodiments by viewing the accompanying drawings, the disclosure, and the appended claims.
[0188] The above content is a further detailed description of the present invention in connection with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A multi-modal image fusion method based on multi-directional gradient filtering, characterized in that, Including: Obtain multiple source images of different modalities to be fused, and perform Laplacian decomposition on each of the source images to obtain multiple Laplacian pyramid sequences; wherein, each Laplacian pyramid sequence includes a low-frequency layer and multiple high-frequency layers; Substitute each of the low-frequency layers into a filtering optimization problem for iterative solution to obtain the filtering result corresponding to the low-frequency layer, and use the filtering result as the saliency layer corresponding to the low-frequency layer to calculate the residual layer corresponding to the low-frequency layer; wherein, the filtering optimization problem is an optimization problem with the filtering effect as the optimization objective, and the optimization objective at least includes the similarity of the brightness distribution of the image before and after filtering and the total gradient in eight directions after filtering; specifically, the filtering optimization problem is: Among them, O represents the filtered image in the iterative process, and O c represents the pixel value of any pixel point c of the filtered image. The coordinates of the pixel point c are (x, y), and I c represents the pixel value of the pixel point with the same coordinates (x, y) in the low-frequency layer I; ||·||2 represents the L2 norm; respectively represent the gradient operators in eight directions: left, right, up, down, upper right, lower right, upper left, and lower left; P c,l , P c,r , P c,u , P c,d , P c,ru , P c,rd , P c,lu , P c,ld are respectively corresponding penalty coefficients; each of the penalty coefficients is calculated according to the corresponding gradient operator; O * is the optimal solution obtained by iterative solution, that is, the filtering result; Calculate the saliency layer fusion factor based on the saliency layer of the target source image, and calculate the residual layer fusion factor based on each residual layer; wherein, the target source image is an infrared image or a SAR image in the multiple source images of different modalities; Perform saliency layer fusion and residual layer fusion on the source images of different modalities respectively based on the saliency layer fusion factor and the residual layer fusion factor, and use the saliency layer fusion result and the residual layer fusion result to obtain the low-frequency layer fusion result; Perform high-frequency layer fusion on the source images of different modalities, and perform image reconstruction using the inverse Laplacian transform method based on the high-frequency layer fusion result and the low-frequency layer fusion result to obtain a multi-modal fusion image.
2. The multimodal image fusion method based on multi-directional gradient filtering according to claim 1, wherein The calculation method of the penalty coefficient is as follows: Among them, "direct" represents directions, namely left, right, up, down, upper right, lower right, upper left, and lower left; α is a preset gradient threshold, β is a preset slope control factor, and P max is a preset maximum penalty coefficient.
3. The multimodal image fusion method based on multi-directional gradient filtering according to claim 1, characterized in that The method for calculating the saliency layer fusion factor based on the saliency layer of the target source image includes: m = mean2(A1), λ = 10 * m, s = std2(A1), R = |A1|, Among them, A1 represents the salient layer of the target source image, A1 represents any pixel point of this salient layer, R is the modulus value of A1, and R max is the maximum modulus value among all pixel points of this salient layer; mean2(·) represents calculating the average value, and std2(·) represents calculating the standard deviation; w is a salient layer fusion factor corresponding to the target source image, and the salient layer fusion factor corresponding to the non-target source image is determined according to w.
4. The multimodal image fusion method based on multi-directional gradient filtering according to claim 1, wherein, The calculation of the residual layer fusion factor based on each residual layer includes: Use the Sobel operator to calculate the Sobel gradient map of each residual layer; Diffuse each Sobel gradient map using a Gaussian filter to obtain the corresponding diffused gradient map; Calculate the residual layer fusion factor corresponding to each residual layer based on each diffused gradient map.
5. The multimodal image fusion method based on multi-directional gradient filtering according to claim 4, wherein The calculation of the residual layer fusion factor corresponding to each residual layer based on each diffused gradient map includes: Among them, represents the diffusion gradient map corresponding to the k-th residual layer, where k ∈ [1, K], and S R represents the sum of all diffusion gradient maps, K is the total number of the multiple source images of different modalities, and K ≥ 2.
6. The multimodal image fusion method based on multi-directional gradient filtering according to claim 4, wherein Both the Gaussian filter template and the standard deviation value of the Gaussian filter are taken as 5.
7. The multimodal image fusion method based on multi-directional gradient filtering according to claim 1, wherein The high-frequency layer fusion of the source images of different modalities includes: Perform high-frequency layer fusion on the source images of different modalities using the maximum absolute value strategy.
8. The multimodal image fusion method based on multi-directional gradient filtering according to claim 1, wherein The multiple source images of different modalities have non-negligible parallax; Before performing Laplacian decomposition on each of the source images, the multi-modal image fusion method further includes: registering and aligning the multiple source images of different modalities obtained.
9. The multimodal image fusion method based on multi-directional gradient filtering according to claim 8, wherein The multiple source images of different modalities include visible light images; The multi-modal image fusion method further includes: Before registering and aligning the multiple source images of different modalities obtained, perform IHS transformation on the visible light image, and extract the I component in the transformation result to form a new image to replace the visible light image as the source image; And, after obtaining the multi-modal fusion image, perform inverse IHS transformation on the multi-modal fusion image.
Citation Information
Patent Citations
Multi-scale transformation decomposition-based image fusion method and system with remarkable target
CN110223265A