A multi-scale decomposition method based on robust self-sparse fuzzy clustering
Through a multi-scale decomposition method of robust self-sparse fuzzy clustering, combined with Gaussian metrics and area density balance strategies, the problem of computational complexity and detail loss in infrared and visible light images is solved to generate clear and efficient fusion images.
Patent Information
- Application Number
- CN202410821786.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-29
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-08-29
AI Technical Summary
The existing infrared and visible light image fusion methods have high computational complexity and are prone to loss of detailed information when processing complex scenes. Deep learning methods require a large number of pre-trained images, and the existing network models have shortcomings in image fusion effect and efficiency.
A multi-scale decomposition method based on robust self-sparse fuzzy clustering is adopted, combined with regularization technology under Gaussian metrics and a connected component filtering algorithm of area density balance strategy, smooth the gradient information of similar and different pixel regions, and fuse infrared and visible light images through morphological heterogeneous processing and different information, and different fusion strategies are used to retain target information and texture details.
Effectively retain infrared target information and visible light texture details, generate a clear and fusion image that conforms to human visual characteristics, while improving computing efficiency and reducing artifacts and information loss.
Smart Images

Figure CN118887099B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly relates to a multi-scale decomposition method based on robust self-sparse fuzzy clustering. Background Art
[0002] Image fusion plays a key role in obtaining more comprehensive, accurate and useful information. Generally speaking, an infrared image is an image generated by an infrared camera or sensor capturing the thermal radiation information released by a target object, presenting the thermal distribution on the object surface. Effective target information can be obtained even in low-light conditions or when the target surface cannot be directly observed, but less information about the appearance features of the object is provided, with the drawback of blurred details. On the contrary, a visible light image is more in line with human vision and can capture appearance features such as the shape, color and texture of an object, but is affected by lighting conditions and may not be able to clearly capture target information in insufficient light or strong light irradiation. Therefore, fusing the two can make full use of their respective advantages, make up for each other's deficiencies, generate more comprehensive and rich image information, and help people obtain information about more scenarios. Currently, infrared and visible light image fusion technologies have been widely applied in various fields such as remote sensing, military, medical, and target detection.
[0003] In the fusion of infrared and visible light images, common fusion methods mainly include multi-scale transformation, sparse representation, subspace, saliency, hybrid model, and deep learning, etc. Among them, the method based on multi-scale transformation has been widely studied and applied because it can effectively utilize the complementary information of infrared and visible light images. It can be divided into spatial domain and frequency domain methods. The spatial domain method directly operates in the pixel space of the image, and realizes multi-scale analysis by changing pixel values or the relative relationship between pixels. It can complete basic image fusion. However, when dealing with images with rich details, the spatial domain method may require more complex algorithms and higher computational costs to capture detail information. In contrast, the frequency domain method transforms the signal from the spatial domain to the frequency domain through transformation, and performs multi-scale analysis in the frequency domain, which is more in line with people's visual trends and helps to retain some important information. However, complex transformation and inverse transformation operations are required during the conversion process between the frequency domain and the spatial domain, which undoubtedly increases the computational complexity and also leads to the loss of potential detail information in the image fusion process. In addition, Dong et al. proposed an adaptive adjustment of the fusion strategy based on PID control technology. In various scenarios, a simple fusion mathematical model can be used for image fusion to reduce the loss of detail information. However, for image fusion in complex scenarios, simple models often perform poorly. With the continuous development of deep learning, scholars have carried out research on image fusion methods based on deep learning. Currently, common deep learning-based image fusion methods are mainly divided into two categories. One is the image fusion method based on Transformer, and the other is the method based on generative adversarial network (GAN). The Transformer network model can capture the dependencies at different positions in the input sequence simultaneously through its self-attention mechanism, realize the capture of local information and long-distance information, effectively maintain the details and edge information of the source image while retaining the global architecture of the image, generate high-quality fusion images, and improve the fusion effect and overall performance. The method based on generative adversarial network (GAN) uses the structure of GAN to achieve image fusion, and through the interactive training of the generator and the discriminator, continuously learns from the source images to generate high-quality fusion images. However, both of the above two network models require a large number of pre-trained images for training to enable the model to learn more image information. Summary of the Invention
[0004] To solve the above problems, the present invention provides a multi-scale decomposition method based on robust self-sparse fuzzy clustering.
[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] A multi-scale decomposition method based on robust self-sparse fuzzy clustering, comprising the following steps:
[0007] S1. Use the Self-Sparse Fuzzy C-Means Clustering Algorithm (SSFCA) combined with the regularization technique under the Gaussian metric to cluster pixels, and smooth the gradient information between similar pixels and different pixel regions;
[0008] S2. Use the Connected Component Filtering Algorithm based on Area Density Balance Strategy (CCFF-ADB) to adaptively merge overly small clustering regions and integrate local information into the feature extraction process;
[0009] S3. Perform morphological heterogeneous processing on the underlying information and use the difference information between infrared and visible light images to enrich the underlying image;
[0010] S4. For feature information at different levels, adopt different fusion strategies to maximize the retention of infrared target information and visible light detail textures, and finally achieve image fusion.
[0011] Furthermore, the objective function of the Self-Sparse Fuzzy C-Means Clustering Algorithm (SSFCA) is as follows:
[0012]
[0013] where γ is a balance factor used to control the member sparsity, c represents the number of image clusters, n represents the number of samples, x j is the j-th sample point in the dataset X = {x1, x2,..., x n}, v i represents the cluster center, u ij represents the membership degree of the sample x j relative to the cluster center v i . By changing the value of γ, the objective function shows different robustness to outliers or noises; Φ(x j |v i , ∑ i ) represents the distance function between x j and v i , and its definition is as follows:
[0014] Φ(x j |vi, ∑ i ) = ln(-ρ(x j |v i , ∑ i )) (2)
[0015] where ρ(x j |v i , ∑ i ) is the Gaussian density function, and its definition is:
[0016]
[0017] Among them, D represents the dimension of the input data, T represents the transpose operation, and ∑ i represents the covariance matrix of the within-class scatter of the i-th class, and x j -v i represents the difference vector between the data point x j and the cluster center v i . Substituting equation (3) into equation (2) gives:
[0018]
[0019] Due to the influence of ln|∑ i |, Φ(x j |v i , ∑ i ) may not satisfy the non-negativity constraint. Therefore, Φ′(x j |v i , ∑ i ) is used to replace Φ(x j |v i , ∑ i ), and the formula is as follows:
[0020]
[0021] Substituting Φ′(x j |v i , ∑ i ) into equation (1) gives the final definition of the objective function:
[0022]
[0023] For each sample x j , it can be divided into c sub-problems, and the constraint condition is 0 ≤ u ij ≤ 1, and we can get:
[0024]
[0025] By simplifying it can be rewritten as:
[0026]
[0027] Among them, h ij = -Φ′(x j |v i , ∑ i ) / 2γ. By adjusting the value of γ, fuzzy memberships with different sparsities can be obtained.
[0028] The present invention has the following beneficial effects:
[0029] 1) The present invention is based on a regularized self-sparse fuzzy clustering algorithm to smooth the gradients between similar and different pixel regions and balance the proportion between the two gradient information. This method not only retains the detailed texture information of infrared and visible light images more comprehensively, but also effectively solves the problem of image imaging differences by using window processing at different scales.
[0030] 2) In the fusion strategy, we adopt morphological heterogeneous processing to balance the brightness and contrast of infrared and visible light, and use the difference information between infrared and visible light images to enrich the underlying images. Then, a linear combination is performed on the high-frequency components in the infrared and visible light images to maximize the retention of infrared target information and visible light detailed texture, obtaining a clear and rich fused image.
[0031] 3) Comparative analysis on public datasets with several advanced fusion methods subjectively and objectively reveals that: the method of the present invention not only effectively retains the significant information of infrared targets and the texture details of visible light, generating a clear fused image that conforms to human visual characteristics, but also provides a fast and effective solution for the field of image fusion technology with its excellent computational efficiency. Description of the Drawings
[0032] Figure 1 It is a comparison chart of the processing results of the method of the present invention with ResNetFusion, MDLatLRR, and CVT;
[0033] In the figure: (a) Visible light; (b) Infrared light; (c) ResNetFusion; (d) MDLatLRR; (e) CVT; (f) The method of the present invention.
[0034] Figure 2 It is the smoothing effect of different smoothing techniques and the corresponding surf graph.
[0035] Figure 3 It is the Surf image result of the clustering algorithm improved by the regularization technique under the Gaussian metric.
[0036] Figure 4 It is the smoothing result of the Surf graph of the method of the present invention.
[0037] Figure 5 It is a schematic diagram of the segmented image;
[0038] In the figure: (a) Visible light image; (b) Visible light underlying information image; (c) Infrared image; (d) Infrared underlying information image.
[0039] Figure 6 It is a schematic diagram of the multi-scale decomposition result of the present invention;
[0040] In the figure: (a) Visible light image; (b) Visible light bright details; (c) Visible light underlying layer; (d) Visible light dark details; (e) Infrared image; (f) Infrared bright details; (g) Infrared underlying layer; (h) Infrared dark details.
[0041] Figure 7 is the fused underlying layer.
[0042] Figure 8 is a schematic diagram of the separation of the pixel-enhanced mixed layer and the weakened mixed layer of the source image;
[0043] In the figure: (a) Enhanced mixed image; (b) Weakened mixed image.
[0044] Figure 9 is I Z effect picture.
[0045] Figure 10 is the fusion effect picture of infrared and visible light bright details.
[0046] Figure 11 is the fusion result of the present invention;
[0047] In the figure: (a) Visible light image; (b) Infrared image; (c) Fusion result.
[0048] Figure 12 The influence of different windows W of the visible light image on the fusion result; n×n dimension;
[0049] In the figure: (a) Visible light image 1; (b) Visible light image 2; (c) Fusion result.
[0050] Figure 13 is the influence of different windows W of the infrared image on the fusion result; n×n dimension;
[0051] In the figure: (a) Infrared image 1; (b) Infrared image 2; (c) Fusion result.
[0052] Figure 14 is the influence of different fusion weight α values on the underlying layer and the fusion result.
[0053] Figure 15 is the influence of different fusion weight β values on the underlying layer and the fusion result.
[0054] Figure 16 is the fusion result of 5 pairs of typical images in the TNO dataset.
[0055] Figure 17For the quantitative comparison of different methods on 21 pairs of images in TNO, the evaluation metrics AG, H, SD, SF, EI, Qab / f, Lab / f, and Nab / f were selected.
[0056] Figure 18 It is the qualitative fusion result obtained from the 8th image frame extracted from the "MainEntrance" video in the INO dataset.
[0057] Figure 19 For the quantitative comparison of different methods on 21 pairs of images in INO, the evaluation metrics AG, H, SD, SF, EI, Qab / f, Lab / f, and Nab / f were selected. Detailed implementation manners
[0058] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made. These all belong to the protection scope of the present invention.
[0059] The present invention provides a multi-scale decomposition method based on robust self-sparse fuzzy clustering. First, it integrates the regularization technology under Gaussian metric into the self-sparse fuzzy C-means clustering algorithm (SSFCA) to cluster pixels, smooth the gradient information between similar pixel and different pixel regions, and adopts a connected component filtering algorithm based on area density balance strategy (CCFF-ADB) to adaptively merge small regions to balance the proportion between similar pixel and different pixel regions, so that the multi-scale decomposition can more comprehensively retain detail textures and target information. Then, in order to make the brightness and contrast of the fused image more in line with the visual trend, morphological heterogeneous processing is performed on the underlying information and the difference information between infrared and visible light images is used to enrich the underlying image. Finally, for different levels of feature information, different fusion strategies are adopted to maximize the retention of infrared target information and visible light detail textures, and finally image fusion is achieved. To more intuitively demonstrate the advantages of the algorithm of the present invention, a pair of typical examples are selected for display. In Figure 1 it, we compare the method of the present invention with the representative ResNetFusion, MDLatLRR, and CVT (the key parts of the detail textures are boxed and enlarged for display). Obviously, although ResNetFusion, MDLatLRR, and CVT retain the human target information in the infrared image, they lose the texture information such as branches in the visible light, resulting in a blurred fused image. In contrast, the fusion result of the method of the present invention not only retains the infrared target but also has the detail texture information of the visible light, and can more clearly display the fused image.
[0060] Specifically, a multi-scale decomposition method based on robust self-sparse fuzzy clustering of the present invention includes the following steps:
[0061] (1) Multi-scale feature extraction
[0062] The present invention extracts image features in the spatial domain using image filtering and clustering algorithms to effectively smooth the image structure, and eliminates the gradient differences between them by clustering similar regions, thereby suppressing or weakening the edge or texture changes between most similar regions. To verify this conjecture, we used classical median filtering and anisotropic diffusion clustering techniques to smooth the images.
[0063] As Figure 2 shown, the Surf graphs processed by the median filtering and anisotropic diffusion clustering algorithms have significantly smoothed the gradient features between similar pixels, effectively reducing the noise and details in the images. The anisotropic diffusion clustering algorithm is superior to the median filtering because it can retain edge information while smoothing the image. However, neither of them can smooth the gradient features between different pixels. To solve this problem, we analyzed existing image clustering algorithms, such as the K-means algorithm and the C-means algorithm, but traditional clustering algorithms often have two problems: First, because of the non-sparsity of fuzzy members, the algorithms are sensitive to outliers. Second, over-clustering causes the algorithms to lose local spatial information of the image.
[0064] Therefore, the present invention designs a regularized self-sparse fuzzy clustering algorithm that can effectively smooth the gradients between similar pixels and different pixels. When processing the gradient features between similar pixels, the regularization technique under the Gaussian metric is integrated into the traditional clustering algorithm to obtain a relatively sparse fuzzy membership degree, reducing the interference of outliers on the clustering results, thereby more effectively smoothing the gradient changes between similar regions. Secondly, to prevent the loss of local information space information caused by over-clustering, especially the imbalance problem that easily occurs between similar pixel and different pixel regions, the present invention adopts a connected component filtering algorithm based on the area density balance strategy (CCFF-ADB) to adaptively merge too small clustering regions, integrating local information into the feature extraction process, thereby avoiding the loss of local details and texture information. The specific clustering algorithm model is as follows:
[0065] The objective function of the regularized self-sparse fuzzy clustering algorithm is as follows:
[0066]
[0067] where γ is a balance factor for controlling the member sparsity, c represents the number of image clusters, n represents the number of samples, and x j is the data set X = {x1, x2,..., xn The j-th sample point in}, v i represents the cluster center, u ij represents the sample x j with respect to the cluster center v i of the membership degree. By changing the γ value, the objective function shows different robustness to outliers or noise; Φ(x j |v i , ∑ i ) represents the distance function between x j and v i , and its definition is as follows:
[0068] Φ(x j |vi, ∑ i ) = ln(-ρ(x j |v i , ∑ i )) (2)
[0069] where ρ(x j |v i , ∑ i ) is the Gaussian density function, and its definition is:
[0070]
[0071] where D represents the dimension of the input data, T represents the transpose operation, ∑ i represents the covariance matrix of the within-class scatter of the i-th class, x j -v i represents the difference vector between the data point x j and the cluster center v i . Substituting equation (3) into equation (2) gives:
[0072]
[0073] Due to the influence of ln|∑ i |, Φ(x j |v i , ∑ i ) may not satisfy the non-negativity constraint. Therefore, we use Φ′(x j |v i , ∑ i ) to represent instead of Φ(x j |v i , ∑ i ), and the formula is as follows:
[0074]
[0075] Substituting Φ′(x j |v i , ∑i ) Substitute into equation (1) to obtain the final objective function definition:
[0076]
[0077] For each sample x j , can be divided into c sub - problems, with the constraint 0 ≤ u ij ≤ 1, we can get:
[0078]
[0079] By simplification can be rewritten as:
[0080]
[0081] where h ij = -Φ′(x j |v i , ∑ i ) / 2γ. By adjusting the value of γ, we can obtain fuzzy memberships with different sparsity levels.
[0082] Similarly, we can divide into n independent sub - problems. By solving we can obtain the cluster center v i . The solution process is as follows:
[0083]
[0084] Solve to get:
[0085]
[0086] In addition, the updated covariance matrix ∑ i can be obtained by solving . The solution process is as follows:
[0087]
[0088] Solve to get:
[0089]
[0090] where ∑ i represents the covariance matrix of the within - class scatter of the i - th class, and x j - v i represents the difference vector between the data point x j and the cluster center v i .
[0091] Although the regularization technology under Gaussian metric is integrated into the traditional clustering algorithm, the interference of non-uniform pixels can be suppressed to a certain extent, and the gradient information between similar regions and different regions can be effectively smoothed, but there is a phenomenon of over-clustering. Figure 4 As shown in the figure, from left to right are the visible light image, the Surf map corresponding to the screenshot of a specific area in the visible light image, and the Surf map obtained by applying the clustering algorithm with regularization technology under Gaussian metric fusion. Comparing these two Surf images, it can be clearly observed that the clustering algorithm integrated with the new technology has achieved remarkable results in smoothing the gradients between similar pixels and different pixel areas. However, after careful observation, we found that there is unevenness in the processing of gradients when processing these two feature extractions. Specifically, the gradient areas between different pixels are too dominant in some cases, resulting in the relatively weakened expression of gradient information in similar areas, which is not conducive to our extraction of target and texture information in infrared and visible light images.
[0092] In order to effectively overcome the above-mentioned gradient unevenness between similar regions and different regions, the present invention adopts a connected component filtering algorithm (CCFF-ADB) based on the area density balance strategy for optimization. First, we use the above-mentioned regularized self-sparse fuzzy clustering algorithm to normalize the connected component image generated by the qth region to obtain the area α q , and use the following formula to calculate the radius, x p α is the center of the circle q The number of q :
[0093]
[0094]
[0095] Where 0≤x p ≤1 and 1≤p≤(K+1),p∈N + , Q represents the number of connected components in the image generated by the above clustering algorithm.
[0096] Then in ε q Mapping is performed to obtain K q :
[0097]
[0098] Finally, K q Normalize and get K′ q , through K′ qThe maximum interval in it can easily calculate the value of the truncated area M. In the connected component image, regions with areas smaller than the M value are merged using the minimum distance, so as to obtain a better gradient distribution of similar pixel regions and different pixel regions, and let the final result obtained be S(x). As Figure 4 shown, while smoothing the gradients of similar pixels and different pixel regions in the Surf image processed by CCF-ADB, the characteristic proportion of the two is also ensured.
[0099] We will merge the same-class pixels of the source image according to the calculated S(x) to obtain the underlying image information, denoted as I(x). The result is as Figure 5 shown: Although the underlying layer obtains the main contour information, providing the overall structure and light and dark distribution of the image, due to the different wavelengths and imaging mechanisms of infrared and visible light images, there is a phenomenon of inconsistent light and dark distribution. This inconsistency often leads to the generation of artifacts in the subsequent image fusion process, affecting the quality of the fused image. Therefore, to improve this problem, the present invention performs morphological image processing on the underlying layer. Specifically, we use the opening operation in morphology (i.e., erode first and then dilate) to obtain the processed underlying layer. The purpose is to smooth the light and dark distribution in the underlying image, make the light and dark contrast of the image more uniform, and facilitate the generation of a fused image with good light and dark contrast in the subsequent fusion process. The opening operation expression is as follows:
[0100]
[0101] Among them, M(x) represents the underlying layer, B represents the structuring element, represents the erosion operation, represents the dilation operation.
[0102] Secondly, based on the underlying layer, we obtain a high-frequency layer containing image detail textures and detailed contour information. We extract the detail parts of infrared and visible light respectively, including dark detail parts and bright detail parts. Among them, the bright detail part is obtained by subtracting the underlying layer from the source image, and the dark detail part is obtained by subtracting the source image from the intermediate frequency component. The expressions are as follows:
[0103] H = I S - M (15)
[0104] L = M - I S (16)
[0105] Among them, I S represents the source image, M represents the underlying layer, H represents the bright detail layer, and L represents the dark detail layer. The multi-scale decomposition result is as Figure 6 shown.
[0106] (2) Multi-scale fusion
[0107] For the underlying layer and high-frequency components obtained by the above-mentioned multi-scale decomposition, the present invention refines the fusion process into two main parts. The first part focuses on processing the underlying layer to construct the approximate contour information of the fused image. The second part focuses on processing the high-frequency detail part to improve the texture details and contour information of the fused image.
[0108] Underlying foundation construction
[0109] First, we fuse the underlying information of the visible light image and the infrared image, taking the minimum value of the underlying information in the visible light and infrared images, so that the finally fused image is more natural and there are no abrupt changes.
[0110] M F = min(M IR (x,y) + M VIS (x,y)) (17)
[0111] M IR represents the underlying information of the infrared image, M VIS represents the underlying information of the visible light image, (x,y) represents the pixel value position, and M F represents the result after fusing the underlying layers. The obtained result is as Figure 7 shown: M F Although the light and dark distribution of the fused image is balanced, the target contour is not prominent. To display the contour of the fused image in the underlying layer, we incorporate the enhanced mixed image I E and the weakened mixed image I D of the source image into the underlying image, so that the underlying image obtains more image information. The expression is as follows:
[0112] I E (x,y) = max(I1(x,y), I2(x,y)) (18)
[0113] I D (x,y) = min(I1(x,y), I2(x,y)) (19)
[0114] where I1 is the result of adaptively linearly adding the source image I after histogram enhancement to the source image, I2 represents the infrared source image, (x,y) represents the pixel value position, and the result of the mixed image is as Figure 8 shown.
[0115] The enhanced mixed image I E and the weakened mixed image I DFuse with the underlying layer. This fusion is weight-based, where the weights of the enhanced and attenuated mixed images are set to α and β respectively. By controlling α and β, it is ensured that while enhancing the texture, the original underlying information is not overly affected. Through this method, we obtain the new underlying image information after texture enhancement, and its fusion result I Z is calculated by the following expression:
[0116]
[0117] Obtain I Z The result is as Figure 9 shown.
[0118] High-frequency detail processing
[0119] After completing the effective fusion of the underlying layer, we follow the principle of texture gradient enhancement and finely integrate the high-frequency detail parts of each original image into the underlying fusion layer in a linearly weighted manner, thus generating the final image fusion result.
[0120] For the bright details extracted from infrared and visible light images, we adopted a linear superposition strategy for meticulous processing. This makes the fused image clearer, while retaining the texture details and edge information in the image, making it more abundant. Its expression is as follows:
[0121] H F = H IR + H VIS (21)
[0122] Here, H IR represents each infrared bright detail, H VIS represents each visible light bright detail part, and H F represents the result after fusion of each image bright detail part. The obtained result is as Figure 10 shown. At this time, although the bright detail part of the high-frequency component in the image has initially shown the characteristics of the fused image in terms of overall structure, upon careful observation, it can be found that the brightness of the image is significantly darker and some detail parts are lost, which to a certain extent affects the clarity and contrast of the image and is not conducive to subsequent more precise image fusion processing.
[0123] To solve the above problems, we use the dark detail part in the high-frequency component for adjustment. Take the minimum value of the dark detail parts of infrared and visible light, min(L IR (x,y), L VIS(x, y)), as the fused dark details, can ensure that the fused image retains the more prominent dark detail parts and avoids information loss during the mixing process. At the same time, to further filter out noise or unimportant detail information in the image, by setting the threshold th and setting to zero the parts where the dark details are less than the threshold, the finally fused image becomes cleaner and clearer.
[0124] L F (i, j) = min(L IR (i, j), L VIS (i, j)) (22)
[0125]
[0126] Among them, L F represents the fused dark detail part, L IR represents the dark detail part of the infrared ray, L VIS represents the dark detail part of the visible light, the threshold th = 5, and (x, y) represents the pixel value position;
[0127] Finally, the fused high-frequency components are combined into the fusion result I Z of the underlying layer through linear addition, thereby obtaining the final image fusion result.
[0128] I F = I Z + H F - L F (24)
[0129] Among them, I F represents the fusion result of the infrared first and the visible light image, I Z represents the information of the underlying layer after texture enhancement, L F represents the fused dark detail part, H F represents the fused bright detail part.
[0130] The fusion result of the present invention is as Figure 11 shown.
[0131] (3) Settings and characteristics of each parameter
[0132] In the research of the present invention, we have particularly focused on the setting problems of a group of key parameters. These parameters have a significant impact on the final research results, and their slight changes may lead to deviations in the results. In order to deeply explore the action mechanism of these parameters, we conduct detailed experiments and analyses on these parameters in this section.
[0133] Determination of the clustering window W n×n and the dimension n
[0134] The size \(n\) of the clustering window has a significant impact on the clustering effect of the source image. Especially when processing the brightness of the underlying layer, it can effectively improve the overall quality and clarity of the image, and avoid artifacts in the subsequent fusion process. In this invention, two sets of fused images are selected to verify the influence of the clustering window size \(n\) on the fusion effect.
[0135] Observation Figure 12 It can be found that when \(n = 0\), although the underlying image obtains the contour of the visible light image, it also brings artifacts. As \(n\) increases, the brightness of the underlying layer gradually tends to be consistent, and the artifacts in the fusion result gradually disappear. When \(n = 250\), the brightness of the underlying layer of the "camouflage car" tends to be consistent, and the artifacts in the fusion result disappear. When \(n = 350\), the underlying layer area of the "ship" is consistent, and the artifacts in the fusion result disappear. Therefore, in this invention, \(n = 350\) is selected as the clustering window size of the visible light image.
[0136] Through observation Figure 13 It can be found that, like the visible light image, although the underlying layer obtains the contour of the infrared image, it also brings the problem of artifacts. As \(n\) increases, the brightness of the underlying layer gradually tends to be consistent, and the artifacts in the fusion result gradually disappear. When \(n = 300\), the brightness of the underlying layers of the "camouflage car" and the "house" tends to be consistent, and the artifacts in the fusion result disappear. Therefore, in this invention, \(n = 300\) is selected as the clustering window size of the infrared image.
[0137] Selection of weights and thresholds
[0138] When constructing the contour, we use \(\alpha\) and \(\beta\) to control the clarity of the control contour. As Figure 14 it can be seen, as the parameter value \(\alpha\) increases, the contour map of the underlying layer becomes clearer and the brightness of the fusion result also increases accordingly. When the parameter value reaches 0.4, we can clearly observe that the fused image has an overexposure situation, and the texture detail information of the clouds in the figure also becomes blurred or even disappears. In order to ensure that the underlying layer can maintain good contour information and the fusion result can maintain good contrast, we comprehensively consider and select the parameter value \(\alpha=0.1\) as the optimal value of this invention.
[0139] Observation Figure 15 it can be seen that as \(\beta\) increases, the contour color of the underlying layer gradually deepens, which is reflected in the image fusion as the camouflage texture information on the "camouflage car" gradually blends with the vehicle color and even disappears. When \(\beta = 0.8\), it can be clearly observed that the camouflage pattern of the vehicle has disappeared. In order to ensure that the fusion effect can meet the requirements of \(\alpha\) and the underlying layer component can obtain good contour information, this invention selects \(\beta = 0.1\) as the optimal value.
[0140] Performance test data:
[0141] To verify the effectiveness and superiority of the method proposed in this experiment, we conducted comparative experiments between the method of the present invention and current mainstream image fusion algorithms, including GANMcC, MDLatLRR, NestFuse, RFN-Nest, CSF, FusionGAN, and SEDRFuse. These algorithms are representative and influential in the field of image fusion. Through comparative experiments, we can more accurately evaluate the performance of the method of the present invention in image fusion tasks and highlight the advantages of the algorithm of the present invention.
[0142] Dataset:
[0143] The present invention uses infrared and visible light images in the TNO dataset and the INO dataset for experiments. The TNO dataset contains a variety of nighttime multispectral (enhanced vision, near-infrared, and long-wave infrared or thermal imaging) images and is registered using different multi-band camera systems. The INO dataset is provided by the National Optics Institute of Canada and contains multiple pairs of visible light and infrared videos captured under different weather conditions, representing different scenarios.
[0144] Evaluation metrics:
[0145] In the experimental design of the present invention, a combination of subjective and objective methods is adopted. Subjectively, we rely on the intuitive judgment of the human eye on image quality and evaluate the visual performance by directly comparing the fusion results generated by different algorithms. Objectively, in order to more accurately evaluate the effect of infrared and visible light image fusion, the present invention selects eight objective evaluation metrics, including average gradient AG, information entropy H, standard deviation SD, spatial frequency SF, edge intensity EI, fusion quantity function Q ab / f , amount of artifacts N ab / f and fusion loss function L ab / f .
[0146] (1) The average gradient (AG) is an important metric for measuring image sharpness and texture features. The larger the AG value, the richer the edges and details in the fused image, and the clearer the image.
[0147] (2) The information entropy (H) is used to measure the amount of information in the fused image. The larger the H value, the richer the information content in the image, and the better the performance of the fusion method in processing information.
[0148] (3) The standard deviation (SD) is an index reflecting image contrast and distribution. Since the human eye is more likely to focus on areas with higher contrast, the larger the SD value, the higher the contrast of the fused image, and thus the better the visual effect achieved.
[0149] (4) Spatial frequency (SF) is an index determined by the gradient distribution, which can effectively display the detail and texture information of an image. The higher the SF value, the higher the image quality, and the more complete the image is in terms of details and texture.
[0150] (5) Edge information (EI) is a measure of the pixel value change in the edge region of an image, reflecting the clarity and sharpness of the object boundaries in the image. The larger the EI value, the clearer the edges in the fused image and the higher the image quality.
[0151] (6) Q ab / f is a function that measures the amount of information of each source image contained in the fused image, ranging from 0 to 1. The larger the Q ab / f value, the more information of each source image is retained in the fused image, and the more ideal the fusion effect is.
[0152] (7) L ab / f is an index that measures the amount of information loss of each source image during the fusion process, ranging from 0 to 1. The smaller the L ab / f value, the less information of the source image is lost during the fusion process, and the more complete the information is retained.
[0153] (8) N ab / f is an index that measures the amount of new information generated during the fusion process of the source images, ranging from 0 to 1. The smaller the N ab / f value, the less new information is generated during the fusion process, which usually means that the fusion process retains more information of the source images rather than generating new information.
[0154] Experimental results on the TNO dataset
[0155] To verify the image fusion method proposed in the present invention, we selected five pairs of typical images from the TNO dataset for qualitative evaluation with the current mainstream fusion algorithms to verify the advantage of the method of the present invention in retaining infrared target information and visible light texture information. Figure 16 It presents in detail a variety of fused images arranged from left to right and their unique detail features. This figure shows the original visible light image, infrared image, and fused effect images obtained by different methods in an orderly manner from top to bottom. To highlight the excellent performance of the fusion method of the present invention, we specifically selected a typical area on each fused image and marked it with a red square, and at the same time provided an enlarged detail display on the right side of the image. From the intuitive visual experience, Figure 16Among the five classic image fusion results, our fusion result is particularly remarkable. It not only successfully highlights the heat source target information but also clearly retains the background texture information, enabling the comprehensive and accurate display of the key information in infrared and visible light images. In the "street scene" experiment that we specifically focused on in the third column, we found that there are varying degrees of problems in the fusion results of methods such as GANMcC, MDLatLRR, RFN-Nest, CSF, FusionGAN, and SEDRFuse, including blurred outlines of infrared target information and darker brightness. Although the Nestfuse method avoids these problems to a certain extent, it still has deficiencies in retaining the detailed texture of visible light images. Specifically, although this method excellently retains the human target information, it fails to effectively retain the detailed information such as the fence next to the person. In contrast, the multi-scale decomposition proposed in the present invention effectively smooths the gradient features between similar and different pixels, performs excellently in retaining infrared target information and visible light detailed texture, has excellent light and dark contrast, and the fusion result is clearer and closer to the visual perception of the human eye.
[0156] In addition, to more prominently demonstrate the advantages of the method of the present invention, we selected the fusion results of 21 typical pairs of images from the TNO dataset for quantitative comparison. As Figure 17 shown, in the comparison of indicators such as average gradient (AG), information entropy (H), standard deviation (SD), spatial frequency (SF), and edge intensity (EI), this algorithm ranks first in all of them, proving that the method of the present invention is superior to other methods in maintaining image clarity, information richness, contrast, and detail capture, and can obtain a relatively clear fusion image. In terms of fusion quality evaluation, we further compared the results of the fusion quantity function (Q ab / f ). Although the method of the present invention only ranks third, this is mainly because this indicator focuses on the fusion quality evaluation of a specific aspect. In fact, our method deeply considers the fusion of infrared and visible light information and retains most of the key information. Therefore, the slight deficiency in the fusion quantity function does not represent a decline in the overall performance. From the evaluation results of the fusion loss function (L ab / f ) and the amount of artifacts (N ab / f ), the method of the present invention achieves the optimal L ab / f loss, which fully proves that our method has significant advantages in dealing with interference information and retains most of the key information. In the evaluation of N ab / f , because we comprehensively consider the information of infrared and visible light, when a certain part of the image is highlighted, there will inevitably be some minor distortions or artifacts in the fusion process. However, these minor flaws do not affect the excellent overall performance of our method, and our evaluation results are still far superior to other methods, which fully demonstrates the significant advantages of this algorithm in the image fusion task.
[0157] Experimental Results on the INO Dataset
[0158] To further study the robustness and performance of the algorithm, we selected the representative video named "MainEntrance" in the INO dataset and divided it into 21 image frames, each frame containing a corresponding pair of infrared and visible light images. Based on these 21 pairs of infrared and visible light images, a qualitative and quantitative comparison was made between the algorithm of the present invention and the current mainstream fusion algorithms. Figure 18 , showing the fusion result of the 8th image frame among the 21 image frames. For easy observation, we intercepted a part and enlarged it below the image. The red box intercepted the target information in the infrared image, and the green box intercepted the visible light texture information. From the overall analysis, the fused image obtained by the method we proposed has a distinct contrast between light and dark and is clearer, bringing a comfortable visual experience. From the perspective of detailed analysis, other methods have lost important information in the original image to varying degrees. Specifically, the target information of the power pole wires in the infrared image (shown by the red box) was lost in the fusion results of GANMcC, MDLatLRR, Nestfuse, RFN-Nest, CSF, FusionGAN, and SEDRFuse. In addition, the texture information of the trees in the visible light image was lost in GANMcC, Nestfuse, RFN-Nest, CSF, and FusionGAN. In contrast, our method effectively retains the key target information in the infrared image and the rich texture information in the visible light image. Figure 19 A quantitative comparison was made for these 21 image pairs. The method we proposed showed good performance in AG, H, SD, SF, EI, and L ab / f metrics and ranked first, fully demonstrating that our fused images are excellent in terms of clarity, contrast, spatial details, etc. Although in Q ab / f and N ab / f metrics, our method did not reach the optimal level, but it is worth mentioning that the average value of our Q ab / f is 0.5126, which is only slightly lower than 0.5363 of Nestfuse. This still shows the strong competitiveness of our method in these metrics. Therefore, the quantitative evaluation results on the INO dataset prove that the algorithm we proposed has significantly improved the robustness and clarity of infrared and visible light image fusion.
[0159] Running Time Comparison
[0160] Table 1: Comparison of the running times of eight different methods for fused images on the TNO and INO datasets on the CPU. Each value represents the average running time of each method on the corresponding dataset, with the unit in seconds (s), and the optimal result is shown in bold black.
[0161]
[0162]
[0163] Since the algorithm of the present invention performs multi-scale decomposition in the spatial domain, its computational complexity is significantly lower than that of other frequency domain algorithms. As shown in Table 1, the method we proposed demonstrated a faster computational speed in comparison with other mainstream fusion algorithms. This not only reflects the high efficiency of the algorithm but also ensures that a fused image can be generated more quickly in practical applications. All in all, our method not only effectively retains the significant information of infrared targets and the texture details of visible light, generating a clear fused image that conforms to human visual characteristics, but also provides a fast and effective solution for the field of image fusion technology with its excellent computational efficiency.
[0164] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various deformations or modifications within the scope of the claims, which do not affect the essence of the present invention.
Claims
1. A multi-scale decomposition method based on robust self-sparse fuzzy clustering, characterized in that: It includes the following steps: S1. Use the self-sparse fuzzy C-means clustering algorithm combined with the regularization technique under the Gaussian metric to cluster pixels, and smooth the gradient information between similar pixels and different pixel regions; S2. Use the connected component filtering algorithm based on the area density balance strategy to adaptively merge too small clustering regions and integrate local information into the feature extraction process; S3. Perform morphological heterogeneous processing on the underlying information and enrich the underlying image by using the difference information between infrared and visible light images; S4. For feature information at different levels, adopt different fusion strategies to maximize the retention of infrared target information and visible light detailed textures, and finally achieve image fusion; specifically: S41. Underlying foundation construction First, take the minimum value of the underlying information in the visible light and infrared images, and fuse the underlying information of the visible light image and the infrared image; M F = min(M IR (x, y)+M VIS (x, y)) (17) Among them, M IR represents the underlying information of the infrared image, M VIS represents the underlying information of the visible light image, (x, y) represents the pixel value position, M F represents the result after the fusion of the underlying layers; Secondly, incorporate the enhanced mixed image I of the source image E and the weakened mixed image I D into the underlying image, so that the underlying image obtains more image information. The expression is as follows: I E (x,y) = max(I1(x,y), I2(x,y)) (18) I D (x,y) = min(I1(x,y), I2(x,y)) (19) Among them, I1 is the result of the source image I after histogram enhancement and adaptively linearly added to the source image, I2 represents the infrared source image, and (x, y) represents the pixel value position; Finally, the enhanced mixed image I of the source image E and the weakened mixed image I D are respectively fused with the underlying layer M F ; among them, the weights of the enhanced mixed image and the weakened mixed image are set as α and β respectively, and the new underlying layer information I after texture enhancement Z The expression is as follows: I Z = M F + α × I E - β × I D (20); Among them, I Z represents the underlying layer information after texture enhancement, M F represents the result after the fusion of the underlying layers, I E represents the enhanced mixed image, I D represents the weakened mixed image, and α and β are weights; S42. High-frequency detail processing After the effective fusion of the underlying layer is completed, the high-frequency detail parts of each original image are finely integrated into the underlying fusion layer in a linearly weighted manner, thereby generating the final image fusion result; among them, for the bright details extracted from the infrared and visible light images, a linear superposition strategy is adopted for detailed processing, and its expression is as follows: H F = H IR + H VIS (21) Among them, H F represents the bright detail part after fusion, and H IR represents the bright detail of the infrared image, and H VIS represents the bright detail part of the visible light image; Take the minimum value min(L IR (x, y), L VIS (x, y)) of the dark detail parts of infrared and visible light as the fused dark detail.
2. The multi-scale decomposition method based on robust self-sparse fuzzy clustering according to claim 1, characterized in that: The objective function of the self-sparse fuzzy C-means clustering algorithm (SSFCA) is as follows: Among them, γ is a balance factor for controlling the member sparsity, c represents the number of image clusters, n represents the number of samples, and x j is the j-th sample point in the data set X = {x1, x2, …, x n}, v i represents the cluster center, and u ij represents the membership degree of the sample x j relative to the cluster center v i . By changing the value of γ, the objective function shows different robustness to outliers or noises; Φ(x j |v i ,∑ i ) represents the distance function between x j and v i , which is defined as follows: Φ(x j |vi,∑ i )=ln(-ρ(x j |v i ,∑ i )) (2) where ρ(x j |v i ,∑ i ) is the Gaussian density function, which is defined as: where D represents the dimension of the input data, T represents the transpose operation, and ∑ i represents the covariance matrix of the within-class scatter of the i-th class, and x j - v i represents the difference vector between the data point x j and the cluster center v i . Substituting equation (3) into equation (2) gives: Due to the influence of ln|∑ i |, Φ(x j |v i ,∑ i ) may not satisfy the non - negative value constraint. Therefore, Φ′(x j |v i ,∑ i ) is used to replace Φ(x j |v i ,∑ i ), and the formula is as follows: Substitute Φ′(x j |v i ,∑ i ) into equation (1) to obtain the final definition of the objective function: For each sample x j , it can be divided into c sub-problems, with the constraint 0 ≤ u ij ≤ 1, and we can obtain: By simplification can be rewritten as: where h ij = -Φ′(x j |v i , ∑ i ) / 2γ, and different sparsity fuzzy members can be obtained by adjusting the value of γ.
3. A multi-scale decomposition method based on robust self-sparse fuzzy clustering as described in claim 2, characterized in that: can be divided into n independent sub-problems, and by solving the clustering center v can be obtained i ; The solution process is as follows: Solve to get: In addition, the updated covariance matrix ∑ i can be obtained by solving and the solution process is as follows: Solve to get: where, ∑ i represents the covariance matrix of the within-class scatter of the i-th class, and x j - v i represents the difference vector between the data point x j and the cluster center v i .
4. A multi-scale decomposition method based on robust self-sparse fuzzy clustering as described in claim 1, characterized in that: The step S2 includes the following steps: Let the connected component image generated for the q-th region in step S1 be normalized by using the self-sparse fuzzy C-means clustering algorithm (SSFCA) described above to obtain an area of α q , and use the following formula to calculate the number ε of α p within the circle with a radius of ε and a center of x q : q : where 0 ≤ x p ≤ 1 and 1 ≤ p ≤ (K + 1), p ∈ N + , Q represents the number of connected components in the image generated by the clustering algorithm described above; Then, for ε q perform mapping to obtain K q : Finally, perform normalization on K q to obtain K′ q . Calculate the value of the truncation area M through the maximum interval in K′ q . Merge the regions with areas smaller than the M value using the minimum distance in the connected component image, so as to obtain a better gradient distribution of similar pixel regions and different pixel regions, and let the obtained final result be S(x).
5. The multi-scale decomposition method based on robust self-sparse fuzzy clustering according to claim 1, wherein: In the step S3, the opening operation in morphology is used to perform morphological heterogeneous processing on the underlying layer, and the opening operation expression is as follows: Among them, M(x) represents the underlying layer, and B represents the structural element. represents the erosion operation. represents the dilation operation.
6. A multi-scale decomposition method based on robust self-sparse fuzzy clustering as described in claim 1, characterized in that: In the step S3, based on the underlying layer, a high-frequency layer containing image detail textures and detailed contour information is obtained; specifically, the detail parts of infrared and visible light are extracted respectively, including dark detail parts and bright detail parts, where the bright detail part is obtained by subtracting the underlying layer from the source image, and the dark detail part is obtained by subtracting the source image from the intermediate frequency component, and the expression is as follows: H = I S -M (15) L = M - I S (16) Among them, I S represents the source image, M represents the underlying layer, H represents the bright detail layer, and L represents the dark detail layer.
7. A multi-scale decomposition method based on robust self-sparse fuzzy clustering as described in claim 1, characterized in that: In the step S42: By setting the threshold th and setting the part where the dark detail is less than the threshold to zero, the finally fused image is made cleaner and clearer; L F (i, j) = min(L IR (i, j), L VIS (i, j)) (22) Among them, L F represents the fused dark detail part, L IR represents the dark detail part of infrared rays, L VIS represents the dark detail part of visible light, the threshold th = 5, and (x, y) represents the pixel value position; Finally, the high-frequency components to be fused are combined into the fusion result I of the underlying layer through linear addition Z to obtain the final image fusion result; I F = I Z + H F - L F (24) Among them, I F represents the fusion result of infrared and visible light images, I Z represents the underlying layer information after texture enhancement, L F represents the dark detail part after fusion, H F represents the bright detail part after fusion.
Citation Information
Patent Citations
Method and system for fusing infrared image and visible light image
CN114549382A