Deep potential low-rank representation infrared and visible light image fusion system and method
By employing a deep latent low-rank representation method and combining it with a weighted summation strategy, the common low-rank matrix of infrared and visible light images is gradually obtained. This solves the problem of insufficient deep-level information in image fusion in existing methods, achieves clearer preservation of edge and texture information, and improves image quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TAISHAN UNIV
- Filing Date
- 2024-02-29
- Publication Date
- 2026-04-17
AI Technical Summary
Existing infrared and visible light image fusion methods fail to effectively consider the common features of salient and basic parts, resulting in insufficient deep information extraction and edge contour clarity in the fused image.
A deep latent low-rank representation method is adopted. After decomposition by latent low-rank representation, pre-fusion is performed. Combined with weighted averaging and summation strategies, the common low-rank matrix of salient and basic images is gradually obtained. Then, multiple rounds of deep decomposition are performed through edge weight fusion and weighted summation strategies to finally obtain a more consistent fused image.
It improves the visual effect and information content of images, especially the preservation of edge and texture information, enhances image clarity and feature information, and solves the problem of insufficient deep information extraction in existing methods.
Smart Images

Figure CN121883267A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image fusion technology, and particularly relates to an infrared and visible light image fusion system and method with deep latent low-rank representation. Background Technology
[0002] Image fusion is a processing technique that improves the quality, clarity, and interpretability of an image by integrating information from multiple sources into a new image. The goal of image fusion is to utilize complementary information from different images to provide more detail and accuracy while removing redundant information. This process typically requires consideration of various image features and fusion methods.
[0003] Common methods for image fusion include:
[0004] Pixel-level fusion: Directly fusing image data at the pixel level, such as averaging, weighting, and high-pass filtering.
[0005] Feature-level fusion: By extracting feature information (such as edges, textures, etc.) from each source image, and then fusing these features, a comprehensive feature map is generated.
[0006] Decision-level fusion: Utilizes high-level information (such as object detection, classification, etc.) to analyze and make decisions on individual source images, and then fuses these decision results.
[0007] Image fusion has a wide range of applications, including but not limited to the following areas:
[0008] Remote sensing imaging: By fusing multispectral images with high-resolution panchromatic images, the spatial resolution and spectral information of the images can be improved.
[0009] Medical imaging: Fusion of medical images from different modalities (such as MRI, CT, PET, etc.) helps improve diagnostic accuracy and reduce misdiagnosis.
[0010] Military and surveillance: Fusion of infrared and visible light images can increase the visibility of targets and improve the ability to detect and identify them.
[0011] Consumer electronics products: such as mobile phone photography, can improve photo quality and dynamic range through multi-camera fusion technology.
[0012] Autonomous driving and robot vision: Integrating image information collected by different sensors (such as cameras, lidar, infrared, etc.) helps improve the accuracy of environmental perception and decision-making.
[0013] In summary, image fusion is a powerful image processing technique that integrates image information from multiple sources to improve image quality and usability. Image fusion plays a crucial role in various practical applications.
[0014] Image fusion, as an image enhancement method, aims to combine images acquired by multiple sensors to create synthetic images with complementary information and significant visual enhancement, thereby revealing more distinct target features and richer background details. Therefore, this technology can provide stable and content-rich image support for advanced vision tasks such as remote sensing imaging, autonomous driving and robotic vision, military and surveillance.
[0015] As a crucial component of image fusion, infrared and visible light image fusion technology combines infrared images (primarily containing thermal information of the target) and visible light images (primarily containing texture and color information of the target) to utilize the complementary information of the two images, thereby obtaining an image with richer information. Latent low-rank representation, due to its superior subspace image decomposition, has wide applications in the field of infrared and visible light image fusion.
[0016] Currently, existing methods based on latent low-rank representation mainly fall into three categories. The first involves decomposing the image using latent low-rank representation and then performing enhancement or multi-scale decomposition on the decomposed image to extract more effective features, such as LatLRR-FCNN and WsLatLRR. The second involves multi-level low-rank decomposition, such as MdLatLRR, which performs continuous decomposition on the initially separated basic parts. The third involves improved latent low-rank representation models, such as the SccLatLRR algorithm based on sparse consistency constraints, which simultaneously performs partial decomposition on the input multimodal image to obtain more salient features and more consistent basic features. Chinese Patent CN114648475A, filed on March 14, 2022, entitled "Infrared and Visible Light Image Fusion Method and System Based on Low-Rank Sparse Representation," discloses the above-mentioned technologies.
[0017] The first method uses only latent low-rank representation for decomposition, and then repeatedly decomposes the resulting base image using convolutional neural networks and multi-scale decomposition to obtain more image feature information. While the SccLatLRR algorithm based on sparse consistency constraints uses two images as input, it essentially applies rank consistency constraints to a single image, obtaining more shared basic features but neglecting the issue of shared salient features. Multi-level low-rank decomposition, such as MdLatLRR, continuously decomposes the base image after decomposition to obtain more salient components. All the above methods only continuously decompose the resulting base image, failing to consider that salient images are composed of their own salient and basic components, and that base images also have their own salient and basic components. To address this issue, a weighted strategy is continuously applied to these unique salient and basic components to mutually constrain them, obtaining more consistent component information, thus obtaining deeper salient and basic features, ultimately resulting in a better-performing fused image. Summary of the Invention
[0018] To overcome the problems existing in related technologies, the present invention discloses an infrared and visible light image fusion system and method based on deep latent low-rank representation. Considering the characteristics of multimodal image information, the present invention uses the pre-fusion result of the first step of latent low-rank representation decomposition as the input for the next step of deep latent low-rank representation. The first step of pre-fusion simply involves averaging the separated basic parts, while the saliency parts are summed using a summation strategy. Then, three rounds of deep decomposition are performed, finally yielding the corresponding result.
[0019] The technical solution is as follows: a method for fusing infrared and visible light images using deep latent low-rank representation, comprising:
[0020] S1, taking infrared and visible light images as task inputs, performs latent low-rank representation decomposition, and uses summation and weighted average strategies to obtain preliminary salient and basic images;
[0021] S2, the obtained salient image and base image are used as input to the deep latent low-rank representation, and the rank weighting strategy is used to gradually fuse them in the decomposition to obtain the common base low-rank matrix and the common salient low-rank matrix of the two images.
[0022] S3, multiply the pre-decomposed basic image and salient image with the common basic low-rank matrix to obtain two basic part images of the first layer; multiply the pre-decomposed basic image and salient image with the common salient low-rank matrix to obtain two salient part images of the first layer;
[0023] S4. Apply an edge weight fusion strategy to the obtained base part image to obtain an edge weight fusion image. Fuse the salient part image using a weighted summation strategy to obtain a weighted summation strategy image. Summate the edge weight fusion image and the weighted summation strategy image to obtain the first fusion image.
[0024] S5 continuously uses the edge weight fusion image and weighted summation strategy image obtained from each depth decomposition as input for the next layer of deep latent low-rank representation decomposition to obtain the fusion image of each decomposition layer.
[0025] In step S1, during the decomposition of the latent low-rank representation method, the latent low-rank representation model is:
[0026]
[0027]
[0028] In the formula, Let represent the column rank matrix and row rank matrix of image Y, respectively. β is a parameter γ that balances the influence of noise to obtain the noise component image for each image. * To solve for the nuclear norm, γ is a parameter to balance the influence of noise, Y is the original image data, ||β||1 is the noise characterized by the L1 norm, and εY is the base image of each image obtained by balancing the influence of noise with parameter γ. The salient portion of each image is obtained by balancing the effect of noise with parameter γ.
[0029] The fused base image and salient image are obtained by using a weighted averaging strategy and a weighted summation strategy.
[0030] In step S2, a rank-weighted strategy is used to gradually fuse the images during decomposition, resulting in a common fundamental low-rank matrix and a common salient low-rank matrix between the two images: the salient image and the fundamental image.
[0031] The salient image and the base image are separated into their respective salient low-rank matrices and base low-rank matrices. A weighted average strategy is applied to the separated base low-rank matrices to obtain the common base low-rank matrix of the two images.
[0032] A weighted summation strategy is applied to the salient low-rank matrices of the two images, the salient image and the base image, to obtain the common salient low-rank matrix of the two images;
[0033] By operating on the common matrix, we obtain the common fundamental image and the common salient image of the salient image and the common fundamental image.
[0034] In step S2, the obtained salient image and basic image are used as inputs to the deep latent low-rank representation. The deep latent low-rank representation decomposition model is as follows:
[0035]
[0036] st
[0037]
[0038]
[0039] In the formula, The low-rank matrices of the salient parts of the first and second input images are decomposed respectively, and the rank of the matrices is approximated by the nuclear norm. These are the low-rank matrices of the basic parts decomposed from the first and second input images, respectively; the rank of the matrix is approximated using the nuclear norm. These are sparse noise components, constrained using the L1 norm; ∥.∥ * For solving the nuclear norm, γ is a parameter to balance the influence of noise, Q1 is the base image after the previous decomposition and fusion, and Q2 is the salient image after the previous decomposition and fusion. Each matrix represents the number of layers; M is the number of model layers, and L is the layer traversal, where L = 1, 2, ..., M; The first input image is decomposed into a low-rank matrix of the salient parts after L-1 level traversal. This is the low-rank matrix of the basic part decomposed from the second type of image after traversing L-1 levels;
[0040] The minimum rank of the input image is obtained by solving the deep latent low-rank representation decomposition model. Part of the deep latent low-rank representation decomposition model is then simplified as follows:
[0041]
[0042]
[0043]
[0044]
[0045] In the formula, Minimize the rank of the first image as input, which is the salient part. Minimize the rank of the second image as input for the salient parts. The minimum rank decomposed from the first image input to the base part. The minimized rank is obtained by decomposing the second image as input to the base part.
[0046] Therefore, the final deep latent low-rank representation decomposition model is obtained as follows:
[0047]
[0048] st
[0049]
[0050]
[0051] Furthermore, the alternating direction multiplier method (ADMM) is used to optimize the transformation of the deep latent low-rank representation decomposition model, constructing an augmented Lagrangian function to decompose the global optimization problem into local subproblems. The optimal solution to the global problem is obtained by coordinating the solution sets of the subproblems. The specific solution process is as follows:
[0052] First, introduce Decompose the problem locally using four variables:
[0053]
[0054] st
[0055]
[0056]
[0057]
[0058]
[0059]
[0060]
[0061] In the formula, The variables in the L-level traversal of the first and second images, which are the inputs for the salient parts, are respectively. These are the minimum rank decomposed from the first image with salient input and the minimum rank decomposed from the second image with salient input, respectively. The Lagrange multipliers in the L-level traversal of the first and second images input into the basic part are respectively used;
[0062] use The problem of obtaining the augmented Lagrange function using six Lagrange multipliers is as follows:
[0063]
[0064] in, This represents the sum of the diagonal elements of the product of the transpose of A and B. For penalty parameters, Let Frobenius norm be the matrix, which represents the square root of the sum of the squares of all elements in the matrix; The Lagrange multipliers are used to iterate through the input image of the base part from level 1 to level 6. The Lagrange multipliers in the L-level traversal of the input image of the base part. Minimize the rank of the decomposition obtained by performing L-level traversal on the input image for the salient parts. The minimum rank is obtained by decomposing the input image into L layers;
[0065] Based on the optimization principle of the Alternating Direction Multiplier Method (ADMM), the variables are updated alternately until the function meets the convergence condition.
[0066] Furthermore, based on the optimization principle of the ADMM algorithm, variables are alternately updated until the function satisfies the convergence condition. The deep latent low-rank representation with M layers is divided into M subproblems. To learn the Lth layer, where L = 1, 2…M, the objective function is defined for a specific layer, expressed as:
[0067]
[0068] The specific variables that are alternately updated include:
[0069] (1) Update
[0070]
[0071]
[0072] In the formula, The closed-form solution can be decomposed using SVD, and the closed-form solution obtained is: The closed-form solution is
[0073] (2) Update
[0074]
[0075]
[0076] In the formula, Q2′Q2 represents the product of the transformation matrix of Q2 and Q2, Q1Q1′ represents the product of the transformation matrices of Q1 and Q1, and (Q1′Q1+I) -1 Let I be the inverse of the matrix, where I is the identity matrix;
[0077] (3) Update
[0078]
[0079]
[0080] In the formula, The closed-form solution can be decomposed using SVD, and the closed-form solution obtained is: The closed-form solution is
[0081] (4) Update
[0082]
[0083]
[0084] (5) Update
[0085]
[0086]
[0087] (6) Update the augmented Lagrange multipliers according to the following conditions:
[0088]
[0089]
[0090]
[0091]
[0092]
[0093]
[0094] In the formula, m is the number of algorithm iterations;
[0095] Update the balance factor according to the following conditions:
[0096]
[0097] Where ρ and μ max The vector is a preset vector; the convergence condition is:
[0098]
[0099]
[0100]
[0101]
[0102]
[0103]
[0104] In the formula, ε is a preset value, ε = 10 -5 ρ = 1.2, where μ is the optimization step size. max =1000.
[0105] In step S3, the final two base images of the first layer are obtained. and The two final salient part images of the first layer are obtained as follows and
[0106] In step S4, the noise portion of the two final base part images and the salient part image obtained in step S3 is processed. and Discard.
[0107] Furthermore, an edge weight fusion strategy is applied to the obtained base image to obtain the edge weight fused image. Based on the weights, the following formula is used for calculation:
[0108]
[0109] in, As the final base part after weight fusion, θ1 and θ2 represent the edge weights of infrared and visible light images, respectively;
[0110] The salient parts of the image are fused using a weighted summation strategy to obtain a weighted summation strategy image, including:
[0111] The weighted summation strategy formula for the separated significant components is as follows:
[0112] and For the separated salient parts, a weighted summation strategy is selected for fusion, as shown in the following formula:
[0113]
[0114] in, The final salient image after fusion;
[0115] The final edge-weighted fused image and the weighted summation strategy image are summed to obtain the first fused image. The image fusion formula using the weighted summation strategy is as follows:
[0116]
[0117] Another object of the present invention is to provide an infrared and visible light image fusion system based on deep latent low-rank representation, the system implementing the infrared and visible light image fusion method based on deep latent low-rank representation, the system comprising:
[0118] The pre-decomposition module is used to decompose infrared and visible light images using a latent low-rank representation method; preliminary salient images and base images are obtained by using a summation strategy and a weighted average strategy, respectively.
[0119] The deep latent low-rank representation decomposition module takes the obtained salient image and the basic image as input to the deep latent low-rank representation. A rank-weighted strategy is used to progressively fuse them during decomposition, resulting in a common basic low-rank matrix and a common salient low-rank matrix between the two images. The basic and salient images obtained from the pre-decomposition are multiplied by the common basic low-rank matrix to obtain the two final basic part images of the first layer. The pre-decomposition of the basic and salient images and the common salient low-rank matrix yields the two final salient part images of the first layer.
[0120] The image fusion module is used to apply an edge weight fusion strategy to the obtained base part image to obtain an edge weight fused image, and to fuse the salient part image through a weighted summation strategy to obtain a weighted summation strategy image. The two images, the final edge weight fused image and the edge weight fused image, are summed to obtain the first fused image. The edge weight fused image and the edge weight fused image obtained from each depth decomposition are continuously used as inputs to the next layer of deep latent low-rank representation decomposition to obtain the fused image of each decomposition layer.
[0121] Combining all the above technical solutions, the advantages and positive effects of this invention are as follows: This invention provides a novel machine learning framework for latent low-rank representation. It uses the basic and salient parts generated after a single decomposition in traditional latent low-rank representation methods as pre-fused images, while simultaneously inputting deep latent low-rank representation. After multiple rounds of deep decomposition, more consistent salient and basic features are obtained. In particular, we employ a novel visual weight map method in the fusion strategy to better highlight the edge parts of the image. Furthermore, this invention can be well applied to the fusion of infrared and visible light images. Compared with existing methods, the fused image obtained by the algorithm of this invention has a significant effect.
[0122] This invention develops a deep latent low-rank representation multimodal image fusion framework by utilizing complementary information from multimodal sensor image data. Unlike previous methods that utilize LatLRR-based decomposition of infrared and visible light images, the proposed method uses latent low-rank representation to decompose a pre-decomposed image. Simultaneously, it performs deep latent low-rank representation fusion decomposition on the salient image and the base image obtained from the pre-decomposition. During the decomposition process, the effective part of the fusion information (the decomposed low-rank matrix) is continuously obtained. In this way, the deep feature information of the image can be obtained by utilizing the fused image.
[0123] This invention proposes a basic part fusion strategy that makes better use of image edge information as the basic part fusion weight, resulting in better visual effects and more feature information. Attached Figure Description
[0124] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure;
[0125] Figure 1 This is a flowchart of the infrared and visible light image fusion method based on deep latent low-rank representation provided in this embodiment of the invention;
[0126] Figure 2 This is a schematic diagram of the infrared and visible light image fusion method based on deep latent low-rank representation provided in this embodiment of the invention.
[0127] Figure 3 This is a schematic diagram of the DeepLatLRR method provided in an embodiment of the present invention;
[0128] Figure 4 This is a schematic diagram of visible light image weighting provided in an embodiment of the present invention;
[0129] Figure 5 This is a schematic diagram of infrared image weighting provided in an embodiment of the present invention;
[0130] Figure 6 This is an infrared image fusion effect diagram provided in an embodiment of the present invention;
[0131] Figure 7 This is a visible light image fusion effect diagram provided in an embodiment of the present invention;
[0132] Figure 8 This is a first-layer fusion image fusion effect diagram provided in an embodiment of the present invention;
[0133] Figure 9 This is a diagram showing the second-layer fusion image fusion effect provided in an embodiment of the present invention;
[0134] Figure 10This is a third-layer fused image effect diagram provided in an embodiment of the present invention. Detailed Implementation
[0135] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0136] To address the shortcomings of current infrared and visible light fusion methods based on latent low-rank representations, this invention designs a deep latent low-rank representation image fusion method. This method employs a pre-fusion mechanism and performs deep latent low-rank representation decomposition, further refining the basic and salient features separated from the previous layer. Fusion is performed layer-by-layer decomposition, which allows for more accurate learning of the image's subspace structure and feature information. Simultaneously, it achieves better results in both subjective and objective evaluation.
[0137] Example 1, such as Figure 1 As shown, the infrared and visible light image fusion method based on deep latent low-rank representation provided in this embodiment of the invention includes the following steps:
[0138] S1, taking infrared and visible light images as task inputs, performs latent low-rank representation decomposition, and uses summation and weighted average strategies to obtain preliminary salient and basic images;
[0139] S2, the obtained salient image and base image are used as input to the deep latent low-rank representation, and the rank weighting strategy is used to gradually fuse them in the decomposition to obtain the common base low-rank matrix and the common salient low-rank matrix of the two images.
[0140] S3, multiply the pre-decomposed basic image and salient image with the common basic low-rank matrix to obtain two basic part images of the first layer; multiply the pre-decomposed basic image and salient image with the common salient low-rank matrix to obtain two salient part images of the first layer;
[0141] S4. Apply an edge weight fusion strategy to the obtained base part image to obtain an edge weight fusion image. Fuse the salient part image using a weighted summation strategy to obtain a weighted summation strategy image. Summate the edge weight fusion image and the weighted summation strategy image to obtain the first fusion image.
[0142] S5 continuously uses the edge weight fusion image and weighted summation strategy image obtained from each depth decomposition as input for the next layer of deep latent low-rank representation decomposition to obtain the fusion image of each decomposition layer.
[0143] In an embodiment of the present invention, Figure 2 This is the principle of the infrared and visible light image fusion method based on deep latent low-rank representation provided in the embodiments of the present invention.
[0144] As can be seen from the above embodiments, the expected benefits and commercial value of the technical solution of the present invention after transformation are as follows:
[0145] Improving the performance of surveillance and security systems: In the field of security surveillance, fusing infrared and visible light images can significantly improve the clarity and reliability of night vision monitoring. The technical solution of this invention can provide more detailed image information, which is of great significance for the improvement of security surveillance systems, and may increase market demand for high-end surveillance equipment, thereby creating greater economic value.
[0146] Enhancing the recognition capabilities of autonomous driving systems: In autonomous driving technology, vehicles need to be able to accurately identify their environment under various lighting conditions. The deep fusion technology provided by this invention can provide richer environmental information at night or in other low-light conditions, thereby helping to improve the reliability and safety of autonomous driving systems.
[0147] Improving Medical Imaging Quality: In the field of medical imaging, combining different types of images can provide more comprehensive diagnostic information. The technology of this invention can also be applied to the fusion of medical images, helping doctors make more accurate diagnoses, which may increase the market value of related medical devices.
[0148] Advanced military applications: Military reconnaissance and target identification often need to be carried out under various environmental conditions, and the image fusion technology provided by this invention can provide clearer images, enhance night vision capabilities and environmental awareness, which has high commercial value for military applications.
[0149] Potential copyright and licensing revenue: This invention can generate copyright royalties by licensing or authorizing other companies to use it, which provides the possibility of creating passive income.
[0150] Technological leadership advantage: Adopting advanced image fusion technology can establish a technological leadership position in related fields, bringing enterprises an increase in brand reputation and market share.
[0151] Value-added services and solutions: As the technology matures, related value-added services and solutions can be developed to provide customers with customized products, thereby creating a differentiated competitive advantage in the market.
[0152] In conclusion, the technical solution of this invention is not only innovative in terms of technology, but also has broad application prospects and huge market potential in terms of commerce. Through integration with related industries, it can bring significant economic benefits and a competitive advantage to enterprises.
[0153] Application of Deep Latent Low-Rank Representation: In existing technologies, infrared and visible light image fusion mainly relies on traditional image processing techniques, such as LatLRR (Latent Low-Rank Representation) decomposition methods. While these methods can handle image fusion problems to a certain extent, they often neglect the deeper information of the images. This invention employs a deep latent low-rank representation method, which can deeply extract the salient features and low-level structural information of the pre-decomposed image, effectively improving the quality and information content of the fused image, thus filling the technological gap in deep information extraction.
[0154] Basic Component Fusion Strategy: Compared to existing technologies, this invention proposes a novel basic component fusion strategy that fully utilizes image edge information as fusion weights. This strategy effectively enhances the edge and texture information of the image, resulting in a clearer visual effect and richer feature information in the fused image. Previous fusion methods often failed to fully consider the importance of edge information; the strategy of this invention fills the technological gap in the field regarding the utilization of edge information.
[0155] Iterative Optimization of Fusion Decomposition: This invention employs an iterative method to obtain the effective portion of the fusion information during the decomposition and fusion process. This method continuously optimizes the low-rank matrix, making each iteration closer to the ideal fusion effect. Most existing methods lack such an iterative optimization mechanism; therefore, this invention fills the gap in existing technologies regarding the continuous improvement of fusion results.
[0156] Comprehensive Utilization of Multimodal Sensor Image Data: This invention leverages the complementary properties of multimodal sensor data, combining fundamental and salient image information through deep latent low-rank representation technology to achieve effective fusion of data from different modalities. Most existing methods fail to effectively combine the complementary information of multimodal data, thus limiting the effectiveness of multimodal image fusion. The method of this invention solves this technical challenge, providing a new and effective approach for multimodal image fusion.
[0157] In summary, this invention proposes a novel multimodal image fusion framework by introducing deep learning and optimization theory into the field of image fusion. This framework not only improves the visual effect and information content of the fused images, but also brings innovative progress to the field of infrared and visible light image fusion in both technical theory and practical application, filling several technical gaps.
[0158] The technical solution of this invention solves a long-standing technical problem that has remained unsolved: it brings significant progress to image fusion technology, particularly in the preservation of texture information and the sharpness of edge contours. In previous techniques, texture details were easily lost during the fusion process, and edge contours often became blurred due to improper fusion. These problems are particularly prominent in fields requiring high-resolution image analysis, such as precision engineering, medical diagnostics, and high-security surveillance.
[0159] By employing a deep latent low-rank representation method, this invention enables more in-depth extraction and analysis of fundamental and salient features in images. This method specifically considers the phase consistency between infrared and visible light images at the feature level, achieving fine-grained learning of the image subspace structure. This in-depth analysis ensures the integrity and consistency of information during image fusion, significantly improving the richness of texture information and the sharpness of edges.
[0160] Meanwhile, this invention introduces an innovative edge weighting calculation mechanism, which optimizes the fusion result by accurately measuring the importance of edge information. This not only enhances edge contours but also largely suppresses the influence of noise. This mechanism provides a new approach to obtaining high-quality fused images, solving a problem that previous techniques could not overcome.
[0161] In summary, this invention provides an efficient and reliable new technology in the field of image fusion, which solves several long-standing problems, and its optimized fusion results have great practical significance and potential commercial value for various image processing applications.
[0162] This invention successfully overcomes some technical biases in the field of image fusion technology through its deep latent low-rank representation method: deep mining of inherent image features: This invention does not rely on the surface feature analysis commonly found in traditional image fusion technology, but instead mines the deep features of images through deep learning technology, thereby extracting richer information content.
[0163] Comprehensive analysis of salient and fundamental components: Unlike traditional methods, this invention comprehensively considers the interrelationship and phase consistency between salient and fundamental images, thereby maintaining the integrity and consistency of image information.
[0164] Precise fusion of edge information: This invention uses an edge weight calculation method to focus on strengthening edge information, which solves the problem of insufficient edge contour in the prior art, and greatly improves the visual effect of the fused image.
[0165] Effective reduction of noise impact: During image fusion, this invention can effectively reduce noise interference, thereby enhancing the accuracy and usability of the images.
[0166] Through these innovative strategies, this invention not only improves the performance of image fusion, but also breaks through the limitations of past technologies, providing a more efficient and reliable new method for image fusion.
[0167] This invention provides an infrared and visible light image fusion system with deep latent low-rank representation, comprising:
[0168] The pre-decomposition module is used to decompose infrared and visible light images using a latent low-rank representation method; preliminary salient images and base images are obtained by using a summation strategy and a weighted average strategy, respectively.
[0169] The deep latent low-rank representation decomposition module takes the obtained salient image and the base image as input to the deep latent low-rank representation; it employs a rank-weighted strategy to progressively fuse the two images during decomposition, obtaining a common base low-rank matrix and a common salient low-rank matrix; it multiplies the pre-decomposed base image and salient image with the common base low-rank matrix to obtain two base part images of the first layer; and it multiplies the pre-decomposed salient image and base image with the common salient low-rank matrix to obtain two salient part images of the first layer.
[0170] The image fusion module is used to apply an edge weight fusion strategy to the obtained base part image to obtain an edge weight fusion image, and to fuse the salient part image through a weighted summation strategy to obtain a weighted summation strategy image. The two images, the final edge weight fusion image and the weighted summation strategy image, are summed to obtain the first fusion image. The edge weight fusion image and the weighted summation strategy image obtained from each depth decomposition are continuously used as inputs for the next layer of deep latent low-rank representation decomposition to obtain the fusion image of each decomposition layer.
[0171] Example 2, as another embodiment of the present invention, the infrared and visible light image fusion system with deep latent low-rank representation provided in this embodiment of the present invention includes three modules: a pre-decomposition module, a deep latent low-rank representation decomposition module, and an image fusion module.
[0172] (1) Pre-decomposition module:
[0173] This module primarily utilizes the latent low-rank representation method to perform image decomposition on infrared and visible light images. Its physical model is as follows:
[0174]
[0175]
[0176] In the formula, γ is a parameter to balance the influence of noise, and Y is the original image data. Let represent the column rank matrix and row rank matrix of image Y, respectively. The original image data includes infrared and visible light images. * To solve for the nuclear norm, ||β||1 represents the noise characterized by the L1 norm, and εY is the parameter γ that balances the influence of noise to obtain the base part image of each image. The salient portion of each image is obtained by balancing the effect of noise with parameter γ, and the noise component image of each image is obtained by balancing the effect of noise with parameter γ.
[0177] (2) Deep latent low-rank representation decomposition module:
[0178] The fused base part image and salient part image obtained from the first module are used as input data for this module. The deep latent low-rank representation decomposition model is shown below:
[0179]
[0180] st
[0181]
[0182]
[0183] In the formula, The low-rank matrices of the salient parts of the first and second input images are decomposed respectively, and the rank of the matrices is approximated by the nuclear norm. These are the low-rank matrices of the basic parts decomposed from the first and second input images, respectively; the rank of the matrix is approximated using the nuclear norm. Q1 represents the sparse noise components, constrained by the L1 norm; Q2 represents the base image after the previous decomposition and fusion, Q1 represents the salient image after the previous decomposition and fusion, and γ is a parameter to balance the influence of noise. Each element is an identity matrix, used to represent the number of layers; ∥.∥ * For solving the nuclear norm, M is the number of model layers, and L is the layer traversal, where L = 1, 2, ..., M; The first input image is decomposed into a low-rank matrix of the salient parts after L-1 level traversal. This is the low-rank matrix of the basic part decomposed from the second type of image after traversing L-1 levels;
[0184] The minimum rank of the input image is obtained by solving the deep latent low-rank representation decomposition model. Part of the deep latent low-rank representation decomposition model is then simplified as follows:
[0185]
[0186]
[0187]
[0188]
[0189] In the formula, Minimize the rank of the first image as input, which is the salient part. Minimize the rank of the second image as input for the salient parts. The minimum rank decomposed from the first image input to the base part. The minimized rank is obtained by decomposing the second image as input to the base part.
[0190] Therefore, the final deep latent low-rank representation decomposition model is obtained as follows:
[0191]
[0192] st
[0193]
[0194]
[0195] In the model optimization stage, this invention employs the Alternating Direction Method of Multipliers (ADMM) for optimization. Specifically, this invention constructs an augmented Lagrangian function, a step that decomposes the large-scale, complex global optimization problem into a series of smaller, more manageable local subproblems. Then, this invention coordinates the solution sets of these subproblems to obtain the optimal solution to the global problem. The specific solution process is as follows:
[0196] First, introduce Decompose the problem locally using four variables:
[0197]
[0198] st
[0199]
[0200]
[0201]
[0202]
[0203]
[0204]
[0205] In the formula, The variables in the L-level traversal of the first and second images, which are the inputs for the salient parts, are respectively. These are the minimum rank decomposed from the first image with salient input and the minimum rank decomposed from the second image with salient input, respectively. The Lagrange multipliers in the L-level traversal of the first and second images input into the basic part are respectively used;
[0206] use The problem of obtaining the augmented Lagrange function using six Lagrange multipliers is as follows:
[0207]
[0208] in,<A,B> =Tr(A ’ B) represents the sum of the diagonal elements of the product of the transpose of A and B; For penalty parameters, Let Frobenius norm be the matrix, which represents the square root of the sum of the squares of all elements in the matrix; The Lagrange multipliers are used to iterate through the input image of the base part from level 1 to level 6. The Lagrange multipliers in the L-level traversal of the input image of the base part. Minimize the rank of the decomposition obtained by performing L-level traversal on the input image for the salient parts. The minimum rank is obtained by decomposing the input image into L layers;
[0209] Then, based on the optimization principle of the Alternating Direction Multiplier Method (ADMM), the variables are updated alternately until the function meets the convergence condition.
[0210] According to the optimization principle of the ADMM algorithm, variables are updated alternately until the function satisfies the convergence condition. A deep latent low-rank representation with M layers is divided into M subproblems. To learn the Lth layer, where L = 1, 2, ..., M, the objective function is defined for a specific layer, expressed as:
[0211]
[0212] The specific variables that are alternately updated include:
[0213] (1) Update
[0214]
[0215]
[0216] In the formula, The closed-form solution can be decomposed using SVD, and the closed-form solution obtained is: The closed-form solution is
[0217] (2) Update
[0218]
[0219]
[0220] In the formula, Q2′Q2 represents the product of the transformation matrix of Q2 and Q2, Q1Q1′ represents the product of the transformation matrices of Q1 and Q1, and (Q1′Q1+I) -1 Let I be the inverse of the matrix, where I is the identity matrix;
[0221] (3) Update
[0222]
[0223]
[0224] In the formula, The closed-form solution can be decomposed using SVD, and the closed-form solution obtained is: Similarly The closed-form solution is
[0225] (4) Update
[0226]
[0227]
[0228] (5) Update
[0229]
[0230]
[0231] (6) Update the augmented Lagrange multipliers according to the following conditions:
[0232]
[0233]
[0234]
[0235]
[0236]
[0237]
[0238] In the formula, m is the number of algorithm iterations;
[0239] Update the balance factor according to the following conditions:
[0240]
[0241] Where ρ and μ max The vector is a preset vector; the convergence condition is:
[0242]
[0243]
[0244]
[0245]
[0246]
[0247]
[0248] In the formula, ε is a preset value, ε = 10 -5 ρ = 1.2, where μ is the optimization step size. max =1000.
[0249] (3) Image fusion module:
[0250] Get the significant portion of the image from the previous step. and Basic Image and For the obtained noise part and Discard.
[0251] (3.1) First, perform basic image fusion:
[0252] In image fusion strategies, this invention proposes a fusion strategy based on edge weight maps, specifically including:
[0253] (3.1.1) Calculate the average value of the image matrix;
[0254] (3.1.2) Obtain the difference between each pixel value and the average value and take the absolute value to obtain matrix S;
[0255] (3.1.3) After obtaining the matrix S from the previous step, find the maximum and minimum values of the matrix;
[0256] (3.1.4) Use S - minimum value / (maximum value - minimum value) to perform image denoising operations;
[0257] (3.1.5) The denoised image is the edge weight image.
[0258] Since low-frequency information in an image represents most of the smooth regions, which contain most of the image energy and typically represent the background, and image intensity information is also reflected in the low frequencies, this invention employs an edge weighting strategy to effectively highlight the background information of the image while maintaining good target feature edges in the fused image. For the fusion of the basic image components, this invention uses the following formula to calculate based on the weights:
[0259]
[0260] in As the final base part after weight fusion, θ1 and θ2 represent the edge weights of infrared and visible light images, respectively.
[0261] (3.2) Significant part fusion:
[0262] and For the separated significant portion, this invention selects a weighted summation strategy for fusion, as shown in the following formula:
[0263]
[0264] in, This represents the salient image after final fusion.
[0265] After obtaining the basal and salient parts of the decomposed image, a weighted sum strategy is used for image fusion, as shown in the following formula:
[0266]
[0267] Example 3: The infrared and visible light image fusion method and system with deep latent low-rank representation provided in this embodiment of the invention can be applied to the fields of security monitoring, medical diagnosis, and military equipment. Taking the field of security monitoring as an example, it improves the monitoring quality in nighttime and low-light environments. By fusing infrared and visible light images to provide more comprehensive image information, it improves nighttime monitoring and target recognition capabilities, helps monitoring personnel detect anomalies, provides a more complete understanding of potential threats, and offers higher security.
[0268] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0269] The information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of the present invention. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0270] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this invention. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments.
[0271] Based on the technical solutions described in the above embodiments of the present invention, the following application examples can be further proposed.
[0272] According to embodiments of this application, the present invention also provides a computer device comprising: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor executes the computer program to implement the steps in any of the above-described method embodiments.
[0273] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps described in the various method embodiments above.
[0274] This invention also provides an information data processing terminal, which, when executed on an electronic device, provides a user input interface to implement the steps described in the above method embodiments. The information data processing terminal is not limited to mobile phones, computers, or switches.
[0275] This invention also provides a computer program product that, when run on an electronic device, enables the electronic device to implement the steps described in the various method embodiments above.
[0276] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks.
[0277] To further demonstrate the positive effects of the above embodiments, the present invention conducts the following experiments based on the above technical solution: The present invention proposes a deep latent low-rank representation method for the first time and applies it to the field of infrared and visible light image fusion. The present invention improves the latent low-rank representation method by revealing deep hidden features and deep structures hidden in the latent subspace through the improved Deep LatLRR method. To improve representation learning ability, the present invention uses a pre-fused image as input, gradually fuses the first few layers, and then fuses the subspace. Specifically, the method proposed in this invention uses the shallow features of the previous layer as input for subsequent layers, then recovers hierarchical information and deeper features. At the same time, different weighted fusion methods are used on the obtained low-rank dictionary to better improve the representation effect of the fused image.
[0278] See details Figures 3-10 ,Right now Figure 3 This is a schematic diagram of the DeepLatLRR method provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of visible light image weighting provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of infrared image weighting provided in an embodiment of the present invention; Figure 6 This is a visible light image from the TNO dataset provided in this embodiment of the invention; Figure 7 This is an infrared image from the TNO dataset provided in this embodiment of the invention; Figure 8 This is a first-layer fusion image fusion effect diagram provided in an embodiment of the present invention; Figure 9 This is a diagram showing the second-layer fusion image fusion effect provided in an embodiment of the present invention; Figure 10This is a diagram illustrating the third-layer fusion image effect provided in this embodiment of the invention. After decomposing the infrared and visible light images into a basic part and a salient part, this invention proposes an edge-weighted fusion strategy for fusing the basic part. Due to the special nature of edge information in image fusion, this fusion strategy improves the subjective evaluation effect of the image to a certain extent. Simultaneously, it provides effective support for subsequent advanced visual learning tasks.
[0279] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention and within the spirit and principles of the present invention should be covered within the scope of protection of the present invention.
Claims
1. A method for fusing infrared and visible light images using deep latent low-rank representation, characterized in that, The method includes: S1, taking infrared and visible light images as task inputs, performs latent low-rank representation decomposition, and uses summation and weighted average strategies to obtain preliminary salient and basic images; S2, the obtained salient image and base image are used as input to the deep latent low-rank representation, and the rank weighting strategy is used to gradually fuse them in the decomposition to obtain the common base low-rank matrix and the common salient low-rank matrix of the two images. S3, multiply the pre-decomposed basic image and salient image with the common basic low-rank matrix to obtain two basic part images of the first layer; multiply the pre-decomposed basic image and salient image with the common salient low-rank matrix to obtain two salient part images of the first layer; S4. Apply an edge weight fusion strategy to the obtained base part image to obtain an edge weight fusion image. Fuse the salient part image using a weighted summation strategy to obtain a weighted summation strategy image. Summate the edge weight fusion image and the weighted summation strategy image to obtain the first fusion image. S5 continuously uses the edge weight fusion image and weighted summation strategy image obtained from each depth decomposition as input for the next layer of deep latent low-rank representation decomposition to obtain the fusion image of each decomposition layer.
2. The infrared and visible light image fusion method based on deep latent low-rank representation according to claim 1, characterized in that, In step S1, during the decomposition of the latent low-rank representation method, the latent low-rank representation model is: In the formula, Let represent the column rank matrix and row rank matrix of image Y, respectively. β is a parameter γ that balances the influence of noise to obtain the noise component image for each image. * To solve for the nuclear norm, γ is a parameter to balance the influence of noise, Y is the original image data, ||β||1 is the noise characterized by the L1 norm, and εY is the base image of each image obtained by balancing the influence of noise with parameter γ. The salient portion of each image is obtained by balancing the effect of noise with parameter γ. The fused base image and salient image are obtained by using a weighted averaging strategy and a weighted summation strategy.
3. The infrared and visible light image fusion method based on deep latent low-rank representation according to claim 1, characterized in that, In step S2, a rank-weighted strategy is used to gradually fuse the images during decomposition, resulting in a common fundamental low-rank matrix and a common salient low-rank matrix between the two images: the salient image and the fundamental image. The salient image and the base image are separated into their respective salient low-rank matrices and base low-rank matrices. A weighted average strategy is applied to the separated base low-rank matrices to obtain the common base low-rank matrix of the two images. A weighted summation strategy is applied to the salient low-rank matrices of the two images, the salient image and the base image, to obtain the common salient low-rank matrix of the two images; By operating on the common matrix, we obtain the common fundamental image and the common salient image of the salient image and the common fundamental image.
4. The infrared and visible light image fusion method based on deep latent low-rank representation according to claim 1, characterized in that, In step S2, the obtained salient image and basic image are used as inputs to the deep latent low-rank representation. The deep latent low-rank representation decomposition model is as follows: st In the formula, The low-rank matrices of the salient parts of the first and second input images are decomposed respectively, and the rank of the matrices is approximated by the nuclear norm. These are the low-rank matrices of the basic parts decomposed from the first and second input images, respectively; the rank of the matrix is approximated using the nuclear norm. These are sparse noise components, constrained using the L1 norm; * For solving the nuclear norm, γ is a parameter to balance the influence of noise, Q1 is the base image after the previous decomposition and fusion, and Q2 is the salient image after the previous decomposition and fusion. Each matrix represents an identity matrix used to represent the number of layers; M is the number of model layers, and L is the layer traversal, where L = 1, 2, ..., M; The first input image is decomposed into a low-rank matrix of the salient parts after L-1 level traversal. This is the low-rank matrix of the basic part decomposed from the second type of image after traversing L-1 levels; The minimum rank of the input image is obtained by solving the deep latent low-rank representation decomposition model. Part of the deep latent low-rank representation decomposition model is then simplified as follows: In the formula, Minimize the rank of the first image as input, which is the salient part. Minimize the rank of the second image as input for the salient parts. The minimum rank decomposed from the first image input to the base part. The minimized rank is obtained by decomposing the second image as input to the base part; Therefore, the final deep latent low-rank representation decomposition model is obtained as follows: st 5. The infrared and visible light image fusion method based on deep latent low-rank representation according to claim 4, characterized in that, The deep latent low-rank representation decomposition model is optimized using the Alternating Directional Multiplier Method (ADMM) to construct an augmented Lagrangian function, thus decomposing the global optimization problem into local subproblems. The optimal solution to the global problem is obtained by harmonizing the solution sets of the subproblems. The specific solution process is as follows: First, introduce Decompose the problem locally using four variables: st In the formula, The variables in the L-level traversal of the first and second images, which are the inputs for the salient parts, are respectively. These are the minimum rank decomposed from the first image with salient input and the minimum rank decomposed from the second image with salient input, respectively. The Lagrange multipliers in the L-level traversal of the first and second images input into the basic part are respectively used; use The problem of obtaining the augmented Lagrange function using six Lagrange multipliers is as follows: in,<A,B> =Tr(A'B), which represents the sum of the diagonal elements of the product of the transpose of A and B; For penalty parameters, Let Frobenius norm be the matrix, which represents the square root of the sum of the squares of all elements in the matrix; The Lagrange multipliers are used to iterate through the input image of the base part from level 1 to level 6. The Lagrange multipliers in the L-level traversal of the input image of the base part. Minimize the rank of the decomposition obtained by performing L-level traversal on the input image for the salient parts. The minimum rank decomposed by performing L-level traversal on the base input image; Based on the optimization principle of the Alternating Direction Multiplier Method (ADMM), the variables are updated alternately until the function meets the convergence condition.
6. The infrared and visible light image fusion method based on deep latent low-rank representation according to claim 5, characterized in that, According to the optimization principle of the ADMM algorithm, variables are updated alternately until the function satisfies the convergence condition. A deep latent low-rank representation with M layers is divided into M subproblems. To learn the Lth layer, where L = 1, 2, ..., M, the objective function is defined for a specific layer, expressed as: The specific variables that are alternately updated include: (1) Update In the formula, The closed-form solution can be decomposed using SVD, and the closed-form solution obtained is: The closed-form solution is (2) Update In the formula, Q2′Q2 represents the product of the transformation matrix of Q2 and Q2, Q1Q1′ represents the product of the transformation matrices of Q1 and Q1, and (Q1′Q1+I) -1 Let I be the inverse of the matrix, where I is the identity matrix; (3) Update In the formula, The closed-form solution can be decomposed using SVD, and the closed-form solution obtained is: The closed-form solution is (4) Update (5) Update (6) Update the augmented Lagrange multipliers according to the following conditions: In the formula, m is the number of algorithm iterations; Update the balance factor according to the following conditions: Where ρ and μ max The vector is a preset vector; the convergence condition is: In the formula, ε is a preset value, ε = 10 -5 ρ = 1.2, where μ is the optimization step size. max =1000.
7. The infrared and visible light image fusion method based on deep latent low-rank representation according to claim 1, characterized in that, In step S3, the final two base images of the first layer are obtained. and The two final salient part images of the first layer are obtained as follows and 8. The infrared and visible light image fusion method based on deep latent low-rank representation according to claim 1, characterized in that, In step S4, the noise portion of the two final base images and the salient image of the first layer obtained in step S3 is processed. and Discard.
9. The infrared and visible light image fusion method based on deep latent low-rank representation according to claim 8, characterized in that, The obtained base image is fused using an edge weight fusion strategy. The resulting edge weight fused image is calculated using the following formula based on the weights: in, This is the final foundational part after weight fusion. and These represent the edge weights of infrared and visible light images, respectively. The salient parts of the image are fused using a weighted summation strategy to obtain a weighted summation strategy image, including: The weighted summation strategy formula for the separated significant components is as follows: and For the separated salient parts, a weighted summation strategy is selected for fusion, as shown in the following formula: in, The final salient image after fusion; The final edge-weighted fused image and the weighted summation strategy image are summed to obtain the first fused image. The image fusion formula using the weighted summation strategy is as follows:
10. A system for fusing infrared and visible light images using deep latent low-rank representation, characterized in that, The system implements the infrared and visible light image fusion method based on the deep latent low-rank representation as described in any one of claims 1-9, and the system comprises: The pre-decomposition module is used to decompose infrared and visible light images using a latent low-rank representation method; preliminary salient images and base images are obtained by using a summation strategy and a weighted average strategy, respectively. The deep latent low-rank representation decomposition module takes the obtained salient image and the basic image as input to the deep latent low-rank representation. A rank-weighted strategy is used to progressively fuse them during decomposition, resulting in a common basic low-rank matrix and a common salient low-rank matrix between the two images. The basic and salient images obtained from the pre-decomposition are multiplied by the common basic low-rank matrix to obtain the two final basic part images of the first layer. The pre-decomposition of the basic and salient images and the common salient low-rank matrix yields the two final salient part images of the first layer. The image fusion module is used to apply an edge weight fusion strategy to the obtained base part image to obtain an edge weight fused image, and to fuse the salient part image through a weighted summation strategy to obtain a weighted summation strategy image. The two images, the final edge weight fused image and the edge weight fused image, are summed to obtain the first fused image. The edge weight fused image and the edge weight fused image obtained from each depth decomposition are continuously used as inputs to the next layer of deep latent low-rank representation decomposition to obtain the fused image of each decomposition layer.