General image fusion method and system based on low rank and sparse prior

By extracting image features based on low-rank decomposition and sparse representation, the problem of insufficient modal feature extraction in the existing image fusion method is solved, and efficient image fusion and network interpretability are achieved.

CN120410892AActive Publication Date: 2025-08-01TIANJIN POLYTECHNIC UNIV

Patent Information

Application Number
CN202510920674.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-08-01
Estimated Expiration
2045-07-04

AI Technical Summary

Technical Problem

The existing image fusion method ignores the unique features of the modal when extracting common features between modals, resulting in high network computing complexity and low fusion efficiency, lack of in-depth exploration of common features between different fusion tasks, and lack of interpretability of the model.

Method used

Common features are extracted by a common feature encoder based on low rank decomposition, and sparse unique features are extracted through a unique feature encoder with sparse representation. Combined with the entropy feature fusion module for fusion, a serial adaptive fusion module and feature reconstruction module are designed to realize feature decoupling and cross-modal interaction.

Benefits of technology

It improves the interpretability and fusion efficiency of the network, fully explores the unique characteristics of the modality, and achieves efficient cross-modal information interaction and fusion image quality improvement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120410892A_ABST
    Figure CN120410892A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image fusion, and provides a general image fusion method and system based on low rank and sparse prior, and the method comprises the steps: extracting common features of two source images through a common feature encoder based on low rank decomposition; respectively extracting sparse unique features of the two source images through a unique feature encoder based on sparse representation; fusing the sparse unique features of the two source images through a serial adaptive fusion module, and fusing the sparse unique features with the common features of the two source images to obtain fused features; and inputting the fused features into a feature reconstruction module to obtain a fused image. According to the method, the image fusion problem is solved by mining the internal relationship of the source image features, and the image features are divided into common features and sparse unique features, so that information flow in the network is clearer, the interpretability of the network is greatly improved, effective cross-modal interaction is realized, and the unique modal features are fully mined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image fusion, and particularly to a general image fusion method and system based on low-rank and sparse priors. Background Art

[0002] Due to the differences in the imaging mechanisms of different sensors, the images obtained under the same scene often contain complementary information. For example, infrared sensors highlight significant targets by capturing thermal radiation, while visible light sensors generate images with rich texture details by reflecting light. In addition, positron emission tomography (PET) can reflect the metabolic activity of tissues, while magnetic resonance imaging (MRI) can provide high-resolution anatomical structure information. A single sensor can only capture part of the information within its dynamic range under different exposure settings. In addition, cameras with different focal lengths can clearly present objects within the depth of field. The goal of image fusion is precisely to integrate the complementary information in multi-source images and generate a fused image containing rich information. Current image fusion tasks mainly include infrared-visible image fusion (IVF), medical image fusion (MIF), multi-exposure image fusion (MEF), and multi-focus image fusion (MFF).

[0003] Although the specific tasks are different, the overall goal of image fusion is the same, that is, to generate a fused image with rich information. Therefore, many studies have begun to focus on general methods for unifying multi-source image fusion tasks. PMGI (Fast Unified Image Fusion Network Based on Gradient and Intensity Ratio Preservation) uses a dual-branch feature extraction network to capture modality-specific information and models multi-source image fusion as a task of simultaneously preserving pixel intensity and gradient details. This method achieves fusion by artificially setting the weight ratios of different tasks. TUFusion (Transformer Network) fuses the Transformer and CNN architectures to build a hybrid encoder, taking into account both the modeling of global representations and the preservation of local details, further enhancing the quality of the fused image. VDMUFusion (General Image Fusion Network Based on Diffusion Model) models image fusion as a pixel-by-pixel weighted average process and introduces a multi-task learning framework to replace the traditional noise prediction network, thereby extending the diffusion model to a general image fusion method.

[0004] Existing fusion methods using a single encoder structure can extract common features between modalities but may ignore modality-unique features; the dual-encoder design emphasizes modality uniqueness but ignores the shared representations of modalities. From the perspective of the internal structure of the network, existing methods usually stack Transformer and convolutional neural network modules, resulting in high network computational complexity and low fusion efficiency. There is a lack of in-depth exploration of the commonalities between different fusion tasks, and the model exhibits "black box" characteristics and lacks interpretability. Summary of the Invention

[0005] The present invention aims to solve at least one of the technical problems existing in the related art. To this end, the present invention provides a general image fusion method and system based on low-rank and sparse priors, extracts common features through a common feature encoder (CFE) based on low-rank decomposition, and extracts sparse unique features of two source images through a unique feature encoder (UFE) based on sparse representation. The sparse unique features of the two source images are fused through an entropy-based feature fusion module and fused with the common features of the two source images to obtain fused features, and the fused features are input into a feature reconstruction module to obtain a fused image, which not only realizes effective cross-modal interaction, but also fully excavates modal unique features; realizes feature decoupling and increases the interpretability of the network.

[0006] The present invention provides a general image fusion method based on low-rank and sparse priors, including: S1: Extract common features of two source images through a common feature encoder based on low-rank decomposition; S11: Extract shallow joint features of two source images through a joint feature extraction module; S12: Perform low-rank approximation on the shallow joint features by using non-negative matrix factorization to obtain a basis matrix and a coefficient matrix; S13: Multiply the basis matrix and the coefficient matrix to obtain common features of two source images; S2: Respectively extract sparse unique features of two source images through a unique feature encoder based on sparse representation; S21: Obtain rough unique features of two source images by subtracting the common features of two source images from the source images; S22: Construct a sparse decomposition model according to the rough unique features of two source images; S23: Solve the sparse decomposition model through a kernel transposed convolution block to obtain sparse unique features of two source images; S3: Fuse the sparse unique features of two source images through a serial adaptive fusion module and fuse them with the common features of two source images to obtain fused features; S4: Input the fused features into a feature reconstruction module to obtain a fused image.

[0007] Further, the joint feature extraction module includes convolution and depthwise separable convolution, and the depthwise separable convolution includes depth convolution and pointwise convolution.

[0008] Further, in step S12, non-negative matrix factorization approximates a non-negative feature matrix by minimizing a reconstruction error, and the calculation expression is: Among them, is the basis matrix, is the coefficient matrix, is the shallow combined feature, is the Frobenius norm of the matrix, is the minimum function.

[0009] Furthermore, in step S22, the calculation expression of the sparse decomposition model is: Among them, is the rough unique feature of the source image , is 's sparse coefficient, is 's dictionary, is the Manhattan norm, is the first learnable parameter, is the source image 's rough unique feature, is 's sparse coefficient, is 's dictionary, is the second learnable parameter, is the minimum function, is the Frobenius norm of the matrix.

[0010] Furthermore, in step S23, Adopt the deep unfolding technology to replace the matrix multiplication operation of the sparse decomposition model with a convolution operation; Unfold the iterative solution process of the sparse decomposition model into a neural network with K stages, and replace the dictionary with a convolutional layer and the transpose of the dictionary with a transposed convolutional layer; Express the solution process of each stage as a kernel transposed convolutional block, and the calculation expression is: Among them, is the th iteration 's sparse coefficient, is the first soft threshold operation, is the th iteration 's sparse coefficient, is the convolutional layer, is the transposed convolutional layer, is the The sub - iteration of the sparse coefficient, is the sub - iteration of the sparse coefficient, is the second soft - threshold operation, and is the convolution operation; Input the sparse coefficient of the K - th stage obtained through the kernel transposed convolution block into the convolution to obtain the sparse unique features of the two source images.

[0011] Furthermore, in step S3, the serial adaptive fusion module includes two entropy - based feature fusion modules, which successively implement the fusion of the sparse unique features of the two source images, and the further fusion of the fused sparse unique features and the common features, including: S31: Extract the details and brightness of the sparse unique features of the two source images through global average pooling and max - pooling, and then obtain the sparse unique fusion features of the two source images through convolution; S32: Calculate the fusion weights of the two source images by calculating the information entropy of the sparse unique fusion features of the two source images; S33: Fuse the sparse unique fusion features of the two source images according to the fusion weights of the two source images to obtain the fused unique features; S34: Fuse the fused unique features and the common features of the two source images to obtain the fused features.

[0012] Furthermore, the feature reconstruction module includes multiple reconstruction convolutional layers, and each of the reconstruction convolutional layers includes convolution and ReLU activation function.

[0013] The present invention also provides a general image fusion system based on low - rank and sparse priors for performing the above - mentioned general image fusion method based on low - rank and sparse priors, including: A low - rank decomposition unit, which extracts the common features of the two source images through a common feature encoder based on low - rank decomposition; A sparse representation unit, which extracts the sparse unique features of the two source images through a unique feature encoder based on sparse representation respectively; A fusion unit, which fuses the sparse unique features of the two source images through a serial adaptive fusion module and fuses them with the common features of the two source images to obtain the fused features; A reconstruction unit, which inputs the fused features into the feature reconstruction module to obtain the fused image.

[0014] One or more of the above technical solutions in the embodiments of the present invention have at least one of the following technical effects: The present invention extracts common features based on a low-rank decomposition-based common feature encoder, and extracts sparse unique features of two source images based on a sparse representation-based unique feature encoder. The sparse unique features of the two source images are fused through an entropy-based feature fusion module, and are fused with the common features of the two source images to obtain fused features. A fused image is obtained through a feature reconstruction module. The present invention divides image features into common features and sparse unique features, making the information flow inside the network clearer, greatly increasing the interpretability of the network, not only realizing effective cross-modal interaction, but also fully mining modal unique features.

[0015] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 is a schematic flow chart of a general image fusion method based on low-rank and sparse priors provided by the present invention.

[0018] Figure 2 is a schematic diagram of the network framework of a general image fusion method based on low-rank and sparse priors provided by the present invention.

[0019] Figure 3 is a schematic diagram of the structure of an entropy-based feature fusion module provided by the present invention.

[0020] Figure 4 is a schematic diagram of the structure of a general image fusion system based on low-rank and sparse priors provided by the present invention.

[0021] Figure 5 is a schematic diagram of a subjective comparison experiment on the fusion result of infrared and visible light images of the M3FD dataset in the embodiments of the present invention.

[0022] REFERENCE SIGNS: 101, low-rank decomposition unit; 102, sparse representation unit; 103, fusion unit; 104, reconstruction unit. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention. The following embodiments are used to illustrate the present invention but cannot be used to limit the scope of the present invention.

[0024] In the description of the embodiments of the present invention, it should be noted that the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. The description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0025] The following combines Figures 1 to 5 to describe a general image fusion method and system based on low-rank and sparse priors of the present invention.

[0026] In the image restoration task, an image can be decomposed into a low-rank matrix representing the background and a sparse matrix representing the abnormal region. In the image fusion task, the source images come from the same scene, so they have similar scene information but different modal features. Therefore, the present invention regards the common features of the source images as the same low-rank matrix and the modal difference features as the sparse matrix of the abnormal region. The present invention constructs a general image fusion network based on low-rank and sparse priors, named (A General Image Fusion Network Based on Low-Rank and Sparse Priors, abbreviated as LRSP-Fusion), to solve the image fusion problem by mining the internal relationship of the source image features. The network framework is as Figure 2 shown.

[0027] As Figure 1 shown, a general image fusion method based on low-rank and sparse priors includes: S1: Extract the common features of two source images through a common feature encoder based on low-rank decomposition; Obtain any two source images to be fused. The source images are infrared-visible light images, medical images, multi-exposure images, multi-focus images, etc. The source images are all from the same scene and have the same basic structure and information, which can be regarded as the common features of the source images and have the same low-rank representation. Their image matrices usually have the property of being low-rank or approximately low-rank. The low-rank matrix represents the main structure and information of the image. Extract the common features of the two source images through CFE, and the calculation expression is: where, is the common feature of the two source images, is the CFE operation, is the splicing operation, is the first source image, is the second source image.

[0028] S11: Extract the shallow joint features of the two source images through the joint feature extraction module (joint feature extraction, abbreviated as JFEB); JFEB includes convolution and depthwise separable convolution. The depthwise separable convolution includes depthwise convolution (DWConv) and pointwise convolution (PWConv); this process is expressed as: where, is the shallow joint feature, is convolution, is depthwise convolution, is pointwise convolution.

[0029] S12: Use non-negative matrix factorization (Non-negative Matrix Factorization, abbreviated as NMF) to perform low-rank approximation on the shallow joint features to obtain the basis matrix and the coefficient matrix; Non-negative matrix factorization is a low-rank approximation model that can extract the latent features of data, and the image matrix itself is also a non-negative matrix. Therefore, the present invention uses NMF technology to perform low-rank decomposition on the image matrix and reconstruct the common features using the obtained basis matrix and coefficient matrix; Non-negative matrix factorization approximates the non-negative feature matrix by minimizing the reconstruction error, and the calculation expression is: where, is the basis matrix, , is of dimension The matrix set, is the column dimension of, is the height, is the width, is the coefficient matrix, , is a matrix set with dimension , is the number of channels, The column dimension of is the same as the row dimension of, and , is the shallow joint feature, is the Frobenius norm of the matrix, is the minimum function, with the constraint that and All elements in are greater than or equal to 0. By iteratively updating and to minimize the loss function, the update formula is as follows: where, is the basis matrix of the th iteration, is the basis matrix of the th iteration, is the Hadamard product, is the transpose of the matrix, is the coefficient matrix of the th iteration, is the coefficient matrix of the th iteration, S13: Multiply the basis matrix and the coefficient matrix to obtain the common feature of the two source images.

[0030] S2: Respectively extract the sparse unique features of the two source images through a unique feature encoder based on sparse representation; The common feature is compressed through a convolution to obtain a common feature map, and then the feature map is subtracted element-wise from the input source image. The calculation expression is: where, is the sparse unique feature of, is the UFE operation, is the convolution operation, is the common feature of the two source images, For sparse unique features.

[0031] S21: By subtracting the common features of two source images from the source images, the rough unique features of the two source images are obtained, and the calculation expression is: where is the rough unique feature of is the rough unique feature of is the common feature of the two source images, is the convolution of, and the common feature is compressed in the channel dimension by using convolution to achieve the integration of the common feature.

[0032] S22: Construct a sparse decomposition model according to the rough unique features of the two source images; The common feature is the main structure and information of the image, while the unique feature represents the modal information of the source image and has sparsity. UFE is constructed based on the sparse representation of the unique feature and is used to solve the following sparse decomposition model: where is the rough unique feature of the source image is the sparse coefficient of is the dictionary of is the Manhattan norm, is the first learnable parameter, is the rough unique feature of the source image is the sparse coefficient of is the dictionary of is the second learnable parameter, is the minimum function, is the Frobenius norm of the matrix.

[0033] The traditional method can be solved by the proximal gradient descent method, and its iterative solution formula is as follows: where For the sparse coefficient of the iteration, is the first soft threshold operation, For the sparse coefficient of the iteration, For the sparse coefficient of the iteration, For the sparse coefficient of the iteration, is the second soft threshold operation, , is the sign function with respect to , is the maximum function, is the absolute value function, is the learnable parameter, can take or .

[0034] S23: Solve the sparse decomposition model through the kernel transposed convolution block to obtain the sparse unique features of the two source images.

[0035] The present invention adopts the deep unfolding technology, replaces the matrix multiplication operation of the sparse decomposition model with a convolution operation; unfolds the iterative solution process of the sparse decomposition model into a neural network including K stages, replaces the dictionary with a convolutional layer, and replaces the transpose of the dictionary with a transposed convolutional layer; As Figure 2 shown, each iteration ( ) corresponds to a step in the solution process. The solution process of each stage is represented as a kernel transposed convolution block (Kernel Transposed Convolution Block, KTCB), and the calculation expression is: Among them, is the sparse coefficient of the iteration , is the first soft threshold operation, is the sparse coefficient of the iteration , is the convolutional layer, is the transposed convolutional layer, is the sparse coefficient of the iteration , For the iteration of the sparse coefficient, is the second soft threshold operation, is the convolution operation; The 0th stage is used to initialize the initial sparse coefficient and the initial sparse coefficient , and the calculation expression is: Among them, the initial sparse coefficient includes and , and the sparse coefficient of the first stage includes the sparse coefficient of the first stage of and the sparse coefficient of the first stage of .

[0036] Input the sparse coefficient of the Kth stage obtained through the kernel transposed convolution block into convolution to obtain the sparse unique features of the two source images, and the calculation expression is: Among them, is the sparse coefficient of the iteration , is the sparse unique feature of, is the iteration of the sparse coefficient, is the sparse unique feature of.

[0037] S3: Fuse the sparse unique features of the two source images through the sequential adaptive fusion module (SAFM), and fuse them with the common features of the two source images to obtain the fused features; As Figure 3 shown, SAFM includes two entropy-based feature fusion blocks (EFFB), which sequentially implement the fusion of the sparse unique features of the two source images, and the further fusion of the fused sparse unique features and the common features. The calculation expression is: Among them, is the fusion feature, is the EFFB operation; The EFFB structure is as Figure 3 shown.

[0038] S31: Extract the details and brightness of the sparse and unique features of the two source images through global average pooling and max pooling, and then through convolution to obtain the sparse and unique fusion features of the two source images. The calculation expression is: Among them, is the sparse and unique fusion feature of, is global average pooling, is the max pooling operation, is the sparse and unique fusion feature of.

[0039] S32: Calculate the fusion weights of the two source images by calculating the information entropy of the sparse and unique fusion features of the two source images; The entropy of an image reflects the average amount of information in the image. The higher the entropy value of the feature map, the richer the features it contains. Therefore, the present invention uses the calculation of entropy to evaluate the amount of information in the feature map, so as to obtain the adaptive weights of the two sets of features. The calculation expression is: Among them, is the weight of, is the weight of, is the information entropy calculation function.

[0040] S33: Fuse the sparse and unique fusion features of the two source images according to the fusion weights of the two source images to obtain the fusion unique features; Among them, is the fusion unique feature.

[0041] S34: Fuse the fusion unique features and the common features of the two source images to obtain the fusion features.

[0042] S4: Input the fusion features into the feature reconstruction module (feature reconstruction module, abbreviated as FRM) to obtain the fusion image.

[0043] The feature reconstruction module includes multiple reconstruction convolutional layers, and each reconstruction convolutional layer includes a convolution and a ReLU activation function. The calculation expression is: where is the fused image, is the FRM operation, is the reconstruction convolutional layer.

[0044] The present invention deeply explores the common problems existing in multi-source image fusion, the low-rank background information and the sparse unique information. On this basis, it unifies the modeling of the multi-source image fusion task from a mathematical perspective, designs an efficient and lightweight fusion network, which not only realizes effective cross-modal interaction, but also fully excavates the unique features of the modality; realizes feature decoupling and increases the interpretability of the network.

[0045] As Figure 4 shown, a general image fusion system based on low-rank and sparse priors is used to execute a general image fusion method based on low-rank and sparse priors, including: The low-rank decomposition unit 101 extracts the common features of two source images through a common feature encoder based on low-rank decomposition; The sparse representation unit 102 extracts the sparse unique features of two source images through a unique feature encoder based on sparse representation respectively; The fusion unit 103 fuses the sparse unique features of two source images through a serial adaptive fusion module, and fuses them with the common features of two source images to obtain fused features; The reconstruction unit 104 inputs the fused features into the feature reconstruction module to obtain a fused image.

[0046] Through the collaborative work of the above units, the common feature encoder based on low-rank decomposition extracts common features, while the unique feature encoder based on sparse representation extracts the sparse unique features of two source images. The sparse unique features of two source images are fused through an entropy-based feature fusion module, and fused with the common features of two source images to obtain fused features. The fused image is obtained through the feature reconstruction module. The present invention divides the image features into common features and sparse unique features, making the information flow inside the network clearer, greatly increasing the interpretability of the network, not only realizing effective cross-modal interaction, but also fully excavating the unique features of the modality.

[0047] The present invention has conducted comparative experiments on quantitative index results and qualitative experimental analyses with a variety of internationally state-of-the-art algorithms on the public dataset M3FD. The internationally advanced algorithms include: MMDRFuse, TUFusion, VDMUFusion, and GIFNet. Among them, MMDRFuse is the 2024 version of the dynamically refreshed distilled micro model, TUFusion is the 2024 version of the transformer network, VDMUFusion is the 2024 version of the general image fusion network based on the diffusion model, and GIFNet is the 2025 version of the three-branch general image fusion network.

[0048] The comparative experiment on the fusion quantitative index results of the M3FD dataset is shown in Table 1.

[0049] Table 1 Comparative experiment on the fusion quantitative index results of the M3FD dataset The higher all the indicators in Table 1 are, the better the fusion effect is. As shown in Table 1, the present invention performs excellently in all evaluation indicators. The qualitative results are as Figure 5 shown. The present invention can effectively retain the texture details of visible light images while extracting significant targets in infrared images.

[0050] To further verify the effectiveness of the present invention in practical applications, the fused infrared-visible light images of the present invention are applied to the target detection task. Specifically, the fused images are used as the input of the target detector, and YOLOv5 is used as the detection model to evaluate the effect of the fused images in improving the detection accuracy. The comparative experiment on the quantitative index results of pedestrian and vehicle detection in the M3FD dataset is shown in Table 2.

[0051] Table 2 Comparative experiment on the quantitative index results of pedestrian and vehicle detection in the M3FD dataset The average precision calculated under the condition that the IoU threshold is 0.5 is used as the evaluation index. As shown in Table 2, the present invention has achieved higher detection accuracy rates and better target recognition effects in both pedestrian and vehicle detection tasks. The quantitative indicators are shown in Table 2 and are comprehensively superior to the comparative methods. The experimental results show that the present invention not only improves the quality of image fusion but also can effectively promote the performance improvement of downstream tasks, and has strong practical value.

[0052] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A general image fusion method based on low-rank and sparse priors, characterized in that Including: S1: Extract the common features of two source images through a common feature encoder based on low-rank decomposition; S11: Extract the shallow joint features of two source images through a joint feature extraction module; S12: Perform low-rank approximation on the shallow joint features using non-negative matrix factorization to obtain a basis matrix and a coefficient matrix; S13: Multiply the basis matrix and the coefficient matrix to obtain the common features of two source images; S2: Extract the sparse unique features of two source images respectively through a unique feature encoder based on sparse representation; S21: Obtain the rough unique features of two source images by subtracting the common features of two source images from the source images; S22: Construct a sparse decomposition model based on the rough unique features of two source images; S23: Solve the sparse decomposition model through a kernel transposed convolutional block to obtain the sparse unique features of two source images; S3: Fuse the sparse unique features of two source images through a serial adaptive fusion module, and fuse them with the common features of two source images to obtain fused features; S4: Input the fused features into a feature reconstruction module to obtain a fused image.

2. A general image fusion method based on low-rank and sparse priors according to claim 1, characterized in that, The combined feature extraction module includes convolution and depthwise separable convolution, and the depthwise separable convolution includes depthwise convolution and pointwise convolution.

3. A general image fusion method based on low-rank and sparse priors according to claim 1, characterized in that In step S12, non-negative matrix factorization approximates a non-negative feature matrix by minimizing the reconstruction error, and the calculation expression is: Among them, is the base matrix, is the coefficient matrix, is the shallow combined feature, is the Frobenius norm of the matrix, is the minimum value function.

4. A general image fusion method based on low-rank and sparse priors according to claim 1, characterized in that In step S22, the calculation expression of the sparse decomposition model is: Among them, is the rough unique feature of the source image , is 's sparse coefficient, is 's dictionary, is the Manhattan norm, is the first learnable parameter, is the rough unique feature of the source image , is 's sparse coefficient, is 's dictionary, is the second learnable parameter, is the minimum value function, is the Frobenius norm of the matrix.

5. A general image fusion method based on low-rank and sparse priors according to claim 4, characterized in that In step S23, Adopt the deep unfolding technology to replace the matrix multiplication operation of the sparse decomposition model with a convolutional operation; Unfold the iterative solution process of the sparse decomposition model into a neural network including K stages, replace the dictionary with a convolutional layer and the transpose of the dictionary with a transposed convolutional layer; Express the solution process of each stage as a kernel transposed convolutional block, and the calculation expression is: Among them, is the sparse coefficient of the n-th iteration, is the first soft threshold operation, is the sparse coefficient of the n-th iteration, is the convolutional layer, is the transposed convolutional layer, is the sparse coefficient of the n-th iteration, is the sparse coefficient of the n-th iteration, is the second soft threshold operation, is the convolutional operation; Input the sparse coefficients of the K-th stage obtained through the kernel transposed convolutional block into the convolution to obtain the sparse unique features of the two source images.

6. A general image fusion method based on low-rank and sparse priors according to claim 1, characterized in that In step S3, the serial adaptive fusion module includes two entropy-based feature fusion modules, which sequentially implement the fusion of the sparse unique features of two source images, and the further fusion of the fused sparse unique features and the common features, including: S31: Extract the details and brightness of the sparse unique features of the two source images through global average pooling and max pooling, and then obtain the sparse unique fusion features of the two source images through convolution; S32: Calculate the fusion weights of two source images by calculating the information entropy of the sparse unique fusion features of two source images; S33: Fuse the sparse unique fusion features of two source images according to the fusion weights of two source images to obtain fused unique features; S34: Fuse the fused unique features with the common features of two source images to obtain fused features.

7. A general image fusion method based on low-rank and sparse priors according to claim 1, characterized in that The feature reconstruction module includes a plurality of reconstruction convolutional layers, and each of the reconstruction convolutional layers includes convolution and a ReLU activation function.

8. A general image fusion system based on low-rank and sparse priors, characterized in that, Used to execute a general image fusion method according to any one of claims 1 to 7, including: A low-rank decomposition unit, which extracts the common features of two source images through a common feature encoder based on low-rank decomposition; A sparse representation unit, which extracts the sparse unique features of two source images respectively through a unique feature encoder based on sparse representation; A fusion unit, which fuses the sparse unique features of two source images through a serial adaptive fusion module, and fuses them with the common features of two source images to obtain fused features; A reconstruction unit, which inputs the fused features into a feature reconstruction module to obtain a fused image.

Citation Information

Patent Citations

  • Multi-source damaged image fusion and recovery joint implementation method based on dictionary learning

    CN112561842A

  • Infrared and visible light image fusion method and system based on low-rank sparse representation

    CN114648475A

  • Real-time instance segmentation method fusing sparse framework and space attention

    CN115100410A

  • Image fusion method and system based on multi-scale transformation and sparse low-rank representation

    CN119991469A

  • Adaptive dynamic magnetic resonance fast imaging method and device based on partially separable function

    WO2024092387A1

Cited By

  • Multi-modal data fusion method and device in enterprise knowledge management, equipment and medium

    CN121030667A

  • A cross-data-domain knee cartilage MRI segmentation method based on sample-level domain routing

    CN122473175A