A deep multi-modal image fusion method and system
By employing a convolutional transformation learning method combined with a deep multimodal image fusion model, the problems of low resolution and poor quality in multimodal image fusion are solved, achieving efficient image fusion and accurate medical diagnostic support.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUILIN UNIV OF ELECTRONIC TECH
- Filing Date
- 2025-01-26
- Publication Date
- 2026-04-24
AI Technical Summary
Existing sparse representation models struggle to effectively fuse multimodal and multiscale image data, especially overexposed and underexposed images, resulting in low image resolution and poor quality. Current technologies also suffer from low storage and computational efficiency when processing large-scale datasets.
The convolutional transform learning method is adopted. By constructing a deep multimodal image fusion model, the convolutional neural network and transform learning techniques are used. Combined with multi-scale transform learning and feature update sub-modules, image features are dynamically learned, sparse features are extracted and fused multiple times to optimize image representation.
It achieves efficient fusion of multimodal images, improves image resolution and quality, provides more comprehensive and accurate lesion information, and supports more precise medical image diagnosis and analysis.
Smart Images

Figure CN120047784B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, specifically to a deep multimodal image fusion method and system. Background Technology
[0002] In today's information age, we face the challenge of acquiring vast amounts of image data from various imaging technologies. This image data typically originates from different imaging modalities, each providing unique information about the imaged object. The goal of multimodal image processing is to integrate this modal image data from different imaging technologies to obtain a more comprehensive and accurate representation of the imaged object, thereby revealing information that a single modality cannot provide. However, image data from different modalities often possess different characteristics and scales, making direct image fusion extremely complex.
[0003] In machine learning, sparse representation models have attracted widespread attention due to their advantages in improving image interpretability, compression, and feature extraction, and can be applied to image fusion techniques. Sparse representation models are mainly divided into traditional dictionary learning and transform learning. Although dictionary learning has achieved success in generating sparse representations, it requires storing and processing a large number of overlapping image patches when dealing with large-scale datasets. Transform learning, as an emerging sparse representation method, directly analyzes image data by learning and analyzing transforms to obtain sparse features, but it still faces challenges when dealing with multimodal and multiscale data. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a deep multimodal image fusion method and system to address the shortcomings of the prior art.
[0005] The technical solution of the present invention to solve the above-mentioned technical problems is as follows:
[0006] A deep multimodal image fusion method includes the following steps:
[0007] S1. Import a set of multiple exposure source images, which includes overexposed images, underexposed images, and standard images;
[0008] S2. Extract features from the overexposed image to obtain overexposed sparse features and overexposed dictionary features; extract features from the underexposed image to obtain underexposed sparse features and underexposed dictionary features.
[0009] S3. The overexposure sparse features and the overexposure dictionary features are updated and calculated using the overall objective function to obtain new overexposure sparse features and new overexposure dictionary features. The underexposure sparse features and the underexposure dictionary features are updated and calculated using the overall objective function to obtain new underexposure sparse features and new underexposure dictionary features.
[0010] S4. Calculate the common features of the new overexposure sparse features and the new underexposure sparse features using a nonlinear function to obtain the overexposure common sparse features and the underexposure common sparse features. Calculate the common features of the new overexposure dictionary features and the new underexposure dictionary features using a nonlinear function to obtain the overexposure common dictionary features and the underexposure common dictionary features.
[0011] S5. Extract features from the overexposed common sparse features and the overexposed common dictionary features to obtain an overexposed fused image; extract features from the underexposed common sparse features and the underexposed common dictionary features to obtain an underexposed fused image.
[0012] S6. Repeat S2-S5 until the preset number of iterations is reached. Perform fusion processing on the overexposed fused image and the underexposed fused image to obtain a reconstructed image. Evaluate and calculate the reconstructed image and the standard image to obtain evaluation indicators.
[0013] Another technical solution of the present invention to solve the above-mentioned technical problems is as follows:
[0014] A deep multimodal image fusion system, comprising:
[0015] An image import unit is used to import a set of images from multiple exposure sources, the set of images from multiple exposure sources including overexposed images, underexposed images, and standard images;
[0016] The feature pre-extraction unit is used to extract features from the overexposed image to obtain overexposed sparse features and overexposed dictionary features, and to extract features from the underexposed image to obtain underexposed sparse features and underexposed dictionary features.
[0017] The feature update unit is used to update the overexposed sparse features and the overexposed dictionary features through the overall objective function to obtain new overexposed sparse features and new overexposed dictionary features, and to update the underexposed sparse features and underexposed dictionary features through the overall objective function to obtain new underexposed sparse features and new underexposed dictionary features.
[0018] The feature extraction unit is used to perform common feature calculation on the new overexposure sparse features and the new underexposure sparse features using a nonlinear function to obtain overexposure common sparse features and underexposure common sparse features, and to perform common feature calculation on the new overexposure dictionary features and the new underexposure dictionary features using a nonlinear function to obtain overexposure common dictionary features and underexposure common dictionary features.
[0019] The feature fusion unit is used to extract features from the overexposed common sparse features and the overexposed common dictionary features to obtain an overexposed fused image, and to extract features from the underexposed common sparse features and the underexposed common dictionary features to obtain an underexposed fused image.
[0020] The image reconstruction unit is used to perform fusion processing on the overexposed fused image and the underexposed fused image to obtain a reconstructed image, and to evaluate and calculate the reconstructed image and the standard image to obtain evaluation indicators.
[0021] The beneficial effects of this invention are as follows: Transform learning is achieved by processing updated features in each iteration based on the overall objective function. By learning and analyzing the transformation, sparse features are directly extracted from image data, effectively preserving the overall structure of the image. Transform learning captures important intrinsic attributes of the image, removes redundant information, and extracts higher-quality sparse features. Updated sparse features and dictionary features are dynamically obtained using a data-driven approach, and common features of various modal images are extracted, making it better suited for image reconstruction and realizing the fusion of multimodal images to provide more accurate support when using image diagnosis and analysis. Attached Figure Description
[0022] Figure 1 This is a general structural diagram of the deep multimodal image fusion model provided in an embodiment of the present invention;
[0023] Figure 2 A flowchart of the deep multimodal image fusion method provided in an embodiment of the present invention;
[0024] Figure 3 This is a structural diagram of the modal multi-scale transformation learning module provided in an embodiment of the present invention;
[0025] Figure 4 This is a structural diagram of the sparse feature update subunit provided in an embodiment of the present invention;
[0026] Figure 5 A structural diagram of the spatial attention module provided in an embodiment of the present invention;
[0027] Figure 6 This is a structural diagram of the dictionary feature update subunit provided in an embodiment of the present invention;
[0028] Figure 7 This is a structural diagram of the fusion module for overexposed image processing provided in an embodiment of the present invention;
[0029] Figure 8 This is a structural diagram of the fusion module for underexposed image processing provided in an embodiment of the present invention;
[0030] Figure 9This is a block diagram of a deep multimodal image fusion system provided in an embodiment of the present invention. Detailed Implementation
[0031] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.
[0032] In sparse representation models, unlike the synthetic framework of dictionary learning, transform learning is analytical. Its key advantage lies in its ability to avoid trivial solutions through analytical transform regularization, which helps to find more meaningful image representations. Furthermore, transform learning exhibits better stability and robustness when processing images because all analytical transforms contribute equally to the analysis of the image.
[0033] Convolutional Dictionary Learning (CDL) combines the ideas of convolutional neural networks with traditional dictionary learning techniques. The core idea of CDL is to learn a set of convolutional kernels that can extract local and global features from the input data, thus forming a sparse representation of the data. In traditional dictionary learning, images are represented as a linear combination of a set of basis vectors, but this is not efficient enough when processing image data with strong spatial correlations. CDL improves computational efficiency by introducing convolution operations to reduce this redundancy, because image data often has strong spatial correlations, and convolution operations can capture the local features of the data during image processing.
[0034] Convolutional Transform Learning (CTL) is an emerging method that combines Convolutional Neural Networks (CNNs) and Transform Learning. Its goal is to make images sparser in the convolutional domain, meaning that most elements in the feature map after convolution are zero, with only a few non-zero elements. This utilizes convolution operations to extract local features from images, making image processing more efficient and natural. Therefore, this invention investigates a technique for fusing multimodal images based on convolutional transform learning.
[0035] This invention, implemented in the field of exposure image processing, represents a significant advancement in modern medical imaging and diagnostics. It enables the provision of more comprehensive and accurate lesion information through multimodal image fusion strategies, thereby supporting more precise diagnosis and analysis. In multi-exposure image fusion technology, the image sequence to be fused often contains large areas of white or black due to overexposure or underexposure. These areas are typically characterized by low contrast, flat colors, and a lack of detail. By fusing overexposed and underexposed images, a single image with a wider dynamic range and richer detail can be obtained.
[0036] like Figure 1 and Figure 2As shown in the figure, a deep multimodal image fusion method provided by an embodiment of the present invention includes the following steps:
[0037] S0. A deep multimodal image fusion model is constructed based on a convolutional neural network. The deep multimodal image fusion model includes multiple cascaded multimodal image fusion modules (MIFB). Each multimodal image fusion module includes a first modal multiscale transformation learning module (MTLBx) and a second modal multiscale transformation learning module (MTLBy) connected in parallel with the fusion module (FB). The first modal multiscale transformation learning module (MTLBx) includes a first convolutional submodule and a first feature update submodule. The second modal multiscale transformation learning module (MTLBy) includes a second convolutional submodule and a second feature update submodule. The fusion module (FB) includes a common feature extraction submodule and a feature fusion submodule. The first modal multiscale transformation learning module (MTLBx) and the second modal multiscale transformation learning module (MTLBy) have the same structure and perform the same image processing operations.
[0038] S1. Import a set of multiple exposure source images, which includes overexposed images, underexposed images, and standard images;
[0039] S2. The overexposed image is subjected to feature extraction through the first convolution submodule to obtain overexposed sparse features and overexposed dictionary features. The underexposed image is subjected to feature extraction through the second convolution submodule to obtain underexposed sparse features and underexposed dictionary features.
[0040] S3. The overexposed sparse features and the overexposed dictionary features are updated and calculated through the first feature update submodule (i.e., the total objective function) to obtain new overexposed sparse features and new overexposed dictionary features. The underexposed sparse features and the underexposed dictionary features are updated and calculated through the second feature update submodule (i.e., the total objective function) to obtain new underexposed sparse features and new underexposed dictionary features.
[0041] S4. The common feature extraction submodule (i.e., nonlinear function) is used to calculate the common features of the new overexposure sparse features and the new underexposure sparse features to obtain the overexposure common sparse features and the underexposure common sparse features. The common feature extraction submodule (i.e., nonlinear function) is used to calculate the common features of the new overexposure dictionary features and the new underexposure dictionary features to obtain the overexposure common dictionary features and the underexposure common dictionary features.
[0042] S5. The overexposed common sparse features and the overexposed common dictionary features are extracted by the feature fusion submodule to obtain the overexposed fused image. The underexposed common sparse features and the underexposed common dictionary features are extracted by the feature fusion submodule to obtain the underexposed fused image.
[0043] S6. Repeat S2-S5 until the preset number of iterations is reached (i.e., multiple cascaded multimodal image fusion modules MIFB process the image). The overexposed fused image and the underexposed fused image are fused to obtain a reconstructed image. The reconstructed image and the standard image are evaluated and calculated to obtain evaluation indicators.
[0044] The set of multiple exposure source images is a pair of multiple exposure source images, which includes overexposed images, underexposed images and standard images. The dataset selected in this embodiment has a total of 440 pairs, of which 340 pairs are training sets and 100 pairs are test sets.
[0045] Specifically, feature extraction for overexposed images in S2 includes decomposing the overexposed image into multiple scales of overexposed images, and extracting the corresponding overexposure sparse features and overexposure dictionary features for each scale of overexposed image. The processing of underexposed images is the same as for overexposed images, extracting the corresponding underexposure sparse features and underexposure dictionary features for each scale of underexposed image. S3-S5 involve corresponding processing of the overexposure sparse features and overexposure dictionary features at multiple scales, as well as the underexposure sparse features and underexposure dictionary features at multiple scales.
[0046] In this embodiment of the invention, multi-scale variation learning is performed on each modality of the input image. By learning dictionaries at multiple scales, feature information of multiple scales of the modal image is extracted, and then multi-scale fusion is performed. Through multiple iterations and feature processing of multiple scales of the image each time, the image resolution is improved from coarse to fine. This gradually optimizes the feature representation of overexposed and underexposed images. Finally, the multi-scale features of each modality are fused, and the fused features of the two modalities are combined into a reconstructed image to overcome the quality problems such as low resolution of a single modal image, resulting in a clearer, higher-quality reconstructed image. The fusion effect is evaluated using evaluation metrics.
[0047] Compared with the fixed constraints (i.e. the overall objective function) of traditional dictionary learning, this invention uses a dynamic learning network to update feature learning constraints in order to find the most suitable prior knowledge. The feature update submodule can adapt to different tasks and input images, thereby improving the accuracy and generalization ability of the model.
[0048] Preferably, the step of extracting features from the overexposed image through the first convolutional submodule to obtain overexposed sparse features and overexposed dictionary features includes:
[0049] The pre-constructed convolution dictionary is initialized, and the initialized convolution dictionary is used as the overexposure dictionary feature. The overexposure image is then convolved using the overexposure dictionary feature to obtain overexposure sparse features.
[0050] Specifically, the first convolutional submodule includes a convolutional layer composed of multiple convolutional kernels; it decomposes the overexposed image into overexposed images at multiple scales, initializes the pre-constructed convolutional dictionary (i.e., multiple convolutional kernels) and sparse coefficient matrix (i.e., adjusts the parameters of the convolutional dictionary to the set dictionary parameters and adjusts the parameters of the sparse coefficient matrix to the set sparse parameters), obtaining the initialized convolutional dictionary and initialized sparse coefficient matrix; it matches the initialized convolutional dictionary with the overexposed images at multiple scales to obtain overexposed dictionary features at multiple scales (i.e., multiple initialized convolutional kernels); it then performs convolution processing on the overexposed images at the corresponding scales using the multiple initialized convolutional kernels to obtain overexposed sparse features at multiple scales, which together form the overexposed sparse features. The overexposed image is represented as:
[0051]
[0052] Where LX is the overexposed image, S is the number of image scales, and x s Let E be the overexposed image at the s-th scale, and E be the number of convolutional dictionaries (i.e., the number of convolutional kernels) for the overexposed image. Let e be the convolutional dictionary (i.e., overexposed dictionary features) at the s-th scale. For from x s pass Extracted sparse features (i.e. overexposure sparse features).
[0053] The process of feature extraction for underexposed images in the second convolutional submodule is the same as that in the first convolutional submodule, and will not be repeated here. The underexposed image is represented as follows:
[0054]
[0055] Where HY represents an underexposed image, and y s Let be the underexposed image at the s-th scale, and F be the number of convolution dictionaries for the underexposed image. Let f be the convolutional dictionary (i.e., underexposed dictionary features) at the s-th scale. For from y s pass The extracted sparse features (i.e., underexposed sparse features) are represented by *, which is a convolution operation.
[0056] It should be understood that the convolution dictionary consists of a set of convolution kernels, each used to extract local features in the image; the initialization process can be random initialization.
[0057] In this embodiment of the invention, dictionary features are used to capture local features of an image to generate sparse features. Sparse features can remove redundant information in an image, which can significantly reduce the amount of data storage and transmission while ensuring image quality.
[0058] Preferably, before the step of updating the overexposed sparse features and the overexposed dictionary features through the first feature update submodule (i.e., the overall objective function), the method further includes:
[0059] Based on the overexposure sparse features and the overexposure dictionary features, the objective function of traditional dictionary learning is improved to obtain the overall objective function, which is:
[0060]
[0061] Where L is the overall objective function, S is the number of image scales, and E is the number of convolution dictionaries. For overexposure sparse features, For overexposure dictionary features, x is the overexposure image. Let λ be the Euclidean norm, λ be the regularization parameter, G(·) be the regularization constraint function for sparse features, β be the regularization parameter, H(·) be the regularization constraint function for dictionary features, and * be the convolution operation.
[0062] Specifically, the overall objective function of traditional dictionary learning is:
[0063]
[0064] in, To reconstruct the error term, For sparse constraint regularization, For dictionary-constrained regularization terms;
[0065] The reconstruction error term has been modified to: The remaining items remain unchanged;
[0066] The overexposed image x will be measured against overexposed sparse features. and overexposure dictionary features The reconstruction error term, which measures the difference between linear combinations, is modified to measure overexposure sparsity features. Overexposure dictionary features The reconstruction error term is the difference between the linear combination of the overexposed image x and the image x.
[0067] The improvement terms of the second feature update submodule (i.e., the overall objective function) are the same as those of the first feature update submodule, and will not be repeated here. The overall objective function is expressed as follows:
[0068]
[0069] In this embodiment of the invention, the improved overall objective function focuses more on the reconstruction error of sparse features, which is beneficial for updating and adjusting sparse features according to the error, so that sparse features can better represent the important information of the image, remove redundant information, keep the sparse constraint regularization term and dictionary constraint regularization term unchanged, and apply regularization constraints to overexposed sparse features and overexposed dictionary features respectively to prevent overfitting.
[0070] Preferably, such as Figure 3 As shown, the overexposure sparse features and the overexposure dictionary features are updated and calculated through multiple first feature update submodules, including:
[0071] For the first input feature x 1 The first output feature is obtained by updating the overexposure sparse features and the overexposure dictionary features at the first scale. (i.e., the first scale update of overexposure sparsity features) And the updated overexposure dictionary features at the first scale ), the first output feature and the second input feature x 2 (That is, the overexposure sparse features at the second scale and the overexposure dictionary features at the second scale) are concatenated and combined, and the combined second input feature is updated to obtain the second output feature. (i.e., the updated overexposure sparsity features at the second scale) and the updated overexposure dictionary features at the second scale ), and so on, for the (s-1)th output feature and the s-th input feature x s By concatenating and combining the input features, the s-th input feature is updated to obtain the s-th output feature. (i.e., the sparse features of overexposure) and new overexposure dictionary features ).
[0072] Preferably, the step of updating the overexposed sparse features and the overexposed dictionary features through the first feature update submodule (i.e., the overall objective function) to obtain new overexposed sparse features and new overexposed dictionary features includes:
[0073] The overall objective function is decomposed into a sparse objective function and a dictionary objective function by using the alternating direction multiplier method;
[0074] New overexposure sparse features are obtained by calculating the overexposure sparse features using the sparse objective function. The sparse objective function is:
[0075]
[0076] Where L1 is the sparse objective function, S is the number of image scales, and E is the number of convolution dictionaries. This refers to overexposure sparse features (i.e., new overexposure sparse features obtained after learning across all scale variations). The * represents the overexposure dictionary features (i.e., the new overexposure dictionary features obtained after learning all scale changes), and * represents the convolution operation. Let λ be the Euclidean norm, λ be the regularization parameter, G(·) be the regularization constraint function for sparse features, and α be the regularization parameter. α For regularization parameters, For overexposure sparse auxiliary variables, x s Let i be the overexposed image that has not undergone dictionary learning (i.e., the overexposed image that remains after dictionary learning at all previous scales), and let i be the i-th iteration.
[0077] The overexposure dictionary features are calculated using the dictionary objective function to obtain new overexposure dictionary features. The dictionary objective function is:
[0078]
[0079] Where L2 is the dictionary objective function, β is the regularization parameter, and α is the regularization parameter. d For regularization parameters, H is an auxiliary variable for the overexposure dictionary, and H(·) is the regularization constraint function for the dictionary features.
[0080] Specifically, the objective function is decomposed into two sub-objectives (i.e., a sparse objective function and a dictionary objective function) using the alternating direction multiplier method. This means the first feature update submodule includes a sparse feature update subunit and a dictionary feature update subunit. The numerical solution to the sparse constraints is solved by the sparse feature update subunit (SRU), such as... Figure 4 and Figure 5 As shown, this sub-unit processes input features through a combination of residual blocks, convolutional layers (Conv), and a spatial attention module, enhancing the expressive power of the features, including:
[0081] Input pre-constructed overexposure sparse auxiliary variables Regularization parameters The input layer contains overexposure sparse features and overexposure dictionary features at multiple scales. The input layer is concatenated with the first residual block set and the first convolutional layer. The output of the first convolutional layer is concatenated with the second residual block set and the second convolutional layer. The output of the second convolutional layer is input to the third residual block set. The outputs of the second and third convolutional layers are element-wise summed and concatenated with the third and fourth convolutional layers. The outputs of the first and fourth convolutional layers are element-wise summed and concatenated with the fourth and fifth convolutional layers. The output of the fifth convolutional layer is element-wise summed with the overexposure sparse auxiliary variable and connected to the spatial attention module and the fifth convolutional layer. The output of the fifth convolutional layer is connected to the output layer, and the output layer outputs new overexposure sparse features.
[0082] Each residual block set includes multiple cascaded residual blocks. The input and output of any residual block are added element-wise and then input into the next residual block. Each residual block includes a cascaded sixth convolutional layer, a first activation function layer (ReLU), a first regularization layer (BN), and a seventh convolutional layer. The spatial attention module includes a parallel-connected max pooling layer (Maxpooling), a pointwise convolutional layer (PWConv), and a first depthwise separable convolutional layer (DSConv). The outputs of the max pooling layer, the pointwise convolutional layer (PWConv), and the depthwise separable convolutional layer (DSConv) are connected to the concatenation layer. The concatenation layer is cascaded with the second depthwise separable convolutional layer and the second activation function layer.
[0083] The numerical solution to the dictionary constraints is solved by the dictionary filter updating (DFU) unit, such as... Figure 6 As shown, this sub-unit processes the input dictionary through a series of convolutional layers (Conv), activation function (ReLU), and regularization layer (BN), and reconstructs the dictionary using a reconstruction layer, enabling the dictionary to adapt to the characteristics of the input data, including:
[0084] Import pre-built overexposure dictionary auxiliary variables Regularization parameters The input layer contains overexposure sparse features at multiple scales and overexposure dictionary features at multiple scales. The input layer is then concatenated with multiple convolutional blocks, with the last convolutional block connected to the reconstruction layer. The output of the reconstruction layer is element-wise added to the overexposure dictionary auxiliary variables and then connected to the output layer. The output layer outputs new overexposure dictionary features.
[0085] Each convolutional block includes a cascaded eighth convolutional layer, a third activation function, and a second regularization layer.
[0086] Furthermore, according to the preset number of iterations, the sparse objective function and the dictionary objective function are solved using the nearest neighbor operator to obtain the overexposed sparse auxiliary variable. Newly exposed sparse features Overexposure dictionary auxiliary variables and new overexposure dictionary features Represented as:
[0087]
[0088] Where i is the iteration number. and It is updated in the i-th iteration. and and It is updated in the i-th iteration. and and The key variable is the variable that, under certain distance constraints, enables the function to achieve a smaller value; α α , α d and (Right now ) are all regularization parameters, x s For the remaining images after dictionary learning at all previous scales, and
[0089]
[0090] It should be understood that the Alternating Direction Method of Multipliers (ADMM) is used to decompose a complex optimization problem into several simpler subproblems, and then gradually approximate the solution to the original problem by alternately solving these subproblems. The nearest neighbor operator is a mathematical tool for solving optimization problems involving non-differentiable functions, updating the solution by combining gradient descent and projection operators.
[0091] The process of updating the underexposed sparse features and the underexposed dictionary features through the second feature update submodule (i.e., the overall objective function) to obtain the new underexposed sparse features and the new underexposed dictionary features is the same as that of the first feature update submodule, and will not be repeated here.
[0092] In this embodiment of the invention, since sparse feature constraints and dictionary constraints are not prior, but are dynamically learned through data-driven methods to obtain the most suitable features after transformation, that is, through each iteration, the updated features are better than the previous features. Therefore, the features used for calculation in the objective function are different and dynamically changing, and finally the feature data that makes the objective function take the minimum value is calculated.
[0093] Preferably, the common feature extraction submodule includes multiple convolutional blocks, each including a ninth convolutional layer, a fourth activation function layer, and a tenth convolutional layer connected sequentially. Feature processing for the first nonlinear function calculation is performed through the first convolutional block, the second convolutional block for the second nonlinear function calculation, the third convolutional block for the third nonlinear function calculation, and the fourth convolutional block for the fourth nonlinear function calculation. Specifically, the first and second convolutional blocks have the same input and output dimensions for image feature processing, and the third and fourth convolutional blocks have the same input and output dimensions. However, the convolutional layer parameters (i.e., bias parameters) of these four convolutional blocks are different, determined by randomly generated convolutional layer parameters (i.e., weights and biases).
[0094] The convolutional layers in each convolutional block are initialized, and random parameters are generated according to the input image features. Specifically, the parameter range of the convolutional layers in the first and second convolutional blocks (when extracting sparse features) is set as follows: nc_x:List[int]=[64,128,256,512], input: nc_x[0]*2, output: nc_x[0]; the parameter range of the convolutional layers in the third and fourth convolutional blocks (when extracting dictionary features) is set as follows: nc_d:List[int]=
[16] , out_nc:int=1, input: out_nc*nc_d[0]*2, output: out_nc*nc_d[0], where * indicates multiplication. This can be understood as follows: When extracting sparse features, the number of input channels is 128 and the number of output channels is 64; when extracting dictionary features, the number of input channels is 32 and the number of output channels is 16. `nc_x` and `nc_d` are lists containing multiple integers representing the number of channels in different layers. `out_nc`: represents the number of output channels for the dictionary; when the input image is used and a convolutional layer is invoked, random values between -1 and 1 are generated to produce the parameters of the convolutional layer.
[0095] Preferably, such as Figure 7 As shown, the common feature extraction submodule (i.e., nonlinear function) calculates the common features of the new overexposure sparse features and the new underexposure sparse features to obtain the overexposure common sparse features and the underexposure common sparse features, including:
[0096] The common sparse features of the new overexposure sparse features and the new underexposure sparse features are calculated using a first nonlinear function to obtain the overexposure common sparse features, which are expressed as follows:
[0097]
[0098] in, For overexposure common sparsity features, For newly overexposed sparse features, Let be the new underexposed sparse feature, s be the s-th scale, E be the number of convolutional dictionaries for the overexposed image, and F be the number of convolutional dictionaries for the underexposed image. For the first nonlinear function of feature processing at the s-th scale in the i-th iteration, it is used to extract sparse features biased towards overexposed images based on the bias parameter.
[0099] The common sparse features of the new overexposure sparse features and the new underexposure sparse features are calculated using a second nonlinear function to obtain the underexposure common sparse features, which are expressed as follows:
[0100]
[0101] in, This is a common sparse feature of underexposure. For the second nonlinear function of feature processing at the s-th scale in the i-th iteration, it is used to extract sparse features biased towards underexposed images based on the bias parameter.
[0102] The nonlinear function is Conv(ReLU(Conv(·))), used to... and Perform common sparse feature calculation. and When processing sparse features, the same dimension is set, and the bias parameter is determined by randomly generated convolutional layer parameters, that is...
[0103]
[0104] In this embodiment of the invention, each modal image does not contain information other than its own modality. Therefore, common features are extracted from different modal features, while retaining their own unique features. That is, common features of two features are extracted based on the weighting parameter, while non-common features of a certain feature are retained. For example, the common feature portion of the overexposed image features is enhanced, but the non-common features of the overexposed image are also retained, resulting in a more informative overexposed image feature representation. Setting different weighting parameters for the nonlinear function effectively fuses complementary information from different modalities, making the common parts more prominent during subsequent image fusion.
[0105] Preferably, such as Figure 8 As shown, the common feature extraction submodule (i.e., nonlinear function) calculates the common features of the new overexposure dictionary features and the new underexposure dictionary features to obtain the overexposure common dictionary features and the underexposure common dictionary features, including:
[0106] The common dictionary features of the new overexposure dictionary features and the new underexposure dictionary features are calculated using a third nonlinear function to obtain the overexposure common dictionary features, which are expressed as follows:
[0107]
[0108] in, To avoid overexposure of common dictionary features, For newly exposed dictionary features, Let be the new underexposure dictionary features, s be the s-th scale, E be the number of convolutional dictionaries (for overexposure), F be the number of convolutional dictionaries (for underexposure), and W be... i s For the third nonlinear function of feature processing at the s-th scale in the i-th iteration, it is used to extract dictionary features biased towards overexposed images based on the bias parameter;
[0109] The common dictionary features of the new overexposure dictionary features and the new underexposure dictionary features are calculated using a fourth nonlinear function to obtain the underexposure common dictionary features, which are expressed as follows:
[0110]
[0111] in, For underexposed common dictionary features, Y i s Let be the fourth nonlinear function for feature processing at the s-th scale in the i-th iteration, used to extract dictionary features biased towards underexposed images based on the bias parameter; the nonlinear function is Conv(ReLU(Conv(·))), used to... and Perform common dictionary feature calculation, W i s and Y i s When processing dictionary features, the same dimension is set, and the bias parameter is determined by randomly generated convolutional layer parameters.
[0112] In this embodiment of the invention, in order to better match common sparse features, the feature dictionary is updated, and the common features of the two features are extracted according to the bias parameter, so that it can adapt to the multimodal image being processed.
[0113] Preferably, such as Figure 7 As shown, the step of extracting features from the overexposure common sparse features and the overexposure common dictionary features through the feature fusion submodule to obtain the overexposure fused image includes:
[0114] The overexposure common sparse features and the overexposure common dictionary features are convolved to obtain the overexposure fused image, which is represented as follows:
[0115]
[0116] in, For the overexposed fused image of the i-th iteration, For the overexposed fusion image at the s-th scale of the i-th iteration, The overexposure common sparse features of the s-th scale image in the i-th iteration are learned by the e-th convolutional dictionary. Let S be the overexposure common dictionary features of the s-th scale image in the i-th iteration learned by the e-th convolution dictionary, where S is the number of image scales, E is the number of convolution dictionaries, and * represents the convolution operation.
[0117] Specifically, the feature fusion submodule includes a convolutional layer composed of multiple convolutional kernels, which integrates the common sparse features of overexposure at the first scale. Common dictionary features of overexposure Perform a convolution operation to obtain the overexposure fused image at the first scale of the first iteration. The second scale of overexposure common sparsity features Common dictionary features of overexposure Perform a convolution operation to obtain the overexposure fused image at the second scale of the first iteration. Similarly, the common sparse features of overexposure at the s-th scale are... Common dictionary features of overexposure Perform a convolution operation to obtain the overexposure fused image at the s-th scale of the first iteration. The overexposed fused images of all scales in the first iteration are summed element-wise to obtain the overexposed fused image of the first iteration.
[0118] like Figure 8 As shown, the process of extracting features from the underexposed common sparse features and the underexposed common dictionary features through the feature fusion submodule to obtain the underexposed fused image is the same as the process of obtaining the overexposed fused image through the feature fusion submodule, and will not be repeated here. The underexposed fused image is represented as follows:
[0119]
[0120] in, For the underexposed fusion image of the i-th iteration, For the underexposed fusion image at the s-th scale of the i-th iteration, The underexposed common sparse features of the s-th scale image in the i-th iteration are learned by the f-th convolutional dictionary. Let S be the underexposure common dictionary features of the s-th scale image in the i-th iteration learned by the f-th convolutional dictionary, where S is the number of image scales and F is the number of convolutional dictionaries.
[0121] In this embodiment of the invention, features at multiple scales are fused through convolution operations to generate a modal image that highlights common features while retaining non-common features of its own modality, thereby reconstructing a clearer image.
[0122] Preferably, the step of fusing the overexposed fused image and the underexposed fused image to obtain the reconstructed image includes:
[0123] The reconstructed image is obtained by weighting the overexposed and underexposed fused images using a weighted fusion expression. The weighted fusion expression is as follows:
[0124]
[0125] Where Z represents the reconstructed image. For the overexposed fused image of the Nth iteration, The image is the underexposed fused image from the Nth iteration, where × represents multiplication, and W... x W is the weighting parameter for overexposed images. y Let W be the weight parameters for the underexposed image, and satisfy W x +W y =1.
[0126] In this embodiment of the invention, images from multiple modalities are fused according to weight parameters to reconstruct a target image for use in assisting diagnosis.
[0127] Preferably, the evaluation calculation of the reconstructed image and the standard image to obtain evaluation indicators includes:
[0128] The reconstructed image and the standard image are calculated using a structural similarity calculation expression to obtain a structural similarity index;
[0129] The peak signal-to-noise ratio (PSNR) is calculated by comparing the reconstructed image with the standard image using a peak signal-to-noise ratio (PSNR) calculation expression.
[0130] It should be understood that the Structural Similarity Index (SSIM) is a metric used to measure the similarity between two images, focusing specifically on the structural information of the images, rather than just pixel-level similarity. The SSIM ranges from 0 to 1, with higher values indicating better image quality. Peak Signal-to-Noise Ratio (PSNR) is an objective metric used to measure the difference between two images, primarily for evaluating the effectiveness of image compression, transmission, or reconstruction algorithms. Higher PSNR values indicate greater similarity between the two images and less quality loss.
[0131] In this embodiment of the invention, the fusion effect of the fusion technology is evaluated by the structural similarity index and the peak signal-to-noise ratio.
[0132] like Figure 9 As shown, an embodiment of the present invention provides a deep multimodal image fusion system, comprising:
[0133] An image import unit is used to import a set of images from multiple exposure sources, the set of images from multiple exposure sources including overexposed images, underexposed images, and standard images;
[0134] The feature pre-extraction unit is used to extract features from the overexposed image to obtain overexposed sparse features and overexposed dictionary features, and to extract features from the underexposed image to obtain underexposed sparse features and underexposed dictionary features.
[0135] The feature update unit is used to update the overexposed sparse features and the overexposed dictionary features through the overall objective function to obtain new overexposed sparse features and new overexposed dictionary features, and to update the underexposed sparse features and underexposed dictionary features through the overall objective function to obtain new underexposed sparse features and new underexposed dictionary features.
[0136] The feature extraction unit is used to perform common feature calculation on the new overexposure sparse features and the new underexposure sparse features using a nonlinear function to obtain overexposure common sparse features and underexposure common sparse features, and to perform common feature calculation on the new overexposure dictionary features and the new underexposure dictionary features using a nonlinear function to obtain overexposure common dictionary features and underexposure common dictionary features.
[0137] The feature fusion unit is used to extract features from the overexposed common sparse features and the overexposed common dictionary features to obtain an overexposed fused image, and to extract features from the underexposed common sparse features and the underexposed common dictionary features to obtain an underexposed fused image.
[0138] The image reconstruction unit is used to perform fusion processing on the overexposed fused image and the underexposed fused image to obtain a reconstructed image, and to evaluate and calculate the reconstructed image and the standard image to obtain evaluation indicators.
[0139] The aforementioned deep multimodal image fusion system can be further described in the above-described implementation details and beneficial effects of a deep multimodal image fusion method, which will not be repeated here.
[0140] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0141] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0142] In the several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0143] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of the present invention, depending on actual needs.
[0144] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A deep multimodal image fusion method, characterized in that, Includes the following steps: S1. Import a set of multiple exposure source images, which includes overexposed images, underexposed images, and standard images; S2. Extract features from the overexposed image to obtain overexposed sparse features and overexposed dictionary features; extract features from the underexposed image to obtain underexposed sparse features and underexposed dictionary features. S3. The overexposure sparse features and the overexposure dictionary features are updated and calculated using the overall objective function to obtain new overexposure sparse features and new overexposure dictionary features. The underexposure sparse features and the underexposure dictionary features are updated and calculated using the overall objective function to obtain new underexposure sparse features and new underexposure dictionary features. S4. Calculate the common features of the new overexposure sparse features and the new underexposure sparse features using a nonlinear function to obtain the overexposure common sparse features and the underexposure common sparse features. Calculate the common features of the new overexposure dictionary features and the new underexposure dictionary features using a nonlinear function to obtain the overexposure common dictionary features and the underexposure common dictionary features. S5. Extract features from the overexposed common sparse features and the overexposed common dictionary features to obtain an overexposed fused image; extract features from the underexposed common sparse features and the underexposed common dictionary features to obtain an underexposed fused image. S6. Repeat S2-S5 until the preset number of iterations is reached. Perform fusion processing on the overexposed fused image and the underexposed fused image to obtain a reconstructed image. Evaluate and calculate the reconstructed image and the standard image to obtain evaluation indicators.
2. The deep multimodal image fusion method according to claim 1, characterized in that, The step of extracting features from the overexposed image to obtain sparse overexposed features and dictionary overexposed features includes: The pre-constructed convolution dictionary is initialized, and the initialized convolution dictionary is used as the overexposure dictionary feature. The overexposure image is then convolved using the overexposure dictionary feature to obtain overexposure sparse features.
3. The deep multimodal image fusion method according to claim 1, characterized in that, Before the step of updating the overexposure sparse features and the overexposure dictionary features using the overall objective function, the method further includes: Based on the overexposure sparse features and the overexposure dictionary features, the objective function of traditional dictionary learning is improved to obtain the overall objective function, which is: , in, Let be the overall objective function. The number of image scales. The number of words in the convolution dictionary. For overexposure sparse features, For overexposure dictionary features, For overexposed images, For the Euclidean norm, For regularization parameters, For sparse features, the regularization constraint function is... For regularization parameters, The regularization constraint function for dictionary features.
4. The deep multimodal image fusion method according to claim 3, characterized in that, The step of updating the overexposure sparse features and the overexposure dictionary features using the overall objective function to obtain new overexposure sparse features and new overexposure dictionary features includes: The overall objective function is decomposed into a sparse objective function and a dictionary objective function by using the alternating direction multiplier method; New overexposure sparse features are obtained by calculating the overexposure sparse features using the sparse objective function. The sparse objective function is: , in, For a sparse objective function, The number of image scales. The number of words in the convolution dictionary. For overexposure sparse features, For overexposure dictionary features, For the Euclidean norm, For regularization parameters, For sparse features, the regularization constraint function is... For regularization parameters, For overexposure sparse auxiliary variables, For overexposed images without dictionary learning, and , For the first iteration For newly exposed dictionary features, This is a sparse feature of overexposure; The overexposure dictionary features are calculated using the dictionary objective function to obtain new overexposure dictionary features. The dictionary objective function is: , in, The objective function is a dictionary. For regularization parameters, For regularization parameters, For overexposure dictionary auxiliary variables, The regularization constraint function for dictionary features.
5. The deep multimodal image fusion method according to claim 1, characterized in that, The step of calculating common features of the new overexposure sparse features and the new underexposure sparse features using a nonlinear function to obtain common overexposure sparse features and common underexposure sparse features includes: The common sparse features of the new overexposure sparse features and the new underexposure sparse features are calculated using a first nonlinear function to obtain the overexposure common sparse features, which are expressed as follows: , in, For overexposure common sparsity features, For newly overexposed sparse features, This is a new underexposed sparse feature. For the s-th scale, and All of these are the number of convolutional dictionaries. The first nonlinear function is used to extract sparse features biased towards overexposed images based on the bias parameter. For the first The next iteration; The common sparse features of the new overexposure sparse features and the new underexposure sparse features are calculated using a second nonlinear function to obtain the underexposure common sparse features, which are expressed as follows: ), in, This is a common sparse feature of underexposure. The second nonlinear function is used to extract sparse features biased towards underexposed images based on the bias parameter; the nonlinear function is Conv(ReLU(Conv( ))), used for and Perform common sparse feature calculation. and When processing sparse features, the same dimension is set, and the bias parameter is determined by randomly generated convolutional layer parameters.
6. The deep multimodal image fusion method according to claim 1, characterized in that, The step of calculating common features of the new overexposure dictionary features and the new underexposure dictionary features using a nonlinear function to obtain common overexposure dictionary features and common underexposure dictionary features includes: The common dictionary features of the new overexposure dictionary features and the new underexposure dictionary features are calculated using a third nonlinear function to obtain the overexposure common dictionary features, which are expressed as follows: , in, To avoid overexposure of common dictionary features, For newly exposed dictionary features, For new underexposed dictionary features, For the s-th scale, Both F and F represent the number of convolutional dictionaries. The third nonlinear function is used to extract dictionary features biased towards overexposed images based on the bias parameter. For the first The next iteration; The common dictionary features of the new overexposure dictionary features and the new underexposure dictionary features are calculated using a fourth nonlinear function to obtain the underexposure common dictionary features, which are expressed as follows: ), in, For underexposed common dictionary features, The fourth nonlinear function is used to extract dictionary features biased towards underexposed images based on the bias parameter; the nonlinear function is Conv(ReLU(Conv( ))), used for and Perform common dictionary feature calculation. and When processing dictionary features, the same dimension is set, and the bias parameter is determined by randomly generated convolutional layer parameters.
7. The deep multimodal image fusion method according to claim 1, characterized in that, The step of extracting features from the common sparse features and the common dictionary features of overexposure to obtain the overexposure fused image includes: The overexposure common sparse features and the overexposure common dictionary features are convolved to obtain the overexposure fused image, which is represented as follows: , in, For the first i The overexposed fused image from the next iteration. For overexposure common sparsity features, To avoid overexposure of common dictionary features, The number of image scales. The number of words in the convolution dictionary.
8. The deep multimodal image fusion method according to claim 1, characterized in that, The process of fusing the overexposed and underexposed images to obtain a reconstructed image includes: The reconstructed image is obtained by weighting the overexposed and underexposed fused images using a weighted fusion expression. The weighted fusion expression is as follows: , in, To reconstruct the image, For the weight parameters of overexposed images, For the overexposed fused image of the Nth iteration, For underexposed images, weight parameters This is the underexposed fusion image from the Nth iteration.
9. The deep multimodal image fusion method according to claim 1, characterized in that, The evaluation calculation of the reconstructed image and the standard image to obtain evaluation indicators includes: The reconstructed image and the standard image are calculated using a structural similarity calculation expression to obtain a structural similarity index; The peak signal-to-noise ratio (PSNR) is calculated by comparing the reconstructed image with the standard image using a peak signal-to-noise ratio (PSNR) calculation expression.
10. A deep multimodal image fusion system, characterized in that, include: An image import unit is used to import a set of images from multiple exposure sources, the set of images from multiple exposure sources including overexposed images, underexposed images, and standard images; The feature pre-extraction unit is used to extract features from the overexposed image to obtain overexposed sparse features and overexposed dictionary features, and to extract features from the underexposed image to obtain underexposed sparse features and underexposed dictionary features. The feature update unit is used to update the overexposed sparse features and the overexposed dictionary features through the overall objective function to obtain new overexposed sparse features and new overexposed dictionary features, and to update the underexposed sparse features and underexposed dictionary features through the overall objective function to obtain new underexposed sparse features and new underexposed dictionary features. The feature extraction unit is used to perform common feature calculation on the new overexposure sparse features and the new underexposure sparse features using a nonlinear function to obtain overexposure common sparse features and underexposure common sparse features, and to perform common feature calculation on the new overexposure dictionary features and the new underexposure dictionary features using a nonlinear function to obtain overexposure common dictionary features and underexposure common dictionary features. The feature fusion unit is used to extract features from the overexposed common sparse features and the overexposed common dictionary features to obtain an overexposed fused image, and to extract features from the underexposed common sparse features and the underexposed common dictionary features to obtain an underexposed fused image. The image reconstruction unit is used to perform fusion processing on the overexposed fused image and the underexposed fused image to obtain a reconstructed image, and to evaluate and calculate the reconstructed image and the standard image to obtain an evaluation index.
Citation Information
Patent Citations
Multi-exposure image fusion method based on low-rank decomposition and sparse representation
CN117291851A
Reconstruction of high-quality images from a binary sensor array
US20170272639A1