Deep multi-modal image fusion method and system

Through the deep multimodal image fusion method, the multi-exposure source image collection is used for feature extraction and updating calculation, which solves the complex problem of multimodal multi-scale image data fusion, and realizes high-quality image fusion and diagnostic support.

CN120047784AActive Publication Date: 2025-05-27GUILIN UNIV OF ELECTRONIC TECH

Patent Information

Application Number
CN202510123737.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-27
Estimated Expiration
2045-01-26

AI Technical Summary

Technical Problem

The prior art is difficult to effectively fusion when processing multimodal multi-scale image data, resulting in complex image fusion and poor effect.

Method used

The deep multimodal image fusion method is adopted, and feature extraction and update calculation is performed by importing the multi-exposure source image set, common features are calculated using nonlinear functions, and through multiple iterations and multi-scale fusion, the reconstruction image is finally obtained.

Benefits of technology

It realizes effective fusion of multimodal images, retains the overall structure of the image, removes redundant information, extracts higher quality sparse features, and supports more accurate image diagnosis and analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047784A_ABST
    Figure CN120047784A_ABST
Patent Text Reader

Abstract

The invention provides a deep multi-modal image fusion method and system, and relates to the technical field of image processing. The method comprises the following steps: respectively extracting features of an overexposure image and an underexposure image, respectively obtaining dictionary features and sparse features of the overexposure image and dictionary features and sparse features of the underexposure image, and respectively updating the dictionary features and sparse features of the underexposure image and the overexposure image through a total objective function, common features are extracted from the updated dictionary features and sparse features through a nonlinear function, feature extraction is carried out according to the extracted common features, an overexposure fusion image and an underexposure fusion image are obtained and fused, a reconstructed image is obtained and compared with a standard image, and an evaluation index is generated. Transformation learning is carried out through total objective function iteration change, sparse features with higher quality are extracted, important internal attributes of the image are captured to reconstruct the image, and fusion of the multi-modal image is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention mainly relates to the field of image processing technology, and specifically to a deep multimodal image fusion method and system. Background Art

[0002] In today's information age, we are faced with the situation of obtaining a large amount of image data from multiple imaging technologies. These image data usually come from different imaging modalities, and each imaging modality can provide unique information about the imaged object. The purpose of multimodal image processing is to integrate these modal image data from different imaging technologies in order to obtain a more comprehensive and accurate representation of the imaged object, thereby revealing information that a single modality cannot provide. However, image data from different modalities often have different characteristics and scales, but direct image fusion is very complicated.

[0003] In machine learning, sparse representation models have attracted widespread attention due to their advantages in improving image interpretability, compression, and feature extraction, and can be applied to image fusion technology. Sparse representation models are mainly divided into traditional dictionary learning and transform learning. Although dictionary learning has been successful in generating sparse representations, it requires storing and processing a large number of overlapping image blocks when processing large-scale data sets. Transformation learning, as an emerging sparse representation method, directly analyzes image data by learning analytical transformations to obtain sparse features, but it still faces challenges when processing multimodal and multi-scale data. Summary of the invention

[0004] The technical problem to be solved by the present invention is to provide a deep multimodal image fusion method and system in view of the deficiencies in the prior art.

[0005] The technical solution of the present invention to solve the above technical problems is as follows:

[0006] A deep multimodal image fusion method comprises the following steps:

[0007] S1. Importing a multi-exposure source image set, wherein the multi-exposure source image set includes an overexposed image, an underexposed image, and a standard image;

[0008] S2, performing feature extraction on the overexposed image to obtain overexposure sparse features and overexposure dictionary features, and performing feature extraction on the underexposed image to obtain underexposure sparse features and underexposure dictionary features;

[0009] S3, updating and calculating the overexposure sparse feature and the overexposure dictionary feature through the total objective function to obtain a new overexposure sparse feature and a new overexposure dictionary feature, and updating and calculating the underexposure sparse feature and the underexposure dictionary feature through the total objective function to obtain a new underexposure sparse feature and a new underexposure dictionary feature;

[0010] S4, performing common feature calculation on the new overexposure sparse feature and the new underexposure sparse feature through a nonlinear function to obtain an overexposure common sparse feature and an underexposure common sparse feature, and performing common feature calculation on the new overexposure dictionary feature and the new underexposure dictionary feature through a nonlinear function to obtain an overexposure common dictionary feature and an underexposure common dictionary feature;

[0011] S5, performing feature extraction on the overexposure common sparse features and the overexposure common dictionary features to obtain an overexposure fused image, and performing feature extraction on the underexposure common sparse features and the underexposure common dictionary features to obtain an underexposure fused image;

[0012] S6. Repeat S2-S5 until a preset number of iterations is reached, fuse the overexposed fused image and the underexposed fused image to obtain a reconstructed image, and evaluate and calculate the reconstructed image and the standard image to obtain an evaluation index.

[0013] Another technical solution of the present invention to solve the above technical problems is as follows:

[0014] A deep multimodal image fusion system, comprising:

[0015] An image importing unit, used for importing a multi-exposure source image set, wherein the multi-exposure source image set includes an overexposed image, an underexposed image, and a standard image;

[0016] A feature pre-extraction unit, configured to perform feature extraction on the overexposed image to obtain overexposure sparse features and overexposure dictionary features, and perform feature extraction on the underexposed image to obtain underexposure sparse features and underexposure dictionary features;

[0017] a feature updating unit, configured to update and calculate the overexposure sparse feature and the overexposure dictionary feature through a total objective function to obtain a new overexposure sparse feature and a new overexposure dictionary feature, and to update and calculate the underexposure sparse feature and the underexposure dictionary feature through a total objective function to obtain a new underexposure sparse feature and a new underexposure dictionary feature;

[0018] a feature extraction unit, configured to perform common feature calculation on the new overexposure sparse feature and the new underexposure sparse feature through a nonlinear function to obtain an overexposure common sparse feature and an underexposure common sparse feature, and perform common feature calculation on the new overexposure dictionary feature and the new underexposure dictionary feature through a nonlinear function to obtain an overexposure common dictionary feature and an underexposure common dictionary feature;

[0019] a feature fusion unit, configured to perform feature extraction on the overexposure common sparse features and the overexposure common dictionary features to obtain an overexposure fused image, and perform feature extraction on the underexposure common sparse features and the underexposure common dictionary features to obtain an underexposure fused image;

[0020] The image reconstruction unit is used to perform a fusion process on the over-exposed fusion image and the under-exposed fusion image to obtain a reconstructed image, and to perform an evaluation calculation on the reconstructed image and the standard image to obtain an evaluation index.

[0021] The beneficial effects of the present invention are as follows: transform learning is realized by processing updated features in each iteration based on the total objective function, sparse features are directly extracted from image data through learning and analyzing transformations, the overall structure of the image is effectively retained, important intrinsic properties of the image are captured through transform learning, redundant information is removed, and higher quality sparse features are extracted, updated sparse features and dictionary features are dynamically obtained in a data-driven manner, and common features of images of various modalities are extracted, which is better suitable for image reconstruction and realizes the fusion of multimodal images to provide more accurate support when using image diagnosis and analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 The overall structure diagram of the deep multimodal image fusion model provided by the embodiment of the present invention;

[0023] Figure 2 A flowchart of a deep multimodal image fusion method provided by an embodiment of the present invention;

[0024] Figure 3 A structural diagram of a modal multi-scale transformation learning module provided in an embodiment of the present invention;

[0025] Figure 4 A structural diagram of a sparse feature updating subunit provided in an embodiment of the present invention;

[0026] Figure 5 A structural diagram of a spatial attention module provided by an embodiment of the present invention;

[0027] Figure 6 A structural diagram of a dictionary feature updating subunit provided in an embodiment of the present invention;

[0028] Figure 7 A structural diagram of a fusion module for overexposed image processing provided by an embodiment of the present invention;

[0029] Figure 8 A structural diagram of a fusion module for underexposed image processing provided by an embodiment of the present invention;

[0030] Fig. 9A module block diagram of a deep multimodal image fusion system provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0031] The principles and features of the present invention are described below in conjunction with the accompanying drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.

[0032] In the sparse representation model, unlike the synthetic framework of dictionary learning, transform learning is analytical, and its key advantage is that it can avoid trivial solutions by regularizing the analytical transformation, which helps to find more meaningful image representations. In addition, transform learning has better stability and robustness when processing images because all analytical transformations contribute equally to the analytical image.

[0033] Convolutional dictionary learning (CDL) combines the idea of ​​convolutional neural networks with traditional dictionary learning techniques. The core idea of ​​convolutional dictionary learning (CDL) is to learn a set of convolution kernels that can extract local and global features from the input data to form a sparse representation of the data. In traditional dictionary learning, images are represented as a linear combination of a set of basis vectors, but this is not effective when processing image data with strong spatial correlation. Convolutional dictionary learning reduces this redundancy by introducing convolution operations, thereby improving computational efficiency. This is because image data usually has strong spatial correlation, and convolution operations can capture local features of data for image processing.

[0034] Convolutional Transform Learning (CTL) is an emerging method that combines convolutional neural networks (CNN) and transform learning. Its goal is to make the image sparse in the convolution domain, that is, most of the elements in the feature map after the convolution operation are zero, and only a few elements are non-zero, that is, the convolution operation is used to extract the local features of the image, making it more efficient and natural when processing the image. Therefore, the present invention studies a technology for fusion of multimodal images based on convolutional transform learning.

[0035] The present invention is implemented in the field of exposure image processing, which can be an important development in the field of modern medical imaging and diagnosis. It can provide more comprehensive and accurate lesion information through multimodal image fusion strategies, thereby supporting more accurate diagnosis and analysis. In multi-exposure image fusion technology, the image sequence to be fused will contain large areas of white or black areas due to overexposure or underexposure. These areas are usually low-contrast, dull colors, and lack of details. By fusing overexposed images and underexposed images, an image with a wider dynamic range and richer details can be obtained.

[0036] like Figure 1 and Figure 2As shown, a deep multimodal image fusion method provided by an embodiment of the present invention includes the following steps:

[0037] S0. Constructing a deep multimodal image fusion model based on a convolutional neural network, the deep multimodal image fusion model includes a plurality of multimodal image fusion modules MIFB connected in series, each of the multimodal image fusion modules includes a first modality multiscale transform learning module MTLBx and a second modality multiscale transform learning module MTLBy connected in parallel with the fusion module FB, wherein the first modality multiscale transform learning module MTLBx includes a first convolution submodule and a first feature update submodule, the second modality multiscale transform learning module MTLBy includes a second convolution submodule and a second feature update submodule, the fusion module FB includes a common feature extraction submodule and a feature fusion submodule, the first modality multiscale transform learning module MTLBx and the second modality multiscale transform learning module MTLBy have the same structure, and the image processing operations are also the same;

[0038] S1. Importing a multi-exposure source image set, wherein the multi-exposure source image set includes an overexposed image, an underexposed image, and a standard image;

[0039] S2. Performing feature extraction on the overexposed image through the first convolution submodule to obtain overexposure sparse features and overexposure dictionary features, and performing feature extraction on the underexposed image through the second convolution submodule to obtain underexposure sparse features and underexposure dictionary features;

[0040] S3, updating and calculating the overexposure sparse feature and the overexposure dictionary feature through the first feature updating submodule (i.e., the overall objective function) to obtain a new overexposure sparse feature and a new overexposure dictionary feature, and updating and calculating the underexposure sparse feature and the underexposure dictionary feature through the second feature updating submodule (i.e., the overall objective function) to obtain a new underexposure sparse feature and a new underexposure dictionary feature;

[0041] S4, performing common feature calculation on the new overexposure sparse feature and the new underexposure sparse feature through a common feature extraction submodule (i.e., a nonlinear function) to obtain an overexposure common sparse feature and an underexposure common sparse feature, and performing common feature calculation on the new overexposure dictionary feature and the new underexposure dictionary feature through a common feature extraction submodule (i.e., a nonlinear function) to obtain an overexposure common dictionary feature and an underexposure common dictionary feature;

[0042] S5. Extracting features from the overexposure common sparse features and the overexposure common dictionary features through a feature fusion submodule to obtain an overexposure fused image, and extracting features from the underexposure common sparse features and the underexposure common dictionary features through a feature fusion submodule to obtain an underexposure fused image;

[0043] S6. Repeat S2-S5 until a preset number of iterations is reached (i.e., image processing by multiple serially connected multimodal image fusion modules MIFB is completed), the overexposed fusion image and the underexposed fusion image are fused to obtain a reconstructed image, and the reconstructed image and the standard image are evaluated and calculated to obtain an evaluation index.

[0044] The multi-exposure source image set is a pair of multi-exposure source images, including over-exposed images, under-exposed images and standard images. The data set selected in this embodiment has a total of 440 pairs, including 340 pairs of training sets and 100 pairs of test sets.

[0045] Specifically, in S2, feature extraction of the overexposed image includes decomposing the overexposed image into overexposed images of multiple scales, and extracting corresponding overexposure sparse features and overexposure dictionary features from the overexposed images of each scale respectively; the processing of the underexposed image is the same as that of the overexposed image, and extracting corresponding underexposure sparse features and underexposure dictionary features from the underexposed images of each scale. In S3-S5, the overexposure sparse features and overexposure dictionary features of multiple scales and the underexposure sparse features and underexposure dictionary features of multiple scales are processed accordingly.

[0046] In an embodiment of the present invention, multi-scale variation learning is performed on each modality of the input image. By learning dictionaries at multiple scales, feature information of multiple scales of the modal image is extracted, and then multi-scale fusion is performed. Through multiple iterations and feature processing of multiple scale images each time, the image resolution is improved from coarse to fine, and then the feature representation of overexposed images and underexposed images is gradually optimized. Finally, after the multiple scale features of each modality are fused, the fused features of the two modalities are combined into a reconstructed image to overcome the quality problems of low resolution of a single modality image, obtain a clearer high-quality reconstructed image, and evaluate the fusion effect through evaluation indicators.

[0047] Compared with the fixed constraints (i.e., the overall objective function) of traditional dictionary learning, the present invention uses a dynamic learning network to update feature learning constraints to find the most appropriate prior knowledge. The feature update submodule can adapt to different tasks and input images, and can improve the accuracy and generalization ability of the model.

[0048] Preferably, the extracting features of the overexposed image by the first convolution submodule to obtain overexposure sparse features and overexposure dictionary features includes:

[0049] The pre-constructed convolution dictionary is initialized, the initialized convolution dictionary is used as an overexposure dictionary feature, and the overexposure image is convolved by the overexposure dictionary feature to obtain an overexposure sparse feature.

[0050] Specifically, the first convolution submodule includes a convolution layer composed of multiple convolution kernels; the overexposed image is decomposed into overexposed images of multiple scales, the pre-constructed convolution dictionary (i.e., multiple convolution kernels) and the sparse coefficient matrix are initialized (i.e., the parameters of the convolution dictionary are adjusted to the set dictionary parameters, and the parameters of the sparse coefficient matrix are adjusted to the set sparse parameters), and the initialized convolution dictionary and the initialized sparse coefficient matrix are obtained, and the initialized convolution dictionary is matched with the overexposed images of multiple scales to obtain the overexposed dictionary features of multiple scales (i.e., multiple initialized convolution kernels), and the overexposed images of corresponding scales are convoluted by multiple initialized convolution kernels to obtain the overexposed sparse features corresponding to the multiple scales, and the overexposed sparse features of multiple scales are composed of the overexposed sparse features. The overexposed image is represented as:

[0051]

[0052] Among them, LX is the overexposed image, S is the number of image scales, x s is the overexposed image of the sth scale, E is the number of convolution dictionaries (i.e., the number of convolution kernels) for the overexposed image, is the e-th convolution dictionary of the s-th scale (i.e., overexposure dictionary feature), From x s pass The extracted sparse features (i.e., overexposed sparse features).

[0053] The process of feature extraction of the underexposed image by the second convolution submodule is the same as that of the first convolution submodule, which will not be repeated here. The underexposed image is represented as:

[0054]

[0055] Among them, HY is the underexposed image, y s is the underexposed image of the sth scale, F is the number of convolutional dictionaries for the underexposed image, is the fth convolution dictionary of the sth scale (i.e., underexposed dictionary feature), From y s pass The extracted sparse features (i.e., underexposed sparse features) are convolution operations.

[0056] It should be understood that the convolution dictionary is composed of a group of convolution kernels, each of which is used to extract local features in the image; the initialization process can be random initialization.

[0057] In the embodiment of the present invention, the dictionary features are used to capture local features of an image to generate sparse features. The sparse features can remove redundant information in the image and can significantly reduce the amount of data storage and transmission while ensuring image quality.

[0058] Preferably, before the step of updating and calculating the overexposure sparse feature and the overexposure dictionary feature through the first feature updating submodule (i.e., the overall objective function), the method further includes:

[0059] The objective function of traditional dictionary learning is improved based on the overexposure sparse feature and the overexposure dictionary feature to obtain a total objective function, which is:

[0060]

[0061] Among them, L is the total objective function, S is the number of image scales, and E is the number of convolution dictionaries. To overexpose sparse features, is the overexposure dictionary feature, x is the overexposure image, is the Euclidean norm, λ is the regularization parameter, G(·) is the regularization constraint function of sparse features, β is the regularization parameter, H(·) is the regularization constraint function of dictionary features, and * is the convolution operation.

[0062] Specifically, the overall objective function of traditional dictionary learning is:

[0063]

[0064] in, is the reconstruction error term, is the sparse constraint regularization term, is the dictionary constraint regularization term;

[0065] The reconstruction error term is modified to The rest of the items remain unchanged;

[0066] The overexposed image x and the overexposed sparse features will be measured and overexposure dictionary features The reconstruction error term is the difference between the linear combinations of , modified to measure the overexposed sparse features With overexposure dictionary feature The reconstruction error term is the difference between the linear combination of the overexposed image x and the overexposed image x.

[0067] The improvement items of the second feature update submodule (i.e., the overall objective function) are the same as those of the first feature update submodule, and will not be repeated here. The overall objective function is expressed as:

[0068]

[0069] In the embodiment of the present invention, the improved total objective function is more focused on the reconstruction error of the sparse features, which is beneficial to updating and adjusting the sparse features according to the errors, so that the sparse features can better represent the important information of the image, remove redundant information, keep the sparse constraint regularization term and the dictionary constraint regularization term unchanged, and perform regularization constraints on the overexposed sparse features and the overexposed dictionary features respectively to prevent overfitting.

[0070] Preferably, if Figure 3 As shown, the overexposure sparse feature and the overexposure dictionary feature are updated and calculated by multiple first feature updating submodules, including:

[0071] For the first input feature x 1 (i.e., the overexposed sparse features of the first scale and the overexposed dictionary features of the first scale) are updated and calculated to obtain the first output feature (i.e., the updated overexposed sparse features of the first scale and the updated overexposed dictionary features of the first scale ), the first output feature and the second input feature x 2 (i.e., the overexposed sparse features of the second scale and the overexposed dictionary features of the second scale) are concatenated and combined, and the combined second input features are updated and calculated to obtain the second output features (i.e., the updated overexposed sparse features of the second scale and the updated overexposed dictionary features of the second scale ), and so on, the s-1th output feature and the sth input feature x s Concatenate and combine, update and calculate the sth input feature after combination, and obtain the sth output feature (i.e., the new overexposed sparse features and new overexposure dictionary features ).

[0072] Preferably, the updating and calculating of the overexposure sparse feature and the overexposure dictionary feature by the first feature updating submodule (i.e., the total objective function) to obtain a new overexposure sparse feature and a new overexposure dictionary feature includes:

[0073] The total objective function is decomposed into a sparse objective function and a dictionary objective function by using the alternating direction multiplier method;

[0074] The overexposure sparse feature is calculated by the sparse objective function to obtain a new overexposure sparse feature, and the sparse objective function is:

[0075]

[0076] Among them, L 1 is the sparse objective function, S is the number of image scales, E is the number of convolution dictionaries, is the overexposed sparse feature (i.e., the new overexposed sparse feature obtained after all scale changes are learned), is the overexposure dictionary feature (i.e., the new overexposure dictionary feature obtained after all scale changes are learned), * is the convolution operation, is the Euclidean norm, λ is the regularization parameter, G(·) is the regularization constraint function of sparse features, α α is the regularization parameter, is the overexposed sparse auxiliary variable, x s is the overexposed image without dictionary learning (i.e., the overexposed image remaining after dictionary learning of all previous scales), i is the i-th iteration;

[0077] The overexposure dictionary feature is calculated by the dictionary objective function to obtain a new overexposure dictionary feature, and the dictionary objective function is:

[0078]

[0079] Among them, L 2 is the dictionary objective function, β is the regularization parameter, α d is the regularization parameter, is the auxiliary variable of the overexposure dictionary, and H(·) is the regularization constraint function of the dictionary feature.

[0080] Specifically, the objective function is decomposed into two sub-objectives (i.e., sparse objective function and dictionary objective function) by alternating direction multiplier method, that is, the first feature updating submodule includes sparse feature updating subunit and dictionary feature updating subunit. The numerical solution of sparse constraint is solved by sparse feature updating subunit (Sparse Representation Updating (SRU) unit), as shown in Figure 4 and Figure 5 As shown in the figure, this subunit processes input features through a combination of residual block set, convolution layer Conv and spatial attention module Spatial Attention module, which enhances the expressiveness of features, including:

[0081] Input pre-built overexposed sparse helper variables Regularization parameter and overexposure sparse features of multiple scales and overexposure dictionary features of multiple scales to the input layer, the input layer is connected in series with the first residual block set and the first convolution layer, the output of the first convolution layer is connected in series with the second residual block set and the second convolution layer, the output of the second convolution layer is connected in series with the input of the third residual block set, the output of the second convolution layer and the output of the third residual block set are added element by element and connected in series with the third convolution layer and the fourth residual block set, the output of the first convolution layer is added element by element with the output of the fourth residual block set, and connected in series with the fourth convolution layer and the fifth residual block set, the output of the fifth residual block set is added element by element with the overexposure sparse auxiliary variable, and connected with the spatial attention module and the fifth convolution layer, the output of the fifth convolution layer is connected with the output layer, and the output layer outputs the new overexposure sparse features

[0082] Among them, each residual block set includes multiple residual blocks connected in series, the input of any residual block is added element by element to its output, and then input into the next residual block, each residual block includes the sixth convolution layer, the first activation function layer ReLU, the first regularization layer BN and the seventh convolution layer connected in series; the spatial attention module includes the maximum pooling layer Maxpooling, the point-by-point convolution layer PWConv and the first depth-separable convolution layer DSConv connected in parallel, the output of the maximum pooling layer Max pooling, the output of the point-by-point convolution layer PWConv and the output of the depth-separable convolution layer DSConv are connected with the splicing layer, and the splicing layer is connected in series with the second depth-separable convolution layer and the second activation function layer.

[0083] The numerical solution of the dictionary constraint is solved by the dictionary filter updating (DFU) unit, such as Figure 6 As shown in the figure, the subunit processes the input dictionary through a series of convolutional layers Conv, activation functions ReLU and regularization layers BN, and reconstructs the dictionary using a reconstruction layer so that the dictionary can adapt to the characteristics of the input data, including:

[0084] Import pre-built overexposure dictionary helper variable Regularization parameter As well as overexposure sparse features of multiple scales and overexposure dictionary features of multiple scales to the input layer, the input layer is connected in series with multiple convolution blocks, the last convolution block is connected to the reconstruction layer, the output of the reconstruction layer is added element by element to the overexposure dictionary auxiliary variable and connected to the output layer, and the output layer outputs the new overexposure dictionary feature

[0085] Each convolution block includes an eighth convolution layer, a third activation function, and a second regularization layer connected in series.

[0086] Furthermore, according to the preset number of iterations, the sparse objective function and the dictionary objective function are solved respectively by the proximity operator to obtain the overexposure sparse auxiliary variable New overexposure sparse features Overexposure dictionary auxiliary variables and new overexposure dictionary features It is expressed as:

[0087]

[0088] Where i is the number of iterations, and is updated in the i-th iteration and and is updated in the i-th iteration and and is the key variable, that is, the variable that can make the function obtain a smaller value under certain distance constraints. α , α d and (Right now ) are regularization parameters, x s is the image remaining after dictionary learning at all previous scales, and

[0089]

[0090] It should be understood that the Alternating Direction Method of Multipliers (ADMM) is used to decompose a complex optimization problem into several simple sub-problems, and then gradually approximate the solution to the original problem by alternately solving these sub-problems. The proximity operator is a mathematical tool used to solve optimization problems involving non-differentiable functions, and the solution is updated by combining the gradient descent method and the projection operator.

[0091] The underexposure sparse features and the underexposure dictionary features are updated and calculated by the second feature updating submodule (ie, the overall objective function), and the processing process of obtaining new underexposure sparse features and new underexposure dictionary features is the same as that of the first feature updating submodule, which will not be repeated here.

[0092] In the embodiment of the present invention, since the sparse feature constraints and dictionary constraints are not a priori, but are dynamically learned in a data-driven manner to obtain the most suitable features after the transformation, that is, through each iteration, the features are updated to be better than the previous features. Therefore, the features used for calculation in the objective function are different and change dynamically, and finally the feature data that minimizes the objective function is calculated.

[0093] Preferably, the common feature extraction submodule includes a plurality of convolution blocks, each of which includes a ninth convolution layer, a fourth activation function layer, and a tenth convolution layer connected one by one. The first convolution block is used to perform feature processing of the first nonlinear function calculation process, the second convolution block is used to perform feature processing of the second nonlinear function calculation process, the third convolution block is used to perform feature processing of the third nonlinear function calculation process, and the fourth convolution block is used to perform feature processing of the fourth nonlinear function calculation process; wherein, the first convolution block and the second convolution block have the same input dimension and output dimension when processing image features, and the third convolution block and the fourth convolution block have the same input dimension and output dimension when processing image features, but the convolution layer parameters (i.e., bias parameters) of the four convolution blocks are different, and are determined by randomly generated convolution layer parameters (i.e., weights and biases).

[0094] The convolutional layers in each convolutional block are initialized and random parameters are generated according to the input image features. Specifically, the parameter value range of the convolutional layers in the first and second convolutional blocks (when extracting sparse features) is set to: nc_x: List[int] = [64, 128, 256, 512], input: nc_x[0]*2, output: nc_x[0]; the parameter value range of the convolutional layers in the third and fourth convolutional blocks (when extracting dictionary features) is set to: nc_d: List[int] =

[16] , out_nc: int = 1, input: out_nc*nc_d[0]*2, output: out_nc*nc_d[0], * indicates multiplication. It can be understood as follows: for sparse feature extraction, the number of input channels is 128 and the number of output channels is 64; for dictionary feature extraction, the number of input channels is 32 and the number of output channels is 16; nc_x and nc_d are lists containing multiple integers, which represent the number of channels in different layers. out_nc: represents the number of output channels of the dictionary; when the image is input and the convolution layer is called, random values ​​are generated between -1 and 1 to generate the parameters of the convolution layer.

[0095] Preferably, if Figure 7 As shown, the common feature extraction submodule (i.e., nonlinear function) is used to calculate the common features of the new overexposure sparse features and the new underexposure sparse features to obtain the overexposure common sparse features and the underexposure common sparse features, including:

[0096] The common sparse feature calculation is performed on the new overexposure sparse feature and the new underexposure sparse feature through the first nonlinear function to obtain an overexposure common sparse feature, and the overexposure common sparse feature is expressed as:

[0097]

[0098] in, is the common sparse feature of overexposure, is the new overexposed sparse feature, is the new underexposed sparse feature, s is the sth scale, E is the number of convolution dictionaries for overexposed images, and F is the number of convolution dictionaries for underexposed images. is the first nonlinear function for processing the features of the s-th scale in the i-th iteration, used for extracting sparse features biased towards the overexposed image according to the bias parameter;

[0099] The new overexposure sparse feature and the new underexposure sparse feature are calculated using a second nonlinear function to obtain an underexposure common sparse feature. The underexposure common sparse feature is expressed as:

[0100]

[0101] in, is the common sparse feature of underexposure, is a second nonlinear function for processing features of the s-th scale in the i-th iteration, used for extracting sparse features biased towards underexposed images according to the bias parameter;

[0102] The nonlinear function is Conv(ReLU(Conv(·))), which is used to and Perform common sparse feature calculations, and The same dimension is set when processing sparse features, and the bias parameters are determined by randomly generated convolution layer parameters, that is,

[0103]

[0104] In the embodiment of the present invention, each modal image does not contain information other than its own modality. Therefore, common features are extracted from different modal features, while retaining their own features, that is, common features of two features are extracted according to the bias parameter, and non-common features of a certain feature are retained, such as enhancing the common feature part of the overexposed image features, but also retaining the non-common features of the overexposed image, to obtain a more informative overexposed image feature representation. Different bias parameters are set for the nonlinear function to effectively fuse complementary information from different modalities, so that the same part is more prominent when subsequent images are fused.

[0105] Preferably, if Figure 8 As shown, the common feature extraction submodule (i.e., nonlinear function) is used to calculate the common features of the new overexposure dictionary features and the new underexposure dictionary features to obtain the overexposure common dictionary features and the underexposure common dictionary features, including:

[0106] A common dictionary feature calculation is performed on the new overexposure dictionary feature and the new underexposure dictionary feature through a third nonlinear function to obtain an overexposure common dictionary feature, and the overexposure common dictionary feature is expressed as:

[0107]

[0108] in, For overexposed common dictionary features, is the new overexposure dictionary feature, is the new underexposure dictionary feature, s is the sth scale, E is the number of convolutional dictionaries (for overexposure), F is the number of convolutional dictionaries (for underexposure), W i s is a third nonlinear function for processing features of the s-th scale in the i-th iteration, used for extracting dictionary features biased towards overexposed images according to the bias parameter;

[0109] A common dictionary feature calculation is performed on the new overexposure dictionary feature and the new underexposure dictionary feature through a fourth nonlinear function to obtain an underexposure common dictionary feature, and the underexposure common dictionary feature is expressed as:

[0110]

[0111] in, is the underexposed common dictionary feature, Y i s is the fourth nonlinear function for processing the features of the sth scale in the i-th iteration, used to extract dictionary features biased towards underexposed images according to the bias parameter; the nonlinear function is Conv(ReLU(Conv(·))), used to and Perform common dictionary feature calculation, W i s and Y i s The same dimension is set when processing dictionary features, and the bias parameters are determined by randomly generated convolution layer parameters.

[0112] In the embodiment of the present invention, in order to better match the common sparse features, the feature dictionary is updated to extract the common features of two features according to the bias parameter so that it can adapt to the multimodal image currently being processed.

[0113] Preferably, if Figure 7 As shown, the feature fusion submodule extracts the overexposure common sparse features and the overexposure common dictionary features to obtain an overexposure fused image, including:

[0114] The overexposure common sparse features and the overexposure common dictionary features are convolved to obtain an overexposure fused image, which is expressed as:

[0115]

[0116] in, is the over-exposed fused image of the i-th iteration, is the overexposed fused image of the sth scale at the i-th iteration, is the overexposed common sparse features of the s-th scale image of the ith iteration learned by the e-th convolution dictionary, is the overexposed common dictionary feature of the s-th scale image of the ith iteration learned by the e-th convolutional dictionary, S is the number of image scales, E is the number of convolutional dictionaries, and * is the convolution operation.

[0117] Specifically, the feature fusion submodule includes a convolution layer composed of multiple convolution kernels, which combines the overexposed common sparse features of the first scale and overexposure common dictionary features Perform a convolution operation to obtain the overexposed fused image of the first scale of the first iteration The overexposed common sparse features of the second scale and overexposure common dictionary features Perform a convolution operation to obtain the overexposed fused image of the second scale of the first iteration Similarly, the overexposed common sparse features of the s-th scale are and overexposure common dictionary features Perform a convolution operation to obtain the overexposed fused image of the sth scale of the first iteration The overexposed fusion images of all scales of the first iteration are added element by element to obtain the overexposed fusion image of the first iteration.

[0118] like Figure 8 As shown, the process of extracting the underexposure common sparse features and the underexposure common dictionary features through the feature fusion submodule to obtain the underexposure fusion image is the same as the process of obtaining the overexposure fusion image through the feature fusion submodule, which will not be repeated here. The underexposure fusion image is expressed as:

[0119]

[0120] in, is the underexposed fused image of the i-th iteration, is the underexposed fused image of the sth scale at the ith iteration, is the underexposed common sparse features of the s-th scale image of the ith iteration learned by the f-th convolutional dictionary, is the underexposed common dictionary feature of the s-th scale image of the ith iteration learned by the f-th convolutional dictionary, S is the number of image scales, and F is the number of convolutional dictionaries.

[0121] In the embodiment of the present invention, features of multiple scales are fused through a convolution operation to generate a modal image that highlights common features while retaining non-common features of its own modality, so as to reconstruct a clearer image.

[0122] Preferably, the fusing the overexposed fused image and the underexposed fused image to obtain a reconstructed image includes:

[0123] The over-exposure fusion image and the under-exposure fusion image are weightedly calculated by a weighted fusion expression to obtain a reconstructed image. The weighted fusion expression is:

[0124]

[0125] Where Z is the reconstructed image, is the over-exposed fused image of the Nth iteration, is the underexposed fusion image of the Nth iteration, × is multiplication, W x is the weight parameter of the overexposed image, W y is the weight parameter of the underexposed image and satisfies W x +W y =1.

[0126] In the embodiment of the present invention, images of multiple modalities are fused according to weight parameters to reconstruct a target image for auxiliary diagnosis.

[0127] Preferably, the evaluating and calculating the reconstructed image and the standard image to obtain an evaluation index includes:

[0128] Calculating the reconstructed image and the standard image using a structural similarity calculation expression to obtain a structural similarity index;

[0129] The peak signal-to-noise ratio is calculated by calculating the peak signal-to-noise ratio of the reconstructed image and the standard image.

[0130] It should be understood that the structural similarity index (SSIM) is an indicator used to measure the similarity of two images, with a special focus on the structural information of the images, rather than just the similarity at the pixel level. The range of SSIM is 0 to 1, and the larger its value, the better the quality of the image. The peak signal-to-noise ratio (PSNR) is an objective indicator used to measure the difference between two images, mainly used to evaluate the effectiveness of image compression, transmission or reconstruction algorithms. The higher the PSNR value, the more similar the two images are and the smaller the quality loss.

[0131] In the embodiment of the present invention, the fusion effect of the present fusion technology is evaluated by the structural similarity index and the peak signal-to-noise ratio.

[0132] like Fig. 9 As shown, an embodiment of the present invention provides a deep multimodal image fusion system, comprising:

[0133] An image importing unit, used for importing a multi-exposure source image set, wherein the multi-exposure source image set includes an overexposed image, an underexposed image, and a standard image;

[0134] A feature pre-extraction unit, configured to perform feature extraction on the overexposed image to obtain overexposure sparse features and overexposure dictionary features, and perform feature extraction on the underexposed image to obtain underexposure sparse features and underexposure dictionary features;

[0135] a feature updating unit, configured to update and calculate the overexposure sparse feature and the overexposure dictionary feature through a total objective function to obtain a new overexposure sparse feature and a new overexposure dictionary feature, and to update and calculate the underexposure sparse feature and the underexposure dictionary feature through a total objective function to obtain a new underexposure sparse feature and a new underexposure dictionary feature;

[0136] a feature extraction unit, configured to perform common feature calculation on the new overexposure sparse feature and the new underexposure sparse feature through a nonlinear function to obtain an overexposure common sparse feature and an underexposure common sparse feature, and perform common feature calculation on the new overexposure dictionary feature and the new underexposure dictionary feature through a nonlinear function to obtain an overexposure common dictionary feature and an underexposure common dictionary feature;

[0137] a feature fusion unit, configured to perform feature extraction on the overexposure common sparse features and the overexposure common dictionary features to obtain an overexposure fused image, and perform feature extraction on the underexposure common sparse features and the underexposure common dictionary features to obtain an underexposure fused image;

[0138] The image reconstruction unit is used to perform a fusion process on the over-exposed fusion image and the under-exposed fusion image to obtain a reconstructed image, and to perform an evaluation calculation on the reconstructed image and the standard image to obtain an evaluation index.

[0139] The above-mentioned deep multimodal image fusion system can refer to the implementation content and beneficial effects of the deep multimodal image fusion method specifically described above, which will not be repeated here.

[0140] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0141] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0142] In the several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0143] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of the present invention.

[0144] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A deep multimodal image fusion method, characterized in that: The steps include: S1. Importing a multi-exposure source image set, wherein the multi-exposure source image set includes an overexposed image, an underexposed image, and a standard image; S2, performing feature extraction on the overexposed image to obtain overexposure sparse features and overexposure dictionary features, and performing feature extraction on the underexposed image to obtain underexposure sparse features and underexposure dictionary features; S3, updating and calculating the overexposure sparse feature and the overexposure dictionary feature through the total objective function to obtain a new overexposure sparse feature and a new overexposure dictionary feature, and updating and calculating the underexposure sparse feature and the underexposure dictionary feature through the total objective function to obtain a new underexposure sparse feature and a new underexposure dictionary feature; S4, performing common feature calculation on the new overexposure sparse feature and the new underexposure sparse feature through a nonlinear function to obtain an overexposure common sparse feature and an underexposure common sparse feature, and performing common feature calculation on the new overexposure dictionary feature and the new underexposure dictionary feature through a nonlinear function to obtain an overexposure common dictionary feature and an underexposure common dictionary feature; S5, performing feature extraction on the overexposure common sparse features and the overexposure common dictionary features to obtain an overexposure fused image, and performing feature extraction on the underexposure common sparse features and the underexposure common dictionary features to obtain an underexposure fused image; S6. Repeat S2-S5 until a preset number of iterations is reached, fuse the overexposed fused image and the underexposed fused image to obtain a reconstructed image, and evaluate and calculate the reconstructed image and the standard image to obtain an evaluation index.

2. The deep multimodal image fusion method according to claim 1, characterized in that: The extracting features of the overexposed image to obtain overexposure sparse features and overexposure dictionary features includes: The pre-constructed convolution dictionary is initialized, the initialized convolution dictionary is used as an overexposure dictionary feature, and the overexposure image is convolved by the overexposure dictionary feature to obtain an overexposure sparse feature.

3. The deep multimodal image fusion method according to claim 1, characterized in that: Before the step of updating and calculating the overexposure sparse features and the overexposure dictionary features by using the total objective function, the method further includes: The objective function of traditional dictionary learning is improved based on the overexposure sparse feature and the overexposure dictionary feature to obtain a total objective function, which is: Among them, L is the total objective function, S is the number of image scales, and E is the number of convolution dictionaries. To overexpose sparse features, is the overexposure dictionary feature, x is the overexposure image, is the Euclidean norm, λ is the regularization parameter, G( · ) is the regularization constraint function of sparse features, β is the regularization parameter, H( · ) is the regularization constraint function of the dictionary feature.

4. The deep multimodal image fusion method according to claim 3, characterized in that: The updating and calculating of the overexposure sparse feature and the overexposure dictionary feature by using the total objective function to obtain a new overexposure sparse feature and a new overexposure dictionary feature includes: The total objective function is decomposed into a sparse objective function and a dictionary objective function by using the alternating direction multiplier method; The overexposure sparse feature is calculated by the sparse objective function to obtain a new overexposure sparse feature, and the sparse objective function is: Among them, L1 is the sparse objective function, S is the number of image scales, and E is the number of convolution dictionaries. To overexpose sparse features, is the overexposure dictionary feature, is the Euclidean norm, λ is the regularization parameter, G(·) is the regularization constraint function of sparse features, α α is the regularization parameter, is the overexposed sparse auxiliary variable, x s is an overexposed image without dictionary learning, and i is the i-th iteration; The overexposure dictionary feature is calculated by the dictionary objective function to obtain a new overexposure dictionary feature, and the dictionary objective function is: Among them, L2 is the dictionary objective function, β is the regularization parameter, α d is the regularization parameter, is the auxiliary variable of the overexposure dictionary, H( · ) is the regularization constraint function of the dictionary feature.

5. The deep multimodal image fusion method according to claim 1, characterized in that: The method of calculating common features of the new overexposure sparse features and the new underexposure sparse features by using a nonlinear function to obtain overexposure common sparse features and underexposure common sparse features includes: The common sparse feature calculation is performed on the new overexposure sparse feature and the new underexposure sparse feature through the first nonlinear function to obtain an overexposure common sparse feature, and the overexposure common sparse feature is expressed as: in, is the common sparse feature of overexposure, is the new overexposed sparse feature, is the new underexposed sparse feature, s is the sth scale, E and F are the number of convolutional dictionaries, is the first nonlinear function, used to extract sparse features biased towards overexposed images according to the bias parameter, and i is the i-th iteration; The new overexposure sparse feature and the new underexposure sparse feature are calculated using a second nonlinear function to obtain an underexposure common sparse feature. The underexposure common sparse feature is expressed as: in, is the common sparse feature of underexposure, is the second nonlinear function, which is used to extract sparse features that are biased towards underexposed images according to the bias parameter; the nonlinear function is Conv(ReLU(Conv(·))), which is used to and Perform common sparse feature calculations, and The same dimension is set when processing sparse features, and the bias parameters are determined by randomly generated convolution layer parameters.

6. The deep multimodal image fusion method according to claim 1, characterized in that: The method of calculating common features of the new overexposure dictionary features and the new underexposure dictionary features by using a nonlinear function to obtain overexposure common dictionary features and underexposure common dictionary features includes: A common dictionary feature calculation is performed on the new overexposure dictionary feature and the new underexposure dictionary feature through a third nonlinear function to obtain an overexposure common dictionary feature, and the overexposure common dictionary feature is expressed as: in, For overexposed common dictionary features, is the new overexposure dictionary feature, is the new underexposed dictionary feature, s is the sth scale, E and F are the number of convolutional dictionaries, W i s is the third nonlinear function, used to extract dictionary features biased towards overexposed images according to the bias parameter, and i is the i-th iteration; A common dictionary feature calculation is performed on the new overexposure dictionary feature and the new underexposure dictionary feature through a fourth nonlinear function to obtain an underexposure common dictionary feature, and the underexposure common dictionary feature is expressed as: in, is the underexposed common dictionary feature, Y i s is the fourth nonlinear function, used to extract dictionary features biased towards underexposed images according to the bias parameter; the nonlinear function is Conv(ReLU(Conv(·))) and Perform common dictionary feature calculation, W i s and Y i s The same dimension is set when processing dictionary features, and the bias parameters are determined by randomly generated convolution layer parameters.

7. The deep multimodal image fusion method according to claim 1, characterized in that: The extracting features of the overexposure common sparse features and the overexposure common dictionary features to obtain an overexposure fused image includes: The overexposure common sparse features and the overexposure common dictionary features are convolved to obtain an overexposure fused image, which is expressed as: in, is the over-exposed fused image of the i-th iteration, is the common sparse feature of overexposure, is the overexposed common dictionary feature, S is the number of image scales, and E is the number of convolutional dictionaries.

8. The deep multimodal image fusion method according to claim 1, characterized in that: The fusing the overexposed fused image and the underexposed fused image to obtain a reconstructed image includes: The over-exposure fusion image and the under-exposure fusion image are weightedly calculated by a weighted fusion expression to obtain a reconstructed image. The weighted fusion expression is: Among them, Z is the reconstructed image, W x is the weight parameter of the overexposed image, is the over-exposed fused image of the Nth iteration, W y is the weight parameter of the underexposed image, is the underexposed fused image of the Nth iteration.

9. The deep multimodal image fusion method according to claim 1, characterized in that: The evaluating and calculating the reconstructed image and the standard image to obtain an evaluation index includes: Calculating the reconstructed image and the standard image using a structural similarity calculation expression to obtain a structural similarity index; The peak signal-to-noise ratio is calculated by calculating the peak signal-to-noise ratio of the reconstructed image and the standard image.

10. A deep multimodal image fusion system, characterized in that: include: An image importing unit, used for importing a multi-exposure source image set, wherein the multi-exposure source image set includes an overexposed image, an underexposed image, and a standard image; A feature pre-extraction unit, configured to perform feature extraction on the overexposed image to obtain overexposure sparse features and overexposure dictionary features, and perform feature extraction on the underexposed image to obtain underexposure sparse features and underexposure dictionary features; a feature updating unit, configured to update and calculate the overexposure sparse feature and the overexposure dictionary feature through a total objective function to obtain a new overexposure sparse feature and a new overexposure dictionary feature, and to update and calculate the underexposure sparse feature and the underexposure dictionary feature through a total objective function to obtain a new underexposure sparse feature and a new underexposure dictionary feature; a feature extraction unit, configured to perform common feature calculation on the new overexposure sparse feature and the new underexposure sparse feature through a nonlinear function to obtain an overexposure common sparse feature and an underexposure common sparse feature, and perform common feature calculation on the new overexposure dictionary feature and the new underexposure dictionary feature through a nonlinear function to obtain an overexposure common dictionary feature and an underexposure common dictionary feature; a feature fusion unit, configured to perform feature extraction on the overexposure common sparse features and the overexposure common dictionary features to obtain an overexposure fused image, and perform feature extraction on the underexposure common sparse features and the underexposure common dictionary features to obtain an underexposure fused image; The image reconstruction unit is used to perform a fusion process on the over-exposed fusion image and the under-exposed fusion image to obtain a reconstructed image, and to perform an evaluation calculation on the reconstructed image and the standard image to obtain an evaluation index.

Citation Information

Patent Citations

  • Real exposure correction method and system guided by camera perception characteristics

    CN116614714A

  • Multi-exposure image fusion method based on low-rank decomposition and sparse representation

    CN117291851A

  • Reconstruction of high-quality images from a binary sensor array

    US20170272639A1

Cited By

  • Learned dictionary based warp blend

    US20260105565A1