A multi-focus image fusion method and system
By employing dense feature extraction, local-global joint attention, and feature fusion modules in a multi-focus image fusion system, the artifacts and blurring problems in fused images of traditional methods are solved, generating high-quality fused images consistent with the source images. This system is suitable for fusion of multi-focus images, medical images, infrared and visible light images, as well as multi-exposure images.
Patent Information
- Application Number
- CN202310760377.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-27
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2043-06-27
AI Technical Summary
In traditional multi-focus image fusion methods, the fused image exhibits artifacts and blurring effects in the focused area and at the defocus boundary. Furthermore, the fused image cannot maintain the same level of realism as the source image, such as differences in brightness and contrast, or even the presence of noise.
A multi-focus image fusion system is adopted, including a dense feature extraction module, a local-global joint attention module, and a feature fusion module. Through dense feature extraction, local-global joint attention, and feature fusion, an accurate decision mapping is generated. The decision map guides the fusion process to avoid distortion, and the feature representation capability of the network is improved through the local-global joint attention module.
The generated fused image is consistent with the source image, with less blurring and artifacts in the focus/defocus boundary areas, resulting in higher realism and clarity. Experiments have verified that it outperforms other methods in the field of multi-focus image fusion.
Smart Images

Figure CN117036874B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image fusion technology, and specifically to a multi-focus image fusion method and system. Background Technology
[0002] In digital photography, multi-focus image fusion combines two or more images with different areas of sharpness into a single, fully sharp composite image. Traditional multi-focus image fusion methods can be broadly classified into two categories: transform domain-based methods and spatial domain-based methods. Transform domain methods typically first process the source image using a decomposition algorithm, transforming it from its original spatial dimension to another abstract spatial dimension. Then, the transform domain coefficients, processed by a certain fusion criterion, are inversely reconstructed to form the final fused image. Spatial domain methods, on the other hand, directly process the source image, applying mathematical operators directly to the pixel values of the source image for fusion.
[0003] However, the fused images obtained by the traditional multi-focus image fusion methods mentioned above have artifacts and blurring effects in the focused area and defocus boundary. In addition, the fused images cannot have the same realism as the source images. For example, the brightness and contrast of the fused images are different from those of the source images, and they may even contain noise. Summary of the Invention
[0004] The purpose of this invention is to provide a multi-focus image fusion method and system that solves the problems of artifacts and blurring effects in the focused area and defocus boundary of the fused image obtained by the traditional multi-focus image fusion method in the prior art. In addition, the fused image cannot have the same realism as the source image. For example, the brightness and contrast of the fused image are different from the source image, and it may even contain noise.
[0005] To achieve the above technical objectives, the technical solution adopted by the present invention is as follows: The present invention provides a multi-focus image fusion system, including a dense feature extraction module, a local-global joint attention module, and a feature fusion module;
[0006] The dense feature extraction module is used to extract dense features from the input image to be fused and to calculate the mean of the extracted dense features;
[0007] The local-global joint attention module takes the mean of dense features as input and outputs a new feature map.
[0008] The feature fusion module is used to perform feature fusion on the feature map to output a focused map, and then transform the output focused map into a decision map.
[0009] The dense feature extraction module has a network structure with nine layers, seven of which are “convolutional layers + batch normalization + activation function” networks, and the other two are tensor concatenation.
[0010] Tensor splicers are embedded between the third and fourth layers of the dense feature extraction module, and between the fourth and fifth layers. The outputs of the second and third layers are respectively input to two tensor splicing modules to achieve the skip connection function.
[0011] The feature fusion module has a network structure consisting of six layers. The first layer is a tensor concatenation layer, the second to fourth layers are all "transposed convolutional layer + batch normalization + activation function" networks, and the last layer is a "convolutional layer + activation function" network.
[0012] A multi-focus image fusion method includes the following steps:
[0013] The dense features of the input images A and B to be fused are extracted by the dense feature extraction module.
[0014] Calculate the mean of the extracted dense features;
[0015] Using the mean of dense features as input, new feature maps are output through two local-global joint attention module paths;
[0016] The feature map is input into the feature fusion module for feature fusion, and finally the focused map is output.
[0017] The focus map output by the feature fusion module is transformed into a decision map after binary segmentation and cell filtering, and finally a fused image is obtained.
[0018] Specifically, the dense features of the input images A and B to be fused are extracted by the dense feature extraction module, including:
[0019] The input image is channel-separated and then fed into six parallel sub-network branches;
[0020] Extract initial features of the image using large convolution kernels in the branch;
[0021] The extracted initial features are sent to four convolutional layers with small kernels for feature analysis and extraction.
[0022] The extracted initial features are sent to four convolutional layers with small kernels for feature analysis and extraction, including:
[0023] Dense connections are introduced between the four convolutional layers to improve the flow of feature information between layers.
[0024] The focus map output by the feature fusion module is transformed into a decision map after binary segmentation and cell filtering, ultimately resulting in a fused image, including:
[0025] By utilizing the weight distribution provided by the decision graph, a weighted fusion strategy is adopted for the source images to obtain the final fused image.
[0026] This invention discloses a multi-focus image fusion method and system. The invention proposes a multi-focus image fusion network (LGAM-Net) with a local-global joint attention module. By utilizing a decision graph to guide the fusion process, it avoids image distortion in the fused image. Simultaneously, the local-global joint attention module enhances the network's feature representation capability, resulting in more accurate decision mappings. Extensive experiments comparing this method with other multi-focus image fusion methods demonstrate that the fused images generated on the Lytro, MFI-WHU, and MFFW test datasets exhibit consistent realism with the source images, exhibiting less blurring and artifacts in the focus / defocus boundary regions. Ablation experiments and module replacement experiments were designed to verify the rationality and effectiveness of the LGAM-Net and LGAM structure design. Results show that with the assistance of LGAM, accurate decision mappings can be generated with fewer misclassified regions. This invention can be applied to other image fusion fields, such as medical image fusion, infrared and visible light image fusion, and multi-exposure image fusion. Attached Figure Description
[0027] The present invention can be further illustrated by the non-limiting embodiments given in the accompanying drawings.
[0028] Figure 1 This is a structural block diagram of the multi-focus image fusion system of the present invention.
[0029] Figure 2 This is a schematic diagram of the network structure of the dense feature extraction module of the present invention.
[0030] Figure 3 This is a flowchart of the algorithm for extracting the mean value of dense features in the dense feature extraction module of the present invention.
[0031] Figure 4 This is a schematic diagram of the local-global joint attention module network structure of the present invention.
[0032] Figure 5 The image obtained is a comparison of the fusion effect of the multi-focus image obtained by the present invention with the fusion effect of 15 existing fusion methods.
[0033] Figure 6 and Figure 7 This is a diagram showing the effect of LGAM-Net of the present invention on focusing and decision mapping compared with four attention modules: SSE, CSE, SCSE and CBAM.
[0034] Figure 8 This is a diagram illustrating the multi-focus source image and multi-quantity fusion effect of the present invention.
[0035] Figure 9This is a flowchart illustrating the steps of the multi-focus image fusion method of the present invention. Detailed Implementation
[0036] To enable those skilled in the art to better understand the present invention, the technical solution of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0037] Please see Figure 1 , Figure 2 and Figure 4 The present invention provides a multi-focus image fusion system, characterized in that it includes a dense feature extraction module, a local-global joint attention module, and a feature fusion module;
[0038] The dense feature extraction module is used to extract dense features from the input image to be fused and to calculate the mean of the extracted dense features;
[0039] The local-global joint attention module takes the mean of dense features as input and outputs a new feature map.
[0040] The feature fusion module is used to perform feature fusion on the feature map to output a focused map, and then transform the output focused map into a decision map.
[0041] For further details, please refer to Figure 2 The network structure of the dense feature extraction module contains nine layers, seven of which are “convolutional layers + batch normalization + activation function” networks, and the other two are tensor concatenation.
[0042] For further details, please refer to Figure 2 Tensor splicers are embedded between the third and fourth layers, and between the fourth and fifth layers. The outputs of the second and third layers are respectively input to two tensor splicing modules to realize the jump connection function.
[0043] For further details, please refer to Figure 4 The network structure of the feature fusion module contains six layers. The first layer is tensor concatenation, the second to fourth layers are all "transposed convolutional layer + batch normalization + activation function" networks, and the last layer is a "convolutional layer + activation function" network.
[0044] In this embodiment, the present invention proposes a multi-focus image fusion network (LGAM-Net) with a local-global joint attention module. The local-global joint attention module generates focus maps. To extract as much initial feature information as possible, the dense feature extraction module (DFEM) employs convolutional kernels of different sizes, max-pooling layers, and skip connections to extract feature information. Then, the features extracted by the DFEM are input into the local-global joint attention module (LGAM), which uses pointwise convolution to extract local information from the feature map. Simultaneously, it uses a spatial pyramid pooling strategy to retain global information from multiple scales. A joint attention map can be obtained by integrating local and global information. Finally, the output of the local-global joint attention module (LGAM) is sent to the feature fusion module (FFM), which can reduce the dimensionality of features and fuse feature maps. Extensive experimental results show that, in terms of both subjective visual effect and objective measurement of the fused image, the method of the present invention outperforms other multi-focus image fusion methods.
[0045] Please see Figure 3 , Figure 5 , Figure 6 , Figure 7 , Figure 8 and Figure 9 A multi-focus image fusion method includes the following steps:
[0046] S100: Extract dense features from the input images A and B to be fused using the dense feature extraction module;
[0047] Specifically, the input image is separated into channels and then fed into six parallel sub-network branches. The initial features of the image are extracted using the large convolutional kernels of the branches. The extracted initial features are then sent to four convolutional layers with small convolutional kernels for feature analysis and extraction. Dense connections are introduced between the four convolutional layers to improve the flow of feature information between layers.
[0048] S101: Calculate the mean of the extracted dense features;
[0049] Extract dense features from the images to be fused.
[0050] The input images to be fused are A and B. Features of the input images are extracted by two Dense Feature Extraction Modules (DFEMs). The two DFEMs are independent but share network weights. The network structure of the DFEM is shown in the figure below. The network structure of the Dense Feature Extraction Module contains nine layers, seven of which are "convolutional layers + batch normalization + activation function" networks, and the other two are tensor concatenation layers. Tensor concatenation layers are embedded between the third and fourth layers, and between the fourth and fifth layers. The outputs of the second and third layers are respectively input to the two tensor concatenation modules to achieve skip connections.
[0051] The parameters of the convolutional network in DFEM are shown in the table below.
[0052] hierarchy Input and output channel dimensions kernel size Step length Conv1 (1,8,256,256) 7 1 Conv2 (8,8,128,128) 3 2 Conv3 (8,16,128,128) 1 1 Conv4 (24,48,128,128) 1 1 Conv5 (72,128,128,128) 1 1 Conv6 (128,128,128,128) 1 1 Conv7 (128,256,128,128) 1 1
[0053] Next, the mean of the extracted dense features is calculated. This mean is calculated based on the outputs of the two dense feature extraction networks DFEM, and the formula is as follows:
[0054]
[0055]
[0056] In this context, FA and FB represent the mean values of two sets of dense feature maps, FA1, FA2, and FA3 represent the features extracted by the three identical parallel sub-networks in the first set, and FB1, FB2, and FB3 represent the features extracted by the three identical parallel sub-networks in the second set.
[0057] S102: Using the mean of dense features as input, new feature maps are output through two local-global joint attention module paths;
[0058] Calculate the feature map where local and global features are blended.
[0059] Using the mean of dense features as input, new feature maps are output through two paths of the Local-Global Attention Module (LGAM). These feature maps exhibit the characteristic of blending local and global features, with the two LGAM channels sharing weight parameters. The network structure of LGAM is shown in the attached figure. Figure 2 As shown.
[0060] The convolutional parameters of each layer in LGAM are shown in the table below.
[0061] hierarchy Input and output channel dimensions kernel size Step length Conv8 (256,256,64,64) 3 1 Conv9 (256,256,64,64) 3 1 PW Conv1_1 (256,16,64,64) 1 1 PW Conv1_2 (16,256,64,64) 1 1 PW Conv2_1 (256,16,64,64) 1 1 PW Conv22 (16,256,64,64) 1 1 PW Conv41 (256,16,64,64) 1 1 PW Conv42 (16,256,64,64) 1 1 Conv10 (768,48,64,64) 1 1 Conv11 (48,256,64,64) 1 1
[0062] For ease of expression, given the input X of LGAM, Taking X as input, we first extract the intermediate feature T(X). Let B(·) denote batch normalization, and l(·) denote the activation function. Both Conv8(·) and Conv9(·) represent convolutions with a kernel size of 3×3 and a stride of 1. The calculation process of the intermediate feature T(X) can then be defined as follows:
[0063] T(X)=B(Conv8(l(B(Conv9(X))))) (3)
[0064] Let P(·,·) denote the adaptive average pooling operation, then P(T(X),k) represents adaptive average pooling of T(X) with an output size of k×k. Therefore, by applying the spatial pyramid pooling strategy to T(X), we can obtain global features G1(T(X),1), G2(T(X),2), and G4(T(X),4) at different scales. The calculation process is as follows:
[0065] G1(T(X),1)=P(T(X),1) (4)
[0066] G2(T(X),2)=P(T(X),2) (5)
[0067] G4(T(X),4)=P(T(X),4) (6)
[0068] Where k = {1, 2, 4}, C represents the number of channels.
[0069] Let local features Where C represents the number of channels, and H and W represent the height and width of the local feature map, respectively. They can be calculated using three independent bottleneck structures:
[0070] L 1(X) =B(PWConv 1_2 (l(B(PWConv 1_1 (T(X)))))) (7)
[0071] L2(X)=B(PWConv 2_2 (l(B(PWConv 2_1 (T(X)))))) (8)
[0072] L4(X)=B(PWConv 4_2 (l(B(PWConv 4_1 (T(X))))) (9)
[0073] To fuse the global features {G1(T(X),1),G2(T(X),2),G4(T(X),4)} with the local features {L1(X),L2(X),L4(X)}, it is necessary to first upsample G1(T(X),1), G2(T(X),2), and G4(T(X),4) to make G... n (T(X),n) matches L n The spatial dimensions are (X) (n = 1, 2, 4). Then, element-wise addition is used to initially fuse the upsampled results and local information to obtain an initial attention feature map. Subsequently, the initial attention feature maps are concatenated along the channel dimension using the Cat(·) concatenation operation. Two convolutional layers are then used to achieve channel dimensionality reduction and feature fusion, thus realizing attention feature fusion. The fusion process is as follows:
[0074] S(X)=Cat(U(G1(T(X),1),G2(T(X),2),G4(T(X),4))) (10)
[0075] w=δ(B(PWConv4(l(B(PWConv3(S(X)))))))) (11)
[0076]
[0077] Where PWConv3(·) and PWConv4(·) represent convolutional layers used for dimensionality reduction, and U(·) represents the upsampling operation. Furthermore, δ(·) and w represent the dot product, the activation function, and the joint attention feature map with a value range of [0,1], respectively. X' is the output of the current LGAM module, which is calculated from the input X via equation (12).
[0078] S103: Input the feature map into the feature fusion module for feature fusion, and finally output the focus map;
[0079] Generate a focus image.
[0080] The feature map obtained from the local and global feature fusion output by LGAM is concatenated with feature tensors and then input into the Feature Fusion Module (FFM) for feature fusion. Finally, a focused map is output. The network structure diagram of FFM is attached. Figure 4 As shown, the network structure of the feature fusion module consists of six layers. The first layer is tensor concatenation, the second to fourth layers are all "transposed convolutional layer + batch normalization + activation function" networks, and the last layer is a "convolutional layer + activation function" network.
[0081] S104: The focus map output by the feature fusion module is transformed into a decision map after binary segmentation and cell filtering, and finally a fused image is obtained;
[0082] Specifically, by utilizing the weight distribution provided by the decision graph, a weighted fusion strategy is adopted for the source images to obtain the final fused image.
[0083] Generate a fused image.
[0084] The feature fusion module consists of four transposed convolutional layers and one convolutional layer connected in series. The dimension of the features decreases layer by layer, which is beneficial for information fusion.
[0085] The focus map output by FFM is transformed into a decision map after binary segmentation and small-region filtering. Using the weight distribution provided by the decision map, a weighted fusion strategy is applied to the source images to obtain the final fused image. The calculation formula for generating the decision map is as follows:
[0086]
[0087] Here, Dmp(x,y) represents the decision map, which is essentially a result of binary segmentation; Fmp(x,y) represents the focus map output by the feature fusion module.
[0088] A small-region filter with a threshold of 0.001×H×W is used to refine Dmp(x,y). Then, a weighted fusion is performed on source image A(x,y) and source image B(x,y) to obtain the fused image I. f (x, y). Symbol If we denote dot product, then the image fusion process is as follows:
[0089]
[0090] To verify that the proposed method outperforms other methods, its performance was evaluated on publicly available test datasets. The proposed method was compared with 15 other methods, including 10 deep learning-based methods and 5 traditional methods.Deep learning methods include Convolutional Neural Networks (CNNs), Ensemble of CNNs (ECNNs), Fusion Densely Connected Networks (FusionDN), Image Fusion Framework Based on Convolutional Neural Networks (IFCNNs), Squeeze and Excitation and Spatial Frequency (SESF), Multi-focus Image Fusion with Generative Adversarial Network (MFF-GAN), Generative Adversarial Network for Multi-focus Image Fusion (MFIF-GAN), Unity Fusion Attention for Fusion (UFA-FUSE), Unified Unsupervised Image Fusion (U2Fusion), and Gradient Aware Cascade Structure Network (GACN); traditional methods are image fusion methods based on image matting. Image fusion methods include Matting (IFM), Guided Filtering (GFF), Multi-scale Weighted Gradient-based Fusion for Multi-focus Image (MWGF), Dense Scale Invariant Feature Transform (DSIFT), and Convolutional Sparse Representation (CSR).
[0091] As attached Figure 5As shown, by magnifying local details in the fused image, the magnified details belong to the focus / defocus boundary region. It can be seen that... Figure 5 The fused images generated by methods such as (j), 5(l), 5(n), and 5(o) exhibit significant artifacts and blurring in the focus / defocus boundary regions. Compared to other fused images, the fusion result obtained using this invention is clearer in the boundary regions with fewer artifacts and blurring. Furthermore, some multi-focus image fusion methods produce fused images that differ significantly in realism from the source images. This difference manifests specifically in variations in brightness and contrast, such as in 5(g). Compared to these methods that compromise the realism of the fused image, this invention offers a significant advantage. First, a focus map is generated using a network model. Then, a decision map is obtained from the focus map and used to guide the fusion process. This ensures that all information in the fused image originates entirely from source images A and B. Therefore, the fusion result of this invention is highly faithful, maintaining the same brightness and contrast as the source images.
[0092] Quantitative evaluation refers to using objective evaluation metrics to measure the performance of fused images. This invention conducted comparative experiments on the Lytro, MFIWHU, and MFFW test datasets. Higher values for the eight metric values used in the experiments indicate higher quality fused images. Detailed comparison results of different methods are shown in Tables 1(a), 1(b), and 1(c).
[0093] It is evident that the method of this invention achieved first place in the index values on the Lytro, MFI-WHU, and MFFW datasets 7, 6, and 8 times, respectively. For a clearer comparison, the index values obtained by different methods on the three datasets were evenly distributed. Table 1 (d) shows the average results of the index values obtained by different methods on three datasets. It can be seen that the method used in this invention is significantly better than the other 15 methods.
[0094] Table 1:
[0095]
[0096]
[0097]
[0098] To verify the rationality and effectiveness of the network architecture design, three ablation experiments were conducted to investigate the impact of LGAM. During the ablation experiments, the learning rate, pruning size, number of training epochs, optimizer, and hardware were kept consistent with those used during the training phase. Table 2 shows the three different LGAM design approaches. In Experiment 1, the original residual blocks were used instead of LGAM. For Experiments 2 and 3, the design structures were "residual blocks + local feature extraction" and "residual blocks + spatial pyramid pooling," respectively, with LGAM used in LGAM-Net.
[0099] Table 2:
[0100] Includes residual blocks Includes local feature extraction Includes spatial pyramid pooling Experiment 1 √ Experiment 2 √ √ Experiment 3 √ √ LGAM-Net √ √ √
[0101] To fully demonstrate the impact of the complete LGAM in the network, objective metrics were used to prove the rationality and effectiveness of the network architecture design. Five metrics were selected to quantify the performance of networks with different design architectures on publicly available test datasets after 240 training epochs. Tables 3(a), 3(b), and 3(c) clearly show that the quantification metrics for LGAM achieved 5, 3, and 2 best times on the Lytro, MFI-WHU, and MFFW test datasets, respectively, and 1 and 3 second-best times on the MFI-WHU and MFFW datasets, respectively. Table 3(d) shows the average quantification metrics for LGAM and other ablation experiments. It is evident that... Table 3 (d) The model with LGAM involved, namely LGAM-Net, achieved first place in Q_CB, Q_NCIE, Q_SG and Q_ABF and second place in SSIM. Therefore, the LGAM design structure is more reasonable and more effective.
[0102] Table 3:
[0103]
[0104]
[0105] Introducing attention modules can enhance the feature representation capabilities of a network. To verify the benefits of the designed LGAM to the network, this invention guided a module replacement experiment based on the idea of controlling variables. Specifically, four commonly used attention modules previously introduced in the above methods—SSE, CSE, SCSE, and CBAM—were embedded into the network, replacing LGAM during the embedding process. Then, the network was retrained and tested on three multi-focus image fusion datasets: Lytro, MFI-WHU, and MFFW.
[0106] To fully demonstrate LGAM's superiority over other attention modules, subjective visual effects and objective metrics were used for evaluation. (See attached document.) Figure 6 and attached Figure 7 These are the 30th image from MFI-WHU and the 4th image from MFFW, respectively. It can be seen that the areas marked with boxes are misclassified regions. (See attached image.) Figure 6 (c),6(d), Figure 6 (e), Figure 6 (f), Appendix Figure 7 (c), Figure 7 (d), Figure 7 (e), Figure 7 As shown in (f), with the introduction of Spatial Squeeze and Channel Excitation (sSE), Channel Squeeze and Spatial Excitation (cSE), Concurrent Spatial and Channel Squeeze and Excitation (scSE), and the Convolutional Block Attention Module (CBAM), the final decision mapping is accompanied by some obvious region misjudgments. (See attached...) Figure 6 (g) and Figure 7 As shown in (g), with the help of LGAM, the quality of the obtained focusing and decision mapping is higher, there are fewer misclassifications, and the boundary regions have no obvious bumps.
[0107] Table 4 presents the objective metrics values of networks with different attention modules on three test sets.
[0108] As can be seen, the network of this invention ranked first 14 times and second 5 times in terms of metric values. Furthermore, LGAM-Net ranked first 6 times in terms of the average of the 8 metrics. Therefore, overall, LGAMNet outperforms models that incorporate other attention modules in terms of metric values.
[0109] Table 4:
[0110]
[0111] To verify that the proposed method is also applicable to cases with more than two multi-focus source images, this invention was tested on a sequence of three multi-focus source images. Specifically, the intermediate result obtained by fusing two of the source images was fused with the third source image to obtain the final fused image. The test results are attached. Figure 8As shown, each column from left to right represents source image 1, source image 2, source image 3, and the fused image, respectively. It is evident that the proposed method can perform sequential multi-focus image fusion very well. The fused image is a fully sharp image that includes the focus areas of the source images.
[0112] This invention discloses a multi-focus image fusion method and system. The invention proposes a multi-focus image fusion network (LGAM-Net) with a local-global joint attention module. By utilizing a decision graph to guide the fusion process, it avoids image distortion in the fused image. Simultaneously, the local-global joint attention module enhances the network's feature representation capability, resulting in more accurate decision mappings. Extensive experiments comparing this method with other multi-focus image fusion methods demonstrate that the fused images generated on the Lytro, MFI-WHU, and MFFW test datasets exhibit consistent realism with the source images, exhibiting less blurring and artifacts in the focus / defocus boundary regions. Ablation experiments and module replacement experiments were designed to verify the rationality and effectiveness of the LGAM-Net and LGAM structure design. Results show that with the assistance of LGAM, accurate decision mappings can be generated with fewer misclassified regions. This invention can be applied to other image fusion fields, such as medical image fusion, infrared and visible light image fusion, and multi-exposure image fusion.
[0113] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.
Claims
1. A multi-focus image fusion system, characterized in that, It includes a dense feature extraction module, a local-global joint attention module, and a feature fusion module; The dense feature extraction module is used to extract dense features from the input image to be fused and to calculate the mean of the extracted dense features; The local-global joint attention module takes the mean of dense features as input and outputs a new feature map. The feature fusion module is used to perform feature fusion on the feature map to output a focused map, and then transform the output focused map into a decision map.
2. The multi-focus image fusion system as described in claim 1, characterized in that, The network structure of the dense feature extraction module contains nine layers, seven of which are "convolutional layers + batch normalization + activation function" networks, and the other two are tensor concatenation.
3. The multi-focus image fusion system as described in claim 2, characterized in that, Tensor splicers are embedded between the third and fourth layers of the dense feature extraction module, and between the fourth and fifth layers. The outputs of the second and third layers are respectively input to two tensor splicing modules to achieve the skip connection function.
4. The multi-focus image fusion system as described in claim 1, characterized in that, The network structure of the feature fusion module consists of six layers. The first layer is tensor concatenation, the second to fourth layers are all "transposed convolutional layer + batch normalization + activation function" networks, and the last layer is a "convolutional layer + activation function" network.
5. A multi-focus image fusion method, applied to the multi-focus image fusion system as described in any one of claims 1 to 4, characterized in that, Includes the following steps: The dense features of the input images A and B to be fused are extracted by the dense feature extraction module. Calculate the mean of the extracted dense features; Using the mean of dense features as input, new feature maps are output through two local-global joint attention module paths; The feature map is input into the feature fusion module for feature fusion, and finally the focused map is output. The focus map output by the feature fusion module is transformed into a decision map after binary segmentation and cell filtering, and finally a fused image is obtained.
6. The multi-focus image fusion method as described in claim 5, characterized in that, The dense feature extraction module extracts dense features from the input images A and B to be fused, including: The input image is channel-separated and then fed into six parallel sub-network branches; Extract initial features of the image using large convolution kernels in the branch; The extracted initial features are sent to four convolutional layers with small kernels for feature analysis and extraction.
7. The multi-focus image fusion method as described in claim 6, characterized in that, The extracted initial features are sent to four convolutional layers with small kernels for feature analysis and extraction, including: Dense connections are introduced between the four convolutional layers to improve the flow of feature information between layers.
8. The multi-focus image fusion method as described in claim 5, characterized in that, The focus map output by the feature fusion module is transformed into a decision map after binary segmentation and cell filtering, ultimately yielding a fused image, including: By utilizing the weight distribution provided by the decision graph, a weighted fusion strategy is adopted for the source images to obtain the final fused image.
Citation Information
Patent Citations
Multi-focus image fusion method based on multi-scale feature interaction network
CN113705675A
Multi-focus image fusion method combining depth context and convolution conditional random field
CN113763300A