A multi-focus image fusion method, system, device and storage medium

By building a fusion network model, using the hollow convolution network and attention mechanism to extract and fuse the features of multi-focus images, the problem of the importance of single and significant features in the prior art is solved, and high-quality multi-focus image fusion is achieved.

CN114627035BActive Publication Date: 2025-06-10NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210116063.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-29
Publication Date
2025-06-10
Estimated Expiration
2042-01-29

AI Technical Summary

Technical Problem

In the existing multi-focus image fusion technology, the importance of extract extract is single and the importance of significant features cannot be effectively highlighted, resulting in insufficient detailed information of the fusion image and poor visual quality.

Method used

A fusion network model is constructed using hollow convolution network and dense convolution neural network based on attention mechanism, multi-scale features are extracted through hollow convolution expansion receptive field, and significant features are adaptively selected through attention mechanism to perform image fusion.

Benefits of technology

Remaining rich detailed information during the multi-focus image fusion process, significantly improving the visual quality of the fusion image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114627035B_ABST
    Figure CN114627035B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-focus image fusion method, system, device and storage medium, belonging to the technical field of multi-focus image fusion. The method includes: acquiring multi-focus images to be fused; inputting the multi-focus images into a pre-trained fusion network model for fusion to obtain a fused image; the fusion network model is constructed by the following method: constructing a fusion network model by using a dilated convolutional network and a densely convolutional neural network based on an attention mechanism; solving the defect in the prior art that the feature extraction of the source image is single and the importance of significant features cannot be effectively highlighted, realizing the retention of rich detail information during multi-focus image fusion, and the visual quality of the fused image obtained by fusion is relatively excellent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multi - focus image fusion method, system, device and storage medium, belonging to the technical field of multi - focus image fusion. Background Art

[0002] During the digital image acquisition process, due to the limitation of the depth of field of the optical sensor lens, it is difficult for the camera system to obtain a panoramic - depth image during imaging. There is often a problem that the scenery within the depth of field is clear, while the scenery outside the depth of field is blurred. However, some unclear areas are not conducive to subsequent image understanding and applications, and may also cause certain errors. The multi - focus image fusion technology can effectively solve this problem by complementarily combining images of the same target scene with different focused regions, effectively solving the problem of clear imaging of the entire scene, making it contain richer information on the same image and meeting the requirements of actual use. At present, the multi - focus image fusion technology has played a crucial role in fields such as machine vision, remote sensing monitoring, and military medicine.

[0003] Existing multi - focus image fusion technologies can be mainly divided into transform - domain methods, spatial - domain methods, and deep - learning methods. The transform - domain methods usually first decompose the original image into different transform coefficients, then fuse these transform coefficients through corresponding fusion rules, and finally perform an inverse transform on the fused coefficients to obtain the fused image. The spatial - domain methods directly perform fusion operations on the pixels or regions of the source images to extract clear pixels in the focused regions. However, both the transform - domain methods and the spatial - domain methods require artificial design of the activity level measurement of significant information and fusion rules, which limits the universality of the fusion algorithm to a certain extent. In recent years, due to the powerful feature extraction and data representation capabilities of deep learning, the multi - focus image fusion technology based on deep learning has also become popular. Currently, the key of most multi - focus image fusion technologies based on deep learning is to use a convolutional neural network to accurately detect the focused regions from multi - focus source images, and then fuse the focused regions from different source images to generate a full - scene clear image. Compared with traditional multi - focus image fusion technologies, the multi - focus image fusion technology based on deep learning has improved the fusion quality to a certain extent, but there are still some limitations, mainly reflected in: 1. The feature extraction scale of the source image is single; 2. It cannot effectively highlight the importance of significant features. Summary of the Invention

[0004] The purpose of the present invention is to provide a multi - focus image fusion method, system, device and storage medium, to solve the defects in the prior art that the feature extraction of the source image is single and cannot effectively highlight the importance of significant features, and to achieve retaining rich detail information during multi - focus image fusion, and the visual quality of the fused image obtained is relatively excellent.

[0005] To achieve the above - mentioned purpose, the present invention is implemented by adopting the following technical solutions:

[0006] In the first aspect, the present invention provides a multi-focus image fusion method, including:

[0007] Obtain multi-focus images to be fused;

[0008] Input the multi-focus images into a pre-trained fusion network model for fusion to obtain a fused image;

[0009] The fusion network model is constructed by the following method: constructing the fusion network model using a dilated convolutional network and a dense convolutional neural network based on an attention mechanism.

[0010] Combined with the first aspect, further, it further includes the step of establishing a training data set, including:

[0011] Extract images from the public data set MS-COCO, crop their sizes into a unified size to obtain labeled images, and form a training data set from the labeled images.

[0012] Combined with the first aspect, further, it further includes the step of preprocessing the training data set, including:

[0013] Perform Gaussian blur processing on different regions of the labeled images in the training data set.

[0014] Combined with the first aspect, further, it further includes the step of setting the loss function of the fusion network model, including:

[0015] Set the following loss function:

[0016] L = L mse + αL ssim + βL per

[0017] L ssim = 1 - SSIM(O, T)

[0018]

[0019] where L is the loss function, L mse is the mean square loss function, L ssim is the structural similarity loss function, L per is the perceptual loss function, α and β are balance parameters, O is the fused image, T is the labeled image, SSIM(O, T) represents the structural similarity between O and T, represents the pixel value at the coordinate (x, y) in the i-th channel of the feature map extracted from the fused image through VGG16, represents the pixel value at the coordinate (x, y) in the i-th channel of the feature map extracted from the labeled image through VGG16, C f 、Hf and W f respectively represent the number of channels, height, and width of any feature map.

[0020] Combined with the first aspect, further, the fusion network model is trained by the following method:

[0021] In Pytorch, the fusion network model is trained using the constructed training dataset, and the batch size during training is set to 8.

[0022] Combined with the first aspect, further, it also includes the step of optimizing the parameters of the fusion network model, including:

[0023] The Adam optimizer is used to optimize the parameters of the fusion network model, where the initial learning rate of the Adam optimizer is set to 0.001.

[0024] In the second aspect, the present invention also provides a multi-focus image fusion system, including:

[0025] An acquisition module: used to acquire multi-focus images to be fused;

[0026] A fusion module: used to input the multi-focus images into the pre-trained fusion network model for fusion to obtain a fused image.

[0027] In the third aspect, the present invention also provides a multi-focus image fusion device, including a processor and a storage medium;

[0028] The storage medium is used to store instructions;

[0029] The processor is used to operate according to the instructions to execute the steps of the method according to any one of the first aspect.

[0030] In the fourth aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the method according to any one of the first aspect.

[0031] Compared with the prior art, the beneficial effects achieved by the present invention are:

[0032] A multi-focus image fusion method, system, device and storage medium provided by the present invention input multi-focus images into a pre-trained fusion network model for fusion to obtain a fused image, realizing the fusion of multi-focus images; among them, the dilated convolutional network expands the receptive field by increasing the dilation rate, and thus more comprehensively extracts multi-scale features in the source image (the multi-focus images to be fused); the dense convolutional neural network can effectively solve the problem of gradient disappearance in deep networks. At the same time, in order to further highlight the importance of significant features, an attention mechanism is introduced into the dense convolutional neural network to adaptively select significant features, thereby improving the fusion performance; in summary, the solution of the present invention can retain rich detail information when fusing multi-focus images, and the visual quality of the fused image obtained by fusion is relatively good. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is one of the flowcharts of a multi-focus image fusion method provided by an embodiment of the present invention;

[0034] Figure 2 is a schematic structural diagram of the fusion network model provided by an embodiment of the present invention;

[0035] Figure 3 is the second flowchart of a multi-focus image fusion method provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be used to limit the protection scope of the present invention.

[0037] Embodiment 1

[0038] As Figure 1 shown, a multi-focus image fusion method provided by an embodiment of the present invention includes:

[0039] S1. Obtain multi-focus images to be fused.

[0040] Obtain multi-focus images with some unclear parts for subsequent fusion.

[0041] S2. Input the multi-focus images into a pre-trained fusion network model for fusion to obtain a fused image.

[0042] Construct a training data set: construct a labeled training data set (multi-focus image data set) based on the public data set MS-COCO.

[0043] In practical applications, it is difficult to obtain multi-focus images and their paired full-depth images. For this reason, the present invention constructs a set of simulated multi-focus image data sets (i.e., training data sets).

[0044] Select 8,000 high-definition natural images from the public dataset MS-COCO, uniformly crop their sizes to 128×128, and then use them as label images.

[0045] Preprocess the training dataset: perform Gaussian blur processing on different regions of the label images. Specifically, Gaussian blur processing can be performed on complementary regions of the label images.

[0046] To simulate the different degrees of blur caused by different depths of field, the 8,000 selected label images in the present invention are evenly divided into 4 groups, and Gaussian blur with Gaussian blur radii of 2, 4, 6, and 8 is respectively used for blur processing.

[0047] Construct a fusion network model using a dilated convolutional network, a convolutional layer, and a densely connected convolutional neural network based on an attention mechanism.

[0048] As Figure 2 shown, the fusion network model includes three parts, namely Feature extraction, Feature fusion, and Image reconstruction.

[0049] The feature extraction part includes two network branches that share weights. Each network branch consists of 1 dilated convolutional network with multiple branches in parallel, 1 1×1 convolutional layer, and 1 densely connected convolutional neural network with an attention mechanism.

[0050] The number of feature channels of the dilated convolutional network with multiple branches in parallel is set to 192, and this network consists of 3 dilated convolutions with a convolutional kernel size of 3×3 and dilation rates of 1, 2, and 3 respectively.

[0051] The input channels and output channels of the densely connected convolutional neural network with an attention mechanism are 64 and 256 respectively, and it consists of 1 dense block and 3 Squeeze-Excitation-Blocks. The dense block contains 3 3×3 convolutional layers, and the output of each layer is cascaded as the input of the next layer.

[0052] The 1×1 convolutional layer is used to adjust the dimension of the feature channels.

[0053] The feature fusion part consists of a "concatenation" operation and 1 1×1 convolutional layer, mainly realizing the fusion of features.

[0054] The feature fusion part performs "concatenation" and 1×1 convolutional operations on the features of the multi-focus images obtained by the feature extraction part to achieve feature fusion and obtain fused features. Among them, the input channels and output channels of the 1×1 convolutional layer are 512 and 64 respectively.

[0055] The image reconstruction part mainly generates a fused image from the fused features. This part consists of 4 convolutional layers of 3×3, with the number of feature channels being 64, 64, 64, and 3 respectively. Each convolutional layer except the last one uses ReLU as the activation function.

[0056] To make the reconstructed image more accurate, the loss function of the above fusion network model is set as follows:

[0057] L = L mse + αL ssim + βL per

[0058] L ssim = 1 - SSIM(O, T)

[0059]

[0060] where L is the loss function, L mse is the mean square loss function, L ssim is the structural similarity loss function, L per is the perceptual loss function, α and β are balance parameters, O is the fused image, T is the label image, SSIM(O, T) represents the structural similarity between O and T, represents the pixel value at the coordinate (x, y) in the i-th channel of the feature map extracted from the fused image by VGG16, represents the pixel value at the coordinate (x, y) in the i-th channel of the feature map extracted from the label image by VGG16, C f 、H f and W f represent the number of channels, height, and width of any feature map respectively.

[0061] In this embodiment, both α and β are 0.5.

[0062] The constructed training dataset is used to train the fusion network model. During the training process, the hyperparameters of the fusion network model include batch size, initial learning rate, number of epochs, and learning rate decay strategy.

[0063] In this embodiment, Pytorch (a scientific computing library based on python) is used to implement the training of the fusion network model. The program running environment is RTX 3080 / 10GB RAM, Intel Core i7 - 10700K@3.80GHz.

[0064] The batch size during training is set to 8. The Adam optimizer is used to optimize the parameters. The initial learning rate of the optimizer is set to 0.001, and the cosine annealing decay method is used to adjust the learning rate. The network is trained for a total of 500 epochs.

[0065] Input the multi-focus images into the pre-trained fusion network model for fusion, and the fused image can be obtained.

[0066] To prove the effectiveness of the fusion technology proposed in the present invention, 8 mainstream multi-focus image fusion methods are selected for comparison with the present invention on the Lytro multi-focus color image dataset, namely the NSCT method, the SR method, the IMF method, the MWGF method, the CNN method, the DeepFuse method, the DenseFuse method (including the DenseFuse-ADD method and the DenseFuse-L1 method), and the IFCNN-MAX method. All comparison methods are tested using the default parameters provided in the literature.

[0067] In this embodiment, four objective indicators are used as quantization indicators, namely the average gradient (AG), the spatial frequency (SF), the visual information fidelity (VIF), and the edge preservation degree (Q AB / F ), Table 1 shows the average index values on the Lytro dataset; it can be seen from Table 1 that the present invention obtains the optimal results in the three indicators of AG, SF, and VIF, and for Q AB / F indicator, it also obtains the sub-optimal result second only to the CNN method; from the results, the present invention is a feasible and efficient multi-focus image fusion method.

[0068] Table 1 Average index values on the Lytro dataset

[0069]

[0070] Example 2

[0071] A multi-focus image fusion system provided by an embodiment of the present invention includes:

[0072] An acquisition module: used to acquire the multi-focus images to be fused;

[0073] A fusion module: used to input the multi-focus images into the pre-trained fusion network model for fusion to obtain a fused image.

[0074] The fusion network model is constructed by the following method: using a dilated convolutional network and a dense convolutional neural network based on an attention mechanism to construct the fusion network model.

[0075] Example 3

[0076] An apparatus for multi-focus image fusion provided by an embodiment of the present invention includes a processor and a storage medium;

[0077] The storage medium is used for storing instructions;

[0078] The processor is used to operate according to the instructions to execute the steps of the following method:

[0079] Obtain multi-focus images to be fused;

[0080] Input the multi-focus images into a pre-trained fusion network model for fusion to obtain a fused image.

[0081] The fusion network model is constructed by the following method: constructing a fusion network model by using a dilated convolutional network and a dense convolutional neural network based on an attention mechanism.

[0082] Embodiment 4

[0083] A computer-readable storage medium provided by an embodiment of the present invention stores a computer program, and when the program is executed by a processor, the steps of the following method are implemented:

[0084] Obtain multi-focus images to be fused;

[0085] Input the multi-focus images into a pre-trained fusion network model for fusion to obtain a fused image.

[0086] The fusion network model is constructed by the following method: constructing a fusion network model by using a dilated convolutional network and a dense convolutional neural network based on an attention mechanism.

[0087] Embodiment 5

[0088] As Figure 3 shown, a multi-focus image fusion method provided by an embodiment of the present invention includes:

[0089] S1. Construct a training data set and preprocess the training data set.

[0090] S2. Construct a fusion network model by using an attention mechanism and a dense convolutional neural network.

[0091] S3. Set a loss function of the fusion network model and optimize network parameters.

[0092] S4. Train the fusion network model by using the training data set to obtain a trained fusion network model.

[0093] S5. Obtain multi-focus images to be fused, and input the multi-focus images into the trained fusion network model to obtain a fused image.

[0094] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0095] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0096] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that realizes the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0097] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0098] The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A multi-focus image fusion method, characterized in that, comprising: Obtaining multi-focus images to be fused; Inputting the multi-focus images into a pre-trained fusion network model for fusion to obtain a fused image; The fusion network model is constructed by the following method: constructing a fusion network model using an atrous convolutional network and a densely convolutional neural network based on an attention mechanism; The fusion network model includes three parts: feature extraction, feature fusion, and image reconstruction; The feature extraction part includes two network branches with shared weights. Each network branch consists of an atrous convolutional network with multiple branches in parallel, a 1×1 convolutional layer, and a densely convolutional neural network with an attention mechanism; The feature channels of the atrous convolutional network with multiple branches in parallel are set to 192, and this network consists of three atrous convolutions with a kernel size of 3×3 and dilation rates of 1, 2, and 3 respectively; The input channels and output channels of the densely convolutional neural network with an attention mechanism are 64 and 256 respectively, and it consists of 1 dense block and 3 Squeeze-Excitation-Blocks. The dense block contains 3 convolutional layers with a size of 3×3, and the output of each layer is cascaded as the input of the next layer; The 1×1 convolutional layer is used to adjust the dimension of the feature channels; The feature fusion part consists of a "concatenation" operation and a 1×1 convolutional layer; The feature fusion part performs concatenation and 1×1 convolutional operations on the features of the multi-focus images obtained by the feature extraction part. Among them, the input channels and output channels of the 1×1 convolutional layer are 512 and 64 respectively; The image reconstruction part generates a fused image from the fused features. This part consists of 4 convolutional layers with a size of 3×3, and the number of feature channels is 64, 64, 64, and 3 respectively. ReLU is used as the activation function for each convolutional layer except the last one.

2. A multi-focus image fusion method according to claim 1, characterized in that, it further includes the step of establishing a training data set, including: Extracting images from the public data set MS-COCO, cropping their sizes into a unified size to obtain label images, and forming a training data set from the label images.

3. A multi-focus image fusion method according to claim 2, characterized in that, it further includes the step of preprocessing the training data set, including: Performing Gaussian blur processing on different regions of the label images in the training data set.

4. A multi-focus image fusion method according to claim 2, characterized in that, it further includes the step of setting the loss function of the fusion network model, including: Setting the following loss function: ; Among them, L is the loss function, L mse is the mean square loss function, L ssim is the structural similarity loss function, L per is the perceptual loss function, 𝛼 and 𝛽 are the balance parameters, is the fused image, is the label image, denotes and the structural similarity between, denotes the pixel value at the coordinate i in the channel of the feature map extracted by VGG16 from the fused image, denotes the pixel value at the coordinate i in the channel of the feature map extracted by VGG16 from the label image, , and respectively denote the number of channels, height, and width of an arbitrary feature map.

5. A multi-focus image fusion method according to claim 2, characterized in that, the fusion network model is trained by the following method: Training the fusion network model using the constructed training data set in Pytorch, and setting the batch size to 8 during the training process.

6. A multi-focus image fusion method according to claim 1, characterized in that, it further includes the step of optimizing the parameters of the fusion network model, including: The parameters of the fusion network model are optimized using the Adam optimizer, and the initial learning rate of the Adam optimizer is set to 0.

001.

7. A multi-focus image fusion system Characterized in that Comprising: An acquisition module: used to acquire multi-focus images to be fused; A fusion module: used to input the multi-focus images into a pre-trained fusion network model for fusion to obtain a fused image; The fusion network model is constructed by the following method: a fusion network model is constructed using a dilated convolutional network and a densely convolutional neural network based on an attention mechanism; The fusion network model includes three parts: feature extraction, feature fusion, and image reconstruction; The feature extraction part includes two network branches with shared weights. Each network branch consists of a dilated convolutional network with multiple branches in parallel, a 1×1 convolutional layer, and a densely convolutional neural network with an attention mechanism; The number of feature channels of the dilated convolutional network with multiple branches in parallel is set to 192, and this network consists of three dilated convolutions with a convolutional kernel size of 3×3 and dilation rates of 1, 2, and 3 respectively; The input and output channels of the densely convolutional neural network with an attention mechanism are 64 and 256 respectively, and it consists of 1 dense block and 3 Squeeze-Excitation-Blocks. The dense block contains 3 convolutional layers with a size of 3×3, and the output of each layer is cascaded as the input of the next layer; The 1×1 convolutional layer is used to adjust the dimension of the feature channels; The feature fusion part consists of a "concatenation" operation and a 1×1 convolutional layer; The feature fusion part performs concatenation and 1×1 convolutional operations on the features of the multi-focus images obtained by the feature extraction part. Among them, the input and output channels of the 1×1 convolutional layer are 512 and 64 respectively; The image reconstruction part generates a fused image from the fused features. This part consists of 4 convolutional layers with a size of 3×3, and the number of feature channels is 64, 64, 64, and 3 respectively. ReLU is used as the activation function for each convolutional layer except the last one.

8. A multi-focus image fusion device Characterized in that Comprising a processor and a storage medium; The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, on which a computer program is stored Characterized in that When the program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.