Multi-focus image fusion method, multi-focus image fusion network model and electronic equipment

By combining multi-branched hollow convolution and multi-scale residual attention module, the problems of high complexity and weak generalization ability of the traditional multi-focus image fusion method are solved, and high-quality image fusion effect is achieved.

CN120339084APending Publication Date: 2025-07-18XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510243276.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The traditional multi-focus image fusion method has the problems of high algorithm complexity, low processing efficiency, weak generalization ability and blurred edge areas. Due to the small scale of the multi-focus image data set, it is difficult for traditional convolutional neural networks to build a powerful fusion model.

Method used

Generative adversarial network (GAN) is used as the basic framework, combined with adversarial training, and through feature extraction modules, feature fusion modules and image reconstruction modules, multi-branch hole convolution and multi-scale residual attention modules are used to extract image features of different scales, and realistic and natural fusion images are generated through adversarial training.

Benefits of technology

The quality and accuracy of image fusion are improved, information loss caused by scale changes is reduced, the accuracy of feature fusion is optimized, and information redundancy is reduced, and a clearer fully focused image is generated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339084A_ABST
    Figure CN120339084A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-focus image fusion method, a multi-focus image fusion network model and electronic equipment, and relates to the field of image processing. The image features of different scales can be extracted, and the image fusion quality is improved. The multi-focus image fusion method comprises the steps that source images of the same scene under different focal lengths are input into a generator, a plurality of feature extraction modules extract features of the source images under the different focal lengths respectively, a multi-scale residual attention module extracts image features of different scales, and a feature map is obtained; the feature fusion module fuses the feature maps of the source images to obtain a fused feature map; the image reconstruction module reconstructs the fusion feature map to obtain a reconstructed fusion image; inputting the gradient map of the fused image into a discriminator to obtain a determination probability output by the discriminator, and calculating a loss function based on the determination probability; and performing adversarial training on the generator and the discriminator through the loss function, and fusing the to-be-fused images under different focal lengths by adopting the trained generator and discriminator.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technologies, and in particular, to a multi-focus image fusion method, a multi-focus image fusion network model, and an electronic device. Background Art

[0002] The multi-focus fusion technology is a key research direction in the field of multi-source image fusion. Its core goal is to overcome the problem of blurred parts of an image caused by the limitation of the lens depth of field during the optical imaging process. This technology generates a fully focused image with clear focus in all regions by combining multiple images taken of the same scene at different focal lengths. This can not only improve the image quality but also provide more accurate information for computer vision systems.

[0003] Traditional processing methods such as pixel-based and region-based fusion have problems such as high algorithm complexity, low processing efficiency, weak generalization ability, and blurred edge regions. With the continuous in-depth research of deep learning technologies in the field of image fusion, their feature extraction and characterization capabilities have made significant progress, effectively breaking through the limitations of traditional algorithms and achieving more accurate and natural image fusion effects. However, due to the generally small scale of multi-focus image datasets and the lack of high-quality reference images for supervised learning, it is difficult for traditional convolutional neural networks to construct a powerful fusion model. Summary of the Invention

[0004] The present application provides a multi-focus image fusion method, a multi-focus image fusion device, and an electronic device, which are based on the framework of Generative Adversarial Networks (GAN) and combine adversarial training to improve the quality of the fused image.

[0005] In a first aspect, the present application provides a multi-focus image fusion method, including:

[0006] Input source images of the same scene at different focal lengths into a generator, where the generator includes a feature extraction module, a feature fusion module, and an image reconstruction module;

[0007] Multiple feature extraction modules respectively extract features of the source images at different focal lengths. The feature extraction module includes multi-branch dilated convolutions, and the multi-branch dilated convolutions extract image features of different scales to obtain feature maps;

[0008] The feature fusion module fuses the feature maps of each source image to obtain a fused feature map;

[0009] The image reconstruction module reconstructs the fused feature map to obtain a reconstructed fused image;

[0010] Input the gradient map of the fused image into the discriminator to obtain the determination probability output by the discriminator, and calculate the loss function based on the determination probability;

[0011] Perform adversarial training on the generator and the discriminator through the loss function, and use the trained generator and discriminator to fuse the images to be fused at different focal lengths.

[0012] According to the multi-focus image fusion method provided in this embodiment, based on the generative adversarial network as the basic framework, the adversarial network includes two parts: a generator and a discriminator. The generator includes multiple feature extraction modules, which are respectively targeted at the source images at different focal lengths. The feature extraction module is a multi-branch dilated convolution structure, which can capture image details and features at different scales, effectively reduce information loss caused by scale changes, optimize the accuracy of feature fusion and reduce information redundancy. Moreover, by adopting the adversarial training strategy of game theory, the generator can generate more realistic and natural fusion results, while the discriminator is responsible for capturing the texture information in the fused image, further improving the quality of the fused image.

[0013] In an exemplary embodiment, the multiple feature extraction modules respectively extract the features of the source images at different focal lengths, including:

[0014] Input the source images with different focal lengths into different feature extraction modules. The feature extraction module includes a first convolutional layer, multiple multi-scale residual attention modules, a first fusion module, and a second convolutional layer connected in sequence;

[0015] The first convolutional layer performs a convolutional operation on the source image to obtain the feature map of the source image and inputs it to the first multi-scale residual attention module;

[0016] Multiple multi-scale residual attention modules are connected in sequence and skip-connected to the first fusion module;

[0017] Each multi-scale residual attention module extracts features from the input feature map, and the first fusion module splices the feature maps extracted by different multi-scale residual attention modules;

[0018] The second convolutional layer performs a convolutional operation on the spliced feature map again to obtain the processed feature map.

[0019] In an exemplary embodiment, the multi-scale residual attention module includes a multi-branch dilated convolution module;

[0020] The multi-branch dilated convolution module includes multiple branches. Each branch includes a convolutional layer with a kernel size of 1. Starting from the second branch, a dilated convolutional layer is connected after the convolutional layer, and the number of dilated convolutional layers in each branch increases sequentially;

[0021] The output of each dilated convolutional layer is connected to the next dilated convolutional layer and the next dilated convolutional layer of the next branch;

[0022] The outputs of the last dilated convolutional layer of each branch are concatenated to obtain a first output feature map, which is used as the output of the multi-branch dilated convolutional module.

[0023] In one exemplary embodiment, the multi-scale residual attention module further includes a dual-channel hybrid attention module connected to the multi-branch dilated convolutional module;

[0024] The dual-channel hybrid attention module includes a channel attention module and a spatial attention module;

[0025] The channel attention module determines channel weights for the first output feature map and weights the first output feature map based on the channel weights to obtain a weighted first feature map;

[0026] The spatial attention module determines spatial weights for the first output feature map and weights the first output feature map based on the spatial weights to obtain a weighted second feature map;

[0027] The first feature map and the second feature map are combined to obtain a second output feature map, which is used as the output of the dual-channel hybrid attention module.

[0028] In one exemplary embodiment, the channel attention module determines channel weights for the first output feature map and weights the first output feature map based on the channel weights to obtain a weighted first feature map, including:

[0029] The channel attention module performs average pooling on the feature maps of each channel in the first output feature map to obtain a first feature vector, performs max pooling on the feature maps of each channel to obtain a second feature vector; and respectively performs dimensionality reduction and dimensionality increase on the first feature vector and the second feature vector in sequence to obtain a processed third feature vector and a fourth feature vector; uses the sigmoid activation function to perform activation processing on the third feature vector and the fourth feature vector respectively to obtain a first enhanced feature map and a second enhanced feature map, and multiplies the first enhanced feature map and the second enhanced feature map with the first output feature map element by element in sequence to obtain a first feature map.

[0030] In one exemplary embodiment, the spatial attention module determines spatial weights for the first output feature map and weights the first output feature map based on the spatial weights to obtain a weighted second feature map, including:

[0031] The spatial attention module performs a convolution operation on the first output feature map, outputs a weight matrix with the same spatial dimension as the first output feature map, normalizes the weight matrix through a sigmoid activation function to generate a spatial attention map, and multiplies the spatial attention map and the first output feature map element by element to obtain a second feature map.

[0032] In an exemplary embodiment, the inputting the gradient map of the fused image into the discriminator to obtain the determination probability output by the discriminator includes:

[0033] Obtain the gradient maps of source images with different focal lengths, and merge the gradient maps of source images with different focal lengths based on the maximum selection method to obtain a joint gradient map;

[0034] Input the gradient map of the fused image and the joint gradient map into the discriminator together to obtain their respective determination probabilities.

[0035] In an exemplary embodiment, the adversarial training of the generator and the discriminator through the loss function includes:

[0036] Calculate the generator loss based on the determination probabilities corresponding to the fused image and the source image respectively;

[0037] Calculate the difference between the light intensity distributions of the source images with different focal lengths and the corresponding fused images to obtain an intensity loss;

[0038] Calculate the difference between the gradient map of the fused image and the gradient maps of the source images with different focal lengths to obtain a gradient loss;

[0039] Calculate the similarity between the fused image and the original image to obtain a structural similarity loss;

[0040] Combine the intensity loss, the gradient loss, and the structural similarity loss to obtain the discriminator loss;

[0041] Perform adversarial training on the generator and the discriminator based on the generator loss and the discriminator loss.

[0042] In a second aspect, the present application provides a multi-focus image fusion network model, including:

[0043] A generator and a discriminator;

[0044] The generator includes a feature extraction module, a feature fusion module, and an image reconstruction module;

[0045] A plurality of feature extraction modules respectively extract features of source images at different focal lengths, wherein the feature extraction module comprises a multi-scale residual attention module, and the multi-scale residual attention module extracts image features of different scales to obtain a feature map;

[0046] The feature fusion module fuses the feature maps of each source image to obtain a fused feature map;

[0047] The image reconstruction module reconstructs the fused feature map to obtain a reconstructed fused image;

[0048] The discriminator predicts the determination probability of the fused image;

[0049] A loss function is calculated based on the determination probability, adversarial training is performed on the generator and the discriminator through the loss function, and the trained generator and discriminator are used to fuse the images to be fused at different focal lengths.

[0050] In a third aspect, the present application provides an electronic device, the electronic device comprising a memory and one or more processors. The memory stores one or more computer programs, the computer programs comprising instructions, and when the instructions are executed by the processor, the electronic device can execute the multi-focus image fusion method as in the first aspect.

[0051] In a fourth aspect, the present application provides a computer-readable storage medium, in which instructions are stored. When the instructions are executed on an electronic device, the electronic device executes the multi-focus image fusion method in the first aspect.

[0052] In a fifth aspect, the present application provides a computer program product. When the computer program product is run on an electronic device, the electronic device executes the multi-focus image fusion method as described in the first aspect.

[0053] It can be understood that the beneficial effects that can be achieved by the application monitoring device, electronic device, computer-readable storage medium, and computer program product provided above can refer to the beneficial effects in the first aspect and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 A schematic diagram of the process of the multi-focus image fusion method provided in the embodiment of the present application;

[0055] Figure 2 Schematic diagram of the model structure in the multi-focus image fusion method provided in the embodiment of the present application Figure 1 ;

[0056] Figure 3 Schematic diagram of the model structure in the multi-focus image fusion method provided in the embodiment of the present applicationFigure 2 ;

[0057] Figure 4 Schematic diagram of the model structure in the multi-focus image fusion method provided by the embodiment of the present application Figure 3 ;

[0058] Figure 5 Schematic diagram of the model structure in the multi-focus image fusion method provided by the embodiment of the present application Figure 4 ;

[0059] Figure 6 Schematic diagram of the structure of the electronic device provided by the embodiment of the present application. Detailed implementation manners

[0060] In order to clearly describe the technical solutions of the embodiments of the present application, in the embodiments of the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and effects. For example, the first chip and the second chip are only used to distinguish different chips, and do not limit their sequence. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and the terms "first" and "second" do not necessarily limit being different. It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner. In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more than two.

[0061] It should be noted that "when... " in the embodiments of the present application can be at the instant when a certain situation occurs, or within a period of time after a certain situation occurs. The embodiments of the present application do not make specific limitations on this.

[0062] The following will describe the implementation manners of this embodiment in detail with reference to the accompanying drawings.

[0063] This embodiment provides a multi-focus image fusion method. Exemplarily, this multi-focus image fusion method can be applied to various electronic devices such as a computer (PC), a tablet computer, a virtual reality / augmented reality device, a wearable device, an industrial computer, and a vehicle-mounted computer; it can also be applied to a server, the cloud, a server cluster, etc. The embodiments of the present application do not make special limitations on this.

[0064] Figure 1 Shows a flowchart of the multi-focus image fusion method provided by the embodiment of the present application.

[0065] AsFigure 1 As shown in the figure, the multi-focus image fusion method may include the following steps:

[0066] Step 10: Input source images of the same scene at different focal lengths into a generator, where the generator includes a feature extraction module, a feature fusion module, and an image reconstruction module.

[0067] In this embodiment, the generator and the discriminator constitute a multi-focus image fusion network model. Among them, the feature extraction module is used to extract the features of the source image; the feature fusion module fuses the features extracted from each source image; and the image reconstruction module restores the image based on the fused features.

[0068] Step 20: Multiple feature extraction modules respectively extract the features of the source images at different focal lengths. The feature extraction module includes a multi-scale residual attention module, and the multi-scale residual attention module extracts image features at different scales to obtain feature maps.

[0069] The multi-scale residual attention (MSRA) module focuses on the information of different regions through residual attention, thereby enhancing the expression ability of features and the stability of the network. Based on this multi-scale residual attention module, the multi-focus image fusion network model in this embodiment can be called a multi-focus image fusion network model based on multi-scale residual attention (MSRA-GAN).

[0070] Specifically, the feature extraction module includes a first convolutional layer, multiple multi-scale residual attention modules, a first fusion module, and a second convolutional layer connected in sequence; input source images with different focal lengths into different feature extraction modules, and the first convolutional layer performs a convolutional operation on the source image to obtain a feature map of the source image and input it into the first multi-scale residual attention module; multiple multi-scale residual attention modules are connected in sequence and skip-connected to the first fusion module; each multi-scale residual attention module extracts features from the input feature map, and the first fusion module splices the feature maps extracted by different multi-scale residual attention modules; the second convolutional layer performs a convolutional operation on the spliced feature map again to obtain a processed feature map.

[0071] The specific structure of the MSRA-GAN model in this embodiment is as Figure 2 shown, referring to Figure 2, where the generator network is specifically as follows: The feature extraction module includes two multi-scale feature extraction sub-networks with the same structure. First, a 3×3 convolutional layer is used to extract the shallow features of the two input images respectively. Then, three MSRA modules are used to further extract the deep features. Subsequently, skip connections are introduced to strengthen the reuse and transmission of features, and at the same time promote the effective flow of gradients in the deep network, improving the stability of training. The feature fusion module performs a simple "concatenation" operation on the extracted features, and then processes the concatenated features through a 1×1 convolutional layer to effectively integrate the image features. The main function of the image reconstruction part is to map the fused features into the final fused image. It consists of four 3×3 convolutional layers. Except for the last convolutional layer, the ReLU activation function is used in the other convolutional layers in the MSRA-GAN model to enhance the nonlinear representation ability of the model, enabling the network to better learn and adapt to complex image fusion tasks.

[0072] The multi-scale residual attention module includes a multi-branch atrous convolution module; the multi-branch atrous convolution module includes multiple branches, and each branch includes a convolutional layer with a kernel size of 1. Starting from the second branch, a dilated convolutional layer is connected after the convolutional layer, and the number of dilated convolutional layers in each branch increases sequentially; the output of each dilated convolutional layer is connected to the next dilated convolutional layer and the next dilated convolutional layer of the next branch; the outputs of the last dilated convolutional layers of each branch are concatenated to obtain the first output feature map, which is used as the output of the multi-branch atrous convolution module.

[0073] The multi-branch atrous convolution (MAC) module includes multiple branches, and each branch includes one or more dilated convolutions. Dilated convolution introduces a "dilation rate" parameter on the basis of the standard convolution operation. The role of this parameter is to expand the receptive field of the convolutional kernel without increasing the number of parameters or the amount of computation. The dilation rate is used to indicate the number of spaces inserted between the elements of the convolutional kernel. By adding spaces in the convolutional kernel, the receptive field of the convolution can be expanded, but the number of parameters used by the convolutional kernel does not increase, and context information can be captured in a larger range.

[0074] The specific structure of the MAC module is as Figure 3 shown, refer to Figure 3, the MAC module can be a network structure consisting of four branches, aiming to effectively capture image features at different scales while reducing the computational load. The first branch achieves dimensionality reduction of the data by using a 1×1 convolutional kernel, effectively reducing the number of data channels; the second branch also starts with a standard 1×1 convolution to reduce the feature dimension, thereby reducing the model parameters and computational burden. Subsequently, a 3×3 dilated convolution with a dilation rate of 1 is introduced, and its key role is to accurately capture the feature information of small-scale textures; the third branch is the same, first using a standard 1×1 convolution to reduce the computational load. Then, a 3×3 dilated convolution with a dilation rate of 2 is used to further extract features, and the purpose of this step is to capture the mid-scale edge contour information of the image. In addition, the features of this branch are combined with the output features of the second branch and then further convolutional processing is performed; the fourth branch continues to follow the strategy of the previous branch processing, that is, using the features from the previous branch to make up for the possible problems caused by the increase in the dilation rate. As the dilation rate increases, the network may over-focus on large-scale structural information and ignore the detailed features. To make up for this deficiency, the fourth branch cleverly combines the features of the previous branch to achieve more comprehensive and balanced feature extraction.

[0075] Specifically, the calculation expressions for the output features of the four branches of the MAC module are respectively

[0076]

[0077] where I represents the input feature map, * represents the convolution operation, Concat represents the concatenation operation, C 1×1 represents a 1×1 convolution, C rate=k represents a dilated convolution with a dilation rate of k, and O1, O2, O3, and O4 respectively represent the output results of the four branches. After the processing of the four branches is completed, the number of channels of the feature map is restored to the initial size by concatenating the output features of each branch, and the calculation expression of the processing process is

[0078] O = Concat(O1, O2, O3, O4) (2)

[0079] The feature map obtained by the multi-branch dilated convolution processing is called the first output feature map, and this feature map will be used as the input of the next layer of the network.

[0080] The multi-scale residual attention module further includes a dual-channel hybrid attention module connected to the multi-branch dilated convolution module; the dual-channel hybrid attention module includes a channel attention module and a spatial attention module; the channel attention module determines channel weights for the first output feature map, and weights the first output feature map based on the channel weights to obtain a weighted first feature map; the spatial attention module determines spatial weights for the first output feature map, and weights the first output feature map based on the spatial weights to obtain a weighted second feature map; the first feature map and the second feature map are combined to obtain a second output feature map, which is used as the output of the dual-channel hybrid attention module.

[0081] The channel attention module determines channel weights for the first output feature map, and weights the first output feature map based on the channel weights to obtain a weighted first feature map, including: the channel attention module performs average pooling on the feature maps of each channel in the first output feature map to obtain a first feature vector, and performs max pooling on the feature maps of each channel to obtain a second feature vector; and respectively performs dimensionality reduction and dimensionality increase on the first feature vector and the second feature vector in sequence to obtain a processed third feature vector and a fourth feature vector; uses the sigmoid activation function to perform activation processing on the third feature vector and the fourth feature vector respectively to obtain a first enhanced feature map and a second enhanced feature map, and multiplies the first enhanced feature map and the second enhanced feature map with the first output feature map element by element in sequence to obtain a first feature map.

[0082] The spatial attention module determines spatial weights for the first output feature map, and weights the first output feature map based on the spatial weights to obtain a weighted second feature map, including: the spatial attention module performs a convolution operation on the first output feature map, outputs a weight matrix with the same spatial dimension as the first output feature map, normalizes the weight matrix through the sigmoid activation function to generate a spatial attention map, and multiplies the spatial attention map and the first output feature map element by element to obtain a second feature map.

[0083] In this embodiment, the generator adopts a new attention mechanism: dual-channel hybrid attention (Double Concurrent Spatial and Channel Squeeze&Excitation, DSCSE). The structural diagram of the DSCSE module is as Figure 4 shown. Refer to Figure 4 , for example, if the input feature map (i.e., the first output feature map output by the multi-branch dilated convolution) is U ∈ R C×H×W, first, the importance weights of each channel are extracted through the channel attention module respectively to obtain the feature map U enhanced by channel attention DcSE ∈R C×1×1 , that is, the first feature map. At the same time, different weights are assigned to each spatial position through the spatial attention module to obtain the feature map U enhanced by spatial attention DsSE ∈R 1×H×W , that is, the second feature map. Subsequently, these two enhanced feature maps are added element by element to obtain the enhanced feature result map U DsCSE ∈R C×H×W . The calculation expression of the specific processing flow of the DSCSE module is

[0084]

[0085] where, represents the element-by-element addition operation.

[0086] Continue to refer to Figure 4 , the channel attention module of the DSCSE mechanism applies average pooling and max pooling in parallel to the input single feature map U ∈ R C×H×W to obtain the compressed feature vectors Z avg ∈R 1×1×C and Z max ∈R 1×1×C , so as to extract the information in the feature map more comprehensively and ensure that different types of features can be effectively utilized. Subsequently, the DSCSE uses a 1×1 convolutional operation to replace the traditional fully connected layer to perform dimensionality reduction and dimensionality increase operations on the feature vectors respectively. In addition, the Sigmoid activation function is used to limit the feature vectors of each channel within the range of (0,1) to obtain two feature maps U avg ∈R C×1×1 and U max ∈R C×1×1 . Finally, they are multiplied element by element with the original input feature map, effectively combining the original features and the key information extracted by the attention module, ensuring that the model can focus on the most informative channel positions, so as to obtain the final feature map U of channel attention DcSE ∈R C×1×1 (that is, the first feature map), and its specific processing flow calculation expression is:

[0087]

[0088] where, σ is the Sigmoid activation function, C 1×1 represents the 1×1 convolutional operation, avgPool(·) represents the average pooling operation, and maxPool(·) represents the max pooling operation.

[0089] Refer toFigure 4 , the spatial attention module of the DSCSE mechanism, for the input feature map U ∈ R C×H×W , first, a 1×1 convolutional layer is used to effectively integrate the feature information and output a weight matrix that is exactly the same as the spatial dimension of the input feature map. Subsequently, the weight matrix is normalized through the Sigmoid activation function to generate a spatial attention map, which can accurately identify the importance of each spatial position. Finally, the spatial attention map is multiplied element-wise with the original input feature map U to obtain the feature map with enhanced spatial attention (i.e., the second feature map). The calculation expression of its specific processing flow is

[0090] U DsSE = σ(C 1×1 (U)) * U (5)

[0091] where σ is the Sigmoid activation function, and C 1×1 represents the 1×1 convolutional operation.

[0092] The structure of the feature extraction module that combines the above multi-branch dilated convolutional module and the dual-channel hybrid attention module is as Figure 5 shown. Refer to Figure 5 , in this embodiment, the MSRA-based feature extraction network extracts image features at different scales by adopting a multi-branch dilated convolutional module, ensuring the integrity and richness of the feature map. Then, using the idea of residual connection, the input feature map U is added to the feature map U MAC processed by the MAC module to optimize the network structure, effectively alleviating the problem of gradient disappearance and significantly improving the training efficiency and performance of the network. In addition, to optimize the accuracy of feature fusion and reduce information redundancy, the MSRA module further integrates the DSCSE attention module, which enhances the network's attention to key information and improves the discriminative ability of features. The calculation expression of its specific processing flow is

[0093]

[0094] where represents the feature map output by the feature extraction module, U represents the input feature map, and U MAC represents the feature map processed by the MAC module.

[0095] Step 30: The feature fusion module fuses the feature maps of each source image to obtain a fused feature map.

[0096] The feature fusion module performs a simple "concatenation" operation on the extracted features, and then processes the concatenated features through a 1×1 convolutional layer to effectively integrate the image features.

[0097] Step 40: The image reconstruction module reconstructs the fused feature map to obtain a reconstructed fused image.

[0098] By splicing each feature map, feature maps with different focal lengths can be fused. The obtained fused feature map includes the texture information of images with different focal lengths. Then, using the fused feature map for image reconstruction can fuse images with multiple focal lengths together to obtain a clearer image.

[0099] Exemplarily, the main function of the image reconstruction module is to map the fused features into the final fused image. Specifically, it can be composed of four 3×3 convolutional layers. Except for the last convolutional layer, the remaining convolutional layers use the ReLU activation function in the MSRA-GAN model to enhance the nonlinear representation ability of the model, enabling the network to better learn and adapt to complex image fusion tasks.

[0100] Step 50: Input the gradient map of the fused image into the discriminator to obtain the determination probability output by the discriminator, and calculate the loss function based on the determination probability. This determination probability is used to indicate the probability that the fused image is real. When the fused image is the same as the source image obtained by shooting, the fused image is real.

[0101] The discriminator network of the used MSRA-GAN model is specifically as follows: one is the joint gradient map generated based on the source image according to the maximum selection principle, and the other is the gradient map of the fused image generated by the generator. The information of these two parts together constitutes the input data of the model, providing a basis for subsequent processing and analysis. First, a 5×5 convolution is used to capture the macroscopic features in the image and perform dimensionality increase processing on the features to more accurately evaluate the information in the image. Subsequently, in order to deeply extract the internal features of the image, 3×3 convolutions are used in the subsequent convolutional layers. All convolutional layers use the ReLU activation function to introduce nonlinearity, and except for the first convolutional layer, batch normalization is used to accelerate the convergence speed of the model.

[0102] Step 60: Perform adversarial training on the generator and the discriminator through the loss function, and use the trained generator and discriminator to fuse the images to be fused under different focal lengths.

[0103] The discriminator predicts the probability that the fused image is real based on the gradient map of the fused image, and also outputs the probability of the source image based on the real source image. Specifically, obtain the gradient maps of the source images with different focal lengths, merge the gradient maps of the source images with different focal lengths based on the maximum selection method to obtain a joint gradient map; input the gradient map of the fused image and the joint gradient map into the discriminator together to obtain their respective determination probabilities.

[0104] The source image is a real image captured of the target scene. The generator uses the source image to reconstruct an image, and the discriminator is used to distinguish between the reconstructed image and the real image. The process of adversarial training can be understood as follows: initially, the generator produces "flawed" images, and the discriminator can distinguish between the flawed reconstructed images and the flawless real images. As the number of training iterations increases, the reconstructed images generated by the generator become increasingly similar to the real images, while the discriminative ability of the discriminator also continuously improves until the two reach a balance and the training is completed. For example, when the number of training iterations reaches a certain value, the discriminator can no longer correctly distinguish between the reconstructed images and the real images, which is equivalent to the discriminator having only a 50% chance of making a determination about the reconstructed images. At the same time, since the generator can no longer obtain positive feedback from the discriminator, the generator no longer needs to be updated, and the discriminator also does not need further optimization. At this point, a balance is reached between the generator and the discriminator, namely the Nash equilibrium.

[0105] In this embodiment, the discriminative loss function of the MSRA-GAN model used is specifically: the fake data input to the discriminator is the gradient map of the fused image generated by the generator The real data input is the gradient map of the fused image obtained according to the principle of maximum pixel selection Their calculation expressions are respectively

[0106]

[0107] Among them, I fused represents the fused image, I1 and I2 represent the input source images, represents the gradient function, abs(·) represents the absolute value function, and max(·) represents the maximum value function. Therefore, the calculation expression of the discriminator loss function is:

[0108]

[0109] Among them, when training the generative adversarial network, N is the number of fused images, b represents the label of the joint gradient map, usually set to 1. c represents the label of the gradient map of the fused image, usually set to 0. D is the determination probability output by the discriminator.

[0110] Exemplarily, a generator loss is calculated based on the determination probabilities corresponding to the fused image and the source image respectively; the difference between the intensity distributions of the source images with different focal lengths and the corresponding fused images is calculated to obtain an intensity loss; the difference between the gradient map of the fused image and the gradient maps of the source images with different focal lengths is calculated to obtain a gradient loss; the similarity between the fused image and the original image is calculated to obtain a structural similarity loss; the discriminator loss is obtained by combining the intensity loss, the gradient loss, and the structural similarity loss; the generator and the discriminator are adversarially trained based on the generator loss and the discriminator loss.

[0111] In the above embodiments, the loss function of the discriminator is introduced. Next, the loss function of the generator is introduced. Specifically, the generator loss function L G is calculated by the following expression

[0112] L G = L adv + αL con + βL SSIM (9)

[0113] where the parameters α and β are balance coefficients used to balance the adversarial loss, the content loss, and the structural similarity loss, and are set to 100 and 0.01 respectively.

[0114] Specifically, the adversarial loss function L adv is an important part of the generator loss function. Its calculation expression is

[0115]

[0116] where N represents the number of fused images in the training process, represents the nth fused image in each batch, a is set to 1, D(·) represents the probability that the fused image generated by the generator is determined to be real by the discriminator, refers to the Laplacian operator, which is used to generate the gradient map of the image. Through adversarial training, the generated image is made more similar to the real image in texture, improving the quality of the generated image.

[0117] The content loss function L con keeps the similarity between the fused image and the source image by combining the intensity loss and the gradient loss, ensuring the retention of important detail information. Its calculation expression is

[0118] L con = γ1L grad + γ2L int (11)

[0119] where γ is a balance factor used to balance the two information losses of intensity and gradient.

[0120] Intensity loss L int Mainly to ensure that the light intensity distribution of the fused image in the focused area is consistent with the source image and retain the key brightness information. Its calculation expression is

[0121]

[0122] where W and H represent the width and height of the image, I1 and I2 represent the input source images, and I fused is the fused image generated by the generator. i and j represent the pixels in the i-th row and j-th column of the gradient map or source image. M1 and M2 represent the feature maps of the source images after being processed by the MSRA module, and their sizes are also W×H.

[0123] Gradient loss L grad The calculation expression is:

[0124]

[0125] where refers to the Laplacian operator, which is used to calculate the gradient values of the fused image I fused and the source images I1 and I2.

[0126] Structural similarity loss function L SSIM The calculation expression is:

[0127]

[0128] where N represents the number of fused images, and SSIM(·) is used to calculate the structural similarity between the fused image I fused and the source image I g .

[0129] The generator loss and discriminator loss are calculated through the above formulas. The parameters of the generator network are updated through the generator loss, and the parameters of the discriminator network are updated through the discriminator loss. The training is repeated until the generator and discriminator reach equilibrium. After training, the generator can fuse the images to be fused and generate a more realistic fused image. Compared with the original images to be fused, the fused image output by the generator combines the focused areas of different images to be fused, thereby improving the quality of the fused image.

[0130] In this embodiment, a generative adversarial network (GAN) is used as the basic framework, and its core module is the multi-scale residual attention (MSRA) module. MSRA-GAN consists of two parts: a generator and a discriminator. First, a multi-branch dilated convolution module is used to capture image details and features at different scales, effectively reducing information loss caused by scale changes, and while maintaining local details, capturing a wider range of context information, thereby enhancing the detail richness and overall consistency of the fused image. At the same time, combined with the idea of residual learning, the generator is further strengthened, improving the performance and stability of the network. In addition, in order to optimize the accuracy of feature fusion and reduce information redundancy, a dual-channel hybrid attention module is integrated into each multi-scale residual unit of the generator. By compressing and exciting the feature maps in the channel and spatial dimensions, the generator can generate feature tensors of different dimensions, thereby enhancing the expression of key channels and spatial features, while effectively suppressing irrelevant information. This design enables the network to more precisely focus on the important parts of the image, further improving the feature expression ability. Finally, an adversarial training strategy based on game theory is adopted, enabling the generator to produce more realistic and natural fusion results, while the discriminator is responsible for capturing rich texture information in the fused image, further improving the quality of the fused image.

[0131] This embodiment also provides a multi-focus image fusion network model for fusing images with different focal lengths and outputting a fused image. The multi-focus image fusion network model specifically includes a generator and a discriminator; the generator includes a feature extraction module, a feature fusion module, and an image reconstruction module; multiple feature extraction modules respectively extract feature maps of source images at different focal lengths; the feature fusion module fuses the feature maps of each source image to obtain a fused feature map; the image reconstruction module reconstructs the fused feature map to obtain a reconstructed fused image; the discriminator predicts the determination probability of the fused image, and the determination probability is used to indicate the probability that the fused image is a real image; a loss function is calculated based on the determination probability, and the generator and the discriminator are subjected to adversarial training through the loss function, and the trained generator and discriminator are used to fuse the images to be fused at different focal lengths.

[0132] The specific details of each module or unit in the above multi-focus image fusion network model have been described in detail in the multi-focus image fusion method, so they will not be elaborated here.

[0133] This application embodiment also provides an electronic device, Figure 6 showing a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure. Figure 6 The electronic device 600 shown is only an example and should not bring any limitation to the functions and usage scopes of the embodiments of the present disclosure.

[0134] As shown Figure 6 in FIG. 2, the electronic device 600 includes a central processing unit (CPU) 601 that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage section 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for system operation are also stored. The CPU 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0135] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read from it can be installed into the storage section 608 as needed.

[0136] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable storage medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 609, and / or installed from the removable medium 611. When the computer program is executed by the central processing unit (CPU) 601, the above functions defined in the embodiments of the present application are executed.

[0137] It should be noted that the computer-readable medium shown in this disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. And in this disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.

[0138] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0139] The units involved in the embodiments described in this disclosure can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation on the unit itself.

[0140] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist alone without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and the one or more programs include instructions that, when executed by the electronic device, cause the electronic device to implement the methods described in the above embodiments.

[0141] It should be noted that although several modules or units of the devices for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0142] The above content is only the specific implementation manners of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A multi-focus image fusion method, characterized in that, Including: Inputting source images of the same scene at different focal lengths into a generator, the generator including a feature extraction module, a feature fusion module, and an image reconstruction module; Multiple feature extraction modules respectively extract features of the source images at different focal lengths, the feature extraction module including a multi-scale residual attention module, the multi-scale residual attention module extracting image features of different scales to obtain feature maps; The feature fusion module fuses the feature maps of each source image to obtain a fused feature map; The image reconstruction module reconstructs the fused feature map to obtain a reconstructed fused image; Inputting the gradient map of the fused image into a discriminator to obtain a determination probability output by the discriminator, and calculating a loss function based on the determination probability; Performing adversarial training on the generator and the discriminator through the loss function, and using the trained generator and discriminator to fuse the images to be fused at different focal lengths.

2. The multi-focus image fusion method according to claim 1, wherein The multiple feature extraction modules respectively extracting features of the source images at different focal lengths includes: Inputting source images of different focal lengths into different feature extraction modules, the feature extraction module including a first convolutional layer, multiple multi-scale residual attention modules, a first fusion module, and a second convolutional layer connected in sequence; The first convolutional layer performs a convolutional operation on the source image to obtain a feature map of the source image and inputs it to the first multi-scale residual attention module; Multiple of the multi-scale residual attention modules are connected in sequence and skip-connected to the first fusion module; Each multi-scale residual attention module extracts features from the input feature map, and the first fusion module splices the feature maps extracted by different multi-scale residual attention modules; The second convolutional layer performs a convolutional operation on the spliced feature map again to obtain a processed feature map.

3. The multi-focus image fusion method according to claim 2, wherein The multi-scale residual attention module includes a multi-branch dilated convolutional module; The multi-branch dilated convolutional module includes multiple branches, each branch including a convolutional layer with a convolutional kernel size of 1. Starting from the second branch, a dilated convolutional layer is connected after the convolutional layer, and the number of dilated convolutional layers in each branch increases sequentially; The output of each dilated convolutional layer is connected to the next dilated convolutional layer and the next dilated convolutional layer of the next branch; The outputs of the last dilated convolutional layer of each branch are spliced to obtain a first output feature map as the output of the multi-branch dilated convolutional module.

4. The multi-focus image fusion method according to claim 3, characterized in that The multi-scale residual attention module further includes a dual-channel hybrid attention module connected to the multi-branch dilated convolutional module; The dual-channel hybrid attention module includes a channel attention module and a spatial attention module; The channel attention module determines channel weights for the first output feature map and weights the first output feature map based on the channel weights to obtain a weighted first feature map; The spatial attention module determines spatial weights for the first output feature map and weights the first output feature map based on the spatial weights to obtain a weighted second feature map; Merge the first feature map and the second feature map to obtain a second output feature map as the output of the dual-channel hybrid attention module.

5. The multi-focus image fusion method according to claim 4, wherein The channel attention module determines channel weights for the first output feature map and weights the first output feature map based on the channel weights to obtain a weighted first feature map, including: The channel attention module performs average pooling on the feature maps of each channel in the first output feature map to obtain a first feature vector, and performs max pooling on the feature maps of each channel to obtain a second feature vector; and respectively performs dimensionality reduction and dimensionality increase on the first feature vector and the second feature vector in sequence to obtain a processed third feature vector and a fourth feature vector; uses a sigmoid activation function to perform activation processing on the third feature vector and the fourth feature vector respectively to obtain a first enhanced feature map and a second enhanced feature map, and multiplies the first enhanced feature map and the second enhanced feature map with the first output feature map element by element to obtain a first feature map.

6. The multi-focus image fusion method according to claim 4, characterized in that The spatial attention module determines spatial weights for the first output feature map and weights the first output feature map based on the spatial weights to obtain a weighted second feature map, including: The spatial attention module performs a convolution operation on the first output feature map, outputs a weight matrix with the same spatial dimension as the first output feature map, normalizes the weight matrix through a sigmoid activation function to generate a spatial attention map, and multiplies the spatial attention map with the first output feature map element by element to obtain a second feature map.

7. The multi-focus image fusion method according to claim 1, wherein Inputting the gradient map of the fused image into the discriminator to obtain the determination probability output by the discriminator, including: Obtain the gradient maps of source images with different focal lengths, and merge the gradient maps of source images with different focal lengths based on the maximum selection method to obtain a joint gradient map; Input the gradient map of the fused image and the joint gradient map into the discriminator together to obtain their respective determination probabilities.

8. The multi-focus image fusion method according to claim 7, wherein Adversarially training the generator and the discriminator through the loss function, including: Calculating a generator loss based on the determination probabilities corresponding to the fused image and the source image respectively; Calculating the difference between the light intensity distributions of the source images with different focal lengths and the corresponding fused images to obtain an intensity loss; Calculating the difference between the gradient map of the fused image and the gradient maps of source images with different focal lengths to obtain a gradient loss; Calculating the similarity between the fused image and the original image to obtain a structural similarity loss; Combining the intensity loss, the gradient loss, and the structural similarity loss to obtain a discriminator loss; Adversarially training the generator and the discriminator based on the generator loss and the discriminator loss.

9. A multi-focus image fusion network model, characterized in that, Including: A generator and a discriminator; The generator includes a feature extraction module, a feature fusion module, and an image reconstruction module; Multiple feature extraction modules respectively extract the features of source images at different focal lengths. The feature extraction module includes a multi-scale residual attention module, and the multi-scale residual attention module extracts image features at different scales to obtain feature maps; The feature fusion module fuses the feature maps of each source image to obtain a fused feature map; The image reconstruction module reconstructs the fused feature map to obtain a reconstructed fused image; The discriminator predicts the determination probability of the fused image; Based on the determination probability, a loss function is calculated, and the generator and the discriminator are adversarially trained through the loss function. The trained generator and discriminator are used to fuse the images to be fused at different focal lengths.

10. An electronic device, characterized in that, It includes a processor and a memory. One or more computer programs are stored in the memory. The one or more computer programs include instructions that, when executed by the electronic device, cause the electronic device to execute the multi-focus image fusion method according to any one of claims 1-8.