Multi-focus image fusion method and device, equipment and storage medium
By performing feature extraction, fusion and enhancement reconstruction of multi-focus images, combined with the feature contribution of different focus areas, the problems of complexity and insufficient accuracy of manual planning in the prior art are solved, and high-quality image fusion effect is achieved.
Patent Information
- Application Number
- CN202510593704.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-05
AI Technical Summary
The existing multi-focus image fusion method requires manual planning in feature extraction, focus area definition and fusion criteria, resulting in high professional knowledge requirements and high uncertainty. Deep learning-based methods lack accuracy when distinguishing focus areas from non-focus areas.
By extracting feature of the fusion multi-focus image pairs, using the preset feature fusion network and the enhanced reconstruction network for feature fusion and image processing, combining the feature contribution of different focus areas, a multi-level deep-supervised convolutional neural network is used for feature-level image fusion and enhanced reconstruction.
Improves the quality and fidelity of image fusion, ensuring the clarity and visual effects of the fusion image, especially in non-focused areas and high-frequency details.
Smart Images

Figure CN120430951A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image fusion technology, and in particular to a multi-focus image fusion method, device, equipment and storage medium. Background Art
[0002] Currently, commonly used multi-focus image fusion methods include traditional manual methods and deep learning-based methods. Traditional multi-focus image fusion mainly includes methods based on manually extracted features, spatial domains, and transform domains. For example, the weighted kernel method based on image gradients focuses on how to construct gradient-based decision graphs and perform image fusion through mathematical morphology; and multi-focus image fusion techniques designed for visual sensor networks utilize sharpening criteria in the wavelet domain. These methods respectively focus on improving image fusion effects through the separation of cartoon and texture content in images and wavelet transforms. In the implementation of most current image fusion technologies, key elements such as feature extraction, focus region definition, and fusion criteria require manual planning and setting. This manual design process itself is considered a highly challenging task, requiring not only a high level of expertise and experience but also facing numerous uncertainties in practice.
[0003] In deep learning-based methods, the network architecture involved essentially follows the design ideas of classification tasks. Specifically, by assigning a single binary label (i.e., 0 or 1) to the training image patch pairs to indicate whether they are focused or not, such a design framework may lead to inaccurate distinctions between the boundaries of focused and unfocused areas.
[0004] It can be seen that how to improve the quality of fused images is a problem to be solved in this field. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a multi-focus image fusion method, device, equipment, and storage medium that can extract image features at different feature levels, fuse image features based on the feature contributions of different focus areas, and process and fuse each fused feature map in the image domain to obtain a final fused image; thus, the overall quality and fidelity of the fused image can be improved. The specific scheme is as follows:
[0006] In a first aspect, the present application provides a multi-focus image fusion method, comprising:
[0007] Perform feature extraction on the fused multi-focus image pair to obtain image features at several feature levels;
[0008] Performing a feature fusion operation on the image features through a preset feature fusion network to obtain a corresponding fusion feature map; the preset feature fusion network includes feature contribution between different focus areas;
[0009] According to the image domain in which each of the fused feature maps is located, each of the fused feature maps is processed by a preset enhancement and reconstruction network, and the corresponding processed images are fused to obtain a fused image corresponding to the multi-focus image pair to be fused; wherein the preset enhancement and reconstruction network is used to perform image enhancement operations, image averaging operations, and image reconstruction operations.
[0010] Optionally, the feature extraction is performed on the multi-focus image pair to be fused to obtain image features at several feature levels, including:
[0011] Performing feature extraction on the multi-focus image pair to be fused by a preset feature extraction network to obtain a first image feature and a second image feature corresponding to two images of the multi-focus image pair to be fused respectively;
[0012] dividing the first image features and the second image features according to the feature levels to obtain image feature groups corresponding to the feature levels;
[0013] The preset feature extraction network includes convolutional layers corresponding to a plurality of feature levels; a single image feature group includes image features of the two images of the multi-focus image pair to be fused at the same feature level.
[0014] Optionally, performing a feature fusion operation on the image features through a preset feature fusion network to obtain a corresponding fusion feature map includes:
[0015] Determine the current image feature group;
[0016] Performing a feature fusion operation on the current image feature group through a network layer corresponding to the current feature level in a preset feature fusion network to obtain a corresponding fused feature map;
[0017] The current feature level is a feature level corresponding to the current image feature group.
[0018] Optionally, during the training process of the preset feature fusion network, the following steps are further included:
[0019] Learning and adjusting the network parameters of the current deep learning network through a back-propagation optimization algorithm to obtain a current adjusted network; the network parameters include parameters for feature contribution between different focus areas;
[0020] A performance test is performed on the current adjusted network, and when a corresponding performance test result indicates that a preset performance condition is satisfied, the current adjusted network is determined as a preset feature fusion network.
[0021] Optionally, processing each fused feature map by using a preset enhanced reconstruction network according to the image domain in which each fused feature map is located includes:
[0022] If the image domain where the fused feature map is located is a pixel space, the fused feature map is directly processed by a preset enhancement reconstruction method;
[0023] If the image domain where the fused feature map is located is a non-pixel space, the fused feature map is projected to the pixel space, and the feature map projected to the pixel space is processed by the preset enhanced reconstruction method.
[0024] Optionally, after processing each of the fused feature maps by using a preset enhanced reconstruction network, the method further includes:
[0025] Supervising each of the fused feature maps based on a first loss function to obtain a corresponding first loss result, so as to optimize the relevant network according to the first loss result;
[0026] Among them, the first loss function is a loss function constructed based on the multi-focus image pair to be fused, the image true value and each of the processed images.
[0027] Optionally, after obtaining the fused image corresponding to the multi-focus image pair to be fused, the method further includes:
[0028] Supervising the fused image based on a second loss function to obtain a corresponding second loss result, so as to optimize the relevant network according to the second loss result;
[0029] The second loss function is a loss function constructed based on the true value of the image and the fused image.
[0030] In a second aspect, the present application provides a multi-focus image fusion device, comprising:
[0031] A feature extraction module is used to extract features from the fused multi-focus image pair to obtain image features at several feature levels;
[0032] A feature fusion module is used to perform a feature fusion operation on the image features through a preset feature fusion network to obtain a corresponding fusion feature map; the preset feature fusion network includes feature contribution between different focus areas;
[0033] An image processing module is used to process each of the fused feature maps through a preset enhancement and reconstruction network according to the image domain in which each of the fused feature maps is located, and to fuse the corresponding processed images to obtain a fused image corresponding to the multi-focus image pair to be fused; wherein the preset enhancement and reconstruction network is used to perform image enhancement operations, image averaging operations, and image reconstruction operations.
[0034] In a third aspect, the present application provides an electronic device, comprising:
[0035] Memory, used to store computer programs;
[0036] A processor is used to execute the computer program to implement the multi-focus image fusion method as described above.
[0037] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements the multi-focus image fusion method as described above.
[0038] It can be seen that the present application first extracts features from the multi-focus image pair to be fused to obtain image features at several feature levels; then performs feature fusion operations on the image features through a preset feature fusion network to obtain corresponding fusion feature maps; the preset feature fusion network includes feature contributions between different focus areas; then, according to the image domain where each fusion feature map is located, each fusion feature map is processed through a preset enhancement and reconstruction network, and the corresponding processed images are fused to obtain a fused image corresponding to the multi-focus image pair to be fused; wherein the preset enhancement and reconstruction network is used to perform image enhancement operations, image averaging operations, and image reconstruction operations. In this way, the present application can extract image features at different feature levels; then fuse the relevant image features in combination with the feature contributions of different focus areas, perform image enhancement, reconstruction, and fusion on each fusion feature map in combination with the image domain to obtain the final fused image; it can improve the quality and fidelity of the fused image as a whole. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0040] Figure 1 This is a flow chart of a multi-focus image fusion method disclosed in this application;
[0041] Figure 2This is a flowchart of a specific multi-focus image fusion method disclosed in this application;
[0042] Figure 3 This is a flowchart of another specific multi-focus image fusion method disclosed in this application;
[0043] Figure 4 This is a schematic structural diagram of a multi-focus image fusion device disclosed in this application;
[0044] Figure 5 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION
[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0046] See also Figure 1 As shown, an embodiment of the present invention discloses a multi-focus image fusion method, comprising:
[0047] Step S11: extract features from the multi-focus image pair to be fused to obtain image features at several feature levels.
[0048] In this application, the multi-focus image pair to be fused is two images with identical content but different focus areas. After feature extraction, image features at several feature levels can be obtained. It is understood that image features include gradually abstract features ranging from simple to complex, such as edges, lines, and textures, as well as shapes and structures. Different types of image features correspond to different feature levels.
[0049] In a specific embodiment, the feature extraction of the multi-focus image pair to be fused to obtain image features at several feature levels may include: performing feature extraction on the multi-focus image pair to be fused using a preset feature extraction network to obtain first image features and second image features corresponding to the two images of the multi-focus image pair to be fused, respectively; dividing the first image features and the second image features according to the feature levels to obtain image feature groups corresponding to the feature levels, respectively; wherein the preset feature extraction network includes convolutional layers corresponding to the several feature levels; and a single image feature group includes image features of the two images of the multi-focus image pair to be fused at the same feature level. Specifically, feature extraction may be performed on the multi-focus image pair to be fused using a preset feature extraction network, the feature extraction network including several convolutional layers corresponding to different feature levels; corresponding first image features and second image features may be extracted for the two images in the multi-focus image pair to be fused, respectively; further, it can be understood that the low-frequency components of an image mainly include large-scale, smooth structures and uniform color areas, while the high-frequency components represent image details, such as edges, textures, and fine structures. The convolutional layer of the low-level feature extraction network, due to its smaller convolution kernel size and simple structure, tends to capture low-frequency information of the image and maintain the general outline and background consistency of the image. As the network level deepens, a larger receptive field (achieved through larger convolution kernels or pooling operations) and more complex filter configurations enable the network to capture higher-frequency details such as edges and textures, while maintaining an understanding of the global structure in high-level features. In this way, for a single input image, the features output by different network layers of the feature extraction network are image features of different feature levels. On this basis, the first image features and the second image features can be divided according to the different feature levels to obtain image feature groups corresponding to each feature level; the single image feature group here contains image features of the same feature level extracted from the input image pair.
[0050] Step S12: performing a feature fusion operation on the image features through a preset feature fusion network to obtain a corresponding fusion feature map; the preset feature fusion network includes feature contribution between different focus areas.
[0051] In the present application, the above steps can be used to obtain image features of different feature levels of the multi-focus image pair to be fused; then, a feature fusion operation can be performed on the image features corresponding to the two images of the multi-focus image pair to be fused; the feature fusion process can be carried out through a preset feature fusion network, which includes feature contributions for different focus areas and can weigh the fusion ratio of information in different focus areas; finally, a fusion feature map of the multi-focus image pair to be fused can be obtained.
[0052] In a specific embodiment, performing a feature fusion operation on the image features through a preset feature fusion network to obtain a corresponding fused feature map may include: determining a current image feature group; performing a feature fusion operation on the current image feature group through a network layer corresponding to the current feature level in the preset feature fusion network to obtain a corresponding fused feature map; wherein the current feature level is the feature level corresponding to the current image feature group. Specifically, an image feature group includes two image features of the same feature level extracted from the input image pair, and a feature fusion operation is performed on the corresponding two image features using the network layer corresponding to the current image feature group in the feature fusion network to obtain a fused feature map corresponding to a single feature level. In a specific embodiment, the two image features corresponding to a single feature level can be connected in series to form a dual-channel feature map for feature fusion.
[0053] In another specific embodiment, the training process of the preset feature fusion network may further include: learning and adjusting network parameters of the current deep learning network using a backpropagation optimization algorithm to obtain a current adjusted network; the network parameters include parameters for feature contribution between different focal regions; performing performance testing on the current adjusted network, and determining the current adjusted network as the preset feature fusion network when the corresponding performance test results indicate that it meets preset performance requirements. Specifically, performing fusion operations on each feature level essentially integrates information hierarchically based on the level of abstraction of the features. Low-level feature fusion helps preserve the general structure and texture of the image, while high-level feature fusion helps capture the precise contours and details of the object. This hierarchical processing ensures comprehensive and efficient utilization of information. When fusing features at different levels, a learning weight allocation mechanism is required to determine the contribution of each feature. This can be automatically learned through the backpropagation and optimization process of the deep learning network, allowing the network to learn to appropriately balance the information fusion ratio between different focal regions. After performance testing, a preset feature fusion network that meets the preset performance requirements is ultimately obtained and used to fuse image features at the same feature level in the multi-focus image pair to be fused.
[0054] In a specific embodiment, natural enhancement methods can be added to the fusion process, such as through network design, the fused image is not only improved in clarity, but also optimized in color, contrast, etc., to further enhance the visual effect.
[0055] Step S13: According to the image domain in which each of the fused feature maps is located, each of the fused feature maps is processed by a preset enhancement and reconstruction network, and the corresponding processed images are fused to obtain a fused image corresponding to the multi-focus image pair to be fused; wherein the preset enhancement and reconstruction network is used to perform image enhancement operations, image averaging operations, and image reconstruction operations.
[0056] In the present application, the above steps can be used to fuse the image features of the same feature level of the multi-focus image pair to be fused to obtain a corresponding fused feature map; then, each fused feature map can be fused through a preset enhancement and reconstruction network. Specifically, each fused feature map can be enhanced and reconstructed according to the image domain in which the fused feature map is located; then, each processed image is fused to obtain a fused image corresponding to the multi-focus image pair to be fused. It is understood that the process of processing the fused feature map by the preset enhancement and reconstruction network may include, but is not limited to, image enhancement operations, image averaging operations, and image reconstruction operations to improve the image display effect.
[0057] In a specific embodiment, the processing of each fused feature map through a preset enhancement and reconstruction network according to the image domain where each fused feature map is located may include: if the image domain where the fused feature map is located is a pixel space, then directly processing the fused feature map through a preset enhancement and reconstruction method; if the image domain where the fused feature map is located is a non-pixel space, then projecting the fused feature map to the pixel space, and processing the feature map projected to the pixel space through the preset enhancement and reconstruction method. Specifically, the enhancement and reconstruction is to further improve the image quality; if the fused feature map is directly located in the image domain, that is, the pixel space, then the enhancement process can be regarded as a direct image enhancement and image averaging operation; and if the fused feature map is in other domains, then the feature is first projected onto the image domain and then enhanced and averaged through the enhancement and reconstruction network.
[0058] In another specific embodiment, after processing each of the fused feature maps through a preset enhanced reconstruction network, the method may further include: supervising each of the fused feature maps based on a first loss function to obtain a corresponding first loss result, so as to optimize the relevant network according to the first loss result; wherein, the first loss function is a loss function constructed based on the multi-focus image pairs to be fused, the image true value, and each of the processed images. Specifically, the reconstruction process can be deeply supervised, and a first loss function can be constructed using the multi-focus image pairs to be fused, the image true value of the multi-focus image pairs to be fused, and the enhanced reconstructed image to characterize the loss of the reconstruction output; the loss result output by the loss function can be used to tune the relevant feature extraction network, feature fusion network, enhanced reconstruction network and other related networks to improve the performance of image fusion and enhancement.
[0059] In another specific embodiment, after obtaining the fused image corresponding to the multi-focus image pair to be fused, the method may further include: supervising the fused image based on a second loss function to obtain a corresponding second loss result, so as to optimize the relevant network based on the second loss result; wherein the second loss function is a loss function constructed based on the true image value and the fused image. Specifically, a second loss function can be constructed through a deep supervision image fusion process using the true image value of the multi-focus image pair to be fused and the fused image to characterize the loss of the fusion output; the loss result output by this loss function can be used to tune the relevant network to optimize network performance.
[0060] It can be seen that the present application can extract image features at different feature levels; then fuse the relevant image features in combination with the feature contribution of different focus areas, and perform image enhancement, reconstruction and fusion of each fused feature map in combination with the image domain to obtain the final fused image, which can improve the visual effect of image fusion; and the loss function can be used to simultaneously supervise the enhanced and reconstructed output and the final fused output, which can improve the overall quality and fidelity of the fused image.
[0061] like Figure 2 As shown, the embodiment of the present application discloses a multi-focus image fusion method, which specifically includes:
[0062] In this embodiment, combined with Figure 3 The multi-focus image fusion method flow chart shown in the figure can be implemented based on a joint multi-level deep supervised convolutional neural network, involving feature extraction, feature fusion, and enhanced reconstruction. Feature extraction involves designing a feature extractor based on deep learning methods for feature extraction, extracting features from the input image pair (source 1 and source 2) to obtain multi-level features. Feature fusion involves fusing features from different images through channel-level concatenation and convolution. Enhanced reconstruction involves deep multi-focus image fusion through multi-level supervision and the fusion of features from different levels and image reconstruction.
[0063] It is understandable that low-level features are good at capturing low-frequency content of images, but may sacrifice high-frequency details; while high-level features, although focused on high-frequency details, are limited in processing low-frequency content. The hierarchical structure of the neural network reflects the hierarchical nature of visual information processing in the brain's visual cortex, that is, the gradual abstraction from simple to complex features. As the network level deepens, larger receptive fields (achieved through larger convolution kernels or pooling operations) and more complex filter configurations enable the network to capture higher-frequency details such as edges and textures, while maintaining an understanding of the global structure in high-level features. The feature extraction module in this embodiment can be composed of a series of Conv+ReLU (convolution layer + activation function) layers, and each feature level can be expressed as:
[0064] ;
[0065] ;
[0066] in, represents the d-th level feature map of the n-th input image; and Represent the convolution filter and bias at the dth level for the nth input image, respectively. This series of operations gradually transitions from low-level features (such as image edges and textures) to more abstract high-level features (such as shape and structure). With each additional layer, the network better captures high-frequency details in the image while also preserving low-frequency content.
[0067] It's understandable that images with different focal regions contain complementary information. That is, the focused portion of one image may be the blurred portion of another. Feature fusion can combine this complementary information to render the details of the entire scene as clearly as possible. For example, if one image has a sharp foreground and a blurred background, while another image has a sharp background and a blurred foreground, fusing the two can yield a single image with both foreground and background in focus. Each focal region has its own distinct sharp areas, and the features of these areas should be preserved as much as possible. Feature fusion strategies should be designed with sufficient sophistication to ensure that while merging information, the details of the sharp portions of the original image are not destroyed, preserving the strengths of each. This requires the fusion algorithm to be able to distinguish which features should be emphasized and which should be suppressed. Performing fusion operations at each feature level essentially integrates information hierarchically based on the level of abstraction. Fusion of low-level features helps preserve the general structure and texture of the image, while fusion of high-level features helps capture the precise contours and details of objects. This hierarchical process ensures comprehensive and efficient utilization of information. When fusing features at different levels, a learned weighting mechanism is often required to determine the contribution of each feature. This can be automatically learned through the backpropagation and optimization process of the deep learning network, allowing the network to learn to appropriately balance the fusion ratio of information between different focus areas. In addition, natural enhancement methods can be added to the fusion process. For example, through network design, the fused image is not only improved in clarity, but also optimized in color, contrast, etc., further improving the visual effect. Furthermore, in this embodiment, given the image features extracted from the input image pair, a feature fusion operation is performed on each level of the feature to fuse the features of the corresponding level; the corresponding feature fusion formula is as follows:
[0068] ;
[0069] in, Represents the d-th level feature Fuse feature maps; and Represent the convolution filter and bias of the d-th level feature respectively; Indicates that the d-th level features of the input image pair are concatenated into a two-channel feature map. In this way, the information of different focus areas can be complemented while retaining their respective advantages.
[0070] Furthermore, in this embodiment, after feature fusion, the enhancement and reconstruction module is intended to further improve the image quality, especially for non-focused areas and high-frequency details. If the fused feature map is directly in the image domain, the enhancement process can be regarded as a direct enhancement and averaging operation. If the representation of the fused feature map is in other domains, the features can be first projected onto the image domain, and then enhanced and averaged. The purpose of the enhancement and reconstruction module is to further improve the quality of the fused image through a series of post-processing operations on the basis of preliminary fusion, especially for those high-frequency details in non-focused areas and images. The implementation of this link can deeply integrate the supervised learning concept in deep learning, ensuring that the fusion results are highly consistent with the real world. Among them, the relevant definition formulas of enhanced reconstruction are as follows:
[0071] ;
[0072] in, represents the reconstructed and enhanced image of the d-th level fusion feature, and and denote the convolution filter and bias of the d-th level fusion feature respectively.
[0073] This step ensures the fusion of low-level and high-level features. In low-level features, the image size is preserved, but the edge details are blurred; in high-level features, the edge details are clear, but information such as the overall image tone is lost. The combination of low-level and high-level features will make the final fused image contain both low-frequency content and high-frequency details. Finally, all reconstructed and enhanced images are combined into a multi-channel feature map, which is then fed into the convolutional layer to obtain the final output:
[0074] ;
[0075] in, Represents the convolutional layer weights fed into the last convolutional layer after the multi-channel feature map is connected. Represents the bias corresponding to the convolutional layer weights of the last convolutional layer.
[0076] Furthermore, it can be understood that combining the above operations can construct a multi-level convolutional neural network (MLCN), in which the convolution weights and biases can be optimized using gradient descent. Through multi-level deep supervised reconstruction, there are D+1 (D feature levels + 1 final output) objectives to minimize: the D outputs in supervised reconstruction include the enhanced and final outputs. For the reconstructed output, the following loss function is proposed:
[0077] ;
[0078] in, and is the input image pair, is the truth value; and represents the set of parameters of the neural network, and is the image reconstructed and enhanced from the dth fused feature. Furthermore, the loss function corresponding to the final output is:
[0079] ;
[0080] in, ; .
[0081] In a specific embodiment, the final loss function can be written as:
[0082] ;
[0083] in, Used to control the balance between two items.
[0084] Furthermore, during the training process, in order to use the second term mainly as regularization, we can use The effect of reducing the loss of reconstructed images is achieved by using a gradually decreasing strategy, as shown below:
[0085] ;
[0086] in, Decays with iteration t, where N is the total number of epochs; the initial value of t can be set to 0.4.
[0087] It can be seen that the present application can realize feature extraction, feature fusion and enhanced reconstruction of multi-focus image pairs through the constructed multi-level convolutional neural network, and finally obtain the corresponding fused image of the multi-focus image pair; image features of different feature levels can be extracted; then the relevant image features are fused in combination with the feature contribution of different focus areas, and the image enhancement reconstruction and fusion of each fused feature map are performed in combination with the image domain to obtain the final fused image; such a combination of multi-level outputs can extract visually distinct features, fuse and enhance; and, the multi-level outputs can be deeply supervised to improve the performance of image fusion and enhancement; especially for input images with common non-focus areas, anisotropic blur and slight misalignment, the quality and fidelity of the fused image can be improved as a whole.
[0088] like Figure 4 As shown, the embodiment of the present application discloses a multi-focus image fusion device, comprising:
[0089] A feature extraction module 11 is used to extract features from the multi-focus image pair to be fused, and obtain image features at several feature levels;
[0090] A feature fusion module 12 is configured to perform a feature fusion operation on the image features through a preset feature fusion network to obtain a corresponding fusion feature map; the preset feature fusion network includes feature contribution between different focus areas;
[0091] The image processing module 13 is used to process each of the fused feature maps through a preset enhancement and reconstruction network according to the image domain in which each of the fused feature maps is located, and to fuse the corresponding processed images to obtain a fused image corresponding to the multi-focus image pair to be fused; wherein the preset enhancement and reconstruction network is used to perform image enhancement operations, image averaging operations, and image reconstruction operations.
[0092] It can be seen that the present application can extract image features at different feature levels; then fuse the relevant image features based on the feature contributions of different focus areas, and perform image enhancement, reconstruction and fusion of each fused feature map in combination with the image domain to obtain the final fused image; it can improve the overall quality and fidelity of the fused image.
[0093] In a specific embodiment, the feature extraction module 11 may include:
[0094] A feature extraction unit is configured to extract features from the multi-focus image pair to be fused using a preset feature extraction network, and obtain a first image feature and a second image feature corresponding to the two images of the multi-focus image pair to be fused respectively;
[0095] A feature division unit is used to divide the first image features and the second image features according to the feature levels to obtain image feature groups corresponding to the feature levels respectively; wherein the preset feature extraction network includes convolution layers corresponding to several feature levels one by one; a single image feature group contains image features of the two images of the multi-focus image pair to be fused at the same feature level.
[0096] In a specific embodiment, the feature fusion module 12 may include:
[0097] A feature group determination unit, configured to determine a current image feature group;
[0098] A feature fusion unit is used to perform a feature fusion operation on the current image feature group through a network layer corresponding to the current feature level in a preset feature fusion network to obtain a corresponding fusion feature map; wherein the current feature level is the feature level corresponding to the current image feature group.
[0099] In a specific embodiment, the device may further include:
[0100] A network parameter adjustment module is used to learn and adjust the network parameters of the current deep learning network through a back-propagation optimization algorithm to obtain the current adjusted network; the network parameters include parameters for feature contribution between different focus areas;
[0101] The performance testing module is used to perform a performance test on the current adjusted network, and when the corresponding performance test result indicates that the current adjusted network meets the preset performance condition, determine the current adjusted network as a preset feature fusion network.
[0102] In a specific embodiment, the image processing module 13 may include:
[0103] A first image processing unit is configured to process the fused feature map directly using a preset enhancement and reconstruction method when the image domain where the fused feature map is located is a pixel space;
[0104] The second image processing unit is used to process the feature map projected to the pixel space by using the preset enhanced reconstruction method when the image domain where the fused feature map is located is a non-pixel space.
[0105] In a specific embodiment, the device may further include:
[0106] A first supervision module is used to supervise each of the fused feature maps based on a first loss function to obtain a corresponding first loss result, so as to optimize the relevant network according to the first loss result; wherein the first loss function is a loss function constructed based on the multi-focus image pair to be fused, the image true value and each of the processed images.
[0107] In another specific embodiment, the device may further include:
[0108] The second supervision module is used to supervise the fused image based on a second loss function to obtain a corresponding second loss result, so as to optimize the relevant network according to the second loss result; wherein the second loss function is a loss function constructed based on the true value of the image and the fused image.
[0109] Furthermore, the embodiment of the present application also discloses an electronic device, Figure 5 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram should not be considered as any limitation to the scope of application of the present application.
[0110] Figure 5This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps of the multi-focus image fusion method disclosed in any of the aforementioned embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0111] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0112] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0113] The operating system 221 is used to manage and control the hardware devices and computer program 222 on the electronic device 20, and can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of implementing the multi-focus image fusion method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 may further include a computer program capable of implementing other specific tasks.
[0114] Furthermore, this application discloses a computer-readable storage medium for storing a computer program; wherein, when executed by a processor, the computer program implements the multi-focus image fusion method disclosed above. The specific steps of this method can be found in the corresponding contents disclosed in the aforementioned embodiments and will not be repeated here.
[0115] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the methods.
[0116] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0117] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0118] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0119] The above is a detailed introduction to the technical solution provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A multi-focus image fusion method, characterized in that: include: Perform feature extraction on the fused multi-focus image pair to obtain image features at several feature levels; Performing a feature fusion operation on the image features through a preset feature fusion network to obtain a corresponding fusion feature map; the preset feature fusion network includes feature contribution between different focus areas; According to the image domain in which each of the fused feature maps is located, each of the fused feature maps is processed by a preset enhancement and reconstruction network, and the corresponding processed images are fused to obtain a fused image corresponding to the multi-focus image pair to be fused; wherein the preset enhancement and reconstruction network is used to perform image enhancement operations, image averaging operations, and image reconstruction operations.
2. The multi-focus image fusion method according to claim 1, characterized in that: The feature extraction is performed on the multi-focus image pair to be fused to obtain image features at several feature levels, including: Performing feature extraction on the multi-focus image pair to be fused by a preset feature extraction network to obtain a first image feature and a second image feature corresponding to two images of the multi-focus image pair to be fused respectively; dividing the first image features and the second image features according to the feature levels to obtain image feature groups corresponding to the feature levels; The preset feature extraction network includes convolutional layers corresponding to a plurality of feature levels; a single image feature group includes image features of the two images of the multi-focus image pair to be fused at the same feature level.
3. The multi-focus image fusion method according to claim 2, characterized in that: The performing a feature fusion operation on the image features through a preset feature fusion network to obtain a corresponding fusion feature map includes: Determine the current image feature group; Performing a feature fusion operation on the current image feature group through a network layer corresponding to the current feature level in a preset feature fusion network to obtain a corresponding fused feature map; The current feature level is a feature level corresponding to the current image feature group.
4. The multi-focus image fusion method according to claim 1, characterized in that: The training process of the preset feature fusion network also includes: Learning and adjusting the network parameters of the current deep learning network through a back-propagation optimization algorithm to obtain a current adjusted network; the network parameters include parameters for feature contribution between different focus areas; A performance test is performed on the current adjusted network, and when a corresponding performance test result indicates that a preset performance condition is satisfied, the current adjusted network is determined as a preset feature fusion network.
5. The multi-focus image fusion method according to claim 1, characterized in that: The processing of each fused feature map by a preset enhanced reconstruction network according to the image domain where each fused feature map is located includes: If the image domain where the fused feature map is located is a pixel space, the fused feature map is directly processed by a preset enhancement reconstruction method; If the image domain where the fused feature map is located is a non-pixel space, the fused feature map is projected to the pixel space, and the feature map projected to the pixel space is processed by the preset enhanced reconstruction method.
6. The multi-focus image fusion method according to any one of claims 1 to 5, characterized in that: After processing each of the fused feature maps by the preset enhanced reconstruction network, the method further includes: Supervising each of the fused feature maps based on a first loss function to obtain a corresponding first loss result, so as to optimize the relevant network according to the first loss result; Among them, the first loss function is a loss function constructed based on the multi-focus image pair to be fused, the image true value and each of the processed images.
7. The multi-focus image fusion method according to claim 6, characterized in that: After obtaining the fused image corresponding to the multi-focus image pair to be fused, the method further includes: Supervising the fused image based on a second loss function to obtain a corresponding second loss result, so as to optimize the relevant network according to the second loss result; The second loss function is a loss function constructed based on the true value of the image and the fused image.
8. A multi-focus image fusion device, characterized in that: include: A feature extraction module is used to extract features from the fused multi-focus image pair to obtain image features at several feature levels; A feature fusion module is used to perform a feature fusion operation on the image features through a preset feature fusion network to obtain a corresponding fusion feature map; the preset feature fusion network includes feature contribution between different focus areas; An image processing module is used to process each of the fused feature maps through a preset enhancement and reconstruction network according to the image domain in which each of the fused feature maps is located, and to fuse the corresponding processed images to obtain a fused image corresponding to the multi-focus image pair to be fused; wherein the preset enhancement and reconstruction network is used to perform image enhancement operations, image averaging operations, and image reconstruction operations.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the multi-focus image fusion method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that Used to store a computer program, which, when executed by a processor, implements the multi-focus image fusion method according to any one of claims 1 to 7.