Unified image fusion method and system based on adaptive distribution difference perception
Through the unified image fusion method of adaptive distribution difference perception, the cascade encoder, decoder and distribution difference perception fusion device are used to solve the problem of insufficient adaptability in complex task scenarios, and high-quality image fusion effect is achieved.
Patent Information
- Application Number
- CN202510926777.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-07-07
AI Technical Summary
The existing multi-source image fusion method is not adaptable enough in complex task scenarios, and it is impossible to effectively capture the correlation between the source images of different fusion tasks, resulting in the model being biased towards specific tasks and unable to generalize to more task scenarios.
A unified image fusion method for adaptive distribution difference perception is designed. By constructing a cascading encoder and decoder, combining the distribution difference perception fusion device and the output layer, a distribution difference perception loss function is used to dynamically adjust the fusion strategy to adapt to the feature differences of different source images.
It realizes robustness and generalization capabilities in complex scenarios, and can effectively integrate multiple fusion tasks such as multimodal, multi-exposure and multi-focus, output high-quality unified fusion images, and preserve target information and texture details of visible images.
Smart Images

Figure CN120410893A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a unified image fusion method and system for adaptive distribution difference perception. Background Art
[0002] Image fusion aims to integrate complementary information in the same scene captured by multiple sensors into one image, and can be widely applied to fields such as medical image enhancement and diagnosis, intelligent transportation environment perception, intelligent security, high-dynamic-range photography, and remote sensing image processing. One of the main challenges faced by current image fusion methods is how to effectively capture the correlation between source images of different fusion tasks and integrate heterogeneous information from different sources. When most existing methods adopt the correlation paradigm of static fusion rules, they ignore the dynamic adaptability of different task modalities, resulting in the model being biased towards specific tasks and unable to generalize to more task scenarios.
[0003] In recent years, researchers have begun to focus on the correlation paradigm of adaptive dynamic fusion to adapt to diverse task scenarios. However, these methods either rely on manually designed optimization objectives, which limits their stability in complex scenarios; or are based on certain assumptions and fail to comprehensively reflect the image distribution relationships of multiple task scenarios. In fact, the scene distributions of different tasks are significantly different, and even for the same task, the distributions from different data sets may also be different. For example, the feature distributions of multi-exposure images vary greatly, and the heterogeneity between infrared and visible light images creates higher scene distribution complexity. Summary of the Invention
[0004] The present invention provides a unified image fusion method and system for adaptive distribution difference perception, and the technical problem to be solved is: the insufficient adaptability of existing multi-source image fusion methods in complex task scenarios.
[0005] To solve the above technical problems, the present invention provides a unified image fusion method for adaptive distribution difference perception, including the steps of: Constructing a unified image fusion network; training the unified image fusion network; applying the trained unified image fusion network to fuse the source image pairs to be fused to obtain a unified fused image; The unified image fusion network includes a cascaded first encoder, a first decoder, a cascaded second encoder, a second decoder, a distribution difference perception fuser, and an output layer; the first encoder and the second encoder have the same structure and both include a plurality of cascaded encoding layers; the first decoder and the second decoder have the same structure and both include a plurality of cascaded decoding layers; the distribution difference perception fuser performs distribution difference perception fusion on the outputs of at least one encoding layer at the same level, and then splices the outputs with the outputs of the two encoding layers at the same level and inputs them into the encoding layer or decoding layer of the next level; the distribution difference perception fuser also performs distribution difference perception fusion on the outputs of at least one decoding layer at the same level, and then splices the outputs with the outputs of the two decoding layers at the same level and inputs them into the decoding layer or output layer of the next level; the output layer is used to obtain a unified fusion image according to the outputs of the first decoder and the second decoder.
[0006] Further, the distribution difference perception fuser performs distribution difference perception fusion on two input features and to obtain high-frequency and low-frequency fusion features , specifically including the steps of: Calculating and to obtain the low-frequency feature distribution difference weight and the high-frequency feature distribution difference weight ; Based on to , the low-frequency features , are adaptively modulated for different brightness channels respectively and then weighted and fused to obtain the low-frequency fusion feature ; Based on to , the high-frequency parts , are edge feature enhanced respectively and then weighted and fused to obtain the high-frequency fusion feature ; To and are weighted and fused to obtain the high-frequency and low-frequency fusion feature .
[0007] Further, the calculation of and to obtain the low-frequency feature distribution difference weight and the high-frequency feature distribution difference weight specifically includes: Extracting the low-frequency components and high-frequency components , and the low-frequency components and high-frequency components ; Calculate the mean and variance between the low-frequency components and and calculate the low-frequency feature distribution difference weight according to the mean and variance ; Calculate the mean and variance between the high-frequency components and and calculate the high-frequency feature distribution difference weight according to the mean and variance .
[0008] Furthermore, the calculation of the low-frequency feature distribution difference weight according to the mean and variance specifically includes: Calculate , take the absolute value of the difference between the means, and obtain the low-frequency mean difference ; Calculate , take the absolute value of the difference between the variances, and obtain the low-frequency variance difference ; Concatenate with to obtain the low-frequency feature distribution difference vector ; Pass through two fully connected layers and the corresponding activation functions to obtain the low-frequency feature distribution difference weight ; Based on the high-frequency components and between the means and variances, use the same process as calculating to calculate .
[0009] Furthermore, the based on for , the low-frequency features , are respectively subjected to adaptive modulation for different brightness channels and then weighted and fused to obtain the low-frequency fused feature , specifically including: Use the mixture of experts module to and respectively perform adaptive brightness normalization to obtain , ; The router weights of the mixture of experts module are obtained by Obtained by applying the Softmax and Top-K strategies; Through the fully connected layer and activation function for 、 Are processed respectively to obtain the corresponding channel attention weights 、 ; Based on For 、 Are weighted and fused to obtain ; Based on For 、 Are weighted and fused to obtain ; After splicing With And passing through two fully connected layers and the corresponding activation functions, the low-frequency fusion weight Is obtained; Based on For 、 Are weighted and fused to obtain the low-frequency fusion feature .
[0010] Furthermore, the high-frequency parts of For 、 Are respectively subjected to edge feature enhancement and then weighted and fused to obtain the high-frequency fusion feature 、 , specifically including: By using a feature design variance filter for 、 To generate the corresponding saliency probability maps And ; After multiplying 、 With their respective And And To obtain the detail enhancement feature maps And ; Based on For 、 Are weighted and summed to obtain the final edge enhancement feature ; Based on For 、 Are weighted and summed to obtain the final edge enhancement feature ; After splicing With After splicing, through two fully connected layers and corresponding activation functions, high-frequency fusion weights are obtained. ; Based on to , are weighted and fused to obtain high-frequency fusion features .
[0011] Furthermore, the weighted fusion of and to obtain high-low frequency fusion features specifically includes: For and , an initial fusion feature is generated through a global average pooling layer and a fully connected layer; For , a low-frequency channel weight and a high-frequency channel weight are obtained through a low-frequency component fully connected layer and a high-frequency component fully connected layer; After splicing , , the Softmax function is used to normalize the splicing result, and then the normalized result is split by channels to obtain a low-frequency channel fusion weight and a high-frequency channel fusion weight ; Based on , to and are weighted and fused to obtain high-low frequency fusion features .
[0012] Furthermore, during the training of the unified image fusion network, the designed loss function is the weighted sum of an intensity loss , a detail loss and a contrast loss . The intensity loss , the detail loss , and the contrast loss are respectively the brightness difference, the high-frequency component difference, and the contrast difference between the source image and the unified fusion image.
[0013] Furthermore, the weights , , and corresponding to the intensity loss are calculated through the following steps: Calculate the absolute value of the mean difference and the absolute value of the difference in standard deviation between pairs of source images ; After and pass through a multi - layer perceptron and Softmax, then perform channel - level separation to obtain weights .
[0014] The present invention also provides a unified image fusion system with adaptive distribution difference perception, which is characterized in that it includes a model construction module, a model training module and a model application module. The model construction module is used to construct a unified image fusion network, and the model training module is used to train the unified image fusion network; the model application module is used to apply the trained unified image fusion network to fuse the source image pairs to be fused, and obtain a unified fused image.
[0015] A unified image fusion method and system with adaptive distribution difference perception provided by the present invention designs a unified image fusion network for the feature distribution differences between different task source images. This network designs a distribution difference perception fusion device to dynamically distinguish the distribution differences of low - frequency and high - frequency features in the image, so as to finely adjust the fusion strategy to adapt to the feature differences of different source images. On this basis, the present invention also proposes a distribution difference perception loss function with weight adaptability, which is used to balance the contributions of different loss terms to the fusion task, and ensure the robustness and generalization ability of the network in complex scenarios. The trained unified image fusion network can not only effectively control the dominant deviation in different fusion tasks, but also integrate multiple fusion tasks such as multi - modality, multi - exposure and multi - focus into a unified framework. Experimental results show that the fusion effect of the present invention in multi - task scenarios is better than that of existing methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is the structural diagram of the unified image fusion network provided by the embodiment of the present invention; Figure 2 is the structural diagram of the distribution difference perception fusion device provided by the embodiment of the present invention; Figure 3 is the qualitative comparison diagram of the multi - exposure image fusion task on the MEFB dataset provided by the embodiment of the present invention; Figure 4 is the qualitative comparison diagram of the multi - focus image fusion task on the MFIFB dataset provided by the embodiment of the present invention; Figure 5 is the qualitative comparison diagram of the infrared and visible light image fusion task on the TNO dataset provided by the embodiment of the present invention; Figure 6 is the qualitative comparison diagram of the medical image fusion task on the Harvard Medical dataset provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The embodiments of the present invention will be specifically described below in conjunction with the accompanying drawings. The given examples are only for illustrative purposes and should not be construed as limiting the present invention. The accompanying drawings are only for reference and illustration and do not constitute a limitation on the scope of patent protection of the present invention, because many changes can be made to the present invention without departing from the spirit and scope of the present invention.
[0018] To achieve the adaptive fusion of paired source images in different fusion task scenarios, an embodiment of the present invention provides a unified image fusion method based on adaptive distribution difference perception, which includes the steps of: Construct a unified image fusion network; Train the unified image fusion network; Apply the trained unified image fusion network to fuse the source image pairs to be fused to obtain a unified fused image.
[0019] Figure 1 FIG. is the structural diagram of the unified image fusion network. As Figure 1 shown, the network includes a cascaded first encoder, a first decoder, a cascaded second encoder, a second decoder, a distribution difference-aware fusion unit (DistributionDifference-Aware Fusion, DDWF), and an output layer; the first encoder and the second encoder have the same structure and share parameters, and both include a cascaded plurality of encoding layers; the first decoder and the second decoder have the same structure, and both include a cascaded plurality of decoding layers; the distribution difference-aware fusion unit performs distribution difference-aware fusion on the outputs of at least one layer of the same-level encoding layer and then splices them with the outputs of the two encoding layers of the same level and inputs them into the next-level encoding layer or decoding layer; the distribution difference-aware fusion unit also performs distribution difference-aware fusion on the outputs of at least one layer of the same-level decoding layer and then splices them with the outputs of the two decoding layers of the same level and inputs them into the next-level decoding layer or output layer; the output layer is used to obtain a unified fused image according to the outputs of the first decoder and the second decoder.
[0020] As Figure 1 shown, given a pair of source images ( represent the height and width of the image respectively, and 3 represents the number of channels), the fusion result is denoted as . In the unified image fusion network, first, the source image pair Each input is fed into a patch embedding layer to extract feature tokens, which are then fed into the first and second encoders, respectively. The entire network is built based on the Vision Transformer (ViT) architecture, with both the encoder and decoder consisting of multiple Transformer blocks. This means that both the encoding and decoding layers utilize Transformer blocks. To perceive and control data distribution differences in multi-task scenarios, this method proposes a distribution difference-aware fusion agent. By introducing a distribution difference sensor, a brightness control module, and an edge control module, this method fine-grainedly adjusts the feature fusion process. Furthermore, this method designs a distribution difference-aware loss function to guide feature extraction and fusion.
[0021] Figure 2 This is the structural diagram of the distribution difference perception fusion. Figure 2 As shown in the figure, the distribution difference perception fusion includes a distribution difference sensor, a brightness control module, an edge control module and a high-low frequency feature modulator (HF-LFFM). Overall, the distribution difference perception fusion is used to transform the two input features into a single pixel. and Perform distribution difference perception fusion to obtain high and low frequency fusion features , specifically including the steps: S1, calculated using distribution difference perceptron and The difference weight of low-frequency feature distribution between and high-frequency feature distribution difference weight ; S2, using brightness control module based on right 、 Low-frequency characteristics 、 After adaptive modulation of different brightness channels, weighted fusion is performed to obtain low-frequency fusion features. ; S3, using edge control module based on right 、 The high frequency part 、 After performing edge feature enhancement, weighted fusion is performed to obtain high-frequency fusion features. ; S4, using high and low frequency characteristic modulator and Perform weighted fusion to obtain high and low frequency fusion features .
[0022] In this embodiment, step S1 (calculating and the weight of the low-frequency feature distribution difference and the weight of the high-frequency feature distribution difference ) specifically includes the steps: S11. Use the feature decoupling module to extract 's low-frequency component and high-frequency component , and 's low-frequency component and high-frequency component ; S12. Use the distribution difference weight generation module to calculate the mean and variance between the low-frequency components and and calculate the low-frequency feature distribution difference weight , calculate the mean and variance between the high-frequency components and and calculate the high-frequency feature distribution difference weight .
[0023] In step S11, in order to achieve dynamic decoupling of the feature map, this method uses learnable low-pass and high-pass filters to form a feature decoupling module to effectively separate the low-frequency and high-frequency components. In addition, by sharing the filters in the group dimension, the model complexity can be effectively reduced while maintaining feature diversity. Specifically, given any input feature , represents the spatial dimension, is the number of channels. After passing through the filter generation layer, a low-pass filter for each group of inputs is generated. The formula of the filter generation layer is as follows: , where is transformed from to , is the kernel size of the low-pass filter, represents the number of groups. BN is batch normalization, W is the convolution parameter, and GAP is global average pooling. Apply the Softmax function to each filter. The group-based operation has fewer parameters and lower complexity compared to generating filters for each pixel. In addition, to obtain the high-pass filter, this method subtracts the low-pass filter from the identity matrix (whose central value is 1 and the rest are 0): .
[0024] Next, the input feature Divided by group, each group of features is denoted as , where is the group index, and , apply a low-pass filter and a high-pass filter (both with a size of ) to each group of features respectively, so as to obtain the corresponding low-frequency and high-frequency components, and their calculation formulas are: , , where, is the channel index, and are the spatial coordinates, represents the offset of the filter kernel.
[0025] In summary, according to the definition of the feature decoupling module, given two source images and , after passing through the feature decoupling module , the low-frequency and high-frequency components are obtained respectively, and the formulas are: , .
[0026] In step S12, after decomposing the feature map into different frequency components, the distribution difference weight generation module generates the corresponding perceptual weights by calculating the mean and variance differences of the low-frequency and high-frequency features. This weight essentially reflects the difference degree of the two source images in key features (such as brightness, edge information, etc.). When there are large differences in these features between the source images, it indicates that more sufficient adjustment and adaptation are required during the fusion process to ensure that the finally obtained fusion result has high quality and consistency. Therefore, this weight can be regarded as a quantization index of the fusion difficulty across tasks and different scenarios. Specifically, this method first constructs a feature distribution difference vector, and by calculating the mean and variance differences of the image in the low-frequency and high-frequency components, quantifies the change degree of the features between the source images, thereby reflecting their differences.
[0027] Specifically, step S12 includes the steps of: S121. Calculate the , mean value of and take the absolute value of the difference to obtain the low-frequency mean difference ; S122. Calculate the , variance of and take the absolute value of the difference to obtain the low-frequency variance difference ; S123. Combine with Splicing is performed to obtain the low-frequency feature distribution difference vector ; S124, will After passing through two layers of fully connected layers and corresponding activation functions, the low-frequency feature distribution difference weight is obtained ; S125, based on high frequency components and The mean and variance between The same process calculation .
[0028] Taking low-frequency features as an example, low-frequency components mainly capture the global structure and brightness information of the image. Its mean represents the overall brightness distribution, while the variance reflects the contrast or texture uniformity of the image. By calculating the mean and variance of two low-frequency feature maps respectively and taking their absolute difference, a low-frequency difference vector can be obtained. : , in, is a channel-level splicing operation, Represents the dimension of the low-frequency feature map.
[0029] Similarly, for high-frequency feature distribution difference vector , its calculation formula is as follows: , in, Represents the dimension of the high-frequency feature map, 、 They represent the high-frequency mean difference and high-frequency variance difference, respectively.
[0030] In order to further generate distribution difference weights, this method designs a dynamic perception module. This module takes the feature difference vectors of high-frequency and low-frequency as input, passes through the fully connected layer and the activation layer, and generates channel-level adaptive weights. Through these weights, the framework can adjust the fine-grained fusion of high-frequency and low-frequency features according to the size of the difference during the fusion process. Specifically, the dynamic perception module is based on 、 Generate low-frequency and high-frequency feature distribution difference weights 、 The process is expressed by the formula: , , in, It is a linear mapping. The value of the Sigmoid activation function is between 0 and 1.
[0031] In this embodiment, step S2 specifically includes the following steps: S21. Use a mixture-of-experts module to and respectively perform adaptive brightness normalization to obtain , ; the router weights of the mixture-of-experts module are obtained by applying the Softmax and Top-K strategies to ; S22. Process , respectively through a fully connected layer and an activation function to obtain the corresponding channel attention weights , ; S23. Based on perform weighted fusion on , to obtain ; based on perform weighted fusion on , to obtain ; S24. After splicing with , pass through two fully connected layers and the corresponding activation functions to obtain the low-frequency fusion weight ; S25. Based on perform weighted fusion on , to obtain the low-frequency fusion feature .
[0032] In step S21, for the low-frequency components, this method uses an adaptive brightness normalization module to adjust the normalization intensity of the brightness channel using the distribution difference weights, aiming to eliminate the interference caused by brightness changes while retaining the robust structural information. The core idea is: based on the distribution difference of the low-frequency features of the input image, determine whether and to what extent to perform brightness normalization adjustment. To achieve this goal, this method designs a normalization module based on the mixture-of-experts (MoE) structure, which borrows the "dynamic routing" and "expert selection" mechanisms in MoE. Specifically, the MoE layer receives two features from the low-frequency part as inputs, and uses the adaptively generated feature distribution weights to allocate them to a group of brightness normalization experts. Each expert processes the input through a predefined brightness normalization function to achieve brightness modulation. The experts are defined as follows: , where and Represents learnable parameters that control the scaling and translation of normalized features, and represent the mean and standard deviation calculated independently in the spatial dimensions of each channel and instance, respectively. Based on the above definitions, given two input low-frequency features and , the process of performing adaptive brightness normalization through the mixture-of-experts module can be defined as: , , where, represents the number of experts, represents the router weights of the low-frequency feature , which are calculated through the Softmax and Top-K strategies, i.e., . This routing mechanism is used to guide the intensity of brightness normalization in the framework to complete adaptive brightness normalization.
[0033] In steps S22 to S25, the present method designs a dynamic fusion module based on the normalized low-frequency features and . This module takes the original low-frequency features and and the corresponding normalized features as joint inputs, and generates dynamic weights through a set of neural networks to achieve adaptive modulation for different brightness channels, and finally fuses the modulated features. Specifically, the module first processes and through a fully connected layer to generate attention weights for each channel. Here, the global average value of each channel is mapped through a fully connected layer, and the output is restricted within the range of through the Sigmoid activation function, so as to obtain the weights . Subsequently, these weights are multiplied by the original low-frequency features respectively and combined through weighted summation to obtain the dynamically modulated channel features . The formula of this module is defined as: , , The design of this module effectively alleviates the problem of information loss during the normalization process. Subsequently, the present method concatenates the two modulated features to form a preliminary fusion feature, and learns the weights of the fusion feature through two fully connected layers and corresponding activation functions. Finally, these weights are used to perform weighted fusion on the two modulated features to generate the final output feature . This process is defined as: , , in, is the sigmoid activation function, is the Relu activation function.
[0034] In this embodiment, step S3 specifically includes the following steps: S31, through the feature design variance filter 、 Generate the corresponding significance probability map and ; S32, will 、 With their respective and Multiply to get the detail enhancement feature map and ; S33, based on right 、 Perform weighted summation to obtain the final edge enhancement feature ;based on right 、 Perform weighted summation to obtain the final edge enhancement feature ; S34, will and After splicing, the high-frequency fusion weight is obtained through two fully connected layers and corresponding activation functions ; S35, based on right 、 Perform weighted fusion to obtain high-frequency fusion features .
[0035] Similar to the design of fusion rules for low-frequency components, high-frequency components primarily contain image details and edge information, reflecting texture, boundaries, and local variations. To meet the need for differentiated processing of high-frequency information in different task scenarios, this method proposes an adaptive control strategy based on distribution difference weights. This strategy consists of an adaptive edge enhancement module and an adaptive edge channel fusion module.
[0036] The adaptive edge enhancement module implements steps S31 to S33. For high-frequency components, the distribution difference weight is used to adjust edge enhancement. The purpose of edge enhancement is to enhance the prominent edge regions in the image, making the details of the image clearer and more prominent. In high-frequency components, edges usually represent the key structures and texture information in the image, playing an important role in the visual perception and analysis of the image. Therefore, precisely adjusting the enhancement intensity of the edge region can significantly improve the image quality. This method uses the distribution difference weight to adjust the intensity of edge enhancement. A large distribution difference between two feature maps indicates a significant difference in texture, structure, or details between them. At this time, the enhancement intensity of the edge should be increased to highlight these detail regions. This enhancement helps to better distinguish subtle structural differences and improve the visibility of details. Conversely, when the distribution difference between two feature maps is small, the intensity of edge enhancement can be appropriately reduced to prevent over-enhancement and mitigate noise interference. Based on this, this method proposes an adaptive edge enhancement module. The input of this module is two high-frequency features and , and the edge information of both is extracted by calculating the saliency edge probability map. Specifically, this module designs variance filters for the two features to generate the saliency edge probability map . To calculate the saliency probability maps and of the two high-frequency features, a saliency measure for each feature is generated based on local variance and global variance. Local variance reflects the variation of the image in the local region, while global variance can provide information about the overall variability of the image. Combining these two variances can effectively extract edge regions or significant details. The calculation formula of the saliency edge probability map is as follows: , , where is the size of the sliding window. The local variance is calculated through a very small receptive field, which results in the inability to accurately distinguish significant edge regions and noise detail images and is prone to the loss of significant edges. To address this problem, this method further calculates the stable global variance for the two high-frequency features. Through the learnable parameter , the framework can automatically adjust the relationship between global and local variances according to the characteristics of the data, improving the adaptability of the framework.
[0037] Next, the two high-frequency features are multiplied by the corresponding saliency edge probability maps respectively to obtain the features with enhanced details. The intensity of enhancement is adjusted using the weight distribution difference, which is multiplied by the enhanced features and added to the original feature maps to obtain the final edge-enhanced features and of the two high-frequency features: , , , , The adaptive edge channel fusion module implements steps S34 and S35. In this method, the enhanced high-frequency features and are concatenated as the input of the fusion module. The concatenated features are processed by two fully connected layers to learn the fusion weights : .
[0038] Subsequently, these weights are used to perform weighted fusion on the two modulated features, and finally the fused high-frequency features are generated: .
[0039] In this embodiment, step S4 specifically includes the steps: S41. For and , an initial fusion feature is generated through a global average pooling layer and a fully connected layer; S42. For , a low-frequency channel weight and a high-frequency channel weight are obtained through a low-frequency component fully connected layer and a high-frequency component fully connected layer; S43. After , are concatenated, the Softmax function is used to normalize the concatenated result, and then the normalized result is split by channels to obtain a low-frequency channel fusion weight and a high-frequency channel fusion weight ; S44. Based on , perform weighted fusion on and to obtain a high-low frequency fusion feature .
[0040] After obtaining the fused high-frequency feature and the low-frequency feature , in order to highlight the truly useful information in the reconstruction process, this method introduces a high-frequency and low-frequency feature modulator (HF-LFFM). Given two feature maps Fusion feature The calculation process is defined as: , where represents the weight parameters of the fully connected layer. To generate the channel weights, this method uses two additional fully connected layers. By concatenating their outputs and normalizing the concatenated result using the Softmax function, the formula is as follows: , where and represent the parameters of the fully connected layers for the low-frequency and high-frequency components respectively, represents the splitting of the features along the channel dimension. Finally, through adaptive fusion , the final low-frequency and high-frequency fusion features are obtained, and this process is defined as: .
[0041] Existing unified image fusion methods are mainly divided into two categories: one category uses fixed optimization objectives to uniformly optimize different task scenarios, and the other category relies on task-related loss functions to achieve data-driven adaptive optimization. However, these two strategies often lead to task biases or ignore the commonalities among multiple tasks, thereby restricting the performance of the fusion model in complex and diverse scenarios. Therefore, it is particularly crucial to design a dynamic general loss function that can simultaneously take into account the requirements of each task and scenario. For this purpose, this method proposes a distribution difference perception loss function, which dynamically adapts to multi-task scenarios by dynamically adjusting the weights of various losses. The formula is as follows: , where the intensity loss , the detail loss , and the contrast loss are the brightness difference, high-frequency component difference, and contrast difference between the source image and the unified fusion image respectively, are the adaptive weights of the intensity loss , the detail loss , and the contrast loss respectively. The intensity loss is used to ensure that the brightness of the fusion image is consistent with the source image, reflecting the commonality in brightness information among tasks. The detail loss uses the Sobel gradient operator to extract high-frequency information (such as edges) and compares the high-frequency components of the fusion image and the source image to meet the personalized requirements of different tasks in terms of detail retention. The contrast loss optimizes its global and local contrast by comparing the histogram distributions of the fusion image, emphasizing the differentiated requirements of specific tasks for the global visual quality.
[0042] In a multi-task image fusion scenario, there are not only features that embody the commonalities of tasks (such as brightness consistency), but also characteristics that reflect task differences (such as detail preservation and contrast enhancement), and the statistical distribution differences among different scenarios within the same task further exacerbate the optimization difficulty. Therefore, this method proposes a weight generation mechanism based on distribution differences, which dynamically generates adaptive weights according to the statistical characteristics of the source images. When the input is the absolute value of the mean difference and the absolute value of the standard deviation , the formula for weight generation is: , where , represents the mean calculation function, represents the standard deviation calculation function. After obtaining a single output through the multi-layer perceptron , and then through the Softmax operation, and then through channel-level separation , the generated weights have non-negativity and normalization.
[0043] It should be noted that the various forms of processes shown above can be used, reordered, steps added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the results expected by the technical solution of the present invention can be achieved. This embodiment does not limit this here.
[0044] To apply the above unified image fusion method, an embodiment of the present invention also provides a unified image fusion system with adaptive distribution difference perception, which includes a model construction module, a model training module, and a model application module. The model construction module is used to construct a unified image fusion network, and the model training module is used to train the unified image fusion network; the model application module is used to apply the trained unified image fusion network to fuse the source image pairs to be fused to obtain a unified fused image. Since the operations implemented by each module in the system have been described in detail above, they will not be elaborated here.
[0045] The embodiments described in the present invention can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0046] The computer programs for implementing the methods and systems of the present invention can be written in any combination of one or more programming languages and stored in a computer-readable storage medium. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0047] A computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a compact disc read-only memory (CD ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0048] In summary, the unified image fusion method and system provided by the embodiments of the present invention design a unified image fusion network for the feature distribution differences between different task source images. The network designs a distribution difference perception fuser to dynamically distinguish the distribution differences of low-frequency and high-frequency features in the image, so as to finely adjust the fusion strategy and adapt to the feature differences of different source images. On this basis, the present invention also proposes a distribution difference perception loss function with adaptive weights, which is used to balance the contributions of different loss terms to the fusion task and ensure the robustness and generalization ability of the network in complex scenarios. The trained unified image fusion network can not only effectively control the dominant deviation in different fusion tasks, but also integrate multiple fusion tasks such as multi-modal, multi-exposure, and multi-focus into a unified framework, which can simultaneously retain the significant information of the target and the texture details of the visible light image, thus achieving a more natural and balanced fusion effect and outputting high-quality super-resolution results.
[0049] The following is an experimental verification.
[0050] In this embodiment, experiments are carried out in four fusion task scenarios, namely multi-exposure image fusion (MEF), multi-focus image fusion (MFF), visible light and infrared image fusion (IVF), and medical image fusion (MMF). The training data set used is composed of a mixture of multiple image fusion task data sets. Specifically, for the multi-exposure image fusion task, 589 pairs of images are selected from the SCIE data set, and the underexposed and overexposed images in this sequence are used as inputs; for the multi-focus image fusion task, 710 pairs of images from the RealMFF data set are used; for the infrared image fusion task, 1000 pairs of images are selected from the LLVIP data set and 2000 pairs of images are selected from the M3FD data set respectively; for the medical image fusion task, 600 pairs of images from the Harvard Medical data set are selected, covering three categories of MRI-CT, MRI-SPECT, and MRI-PET.
[0051] To comprehensively evaluate the performance of the proposed method, this embodiment conducts tests on multiple existing benchmark data sets, covering four types of tasks: MEF, MFF, IVF, and MMF. Specifically, the MEF task is evaluated using the MEFB data set; the MFF task follows the MFIFB benchmark setting; for the IVF task, it is tested on the TNO and RoadScene data sets; and the MMF task is evaluated on the Harvard Medical data set for three categories of MRI-CT, MRI-PET, and MRI-SPECT respectively.
[0052] In this embodiment, six metrics are selected from four major categories of evaluation metrics to quantitatively measure the performance of the fusion result. These metrics are: AG (Average Gradient, which measures the image sharpness or edge sharpness), EN (Entropy, which reflects the texture complexity or information content of the image), Q abf (edge Preservation Information Transfer Factor, which describes the degree of edge information retained in the fused image), SF (Spatial Frequency, which measures the texture and detail performance of the fused image), MS-SSIM (Multi-Scale Structural Similarity, which evaluates the structural similarity between two images), and FMI (Feature Mutual Information, which measures the correlation between two images by calculating the mutual information of features such as edges and textures in the two images).
[0053] In this embodiment, two experimental settings are used to comprehensively compare the proposed method: one is the unified image fusion method, and the other is the task-specific fusion method. Specifically, for the unified fusion method, the baselines include U2Fusion [1] 、SDNet [1] 、SwinFusion [1] 、DeFusion [1] 、MUFusion [2] 、DDBFusion [1] 、TC-MoA [3] ; while in the comparison of the task-specific fusion method, for the MFF task, MFF-GAN [4] 、ZMFF [5] and DB-MFIF [6] are selected as the comparison baselines; for the MEF task, TransMEF [7] 、BHFMEF [8] and HSDS-MEF [9] are used as the comparison methods; for the IVF task, the proposed method is compared with LRRNet [1] 、DDFM [1] and EMMA [1] ; and in the MMF task, MATR
[10] 、DDFM and EMMA are selected as the comparison methods.
[0054] The experiment was conducted on a high-performance server equipped with four NVIDIA GeForce RTX 3090 GPUs. For different image fusion tasks, all training samples were uniformly and randomly cropped to a size of 224×224 in the preprocessing stage. This operation helps to standardize the input images and reduce task biases that may be caused by differences in image sizes. During the training process, the model underwent a total of 60 epochs to ensure sufficient training and convergence. To optimize the training process, the batch size was set to 2, and the AdamW optimizer was used. The initial learning rate was set to 8.0×10 -5 , and as the training progressed, the learning rate was gradually reduced to 1.0×10 -6 . Through this strategy of dynamically adjusting the learning rate, overfitting can be effectively prevented and the model can be helped to more finely adjust the weights in the later stage.
[0055] The model incorporates the Transformer architecture to enhance the ability to capture image features and the fusion effect. In the Transformer blocks of the fourth to eighth layers of the encoder and the first to fifth layers of the decoder, a distribution difference-aware fusion module is embedded. The introduction of this module enables the model to dynamically capture the distribution characteristics between different source images and flexibly adjust the fusion strategy by deeply analyzing the differences between each pair of source images in the feature space.
[0056] The experiment comprehensively evaluated the fusion performance of the proposed method through six quantitative metrics. The specific results of this method and other comparison methods on the MEFB dataset are shown in Table 1 and Figure 3 . Figure 3 . In (a), overexposed image; (b), underexposed image; (c), TransMEF; (d), BHFMEF; (e), HSDS-MEF; (f), U2Fusion; (g), SDNet; (h), DeFusion; (i), SwinFusion; (g), MUFusion; (k), DDBFusion; (l), TC-MoA; (m), this method.
[0057] Table 1: Quantitative comparison of multi-exposure image fusion tasks on the MEFB dataset
[0058] As can be seen from Table 1, in comparison with the current state-of-the-art unified image fusion methods, our method performs best, demonstrating excellent compatibility in various fusion tasks. Especially in information-theoretic metrics (such as EN, MS-SSIM), our method has obvious advantages, indicating that the fused images generated by it can retain more source image information and are more in line with the characteristics of human visual perception. Although many image fusion methods specifically designed for single tasks adopt complex task-specific strategies, our method also achieves excellent results in the competition with the task-specific multi-exposure fusion method HSDS-MEF. In addition, the good performance of our method in metrics such as MS-SSIM and FMI further proves its advantages in retaining structural and gradient information. As Figure 3 shown, our model is significantly superior to other methods in visual quality: for example, the outlines of the trees outside the window are clearer and the background textures are finer. It is worth noting that our model can not only directly generate natural color images, but also achieve color reconstruction through post-processing of grayscale images.
[0059] Our method has been systematically compared and evaluated with three task-specific methods (MFF-GAN, ZMFF, and DB-MFIF) and eight unified image fusion methods in the multi-focus image fusion task. On the MFIFB dataset, we used six quantitative metrics to evaluate the performance of each method in detail, and the evaluation results are shown in Table 2 and Figure 4 shown. Figure 4 In it, (a) near-focus image, (b) far-focus image, (c) MFF-GAN, (d) ZMFF, (e) DB-MFIF, (f) U2Fusion, (g) SDNet, (h) DeFusion, (i) SwinFusion, (g) MUFusion, (k) DDBFusion, (l) TC-MoA, (m) our method.
[0060] Table 2: Quantitative comparison of multi-focus image fusion tasks on the MFIFB dataset
[0061] Table 2 and Figure 4 The experimental results show that our method is highly competitive in most metrics, especially achieving excellent performance in unified image fusion tasks. At the same time, our method also performs well in the comparison with the latest task-specific multi-focus image fusion method DB-MFIF, which indicates that our method has obvious advantages in retaining the unique details of the source images and presents higher fusion quality in human visual perception. The qualitative comparison results (see Figure 4)This conclusion was further verified: The fused images generated by this method are superior to other methods in terms of both texture and color consistency. U2Fusion has color deviation in the far-focus region, while DDBFusion shows blurring in the near-focus region. At the same time, it is difficult for these methods to effectively maintain font clarity and color information.
[0062] On the TNO dataset, six quantitative metrics were used to comprehensively evaluate the proposed model and each comparative method in the infrared and visible image fusion task. The experimental results are shown in Table 3 and Figure 5 as follows. Figure 5 In it, (a) visible image, (b) infrared image, (c) LRRNet, (d) DDFM, (e) EMMA, (f) U2Fusion, (g) SDNet, (h) DeFusion, (i) SwinFusion, (g) MUFusion, (k) DDBFusion, (l) TC-MoA, (m) this method.
[0063] Table 3: Quantitative comparison of infrared and visible image fusion tasks on the TNO dataset
[0064] The experimental results in Table 3 show that the fusion framework proposed by this method not only matches or even surpasses the existing unified image fusion methods and task-specific methods in terms of performance, but is also significantly superior to other latest methods. Figure 5 It is further proved that in low-light environments, this method can simultaneously retain the significant information of the target and the texture details of the visible image, thus achieving a more natural and balanced fusion effect; while other methods often suffer from problems such as over-dark scenes, loss of significant information or texture details, and it is difficult to achieve an ideal balance between details and significance.
[0065] In the SPECT-MRI fusion task on the Harvard Medical dataset, this method systematically evaluated the performance of eleven methods using six quantitative metrics. The experimental results are shown in Table 4 and Figure 6 as follows. Figure 6 In it, (a) SPECT image, (b) MRI image, (c) MATR, (d) DDFM, (e) EMMA, (f) U2Fusion, (g) SDNet, (h) DeFusion, (i) SwinFusion, (g) MUFusion, (k) DDBFusion, (l) TC-MoA, (m) this method.
[0066] Table 4: Quantitative comparison of medical image fusion tasks on the Harvard Medical dataset
[0067] The experimental results in Table 4 show that the proposed method achieved the best performance in all metrics, fully demonstrating its excellent effectiveness in preserving edge details, maintaining source image information, and capturing key features. As Figure 6 shown, although U2Fusion and SDNet performed well in processing functional and structural information, they tended to introduce noise; SwinFusion could make full use of functional information but often ignored key structural details; MUFusion tried to balance between structure and function but still had obvious artifacts; while DDBFusion showed obvious color deviations. In contrast, the proposed method demonstrated better visual quality, effectively retaining the complementary information of the source image and exhibiting more excellent fusion performance.
[0068] The experimental results indicate that the proposed invention has a better fusion effect in multi-task scenarios than existing methods, achieving a more natural and balanced fusion effect.
[0069] Cited references: 1. Liu J, Wu G, Liu Z, et al. Infrared and visible image fusion: From data compatibility to task adaption[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024; 2. Cheng C, Xu T, Wu X J. MUFusion: A general unsupervised image fusion network based on memory unit[J]. Information Fusion, 2023, 92: 80 - 92; 3. Zhu P, Sun Y, Cao B, et al. Task-customized mixture of adapters for general image fusion[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2024: 7099 - 7108; 4. Zhang H, Le Z, Shao Z, et al. MFF-GAN: An unsupervised generative adversarial network with adaptive and gradient joint constraints for multi-focus image fusion[J]. Information Fusion, 2021, 66: 40-53; 5. Hu X, Jiang J, Liu X, et al. ZMFF: Zero-shot multi-focus image fusion[J]. Information Fusion, 2023, 92: 127-138; 6. Zhang J, Liao Q, Ma H, et al. Exploit the best of both end-to-end and map-based methods for multi-focus image fusion[J]. IEEE Transactions on Multimedia, 2024, 26: 6411-6423; 7. Qu L, Liu S, Wang M, et al. TransMEF: A transformer-based multi-exposure image fusion framework using self-supervised multi-task learning[C] / / Proceedings of the AAAI conference on artificial intelligence. 2022, 36(2): 2126-2134; 8. Mu P, Du Z, Liu J, et al. Little strokes fell great oaks: Boosting the hierarchical features for multi-exposure image fusion[C] / / Proceedings of the 31st ACM International Conference on Multimedia. 2023: 2985-2993; 9. Wu G, Fu H, Liu J, et al. Hybrid-supervised dual-search: Leveraging automatic learning for loss-free multi-exposure image fusion[C] / / Proceedings of the AAAI conference on artificial intelligence. 2024, 38(6): 5985-5993; 10. Tang W, He F, Liu Y, et al. MATR: Multimodal medical image fusion via multiscale adaptive transformer[J]. IEEE Transactions on Image Processing, 2022, 31: 5134-5149。
[0070] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. An adaptive distribution difference-aware unified image fusion method, characterized in that, Including the steps: Construct a unified image fusion network; train the unified image fusion network; apply the trained unified image fusion network to fuse the source image pairs to be fused to obtain a unified fused image; The unified image fusion network includes a cascaded first encoder, a first decoder, a cascaded second encoder, a second decoder, a distribution difference perception fusion unit and an output layer; the first encoder and the second encoder have the same structure and both include a plurality of cascaded encoding layers; the first decoder and the second decoder have the same structure and both include a plurality of cascaded decoding layers; the distribution difference perception fusion unit performs distribution difference perception fusion on the outputs of at least one layer of the same-level encoding layers and then splices them with the outputs of the two encoding layers of the same level respectively and inputs them into the next-level encoding layer or decoding layer; the distribution difference perception fusion unit also performs distribution difference perception fusion on the outputs of at least one layer of the same-level decoding layers and then splices them with the outputs of the two decoding layers of the same level respectively and inputs them into the next-level decoding layer or output layer; The output layer is used to obtain a unified fused image according to the outputs of the first decoder and the second decoder.
2. The unified image fusion method with self-adaptive distribution difference perception according to claim 1, wherein The distribution difference-aware fusion unit performs distribution difference-aware fusion on the two input features and to obtain high-low frequency fusion features , which specifically includes the steps: Calculation and the weight of the low-frequency feature distribution difference and the weight of the high-frequency feature distribution difference ; Based on for 、 low-frequency features 、 perform adaptive modulation for different luminance channels respectively and then perform weighted fusion to obtain low-frequency fusion features ; Based on For 、 high-frequency part 、 After edge feature enhancement and weighted fusion are performed respectively, a high-frequency fusion feature is obtained ; Pair and are weighted and fused to obtain the high-low frequency fused feature .
3. An adaptive distribution difference-aware unified image fusion method according to claim 2, characterized in that The calculation and the weight of the difference in the distribution of low-frequency features and the weight of the difference in the distribution of high-frequency features , specifically including: Extract low-frequency components of and high-frequency components of , as well as low-frequency components of and high-frequency components of ; Calculate the low-frequency components and Calculate the mean and variance between them and calculate the difference weight of the low-frequency feature distribution based on the mean and variance ; Calculate high-frequency components and Calculate the mean and variance between them, and calculate the high-frequency feature distribution difference weight based on the mean and variance .
4. An adaptive distribution difference-aware unified image fusion method according to claim 3, characterized in that, Calculating the difference weight of the low-frequency feature distribution based on the mean value and variance Specifically, it includes: Calculation , calculate the mean value and take the absolute value of the difference to obtain the low-frequency mean difference ; Calculate , to obtain the low-frequency variance difference by calculating the variance of and taking the absolute value of their difference; Concatenate with to obtain the low-frequency feature distribution difference vector ; By passing through two fully connected layers and corresponding activation functions, the weight of the low-frequency feature distribution difference is obtained ; Based on the mean and variance between the high-frequency components and , calculate using the same process as calculating . .
5. An adaptive distribution difference-aware unified image fusion method according to claim 2, wherein The one based on pair , low-frequency features , are adaptively modulated for different luminance channels respectively and then weighted and fused to obtain low-frequency fusion features , specifically including: Use a mixture of experts module to and respectively perform adaptive brightness normalization to obtain 、 ; The router weights of the mixture of experts module are obtained by applying the Softmax and Top-K strategies to ; Process and respectively through the fully connected layer and the activation function to obtain the corresponding channel attention weights and ; , respectively through the fully connected layer and the activation function to obtain the corresponding channel attention weights , ; Based on Perform weighted fusion on and to obtain ; Based on Perform weighted fusion on and to obtain ; After splicing with , the low-frequency fusion weight is obtained through two fully-connected layers and corresponding activation functions; Based on pair and perform weighted fusion to obtain low-frequency fusion features .
6. An adaptive distribution difference-aware unified image fusion method according to claim 2, wherein The said based on pair , high-frequency part of , After respectively performing edge feature enhancement and then performing weighted fusion, a high-frequency fusion feature is obtained, specifically including: By designing a variance filter for features , generate corresponding saliency probability maps and ; Multiply , with their respective and to obtain the detail-enhanced feature maps and ; Based on Perform weighted summation on and to obtain the final edge enhancement feature ; Based on Perform weighted summation on and to obtain the final edge enhancement feature ; After splicing and , high-frequency fusion weights are obtained through two fully connected layers and corresponding activation functions; Based on pair , perform weighted fusion to obtain high-frequency fusion features .
7. An adaptive distribution difference perception-based unified image fusion method according to claim 6, characterized in that, The pair of and are weighted and fused to obtain the high-low frequency fused feature , specifically including: Pair And Generate initial fusion features through a global average pooling layer and a fully connected layer ; Pair The low-frequency channel weights are obtained through the low-frequency component fully-connected layer and the high-frequency component fully-connected layer and the high-frequency channel weights ; After splicing and , the Softmax function is used to normalize the splicing result, and then the normalized result is split by channels to obtain the low-frequency channel fusion weight and the high-frequency channel fusion weight ; Based on , , perform and weighted fusion to obtain high-low frequency fusion features .
8. An adaptive distribution difference perception-based unified image fusion method according to any one of claims 1 to 7, characterized in that: During the training of the unified image fusion network, the loss function adopted is designed as the intensity loss , the detail loss and the contrast loss weighted sum. The intensity loss , the detail loss , and the contrast loss are the brightness difference, the high-frequency component difference, and the contrast difference between the source image and the unified fusion image, respectively.
9. An adaptive distribution difference perception-based unified image fusion method according to claim 8, characterized in that, Strength loss , Detail loss and Contrast loss The corresponding weights , are calculated through the following steps: Calculate the absolute value of the mean difference between source image pairs and the absolute value of the difference in standard deviations ; After and pass through a multi-layer perceptron and Softmax, channel-level separation is then performed to obtain the weight .
10. An adaptive distribution difference-aware unified image fusion system, characterized in that: Including a model construction module, a model training module and a model application module, the model construction module is used to construct a unified image fusion network, the model training module is used to train the unified image fusion network; the model application module is used to apply the trained unified image fusion network to fuse the source image pairs to be fused to obtain a unified fused image.
Citation Information
Patent Citations
PET-MRI image fusion method based on adaptive generative adversarial network
CN115457359A
Infrared image and visible light image fusion method and system based on deep learning
CN120107089A
Image fusion method and system based on two-stage adversarial training and edge perception
CN120163719A
Method and apparatus of fusing image, method and apparatus of training image fusion model, electronic device, storage medium and computer program
JP2023001926A