Three-modal image fusion method and system based on cross attention mechanism
Through the three-modal image fusion method based on the cross attention mechanism, the image fusion problem in multi-source heterogeneous data and high-noise environment is solved, and high-precision all-weather detection of aerial images of drones is realized, taking into account real-time and computing efficiency.
Patent Information
- Application Number
- CN202510364804.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-25
AI Technical Summary
The prior art is difficult to achieve the accuracy and real-time requirements of all-weather detection and reconnaissance when processing multi-source heterogeneous data, dynamic scenes and high-noise environments.
A three-modal image fusion method based on the cross attention mechanism is adopted to simulate rainy days and noise effects through image enhancement neural networks, combining the complementary features of visible light, long-wave infrared and short-wave infrared, and feature extraction and fusion are used for cross attention mechanism, and a comprehensive loss function is designed to improve the fusion accuracy.
It improves the fusion accuracy of drone aerial images, takes into account real-time and accuracy, simplifies the calculation amount, and reduces the risk of slow response.
Smart Images

Figure CN120374414A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image fusion, and in particular, to a three-modal image fusion method and system based on a cross-attention mechanism. Background Art
[0002] Image fusion is a technology that combines image information from different sources and is widely used in multiple fields such as remote sensing, medical imaging, security monitoring, robot vision, and autonomous driving. Its core goal is to enhance the available information of images and provide more comprehensive and accurate decision-making support. Image fusion generally follows three main methods: Pixel-level fusion generates a fused image by performing operations such as weighted averaging or selecting the maximum value of pixels; Feature-level fusion performs fusion after extracting image features and is often used for target recognition in complex scenes; Decision-level fusion integrates the outputs of different models or sensors to obtain more accurate results.
[0003] In recent years, with the progress of technology, multi-scale decomposition techniques (such as wavelet transform) and deep learning methods (such as convolutional neural networks) have been widely used in image fusion and can automatically extract and optimize image features. In addition, image registration technology ensures the geometric alignment of different source images and improves the fusion effect. To evaluate the fusion effect, common evaluation metrics include entropy, mean square error (MSE), and structural similarity (SSIM). Although image fusion technology has achieved remarkable results in multiple fields, it still faces challenges in dealing with multi-source heterogeneous data, dynamic scenes, and high-noise environments. Summary of the Invention
[0004] The present application provides a three-modal image fusion method and system based on a cross-attention mechanism, and its technical purpose is to effectively integrate the color information of visible light, the temperature data of long-wave infrared, and the anti-fog characteristics of short-wave infrared, and utilize the complementary information between multi-modalities to improve the fusion accuracy of UAV aerial images.
[0005] The above technical purpose of the present application is achieved through the following technical solutions:
[0006] A three-modal image fusion method based on a cross-attention mechanism, comprising:
[0007] Step S1: Add simulated raindrop effects and noise effects to the collected three-modal data through an image enhancement neural network to obtain enhanced data of different modalities;
[0008] Step S2: Extract the information feature maps of the enhanced data of different modalities through a feature extraction network, and then cascade and output the information feature maps of each modality;
[0009] Step S3: Calculate the global standard feature map representing all modal feature information, perform cross-attention calculation on the information feature maps of each modality and this global standard feature map, and obtain the cross-attention mechanism feature maps of each modality;
[0010] Step S4: Perform multi-layer convolution on the cross-attention mechanism feature maps of each modality to restore the channels, and generate a fused image containing the information of three modalities.
[0011] A three-modal image fusion system based on cross-attention mechanism, which is used for a fusion method, including:
[0012] A data enhancement module, which adds simulated raindrop effects and noise effects to the collected three-modal data through an image enhancement neural network to obtain enhanced data of different modalities;
[0013] A feature extraction module, which extracts the information feature maps of the enhanced data of different modalities through a feature extraction network, and then cascades and outputs the information feature maps of each modality;
[0014] A feature cross module, which calculates the global standard feature map representing all modal feature information, performs cross-attention calculation on the information feature maps of each modality and this global standard feature map, and obtains the cross-attention mechanism feature maps of each modality;
[0015] A fused image reconstruction module, which performs multi-layer convolution on the cross-attention mechanism feature maps of each modality to restore the channels, and generates a fused image containing the information of three modalities.
[0016] The beneficial effects of this application are as follows: For the three-modal image fusion method and system based on cross-attention mechanism described in this application, first, the input UAV aerial images are enhanced, and the rainy day and noise effects are simulated to enrich the data set and improve its diversity; secondly, the cross-attention mechanism is used to fully extract and fuse the complementary features in the multi-modal images. The multi-modal cross-attention mechanism can ensure the cross-supervision effect during the feature extraction process, greatly simplify the sudden increase in computational complexity caused by the increase in modalities, and reduce the possible slowdown in response caused by the introduction of the attention mechanism. In addition, a comprehensive loss function is designed, which combines the structure, texture, and intensity information of the source images and can well learn the complementary information of multi-modal images to improve the fusion accuracy. Compared with the existing fusion algorithms, the method described in this application innovatively combines the complementary advantages of visible light, long-wave infrared, and short-wave infrared, taking into account both real-time performance and accuracy. Description of the Drawings
[0017] Figure 1 It is the schematic diagram of the image enhancement neural network in the embodiment of this application;
[0018] Figure 2 Flow chart of the three - modal image fusion method based on the cross - attention mechanism in the embodiment of the present application;
[0019] Figure 3 Schematic diagram of the calculation of the global standard feature map in the embodiment of the present application;
[0020] Figure 4 Principle diagram of the multi - modal cross - attention mechanism in the embodiment of the present application;
[0021] Figure 5 Effect diagram of the three - modal image fusion method in the embodiment of the present application. Specific implementation manner
[0022] The technical solution of the present application will be described in detail below with reference to the accompanying drawings. It should be understood that the following specific implementation manners are only used to illustrate the present invention and not to limit the scope of the present invention.
[0023] As Figure 2 shown, the three - modal image fusion method based on the cross - attention mechanism described in the present application includes:
[0024] Step S1: Add simulated raindrop effects and noise effects to the collected three - modal data through an image enhancement neural network to obtain enhanced data of different modalities.
[0025] Preferably, as Figure 1 shown, the image enhancement neural network includes an input layer, an encoding layer, a feature extraction and analysis layer, a decoding layer, and an output layer. The input layer is used to receive the original collected three - modal data. At the same time, to improve the efficiency of the network, a variety of noise feature images are also provided to the network to simulate weather conditions such as rainy days. The encoding layer includes a down - sampling process, which is used to transfer the information of the collected three - modal data and noise features from the spatial dimension to the channel dimension for subsequent feature analysis. The feature extraction and analysis layer includes an up - sampling process and a down - sampling process, which are used to train the network's ability to fuse noise features onto the original image. The decoding layer includes an up - sampling process, which is used to transfer the image information from the channel dimension back to the spatial dimension. The output layer outputs the enhanced data fused with raindrop effects and noise effects.
[0026] Step S2: Extract the information feature maps of the enhanced data of different modalities through a feature extraction network, and then cascade and output the information feature maps of each modality.
[0027] Preferably, the feature extraction network includes five convolutional layers. First, a 1×1 convolutional layer is designed to reduce the differences in different modality information. Then, four convolutional layers are used for the deep feature extraction of visible light images, long-wave infrared images, and short-wave infrared images. It should be noted that the outputs of the second, third, and fourth layers are all added element-wise to the output of the fusion image reconstruction module to exchange modality complementary features. Except for the first layer, the kernel size of all convolutional layers is 3×3. All convolutional layers of the feature extraction network use Leaky Relu as the activation function.
[0028] Step S3: Calculate the global standard feature map representing all modality feature information, and perform cross-attention on the information feature maps of each modality and this global standard feature map to obtain the cross-attention mechanism feature maps of each modality.
[0029] Specifically, as Figure 4 shown, the step S3 includes:
[0030] Step S31: As Figure 3 shown, perform pixel-wise addition on the information feature maps of each modality at the current layer to obtain the global standard feature map of the current layer.
[0031] Step S32: Perform pixel-wise subtraction on the global standard feature map and the information feature map of the corresponding modality respectively, and perform pixel-wise mean processing to achieve the cross of the information feature map of each modality and the global standard feature map, expressed as:
[0032]
[0033] where, F cross represents the cross feature map after cross processing; u standard (i,j) represents the pixel unit in the global standard feature map; u vis (i,j) represents the pixel unit of the information feature map in the visible light modality; i and j respectively represent the unit pixels in the x and y axis directions.
[0034] Step S33: Introduce an attention mechanism to the cross feature map through the squeeze-and-excitation method, and output the channel feature map with attention added.
[0035] Preferably, the implementation process of the attention mechanism is roughly divided into three steps. The specific process is shown in Figure 4 . The step S33 includes:
[0036] Step S331: Perform global average pooling on the cross feature map to generate a list vector. In a macroscopic sense, each element in the list vector represents the feature information of the corresponding channel to a certain extent. This list vector is expressed as:
[0037]
[0038] Among them, z c represents the list vector; u c represents the value at the corresponding position in the input cross-feature map; F sq represents the squeeze operation; H and W respectively represent the height and width of the cross-feature map; i and j respectively represent the unit pixels in the x and y axis directions;
[0039] Step S332: Generate a weight information list W through two fully connected layers. The weight information list W is obtained through learning and is used to display feature correlations. The channel weight value is obtained according to the weight information list W, expressed as:
[0040] s c = F ex (z c , W) = σ(g(z c , W)) = σ(W2δ(W1z c ));
[0041] Among them, s c represents the channel weight value; F ex represents the extension operation; δ represents the sigmoid activation function; W1 and W2 respectively represent two fully connected layers;
[0042] Step S333: Multiply the channel weight value with the corresponding channel to obtain the channel feature map with added attention, expressed as:
[0043] u se = F scale (u c , s c ) = s c u c ;
[0044] Among them, F scale represents the mapping operation; s c represents each channel weight value, u c represents the corresponding channel.
[0045] Step S34: Add the channel feature map and the information feature map to obtain the cross-attention mechanism feature map of each modality, expressed as:
[0046]
[0047] Among them, u se represents the channel feature map.
[0048] Step S4: Perform multi-layer convolution on the cross-attention mechanism feature map of each modality to restore the channels and generate a fused image containing three-modal information. The final fusion effect is as Figure 5 shown.
[0049] Preferably, in the step S4, the four convolutional layers of the multi-layer convolutional reduction channel, the kernel sizes of the first three convolutional layers are 3×3, which are responsible for completing the fusion of spatial information, and the kernel of the last convolutional layer is 1×1, which is responsible for completing the fusion of channel information.
[0050] Preferably, the last convolutional layer of the multi-layer convolutional reduction channel uses the Tanh activation function, and the remaining convolutional layers use the Leaky Relu activation function.
[0051] Preferably, the loss function of the multi-layer convolutional reduction channel is expressed as:
[0052] L all = λ1L st + λ2L text + λ3L ssim ;
[0053] where λ1, λ2, and λ3 all represent weight parameters; L st represents the intensity loss, which characterizes the retention of the intensity information of the significant target features in the fused image; L text represents the texture loss, which characterizes the retention of the rich texture detail information in the fused image; L ssim represents the similarity loss, which is used to maintain the similarity of the fused image and the source image in terms of structure, color, and brightness.
[0054] The intensity loss is expressed as:
[0055]
[0056] where H and W respectively represent the height and width of the input image; I fus , I vis , I l_inf , I s_inf respectively represent the fused image, the input visible light image, the long-wave infrared image, and the short-wave infrared image; ||·||1 represents the first-order norm calculation.
[0057] The texture loss is expressed as:
[0058]
[0059] where represents the gradient operator, which is used to measure the texture information of the image;
[0060] L ssim = w1(1 - ssim(I fus , I vis )) + w2(1 - ssim(I fus , I l_inf))
[0061] +w3(1 - ssim(I fus , I s_inf ))
[0062] Among them, the result obtained by the ssim(·) function represents the structural similarity of two images; w1, w2, and w3 respectively represent the weights of the three modalities in this loss.
[0063] The three - modality image fusion system based on the cross - attention mechanism described in this application includes a data enhancement module, a feature extraction module, a feature cross - module, and a fused image reconstruction module.
[0064] The data enhancement module adds simulated raindrop effects and noise effects to the collected three - modality data through an image enhancement neural network to obtain enhanced data of different modalities.
[0065] The feature extraction module extracts the information feature maps of the enhanced data of different modalities through a feature extraction network, and then cascades and outputs the information feature maps of each modality.
[0066] The feature cross - module is used to calculate the global standard feature map representing the feature information of all modalities, and perform cross - attention on the information feature maps of each modality and the global standard feature map to obtain the cross - attention mechanism feature maps of each modality.
[0067] The fused image reconstruction module is used to perform multi - layer convolution to restore the channels on the cross - attention mechanism feature maps of each modality to generate a fused image containing the information of the three modalities.
[0068] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.
Claims
1. A three-modal image fusion method based on cross-attention mechanism, characterized in that, Including: Step S1: Add simulated raindrop effects and noise effects to the collected three-modal data through an image enhancement neural network to obtain enhanced data of different modalities. Step S2: Extract the information feature maps of the enhanced data of different modalities through a feature extraction network, and then cascade and output the information feature maps of each modality. Step S3: Calculate the global standard feature map representing the feature information of all modalities, and perform cross-attention calculation between the information feature maps of each modality and the global standard feature map to obtain the cross-attention mechanism feature maps of each modality. Step S4: Perform multi-layer convolution to restore the channels on the cross-attention mechanism feature maps of each modality to generate a fused image containing the information of three modalities.
2. The three-modal image fusion method according to claim 1, wherein In the step S1, the image enhancement neural network includes an input layer, an encoding layer, a feature extraction and analysis layer, a decoding layer, and an output layer; the input layer is used to receive the original collected three-modal data; the encoding layer includes a downsampling process for transferring the information of the collected three-modal data and the noise features from the spatial dimension to the channel dimension; the feature extraction and analysis layer includes an upsampling process and a downsampling process for training the network's ability to fuse the noise features onto the original image; the decoding layer includes an upsampling process for transferring the image information from the channel dimension back to the spatial dimension; the output layer outputs the enhanced data fused with raindrop effects and noise effects.
3. The three-modal image fusion method according to claim 1, wherein In the step S2, the feature extraction network includes five convolutional layers. The kernel size of the first convolutional layer is 1×1, and the kernel sizes of the remaining convolutional layers are 3×3. The activation function of all convolutional layers is the Leaky Relu function.
4. The three-modal image fusion method according to claim 1, wherein The step S3 includes: Step S31: Perform pixel-wise addition on the information feature maps of each modality in the current layer to obtain the global standard feature map of the current layer. Step S32: Perform pixel-wise subtraction between the global standard feature map and the information feature maps of the corresponding modalities respectively, and perform pixel-wise mean processing to realize the cross between the information feature maps of each modality and the global standard feature map, expressed as: Among them, F cross represents the cross feature map after cross processing; u standard (i,j) represents the pixel unit in the global standard feature map; u vis (i,j) represents the pixel unit of the information feature map in the visible light modality; i and j respectively represent the unit pixels in the x and y axis directions. Step S33: Introduce an attention mechanism to the cross feature map through the squeeze-and-excitation method, and output the channel feature map with added attention. Step S34: Add the channel feature map and the information feature map to obtain the cross-attention mechanism feature maps of each modality, expressed as: Among them, u se represents the channel feature map.
5. The three-modal image fusion method according to claim 4, wherein The step S33 includes: Step S331: Perform global average pooling on the cross feature map to generate a list vector, expressed as: Among them, z c represents the list vector; u c represents the value at the corresponding position in the input cross feature map; F sq represents the squeeze operation; H and W respectively represent the height and width of the cross feature map; i and j respectively represent the unit pixels in the x and y axis directions; Step S332: Generate a weight information list W through two fully connected layers, and obtain the channel weight values according to the weight information list W, expressed as: s c = F ex (z c , W) = σ(g(z c , W)) = σ(W2δ(W1z c )); Among them, s c represents the channel weight value; F ex represents the extension operation; δ represents the sigmoid activation function; W1 and W2 respectively represent two fully connected layers; Step S333: Multiply the channel weight values with the corresponding channels to obtain the channel feature map with added attention, expressed as: u se = F scale (u c , s c ) = s c u c ; Among them, F scale represents a mapping operation; s c represents the weight values of each channel, and u c represents the corresponding channel.
6. The three-modal image fusion method according to claim 5, wherein In the step S4, the multi-layer convolution for restoring channels includes four convolutional layers. The kernel sizes of the first three convolutional layers are 3×3, and the kernel of the last convolutional layer is 1×1.
7. The three-modal image fusion method according to claim 6, wherein The activation function of the last convolutional layer of the multi-layer convolution for restoring channels is the Tanh function, and the activation functions of the remaining convolutional layers are the Leaky Relu functions.
8. The three-modal image fusion method according to claim 7, wherein The loss function of the multi-layer convolution for restoring channels is expressed as: L all = λ1L st + λ2L text + λ3L ssim ; Among them, λ1, λ2, and λ3 all represent weight parameters; L st represents the intensity loss, which characterizes the retention of the intensity information of the significant target features in the fused image; L text represents the texture loss, which characterizes the retention of the rich texture detail information in the fused image; L ssim represents the similarity loss, which is used to maintain the similarity between the fused image and the source image in terms of structure, color, and brightness; The strength loss is expressed as: where H and W represent the height and width of the input image, respectively; I fus , I vis , I l_inf , I s_inf represent the fused image, the input visible light image, the long-wave infrared image, and the short-wave infrared image, respectively; ||·||1 represents the first-order norm calculation; The texture loss is expressed as: Among them, represents a gradient operator used to measure image texture information; L ssim = w1(1 - ssim(I fus , I vis )) + w2(1 - ssim(I fus , I l_inf )) + w3(1 - ssim(I fus , I s_inf )) Among them, the result obtained by the ssim(·) function characterizes the structural similarity of two images; w1, w2, and w3 respectively represent the weights of the three modalities in this loss.
9. A three-modal image fusion system based on a cross-attention mechanism, the fusion system being used for the fusion method according to any one of claims 1-8, characterized in that, It includes: A data augmentation module that adds simulated raindrop effects and noise effects to the collected three-modal data through an image enhancement neural network to obtain augmented data of different modalities; A feature extraction module that extracts the information feature maps of the augmented data of different modalities through a feature extraction network, and then cascades and outputs the information feature maps of each modality; A feature cross module that calculates the global standard feature map representing the feature information of all modalities, and performs cross-attention calculation on the information feature maps of each modality and this global standard feature map to obtain the cross-attention mechanism feature maps of each modality; A fused image reconstruction module that performs multi-layer convolution to restore the channels of the cross-attention mechanism feature maps of each modality to generate a fused image containing the information of the three modalities.