Multi-modal image fusion method based on modal self-adaption and modal interaction compensation
By employing a multimodal image fusion method that combines modal adaptation and modal interaction compensation, and utilizing a shared weight transformer network and a self-attention mechanism, along with differential feature extraction and modal interaction compensation modules, the problem of modal differences and correlation characteristics between infrared and visible light images in different scenarios is solved, achieving a more efficient multimodal image fusion effect.
Patent Information
- Application Number
- CN202511003843.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-07
AI Technical Summary
Existing technologies struggle to effectively coordinate the modal differences and dynamic correlation characteristics of infrared and visible light images in different scenarios, resulting in poor multimodal image fusion performance.
A multimodal image fusion method based on modal adaptation and modal interaction compensation is adopted. Features are extracted synchronously through a shared weight transformer network, combined with self-attention mechanism and differential feature extraction strategy, and a modal interaction compensation module is added to achieve dynamic feature balancing.
It significantly improves the accuracy and adaptability of multimodal image fusion, solves the problems of asynchronous feature extraction and weak correlation in traditional methods, and achieves better fusion results.
Smart Images

Figure CN120912448A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer vision, in particular to a multi-modal image fusion method based on modal adaptation and modal interaction compensation. BACKGROUND
[0002] In recent years, multi-modal image fusion has attracted extensive research attention in the academic and industrial fields of artificial intelligence, and it plays an important role in intelligent transportation, remote sensing monitoring, medical image analysis, and military reconnaissance. Single-modal images (such as relying only on visible light or infrared imaging) often cannot provide comprehensive enough information, while multi-modal image fusion aims to integrate the advantages of different imaging mechanisms to achieve more robust and rich visual representation. The key challenges of multi-modal image fusion are: (1) significant modal difference: infrared images mainly reflect the radiation characteristics of the target and are not sensitive to light changes, but lack texture details; while visible light images contain rich spatial structure and color information, but are prone to failure in low light or complex environments. (2) Dynamic correlation characteristics: the information contribution of different modalities varies in different scenes (such as infrared modality dominance at night and visible light modality dominance during the day), which requires adaptive balancing of the fusion strategy.
[0003] Benefiting from the powerful perceptual ability of deep neural networks, image recognition, image retrieval, and target detection tasks have achieved widespread success, among which the use of visual features has made significant progress in image fusion visual tasks. Since different modal images have both differences and correlations in feature representation, an adaptive mechanism needs to be designed to balance their feature contributions to achieve better fusion results. Therefore, the key to multi-modal image fusion is to coordinate the complementary characteristics of infrared and visible light images. SUMMARY
[0004] The present application proposes a multi-modal image fusion method based on modal adaptation and modal interaction compensation, which can accurately and effectively perform multi-modal image fusion.
[0005] The present application adopts the following technical solutions.
[0006] The multi-modal image fusion method based on modal adaptation and modal interaction compensation comprises the following steps:
[0007] Step S1: Obtain a multi-modal image fusion dataset, and through automatic preprocessing of infrared and visible light images, obtain a dataset for training a multi-modal image fusion model
[0008] Step S2: Analyze the modal difference characteristics of infrared and visible light images, and dynamically evaluate the correlation characteristics of image pairs in different scenes by combining JS divergence and K-means clustering algorithm;
[0009] Step S3: The features of the infrared and visible light images are synchronously extracted by using a transformer network with shared weights, and the cross-modal feature dependency relationship is efficiently captured through a self-attention mechanism.
[0010] Step S4: A differential feature extraction strategy is adopted for the difference in information sparsity between the infrared image and the visible light image, and the model parameters are optimized through iterative training, and the final model is used to complete the multi-modal image fusion.
[0011] Step S5: A modal interaction compensation module is added, the common features and the modal specific features output by step S4 are dynamically balanced through an adaptive weight fusion unit, and the cross-modal feature compensation mechanism is used to realize feature complementation in the channel and spatial dimensions, and then the parameters of the modal interaction compensation module are optimized through iterative training, so that the model can adaptively learn the optimal fusion weight of the multi-modal features, and the final model is used to realize the generation optimization of the multi-modal fusion image.
[0012] In step S1, the automatic preprocessing of the infrared and visible light images includes block cutting and low contrast screening.
[0013] Step S1 specifically includes the following steps:
[0014] Step S11: Obtain a public multi-modal image fusion dataset, perform color space conversion on the visible light image to extract the luminance component, and perform pixel value normalization on the infrared image with a wavelength of 814 μm to complete data standardization;
[0015] Step S12: Use a sliding window segmentation algorithm to divide the input image of 640x480 into fixed-size 128x128 images, set an appropriate sliding step to ensure image coverage integrity, and use a double-percentile method to calculate the contrast index of each image block;
[0016] Step S13: Screen the image blocks according to the preset contrast threshold, remove low-contrast samples, and store the qualified image blocks in an HDF5 hierarchical format, while preserving the original spatial dimensions and data types;
[0017] Step S14: Perform integrity verification on the constructed dataset to ensure the spatial alignment accuracy and intensity distribution consistency of the infrared and visible light image blocks, and finally generate a training dataset that meets the requirements.
[0018] In step S2, the modal difference characteristics of the infrared and visible light images are quantitatively analyzed based on multi-dimensional sparsity indicators such as gradient change rate, entropy value, L1 norm, and Laplacian variance.
[0019] Step S2 specifically includes the following steps:
[0020] Step S21: Calculate the gradient change rate, entropy value, L1 norm and Laplacian variance index of the infrared and visible light image blocks respectively, and establish a multi-dimensional sparsity difference quantization model; wherein the gradient change rate reflects the richness of image edge information, the entropy value represents the uncertainty of image information quantity, the L1 norm measures the sparsity of image features, and the Laplacian variance is used to evaluate the overall definition of the image. The quantization model constructed by the four complementary indexes can comprehensively depict the essential difference in information sparsity of different modal images;
[0021] The specific calculation method is: for each image pair (IR, VIS), the infrared image index value F IR and the visual image index value F VIS are calculated respectively, and on this basis, the difference value ΔF = F IR -F VIS is calculated, and then the proportion distribution of samples with ΔF≥0 and ΔF<0 in the whole data set is counted to reveal the systematic difference characteristics of different modalities in information sparsity; To clarify the calculation process of the above indexes, the calculation method of each quantization index is defined as follows:
[0022] For the gradient change rate, the horizontal and vertical gradients (G x ,G y ) of the image block are calculated by using the Sobel operator, and then the arithmetic mean of the gradient amplitudes of all pixels in the image block is counted as the gradient change rate index. The specific calculation method is as follows:
[0023]
[0024] Where N is the total number of pixels, reflecting the average edge strength;
[0025] For the entropy value, the Shannon entropy of the image block gray histogram is calculated by using the probability distribution discretization method, which can effectively avoid the floating point operation error and improve the numerical stability of the modal difference analysis. The specific calculation method is as follows:
[0026] Where p(k) represents the frequency of pixels with gray value k in the image block, i.e. it needs to satisfy The base of logarithm is 2;
[0027] For the L1 norm, the sparsity feature strength of the image block is calculated by using the absolute value and mean value method, which can effectively quantify the information density difference between modalities. The specific calculation method is as follows:
[0028]
[0029] Where I(i) represents the normalized gray value of the i-th pixel in the image, and N represents the total number of pixels;
[0030] For Laplacian variance, the second-order differential operator is combined with statistical dispersion measurement to effectively represent the modal difference characteristics of image local contrast, and the specific calculation method is as follows:
[0031]
[0032] Where L represents the Laplacian convolution result, μ L represents the mean of the Laplacian response, and N represents the total number of pixels.
[0033] Step S22: The JS divergence is combined with the K-means clustering algorithm to evaluate the modal difference through symmetry measurement and automatically identify the scene correlation strength to provide adaptive basis for feature fusion.
[0034] For JS divergence, the statistical characteristic difference of infrared and visible light image blocks is calculated by using the probability distribution similarity measurement method, which can effectively evaluate the correlation characteristics between modes, and the specific calculation method is as follows:
[0035]
[0036] Where P, Q represent the normalized gray histogram distribution of infrared and visible light image blocks respectively, M is the intermediate distribution, and D KL is the KL divergence.
[0037] For K-means clustering, an unsupervised learning method is used to automatically classify the image modal correlation characteristics to identify the modal correlation mode in different scenes, and the specific calculation method is as follows:
[0038]
[0039] Where S represents k cluster classes, μ i represents the center point of the i-th cluster, and x is the modal correlation characteristic value calculated by JS divergence.
[0040] Step S23: The sparse difference quantization result and the scene correlation characteristic analysis result are used as priori knowledge to provide feature distribution basis for the subsequent model framework including the difference feature extraction stage and the modal interaction compensation stage.
[0041] Step S3 specifically includes the following steps:
[0042] Step S31: The embedding projection of the input image is realized through the 3x3 convolution layer of the transformer network, the step size is 1, and the single-pixel zero padding is used to keep the feature map space dimension unchanged. The linear mapping layer converts the original input into deep feature representation through the learnable weight matrix, and its mathematical form is represented as:
[0043] X out = conv(X in )
[0044] where conv denotes the convolution kernel, X in denotes the batch input of the structured training dataset generated by step S1, X out denotes the resulting deep feature representation;
[0045] Step S32: Standardize the deep feature representation X out The layer normalization module with configurable bias is used to standardize the input features and eliminate the scale difference of feature distribution between modalities, which is mathematically expressed as:
[0046]
[0047] where μ and σ represent the feature mean and standard deviation respectively, γ and β represent the learnable affine parameters, and ∈ represents the numerical stability term;
[0048] Step S33: Process X norm and generate query (Q), key (K), and value (V) vectors, on which multi-head temperature adjustment attention calculation is performed, which is mathematically expressed as:
[0049]
[0050] where τ represents the learnable temperature parameter, d k denotes the key vector dimension, and Softmax represents the standardization exponential function;
[0051] Step S34: Fuse the attention features generated in step S33 with the original input features through residual connection, and perform layer normalization, which is mathematically expressed as:
[0052] X norm = LayerNorm(X out + Attn)
[0053] where + denotes the residual connection operation of element-wise addition, X out denotes the deep feature representation obtained in step S31, and LayerNorm(·) denotes the normalization function, which is calculated in the same way as step S32, using the same normalization strategy and parameter configuration;
[0054] Step S35: Perform gated dynamic feedforward transformation on the normalized features and fuse them through residual connection, which is mathematically expressed as:
[0055]
[0056] where W out denotes the shrinkage projection matrix, and GelU denotes the activation function, and both represent extended projection matrices, and ⊙ represents Hadamard product.
[0057] In step S4, the infrared image is effectively extracted by the neighborhood grouping matching model, and the visible light image is fully captured by the context perception modeling network.
[0058] Step S4 specifically includes the following steps:
[0059] Step S41: The visual features output in step S3 are processed by the channel expansion projection module. First, a 1x1 convolution is used for channel expansion projection to increase the number of feature channels to twice the original dimension; then a 7x7 dilated convolution is used for multi-scale spatial feature extraction, where the dilated rate is dynamically adjusted according to the feature map resolution to maintain the optimal receptive field; finally, a 1x1 convolution is used for channel compression projection to restore the feature dimension to the original number of channels, and the processed features are added to the input features through a residual connection;
[0060] The processing process through the "expansion-transformation-compression" cascade structure effectively enhances the expression ability and spatial context information of the features while maintaining the consistency of the feature dimension, which is mathematically expressed as:
[0061]
[0062] where conv 1×1 represents a 1x1 convolution kernel, dc 7×7 represents a 7x7 dilated convolution kernel;
[0063] Step S42: The features output in step S41 are processed by the hierarchical feature enhancement module, which specifically includes: first, the input features are standardized by layer normalization; then, the features are transformed by a bottleneck structure constructed by two levels of point convolution: the first level uses a 1x1 point convolution to realize channel expansion projection, which expands the number of feature channels to four times the original dimension, and the second level uses a 1x1 point convolution to realize channel contraction projection, which restores the number of feature channels to the original dimension; finally, the transformed features are added to the original input features through a residual connection; this processing process realizes efficient enhancement of the features through the expansion-contraction bottleneck structure, which improves the expression ability of the features while effectively preserving the key information of the original features, which is mathematically expressed as:
[0064]
[0065] where pc 1×1 represents a 1x1 point convolution kernel, and LayerNorm(·) represents a normalization function, which is calculated in the same way as step S32, using the same normalization strategy and parameter configuration;
[0066] Step S43: projecting the infrared feature output in step S3 to the spherical space, and then generating an index matrix S through the local sensitive hashing mapping, for locating the TopK associated features in the feature space, and finally using the index matrix to synchronously group the left / right branch features, the mathematical expression is as follows: First, two groups of parameter-shared 1x1 convolution layers are used to generate left / right branch features, and the mathematical expression is as follows:
[0067]
[0068] Wherein, conv 1×1 represents a 1x1 convolution kernel, and the above convolution kernel uses the same parameter configuration;
[0069] Step S44: projecting the right branch feature output in step S43 to the spherical space, and then generating an index matrix S through the local sensitive hashing mapping, for locating the TopK associated features in the feature space, and finally using the index matrix to synchronously group the left / right branch features, the mathematical expression is as follows:
[0070]
[0071] Wherein, P represents the spherical projection feature, ||·||2 represents the L2 norm, R represents a random rotation matrix, S represents a hash index matrix, represents the maximum index taken along the hash bucket dimension b, sign(·) represents a binary sign function, and U represents a set aggregation operation of the features according to the index S;
[0072] Step S45: based on the grouped features obtained in step S44, first calculating the intra-group feature similarity, then using the feature similarity to perform feature aggregation, and finally adding and fusing the aggregation result with the infrared feature output in step S3, the mathematical expression is as follows:
[0073]
[0074] Step S46: calculating the overall loss function Loss of the network, and the calculation formula is as follows:
[0075] Loss=α1Loss ssim +α2Loss mse +α3Loss int +α4Loss grad
[0076] The overall optimization target is composed of four parts, and α1, α2, α3, and α4 are hyperparameters of the loss functions, and the calculation methods of the loss functions are defined as follows:
[0077] For the structural similarity loss Loss ssim , it is used to measure the fusion image I F , the infrared image I IR , and the visible light image I VISThe structural similarity difference between them is calculated as follows:
[0078] Loss ssim = 1-SSIM(I F ,I IR )+1-SSIM(I F ,I VIS )
[0079] Wherein SSIM(·) represents the structural similarity index;
[0080] For the mean square error loss Loss mse , used to constrain the overall brightness distribution consistency of the fusion result and the input image, the specific calculation formula is as follows:
[0081]
[0082] Wherein is the square of L2 norm;
[0083] For the intensity preservation loss Loss int , the fused image is forced to retain the significant intensity features in the source image, and the specific calculation formula is as follows:
[0084]
[0085] Wherein HW respectively represents the image width and height, ||·||1 is L1 norm, and max(·) takes the maximum of the two for each pixel position of the comparison image;
[0086] For the gradient consistency loss Loss grad , used to ensure that the fused image inherits the clearest edge information in the source image, the specific calculation formula is as follows:
[0087]
[0088] Wherein represents the Sobel operator;
[0089] Step S47: Gradient update and parameter adjustment according to the overall loss value of step S46, performance evaluation on the validation set according to the preset interval period and reservation of the optimal model, automatic termination when the training round reaches the maximum value, and finally saving the best performance fusion model.
[0090] In step S5, based on the dynamic correlation characteristics between infrared and visible light modalities, a modal interaction compensation module is added on the basis of freezing the parameters of the S4 module.
[0091] Step S5 specifically includes the following steps:
[0092] Step S51: First, the visible features output in step S3 and infrared features are added to generate a mixed feature X sum , whose calculation formula is as follows:
[0093]
[0094] Step S52: Based on the mixed feature X sum obtained in step S51, a spatial attention module is constructed to take the mixed feature as input to generate a spatial attention mask SA(X sum ), SA(·) represents a spatial attention block, and the mixed feature X sum is input into global average pooling and global maximum pooling at the same time, and then the two pooling results are concatenated in the channel dimension and processed through a 7x7 convolution layer to obtain the spatial attention mask SA(X sum ); the specific calculation method is as follows:
[0095] SA(X sum )=conv 7×7 (Cat(GAP(X sum ),GMP(X sum )))
[0096] Where Cat(·) represents a channel dimension concatenation operation, GAP(·) represents a global average pooling operation, and GMP(·) represents a global maximum pooling operation.
[0097] Step S53: Based on the mixed feature X sum obtained in step S51, a channel attention module is constructed to take the mixed feature as input to generate a channel attention mask CA(X sum ), CA(·) represents a channel attention block, and the mixed feature X sum is input into global average pooling, and then the pooling result is sent into two layers of cascaded 1x1 convolution layers for processing and obtaining the channel attention mask CA(X sum ); the specific calculation method is as follows:
[0098] CA(X sum )=conv 1×1 (conv 1×1 (GAP(X sum )))
[0099] Step S54: Based on the spatial attention mask generated in step S52 and the channel attention mask obtained in step S53, the features are first fused through element-wise addition, and then the fused attention features are concatenated with the original mixed feature in step S51 in the channel dimension to realize comprehensive enhancement of multi-scale features; the specific calculation method is as follows:
[0100] X s54 = Cat(SA(X sum )+CA(X sum ),X sum )
[0101] Step S55: Based on the output feature X s54 of step S54, first feature extraction is performed through a 7*7 hollow convolution layer to expand the receptive field and maintain the spatial resolution, and then a sigmoid activation function is applied to generate a gating weight map; the specific calculation method is as follows:
[0102] W = sigma (dc 7×7 (X s54 ))
[0103] Where sigma represents the sigmoid function;
[0104] Step S56: Based on the gating weight map W of S55, the visible feature and the infrared feature obtained by S3 are dynamically weighted and fused, and the weighted results are respectively connected with the residual connection of the output features X camn and X nngm of S42 and S45, and finally the enhanced feature after modal compensation is obtained; the specific calculation formula is as follows:
[0105]
[0106] Where * represents matrix element multiplication;
[0107] Step S57: Gradient back propagation and parameter adjustment are performed based on the overall loss function Loss defined in step S46 (its composition structure and calculation method are exactly the same as the foregoing); during the process, the model performance is evaluated on the validation set according to the preset iteration interval, and the optimal weight is dynamically saved, and when the training round reaches the preset maximum value, the process is automatically terminated, and finally the best fusion model after complete optimization is output.
[0108] Compared with the prior art, the present application has the following beneficial effects:
[0109] 1、The shared weight transformer network proposed in the present application combines the self-attention mechanism, synchronously extracts and associates the infrared and visible light image features, effectively solves the problem of asynchronous double-modal feature extraction and weak correlation in the traditional method, and significantly improves the accuracy of cross-modal feature fusion.
[0110] 2. According to the information characteristic difference of infrared and visible light images, the differential feature extraction strategy (the infrared adopts a neighborhood grouping matching model, and the visible light adopts a context perception modeling network) innovatively designed by the application can realize the targeted extraction of different modal characteristics, and overcome the limitations of the traditional single network in extracting multi-modal characteristics.
[0111] The modal interaction compensation module provided by the application realizes dynamic feature balance on the basis of freezing the backbone network through the adaptive weight fusion unit and the cross-modal feature compensation mechanism, and solves the problem that the traditional fixed weight fusion method is difficult to adapt to different scene requirements. BRIEF DESCRIPTION OF DRAWINGS
[0112] The application will be further described in detail below in combination with the drawings and specific embodiments:
[0113] The accompanying drawings are schematic diagrams of the principles of the application. Figure 1 DETAILED DESCRIPTION
[0114] The application will be further described in detail below in combination with the drawings and specific embodiments:
[0115] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise indicated, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs.
[0116] It should be noted that the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, the singular form is intended to include the plural form unless the context clearly indicates otherwise, and it should also be understood that when the terms "comprise" and / or "include" are used in the specification, there is a feature, step, operation, device, component and / or combination thereof.
[0117] As shown in FIG. 1, the multi-modal image fusion method based on modal adaptation and modal interaction compensation includes the following steps: Figure 1
[0118] Step S1: Obtain a multi-modal image fusion dataset, and obtain a dataset for training a multi-modal image fusion model by automatic pre-processing of infrared and visible light images
[0119] Step S2: Analyze the modal difference characteristics of infrared and visible light images, and dynamically evaluate the correlation characteristics of image pairs in different scenes in combination with JS divergence and K-means clustering algorithm;
[0120] Step S3: The features of the infrared and visible light images are synchronously extracted by using a transformer network with shared weights, and the cross-modal feature dependency relationship is efficiently captured through a self-attention mechanism;
[0121] Step S4: A differential feature extraction strategy is adopted for the difference in information sparsity between the infrared image and the visible light image, and the parameters of the multi-modal image fusion model are optimized through iterative training, and the final model is used to complete the multi-modal image fusion;
[0122] Step S5: A modal interaction compensation module is added, the common features and the modal specific features output by step S4 are dynamically balanced through an adaptive weight fusion unit, and the cross-modal feature compensation mechanism is used to realize feature complementation in the channel and spatial dimensions, and the parameters of the modal interaction compensation module are optimized through iterative training, so that the model adaptively learns the optimal fusion weight of the multi-modal features, and the final model is used to realize the generation optimization of the multi-modal fusion image.
[0123] In step S1, the automatic preprocessing of the infrared and visible light images includes block cutting and low contrast screening.
[0124] Step S1 specifically includes the following steps:
[0125] Step S11: Obtain a public multi-modal image fusion dataset, perform color space conversion on the visible light image to extract the luminance component, and perform pixel value normalization on the infrared image with a wavelength of 814 μm to complete data standardization;
[0126] Step S12: Use a sliding window segmentation algorithm to divide the input image of 640x480 into fixed-size 128x128 images, set an appropriate sliding step to ensure image coverage integrity, and use a double-percentile method to calculate the contrast index of each image block;
[0127] Step S13: Screen the image blocks according to the preset contrast threshold, remove low-contrast samples, store the qualified image blocks in HDF5 hierarchical format, and completely retain the original spatial dimensions and data types;
[0128] Step S14: Perform integrity check on the constructed dataset to ensure the spatial alignment accuracy and intensity distribution consistency of the infrared and visible light image blocks, and finally generate a training dataset that meets the requirements.
[0129] In step S2, the modal difference characteristics of the infrared and visible light images are quantitatively analyzed based on gradient change rate, entropy value, L1 norm, and Laplacian variance.
[0130] Step S2 specifically includes the following steps:
[0131] Step S21: Calculate the gradient change rate, entropy value, L1 norm and Laplacian variance index of the infrared and visible light image blocks respectively, and establish a multi-dimensional sparsity difference quantization model; wherein the gradient change rate reflects the richness of image edge information, the entropy value represents the uncertainty of image information quantity, the L1 norm measures the sparsity of image features, and the Laplacian variance is used to evaluate the overall definition of the image. The quantization model constructed by the four complementary indexes can comprehensively depict the essential difference in information sparsity of different modal images;
[0132] The specific calculation method is: for each image pair (IR, VIS), the infrared image index value F IR and the visible image index value F VIS are calculated respectively, and on this basis, the difference value ΔF = F IR -F VIS is calculated, and then by counting the sample proportion distribution of ΔF≥0 and ΔF<0 in the whole data set, the systematic difference characteristics of different modalities in information sparsity are revealed; in order to clarify the calculation process of the above indexes, the calculation method of each quantization index is defined as follows:
[0133] For the gradient change rate, Sobel operator is used to calculate the horizontal and vertical gradients (G x ,G y ) of the image block, and then the arithmetic average value of the gradient amplitude of all pixels in the image block is calculated as the gradient change rate index. The specific calculation method is as follows:
[0134]
[0135] Where N is the total number of pixels, reflecting the average edge strength;
[0136] For the entropy value, the Shannon entropy of the image block gray histogram is calculated by using the probability distribution discretization method, which can effectively avoid the floating point operation error and improve the numerical stability of the modal difference analysis. The specific calculation method is as follows:
[0137] Where p(k) represents the frequency of pixels with gray value k in the image block, that is, it needs to satisfy The base of logarithm is 2;
[0138] For the L1 norm, the absolute value and mean value method is used to calculate the sparsity feature strength of the image block, which can effectively quantify the information density difference between modalities. The specific calculation method is as follows:
[0139]
[0140] Where I(i) represents the normalized gray value of the i-th pixel in the image, and N represents the total number of pixels;
[0141] For Laplacian variance, the second-order differential operator is combined with statistical dispersion measurement to effectively represent the modal difference characteristics of image local contrast, and the specific calculation method is as follows:
[0142]
[0143] Where L represents the Laplace convolution result, μ L represents the mean of the Laplace response, and N represents the total number of pixels.
[0144] Step S22: The JS divergence is combined with the K-means clustering algorithm to evaluate the modal difference through symmetry measurement and automatically identify the scene correlation strength to provide adaptive basis for feature fusion.
[0145] For JS divergence, the statistical characteristic difference of infrared and visible light image blocks is calculated by using the probability distribution similarity measurement method, which can effectively evaluate the correlation characteristics between modes, and the specific calculation method is as follows:
[0146]
[0147] Where P and Q represent the normalized gray histogram distribution of infrared and visible light image blocks respectively, M is the intermediate distribution, and D KL is the KL divergence.
[0148] For K-means clustering, an unsupervised learning method is used to automatically classify the image modal correlation characteristics to identify the modal correlation mode in different scenes, and the specific calculation method is as follows:
[0149]
[0150] Where S represents k cluster classes, μ i represents the center point of the i-th cluster, and x is the modal correlation characteristic value calculated by JS divergence.
[0151] Step S23: The sparse difference quantization result and the scene correlation characteristic analysis result are used as priori knowledge to provide feature distribution basis for the subsequent model framework including the difference feature extraction stage and the modal interaction compensation stage.
[0152] Step S3 specifically includes the following steps:
[0153] Step S31: The 3×3 convolution layer of the transformer network is used to realize the embedding projection of the input image, the step size is 1, and the single-pixel zero padding is used to keep the feature map space dimension unchanged. The linear mapping layer uses the learnable weight matrix to convert the original input into deep feature representation, and its mathematical form is represented as:
[0154] X out = conv(X in )
[0155] where conv denotes the convolution kernel, X in denotes the batch input of the structured training dataset generated by step S1, X out denotes the obtained deep feature representation;
[0156] Step S32: The deep feature representation X out The layer normalization module with configurable bias is used to standardize the input features and eliminate the scale difference of feature distribution between modalities, and its mathematical expression is:
[0157] where μ and σ represent the feature mean and standard deviation respectively, γ and β represent the learnable affine parameters, and ∈ represents the numerical stability term;
[0158] Step S33: The 3x3 deep separable convolution is used to process X norm and generate query (Q), key (K), and value (V) vectors, on which the multi-head temperature adjustment attention calculation is performed, and its mathematical form is expressed as:
[0159]
[0160] where τ represents the learnable temperature parameter, d k denotes the key vector dimension, and softmax denotes the standardization exponential function; Step S34: The attention features generated in step S33 are fused with the original input features through residual connection, and layer normalization processing is performed, and its mathematical expression is:
[0161] X norm = LayerNorm(X out + Attn)
[0162] where + denotes the residual connection operation of element-wise addition, X out denotes the deep feature representation obtained in step S31, and LayerNorm(·) denotes the normalization function, which is calculated in the same way as step S32, using the same normalization strategy and parameter configuration;
[0163] Step S35: The normalized features are executed through the gated dynamic feedforward transformation and fused through the residual connection, and its mathematical expression is:
[0164]
[0165] where W out denotes the shrinkage projection matrix, GeLU denotes the activation function, and both denote the expansion projection matrix, and ⊙ denotes the Hadamard product.
[0166] In step S4, the infrared image is effectively extracted by the neighborhood grouping matching model, and the visible light image is fully captured by the context perception modeling network.
[0167] Step S4 specifically includes the following steps:
[0168] Step S41: The visual feature output in step S3 is processed by a channel expansion projection using a 1×1 convolution to increase the number of feature channels to twice the original dimension. First, a 1×1 convolution is used for channel expansion projection to increase the number of feature channels to twice the original dimension. Then, a 7×7 dilated convolution is used for multi-scale spatial feature extraction, where the dilated rate is dynamically adjusted according to the feature map resolution to maintain the optimal receptive field. Finally, a 1×1 convolution is used for channel compression projection to restore the feature dimension to the original channel number, and the processed feature is added to the input feature through residual connection.
[0169] The processing process uses a cascade structure of "expansion-transformation-compression" to effectively enhance the expression ability of the feature and the spatial context information while maintaining the consistency of the feature dimension, which can be mathematically expressed as:
[0170]
[0171] where conv 1×1 represents a 1×1 convolution kernel, dc 7×7 represents a 7×7 dilated convolution kernel.
[0172] Step S42: The feature output in step S41 is processed by a hierarchical feature enhancement process, which specifically includes: first, the input feature is standardized by layer normalization; then, a bottleneck structure constructed by two levels of point convolution is used for feature transformation: the first level uses a 1×1 point convolution to realize channel expansion projection, which expands the number of feature channels to four times the original dimension, and the second level uses a 1×1 point convolution to realize channel contraction projection, which restores the number of feature channels to the original dimension; finally, the transformed feature is added to the original input feature through residual connection; this processing process realizes efficient enhancement of the feature through an expansion-contraction bottleneck structure, which not only improves the expression ability of the feature but also effectively preserves the key information of the original feature, which can be mathematically expressed as:
[0173] X camn = pc 1×1 (pc 1×1 (LayerNorm(X s41 )))+X s41
[0174] where pc 1×1 represents a 1×1 point convolution kernel, LayerNorm(·) represents a normalization function, and its calculation method is consistent with step S32, using the same normalization strategy and parameter configuration.
[0175] Step S43: Projecting the right branch feature outputted in step S43 to the spherical space, and then generating an index matrix S through the local sensitive hashing mapping, which is used to locate the TopK related features in the feature space, and finally using the index matrix to synchronize the grouping of the left / right branch features, the mathematical expression is as follows: First, two groups of parameter-shared 1x1 convolution layers are used to generate left / right branch features, and the mathematical expression is as follows:
[0176]
[0177] Where conv 1×1 represents a 1x1 convolution kernel, and the above convolution kernel uses the same parameter configuration;
[0178] Step S44: Projecting the right branch feature outputted in step S43 to the spherical space, and then generating an index matrix S through the local sensitive hashing mapping, which is used to locate the TopK related features in the feature space, and finally using the index matrix to synchronize the grouping of the left / right branch features, the mathematical expression is as follows:
[0179]
[0180] Where P represents the spherical projection feature, ||·||2 represents the L2 norm, R represents the random rotation matrix, S represents the hash index matrix, represents the maximum index along the hash bucket dimension b, sign(·) represents the binary sign function, and U represents the set aggregation operation of the features aggregated according to the index S;
[0181] Step S45: Based on the grouped features obtained in step S44, first calculate the intra-group feature similarity, then use the feature similarity to aggregate the features, and finally add and fuse the aggregation result with the infrared feature outputted in step S3, the mathematical expression is as follows:
[0182]
[0183] Step S46: Calculate the overall loss function Loss of the network, and the calculation formula is as follows:
[0184] Loss=α1Loss ssim +α2Loss mse +α3Loss int +α4Loss grad
[0185] The overall optimization target is composed of four parts, and α1, α2, α3, α4 are the hyperparameters of each part of the loss function, and the calculation method of each loss function is defined as follows:
[0186] For structural similarity loss Loss ssim , it is used to measure the similarity between the fused image I F and the infrared image I IR , the visible light image IVIS The structural similarity difference between them is calculated as follows:
[0187] Loss ssim = 1-SSIM(I F ,I IR )+1-SSIM(I F ,I VIS )
[0188] Wherein SSIM(·) represents the structural similarity index;
[0189] For the mean square error loss Loss mse , used to constrain the overall brightness distribution consistency of the fusion result and the input image, the specific calculation formula is as follows:
[0190]
[0191] Wherein is the square of L2 norm;
[0192] For the intensity preservation loss loss int , the fusion image is forced to retain the significant intensity features in the source image, and the specific calculation formula is as follows:
[0193]
[0194] Wherein HW respectively represents the width and height of the image, ||·||1 is L1 norm, and max(·) takes the maximum of the two for each pixel position of the comparison image;
[0195] For the gradient consistency loss Loss grad , used to ensure that the fusion image inherits the clearest edge information in the source image, the specific calculation formula is as follows:
[0196]
[0197] Wherein represents the Sobel operator;
[0198] Step S47: Gradient update and parameter adjustment according to the overall loss value of step S46, evaluate the performance on the validation set according to the preset interval period and reserve the optimal model, automatically terminate when the training round reaches the maximum value, and finally save the best performance fusion model.
[0199] In step S5, based on the dynamic correlation characteristics between infrared and visible light modalities, a modal interaction compensation module is added on the basis of freezing the parameters of the S4 module.
[0200] Step S5 specifically includes the following steps:
[0201] Step S51: First, the visible features output in step S3 and infrared features are added to generate a mixed feature X sum , whose calculation formula is as follows:
[0202]
[0203] Step S52: Based on the mixed feature X sum obtained in step S51, a spatial attention module is constructed to take the mixed feature as input to generate a spatial attention mask SA(X sum ), SA(·) represents a spatial attention block, and the mixed feature X sum is simultaneously input into global average pooling and global maximum pooling, and then the two pooling results are concatenated in the channel dimension and processed through a 7x7 convolution layer to obtain the spatial attention mask SA(X sum ); the specific calculation method is as follows:
[0204] SA(X sum )=conv 7×7 (Cat(GAP(X sum ),GMP(X sum )))
[0205] where Cat(·) represents a channel dimension concatenation operation, GAP(·) represents a global average pooling operation, and GMP(·) represents a global maximum pooling operation;
[0206] Step S53: Based on the mixed feature X sum obtained in step S51, a channel attention module is constructed to take the mixed feature as input to generate a channel attention mask CA(X sum ), CA(·) represents a channel attention block, and the mixed feature X sum is input into global average pooling, and then the pooling result is input into two cascaded 1x1 convolution layers for processing to obtain the channel attention mask CA(X sum ); the specific calculation method is as follows:
[0207] CA(X sum )=conv 1×1 (conv 1×1 (GAP(X sum )))
[0208] Step S54: Based on the spatial attention mask generated in step S52 and the channel attention mask obtained in step S53, the features are first fused through element-wise addition, and then the fused attention features are concatenated with the original mixed feature in step S51 in the channel dimension to realize comprehensive enhancement of multi-scale features; the specific calculation method is as follows:
[0209] X s54 = (Cat(SA(X sum )+CA(X sum ), X sum )
[0210] Step S55: Based on the output feature X s54 of step S54, first feature extraction is performed through a 7x7 hollow convolutional layer to expand the receptive field and maintain spatial resolution, and then a sigmoid activation function is applied to generate a gating weight map; the specific calculation method is as shown below:
[0211] W = σ(dc 7×7 (X s54 ))
[0212] Wherein σ represents the sigmoid function;
[0213] Step S56: Based on the gating weight map W of S55, the visual feature and the infrared feature obtained by S3 are dynamically weighted and fused, and the weighted results are respectively connected with the residual connection of the output features X camn and X nngm of S42 and S45, and finally the enhanced features after modal compensation are obtained; the specific calculation formula is as follows:
[0214]
[0215] Wherein * represents matrix element multiplication;
[0216] Step S57: Gradient back propagation and parameter adjustment are performed based on the overall loss function Loss defined in step S46 (its composition structure and calculation method are exactly the same as the foregoing); during the process, the model performance is evaluated on the validation set according to the preset iteration interval, and the optimal weight is dynamically saved, and when the training round reaches the preset maximum value, the process is automatically terminated, and finally the best fusion model after complete optimization is output.
[0217] The example proposes a multi-modal image fusion method based on modal adaptation and modal interaction compensation. First, a multi-modal image fusion data set is obtained, and training data set is obtained by automatic preprocessing of infrared and visible light image pairs. Then, the modal difference characteristics of infrared and visible light images are quantitatively analyzed based on multi-dimensional sparsity indicators, and the correlation characteristics of image pairs in different scenes are dynamically evaluated combined with JS divergence and K-means clustering algorithm. Then, the transformer network with shared weights is used to synchronously extract the features of the two modal images, and the cross-modal feature dependency is captured through the self-attention mechanism. Then, the differential feature extraction strategy is implemented according to the difference characteristics of infrared and visible light images, in which the sparse thermal radiation features of infrared images are extracted through the neighborhood grouping matching model, and the spatial detail features of visible light images are captured through the context perception modeling network, and the model parameters are optimized through iterative training. Finally, based on the dynamic correlation characteristics between modalities, a modal interaction compensation module is added, the common features and modal exclusive features are balanced through the adaptive weight fusion unit, and the feature complementation is realized through the cross-modal feature compensation mechanism. The model parameters are optimized through iterative training to make the model adaptively learn the optimal fusion weight, and the final model is used to realize the generation and optimization of multi-modal fusion images.
[0218] The embodiment is completely driven by image data without any manual annotation information. For the core challenge of multi-modal image fusion task, the application innovatively automatically mines the correlation characteristics between modalities as prior knowledge through JS divergence and K-means clustering algorithm. Most existing methods rely on fixed fusion rules or manually set weights, while the application dynamically evaluates the correlation characteristics of image pairs in different scenes to realize completely adaptive feature fusion. For the characteristic difference between infrared and visible light images, based on the quantitative analysis results of gradient change rate, entropy and other multi-dimensional sparsity indicators, the application designs a differential feature extraction strategy, in which the sparse thermal radiation features of infrared images extracted through the neighborhood grouping matching model and the spatial detail features of visible light images captured through the context perception modeling network form complementary advantages. The modal interaction compensation module constructed by the application fully utilizes the prior knowledge obtained through the foregoing analysis, dynamically balances the modal features through the adaptive weight fusion unit, and finally realizes the optimal fusion effect of preserving thermal radiation information and spatial details.
[0219] The above only describes the preferred embodiments of the application, and any equivalent changes and modifications made within the scope of the application should be included in the scope of the application.
Claims
1. A multi-modal image fusion method based on modal adaptation and modal interaction compensation, characterized in that: The method comprises the following steps: Step S1: Obtain a multi-modal image fusion dataset, and obtain a dataset for training a multi-modal image fusion model through automatic preprocessing of infrared and visible light images Step S2: Analyze the modal difference characteristics of infrared and visible light images, and dynamically evaluate the correlation characteristics of image pairs in different scenes in combination with a JS divergence and a K-means clustering algorithm Step S3: Synchronously extract features of infrared and visible light images by using a transformer network with shared weights, and efficiently capture cross-modal feature dependency relationships by using a self-attention mechanism Step S4: In view of the difference characteristics of infrared images and visible light images in information sparsity, a differential feature extraction strategy is used, and parameters of the multi-modal image fusion model are optimized through iterative training, and the model is used to complete multi-modal image fusion Step S5: A modal interaction compensation module is added, common features and modal exclusive features output by step S4 are dynamically balanced by using an adaptive weight fusion unit, and a cross-modal feature compensation mechanism is used to realize feature complementation in channel and spatial dimensions, and parameters of the modal interaction compensation module are optimized through iterative training, so that the model adaptively learns optimal fusion weights of multi-modal features, and the model is used to realize generation optimization of multi-modal fusion images.
2. The multi-modal image fusion method based on modal adaptation and modal interaction compensation according to claim 1, characterized in that: In step S1, the automatic preprocessing of the infrared and visible light images includes block cutting and low-contrast screening.
3. The multi-modal image fusion method based on modal adaptation and modal interaction compensation according to claim 2, characterized in that: Step S1 specifically comprises the following steps; Step S11: Obtain a public multi-modal image fusion dataset, perform color space conversion on the visible light image to extract a luminance component, and perform pixel value normalization processing on the infrared image with a wavelength of 814 μm, to complete data standardization Step S12: Divide the input image of 640*480 into fixed-size 128*128 images by using a sliding window segmentation algorithm, set an appropriate sliding step to ensure image coverage integrity, and calculate the contrast indicators of each image block by using a double-percentile method Step S13: Screen the image blocks according to a preset contrast threshold, remove low-contrast samples, store the qualified image blocks in an HDF5 hierarchical format, and completely retain the original spatial dimensions and data types Step S14: Perform integrity checking on the constructed dataset to ensure the spatial alignment accuracy and intensity distribution consistency of the infrared and visible light image blocks, and finally generate a required training dataset.
4. The multi-modal image fusion method based on modal adaptation and modal interaction compensation according to claim 1, characterized in that: In step S2, the modal difference characteristics of the infrared and visible light images are quantitatively analyzed based on multi-dimensional sparsity indicators including a gradient change rate, an entropy value, an L1 norm and a Laplacian variance.
5. The multi-modal image fusion method based on modal adaptation and modal interaction compensation according to claim 4, characterized in that: Step S2 specifically comprises the following steps; Step S21: Calculate the gradient change rate, the entropy value, the L1 norm and the Laplacian variance indicators of the infrared and visible light image blocks respectively, and establish a multi-dimensional sparsity difference quantization model; wherein the gradient change rate reflects the richness of image edge information, the entropy value represents the uncertainty of image information, the L1 norm measures the sparsity of image features, and the Laplacian variance is used to evaluate the overall sharpness of the image, and the quantization model constructed by the four complementary indicators comprehensively describes the essential difference in information sparsity between different modal images The specific calculation method is: for each image pair (IR, VIS), the infrared image index value F IR and the visible image index value F VIS are calculated respectively, and on this basis, the difference value ΔF = F IR -F VIS is calculated, and then by counting the sample proportion distribution of ΔF≥0 and ΔF<0 in the whole data set, the systematic difference characteristics of different modalities in information sparsity are revealed; the calculation methods of each quantitative index are defined as follows: For the gradient variation rate, the Sobel operator is used to calculate the horizontal and vertical gradients (G x ,G y ) of the image block, and then the arithmetic mean of the gradient amplitudes of all pixels in the image block is calculated as the gradient variation rate indicator. The specific calculation method is as follows: Wherein N is the total number of pixels, and reflects the average edge intensity The Shannon entropy of the image block gray scale histogram is calculated by using the probability distribution discretization method to improve the numerical stability of the modal difference analysis, and the specific calculation method is as follows: where p(k) represents the frequency of pixels with gray value k in the image block, i.e. needs to satisfy Logarithm base 2; For the L1 norm, the absolute value and mean value method is used to calculate the sparse feature strength of the image block, and the specific calculation method is as follows: Where I(i) represents the normalized gray value of the i-th pixel in the image, and N represents the total number of pixels. For Laplacian variance, the second-order differential operator is combined with statistical dispersion measurement, and the specific calculation method is as follows: wherein L represents a Laplacian convolution result, μ L represents a Laplacian response mean value, and N represents a total number of pixels; in step S22, a JS divergence is combined with a K-means clustering algorithm to evaluate a modality difference through symmetry measurement and to automatically identify a scene correlation strength, thereby providing an adaptive basis for feature fusion; For JS divergence, the probability distribution similarity measurement method is used to calculate the statistical characteristic difference of the infrared and visible light image blocks, and the specific calculation method is as follows: where P, Q represent the normalized gray histogram distribution of the infrared and visible image blocks respectively, M is the intermediate distribution, D KL is the KL divergence; For K-means clustering, an unsupervised learning method is used to automatically classify the image modal correlation characteristics, identify the modal correlation mode in different scenes, and the specific calculation method is as follows: where S denotes k cluster classes, μ i represents the center point of the i-th cluster, and x is the modal correlation eigenvalue calculated by JS divergence. Step S23: The sparse difference quantization result and the scene correlation characteristic analysis result are used as prior knowledge to provide feature distribution basis for the subsequent model framework including the difference feature extraction stage and the modal interaction compensation stage.
6. The multi-modal image fusion method based on modal adaptation and modal interaction compensation according to claim 1, characterized in that: Step S3 specifically includes the following steps: Step S31: The embedding projection of the input image is realized through the 3x3 convolution layer of the transformer network, the step size is 1, and the single pixel zero padding is used to keep the feature map space dimension unchanged. The linear mapping layer uses the learnable weight matrix to convert the original input into a deep feature representation, and its mathematical form is represented as: X out = conv(X in ) where conv denotes a convolution kernel, X in denotes a batch input of the structured training dataset generated by step S1, X out denotes the resulting deep feature representation; Step S32: performing depth feature representation X out The layer normalization module with configurable bias is adopted to standardize the input features and eliminate the scale difference of the feature distribution between modalities, and the mathematical expression is as follows: Where μ and σ represent the feature mean and standard deviation, γ and β represent the learnable affine parameters, and ∈ represents the numerical stability term. Step S33: X is processed by 3x3 depth separable convolution norm and generate query Q, key K, value V vectors, on which multi-head temperature regulation attention calculation is performed, which is mathematically expressed as: where τ denotes a learnable temperature parameter, d k denotes the key vector dimension, and softmax denotes the normalized exponential function; Step S34: The attention features generated in step S33 are fused with the original input features through residual connection, and layer normalization processing is performed, which is mathematically expressed as: X norm = LayerNorm(X out + Attn) where + denotes an element-wise addition of residual connections, X out denotes the depth feature representation obtained in step S31, and LayerNorm(·) denotes a normalization function, which is calculated in the same way as step S32, using the same normalization strategy and parameter configuration. Step S35: The normalized features are subjected to gated dynamic feedforward transformation and fused through residual connection, which is mathematically expressed as: where W out denotes a shrinkage projection matrix, GeLU denotes an activation function, and both denote an expansion projection matrix, and denotes a Hadamard product.
7. The multi-modal image fusion method based on modal adaptation and modal interaction compensation according to claim 1, characterized in that: In step S4, the infrared image effectively extracts sparse thermal radiation features through the neighborhood grouping matching model, The visible light image fully captures the spatial detail features through the context perception modeling network.
8. The multi-modal image fusion method based on modal adaptation and modal interaction compensation according to claim 7, characterized in that: Step S4 Specific includes the following steps; Step S41 : visual features outputted in step S3 First, a 1x1 convolution is used for channel expansion projection, which increases the number of feature channels to twice the original dimension. Then, a 7x7 empty convolution is used for multi-scale spatial feature extraction, in which the empty rate is dynamically adjusted according to the feature map resolution to maintain the optimal receptive field. Finally, a 1x1 convolution is used for channel compression projection to restore the feature dimension to the original channel number, and the processed features are added to the input features through a residual connection. The processing process enhances the expression ability and spatial context information of the features through the cascade structure of "expansion-transformation-compression", and its mathematical expression is as follows: where conv 1×1 represents a 1x1 convolution kernel, dc 7×7 represents a 7x7 dilated convolution kernel; Step S42: The features output by step S41 are subjected to hierarchical feature enhancement processing, which specifically includes: first, the input features are standardized by layer normalization; then, the features are transformed through the bottleneck structure constructed by two levels of point convolution: the first level uses 1x1 point convolution to realize channel expansion projection, which expands the feature channel number to 4 times the original dimension, and the second level uses 1x1 point convolution to realize channel contraction projection, which restores the feature channel number to the original dimension; finally, the transformed features are added to the original input features through residual connection; this processing process realizes efficient enhancement of the features through the expansion-shrinkage bottleneck structure, improves the feature expression ability, and preserves the key information of the original features, which is mathematically expressed as: X camn = pc 1×1 (pc 1×1 (LayerNorm(X s41 )))+X s41 where pc 1×1 represents a 1x1 pointwise convolution kernel, and LayerNorm(·) represents a normalization function, which is calculated in the same manner as step S32, and the same normalization strategy and parameter configuration are adopted. Step S43: generating the infrared feature for the output of step S3 First, two groups of parameters are shared to generate left / right branch features by using 1x1 convolution layers, and the mathematical expression is as follows: where conv 1×1 denotes a 1x1 convolution kernel, which uses the same parameter configuration as above. Step S44: Based on the right branch feature projection output in step S43, the index matrix S is generated by local sensitive hashing mapping for positioning TopK associated features in the feature space, and finally the index matrix is used to synchronize the grouping of left / right branch features, which is mathematically expressed as follows: where P denotes a spherical projection feature, ||·||2 denotes an L2 norm, R denotes a random rotation matrix, S denotes a hash index matrix, denotes taking the maximum value index along the hash bucket dimension b, sign(·) denotes a binary sign function, and U denotes a set-wise aggregation operation of features aggregated by index S; Step S45: Based on the grouped features obtained in step S44, the intra-group feature similarity is first calculated, then the feature similarity is used for feature aggregation, and finally the aggregation result is added and fused with the infrared feature output in step S3, which is mathematically expressed as follows: Step S46: Calculate the overall loss function Loss of the network, and the calculation formula is as follows: Loss = a1Loss ssim + a2Loss mse + a3Loss int + a4Loss grad The overall optimization target is composed of four parts, and α1, α2, α3, α4 are the hyperparameters of each part of the loss function, and the calculation method of each loss function is defined as follows: Structural similarity loss Loss ssim , for measuring the structural similarity difference between the fused image I F and the infrared image I IR , the visible light image I VIS , the specific formula is as follows: Loss ssim = 1 - SSIM(I F , I IR ) + 1 - SSIM(I F , I VIS ) Where SSIM(·) represents the structural similarity index; mean square error loss mse , for constraining the consistency of the fusion result with the overall brightness distribution of the input image, the specific calculation formula is as follows: wherein is the L2 norm square; For intensity preservation loss Loss int The significant intensity features in the source image are preserved by the forced fusion image, and the specific calculation formula is as follows: Where HW represents the image width and height respectively, ||·||1 is the L1 norm, and max(·) takes the maximum value of each pixel position of the two compared images; Gradient consistency loss Loss grad To ensure that the fused image inherits the sharpest edge information in the source image, the specific calculation formula is as follows: wherein represents a Sobel operator; Step S47: Gradient update and parameter adjustment are performed according to the overall loss value in step S46, and the performance is evaluated on the validation set at a preset interval and the optimal model is retained, and the training is automatically terminated when the maximum value is reached, and finally the best performance fusion model is saved.
9. The multi-modal image fusion method based on modal adaptation and modal interaction compensation according to claim 8, characterized in that: In step S5, based on the dynamic association characteristics between infrared and visible light modalities, a modal interaction compensation module is added based on the frozen parameters of the S4 module.
10. The multi-modal image fusion method based on modal adaptation and modal interaction compensation according to claim 9, characterized in that: Step S5 specifically includes the following steps: Step S51: First, the visible features outputted in step S3 are added to the infrared features outputted in step S4 to generate the mixed features X sum The calculation formula is as follows: Step S52: based on the mixed features X obtained in step S51 sum , a spatial attention module is constructed to take the mixed features as input to generate a spatial attention mask SA(X sum ), SA(·) represents a spatial attention block, and the mixed features X sum are input into a global average pooling layer and a global maximum pooling layer at the same time. After the two pooling results are concatenated in the channel dimension and processed by a 7x7 convolution layer, the spatial attention mask SA(X sum ) is obtained; the specific calculation method is as shown below: SA(X sum ) = conv 7×7 (Cat(GAP(X sum ), GMP(X sum ))) Where Cat(·) represents the channel dimension concatenation operation, GAP(·) represents the global average pooling operation, and GMP(·) represents the global maximum pooling operation; Step S53: based on the mixed features X obtained in step S51 sum , a channel attention module is constructed to take the mixed features as input to generate a channel attention mask CA(X sum ), CA(·) represents a channel attention block, and the mixed features X sum are input into a global average pooling layer, and then the pooled results are sent into a two-layer cascaded 1x1 convolution layer for processing to obtain the channel attention mask CA(X sum ); the specific calculation method is as shown below: CA(X sum ) = conv 1×1 (conv 1×1 (GAP(X sum ))) Step S54: Based on the spatial attention mask generated in step S52 and the channel attention mask obtained in step S53, first, the features are fused by element-wise addition, and then the fused attention features are concatenated with the original mixed features in step S51 in the channel dimension to realize the comprehensive enhancement of multi-scale features; The specific calculation method is as follows: X s54 = Cat(SA(X sum )+ CA(X sum ), X sum ) Step S55: output features X based on step S54 s54 First, feature extraction is performed by a 7x7 empty convolution layer to expand the receptive field and maintain spatial resolution, and then a sigmoid activation function is applied to generate a gating weight map. The specific calculation method is as follows: W = σ(dc 7×7 (X s54 )) Where σ represents the sigmoid function; Step S56: Based on the gated weight map W obtained in S55, for the visual features obtained in S3 and infrared features Perform dynamic weighted fusion, and then combine the weighted results with the output features X of S42 and S45 respectively. camn X nngm Residual connections are performed to ultimately obtain the enhanced features after modal compensation; the specific calculation formula is as follows: Where * represents matrix element multiplication; Step S57: Gradient backpropagation and parameter adjustment are performed based on the overall loss function Loss defined in step S46, and during the process, the model performance is evaluated on the validation set at a preset iteration interval and the optimal weights are dynamically saved, and the process is automatically terminated when the training round reaches the preset maximum value, and finally the best fusion model optimized is output.
Citation Information
Cited By
Port large model training data image segmentation method and system
CN121147935A
Port large model training data image segmentation method and system
CN121147935B
Infrared light-visible light image data fusion method
CN121330445A
Ecological quality multi-source data fusion method and system based on air-space-ground cooperation
CN121637432A
Superlens element atomic structure reverse design method and electronic equipment
CN121659811A