Multi-modal frequency domain cross image fusion method
By building a multimodal image fusion network, extracting and fusing different frequency domain features of infrared and visible images, the problem of poor results in the existing technology in low light or complex environments is solved, and the image fusion effect with clearer details and richer information is achieved.
Patent Information
- Application Number
- CN202510004454.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-02
AI Technical Summary
The existing infrared and visible image fusion technology has significantly reduced its effect under low light or complex environmental conditions, and it is difficult to effectively process and interpret multimodal perceptual information.
The multimodal frequency domain cross-image fusion method is adopted to build a multimodal image fusion network, including an encoder network, a cross-modal cross-fusion network and a decoder, extract different frequency domain features and realize feature fusion through cross-layer connection and cross-connection.
The fusion effect of multimodal images is improved, so that the fusion results have clearer details and richer information, adapting to low-light or complex environmental conditions.
Smart Images

Figure CN119992266A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image fusion, and in particular to a multi-modal frequency domain cross-image fusion method. Background Art
[0002] In modern multi-sensor systems, the fusion technology of infrared images and visible light images is developing rapidly and has become a key tool in many applications. This technology improves the overall perception ability by combining information obtained in different spectral ranges, which not only enhances the details and features of the image, but also expands the possibility of application. It has important application value in the fields of target detection, monitoring, navigation and assisted driving. Infrared imaging equipment can detect objects in low light or night conditions by capturing thermal radiation, but usually has problems of low resolution and insufficient contrast. Visible light imaging relies on ambient light and obtains the structure and color information of objects through reflected light. Its image resolution is high and detailed, but the image effect will be significantly reduced under low light or complex environmental conditions. The fusion image of infrared image and visible light image combines the advantages of both, improves the overall performance of the system, and the fusion result shows obvious advantages in scenarios such as target recognition ability, environmental adaptability and reliability.
[0003] There are three main types of infrared and visible light image fusion methods. Image-level fusion directly synthesizes the two images at the pixel level and generates a fused image through transformation and reconstruction. The method is simple and intuitive, but it may introduce data noise and distortion from different sources. Feature-level fusion extracts the features of the two images, fuses the feature information, and then makes a decision. It can achieve significant results in information compression and feature enhancement, but it also needs to solve the complexity of feature extraction. Decision-level fusion fuses the decision results from different sensors based on high-level information to improve the accuracy of judgment. It is usually used in application scenarios that only care about the final decision result. With the increasing maturity of artificial intelligence and deep learning technologies, the future of infrared and visible light image fusion technology may rely more and more on data-driven intelligent models. Through a large amount of training data and complex network structures, deep learning methods are likely to significantly improve the automation level and fusion effect of image fusion. At the same time, how to effectively process and interpret multimodal perception information and provide efficient fusion algorithms in real-time applications are still important challenges we will face. Summary of the invention
[0004] The embodiment of the present invention provides a multimodal frequency domain cross-image fusion method, which can improve the fusion effect of multimodal images and make the fusion result have clearer details and richer information. The technical solution is as follows:
[0005] The embodiment of the present invention provides a multi-modal frequency domain cross-image fusion method, comprising:
[0006] S1. Construct a multimodal image fusion network; wherein the multimodal image fusion network includes: an encoder network, a cross-modal cross-fusion network and a decoder; the execution steps of the multimodal image fusion network include:
[0007] Use the encoder network to extract different frequency domain features from each input modality image;
[0008] A cross-modal cross-fusion network is used to achieve the fusion of frequency domain features of different modal images through cross-layer connection and cross-connection.
[0009] Input the fused features into the decoder, decode the fused features, and output the fused image;
[0010] S2. Use multimodal images to train the constructed multimodal image fusion network and output a fused image containing the features of each modality image.
[0011] Further, the encoder network includes: a shared feature encoder, a basic conversion network encoder, and a detail convolutional neural network encoder;
[0012] Modal images include: infrared images and visible light images;
[0013] The method of extracting different frequency domain features from each input modality image using an encoder network includes:
[0014] Cross-modal shallow feature extraction is achieved through a shared feature encoder to obtain shallow features of infrared images and visible light images in, Represent the shallow features of infrared image and visible light image respectively;
[0015] The basic conversion network encoder is used to further extract the shallow features to obtain low-frequency basic features. in, Represent the low-frequency basic features of infrared images and visible light images respectively;
[0016] The detailed convolutional neural network encoder is used to further extract the shallow features to obtain high-frequency detail features. in, Represent the high-frequency detail features of infrared images and visible light images respectively.
[0017] Further, the shared feature encoder comprises: 2 transformer layers based on a recovery transformer block;
[0018] Each shared feature encoder is composed of two stacked transformer layers based on the recovery transformer block, and the parameters between the two shared feature encoders are independent of each other;
[0019] The formula of the shared feature encoder is expressed as:
[0020]
[0021] Among them, S(·) represents the process of extracting features by the shared feature encoder, and {I, V} represents the input paired infrared image and visible light image, respectively.
[0022] Further, the base conversion network encoder includes a conversion network module based on a lightweight converter for extracting low-frequency basic features from shallow features;
[0023] Each conversion network module includes: 2 Norm layers, 1 basic Attention layer and 1 MLP layer; Norm is normalization, Attention is attention, and MLP is a multi-layer perceptron;
[0024] The formula of the basic transformation network encoder is expressed as:
[0025] Norm(Attention(Norm(MLP(·))))=B(·)
[0026]
[0027] Among them, B(·) represents the process of extracting features by the basic transformation network encoder.
[0028] Furthermore, the detail convolutional neural network encoder includes three INNs for extracting high-frequency detail features from shallow features; wherein the INN is a reversible neural network;
[0029] In each reversible neural network processing, the transformation process is:
[0030]
[0031] Among them, ⊙ represents the Hadamard product operation, represents the 1st to cth channels in the kth reversible layer input feature, k = 1, 2, ..., K, K represents the number of reversible layer processing, C represents the number of input feature channels, CAT(·) is the channel connection operation, I i is an arbitrary mapping function, i=1,...,3, i represents different mapping function identifiers. After the reversible neural network is repeatedly processed for 3 times, the detail convolutional neural network encoder feature extraction result, i.e., the high-frequency detail feature, is output;
[0032] The formula of the entire detail convolutional neural network encoder is expressed as:
[0033]
[0034] Where D(·) represents the process of extracting features by the detail convolutional neural network encoder.
[0035] Furthermore, the cross-modal cross-fusion network includes high-frequency and low-frequency branches and a channel attention module, which are used to fuse the frequency domain features extracted by the encoder; wherein each high-frequency and low-frequency branch is composed of three layer feature extraction modules representing different depths;
[0036] The layer feature extraction process of the layer feature extraction module is expressed as:
[0037] LFC(Θ B ') = Θ B ″,LFC(Θ B ′,Θ B ″)=Θ B ″′,
[0038] LFC(Θ D ') = Θ D ″,LFC(Θ D ′,Θ D ″)=Θ D ″′
[0039] in, Represent the low-frequency basic features of infrared images and visible light images respectively, They represent the high-frequency detail features of infrared images and visible light images respectively. LFC(·) stands for layer feature extraction, which is used to extract high- and low-frequency features of different levels. Different stage features from shallow to deep can be generated according to the number of processing times. B ′,Θ B ″,Θ B ″′} represents the characteristics of each stage of low frequency; {Θ D ′,Θ D ″,Θ D ″′} represents the characteristics of each high frequency stage.
[0040] Furthermore, the cross-modal cross-fusion network is used to achieve the fusion of frequency domain features of different modal images by means of cross-layer connection and cross-connection, including:
[0041] For the obtained frequency domain features of different modality images, the layer feature extraction module is used to extract the features of different stages from shallow to deep;
[0042] In the process of processing stage features, cross-layer connection is used to link features at different levels and transfer shallow features to each deep feature.
[0043] Based on the requirements of processing different frequency domain feature information, frequency domain feature fusion is completed, and the cross connection method is adopted to extract primary features {F B ,F D} and linked to the cross-layer processing results to obtain the heteromodal same-frequency fusion features {Θ B ,Θ D}; The process of cross-layer connection and cross-connection is expressed as:
[0044]
[0045] LFC(Θ B ′,Θ B ″,Θ B ″′,F D )=Θ B ,LFC(Θ D ′,Θ D ″,Θ D ″′,F B )=Θ D
[0046] Among them, F(·) base and F(·) detail Respectively represent the primary basic feature and primary detail feature extraction processes;
[0047] The heteromodal and same-frequency fusion features are sent to the channel attention module for fusion to obtain the final fusion feature map.
[0048] Furthermore, a mixed gradient loss is added in the training process, where the mixed gradient loss is composed of a surface layer and a gradient layer, ensuring that the multimodal image fusion network training process can learn more complete fusion gradient information.
[0049] Furthermore, the mixed gradient loss converts the visible light image into the Lab color space, extracts the L brightness channel as the visible light processing channel, and mixes the visible light image L brightness channel with the infrared image channel to generate a surface layer loss result. At the same time, the Sobel operator is used to extract the gradient of the visible light image L brightness channel and the infrared image channel, and the extracted gradient map is subjected to gradient mixing processing MixGrad(·) to generate the gradient layer loss result. The obtained surface layer loss result and gradient layer loss result are summed according to the weights to obtain the final mixed gradient loss Loss Mixed_Grad , expressed as:
[0050]
[0051] Among them, {V L, I, F} represent the L brightness channel of the visible light image, the infrared image and the fused image respectively. They represent the L brightness channel gradient map of the visible light image, the infrared image gradient map and the fused image gradient map respectively, and λ is the sum weight of the surface layer loss and the gradient layer loss.
[0052] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement any one of the above-mentioned multimodal frequency domain cross-image fusion methods.
[0053] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0054] In this embodiment, a multimodal image fusion network is constructed, and the multimodal image fusion network includes: an encoder network, a cross-modal cross-fusion network and a decoder; the encoder network is used to extract different frequency domain features from each input modal image, and the high and low frequency feature extraction process of different modal images is realized by designing encoder networks for different frequency domain features; an efficient cross-modal cross-fusion network is used to realize the fusion of each frequency domain feature of different modalities, and the effect of effectively utilizing and efficiently transmitting high and low frequency information of different levels is achieved through cross-layer connection and cross-connection, thereby improving the transmission efficiency and utilization rate of information; the fused features are input into the decoder, the fused features are decoded, and a fused picture is output; the entire multimodal image fusion network is trained, and a fused picture containing the features of each modal image is output. The present invention can improve the fusion effect of multimodal images, so that the fusion result has clearer details and richer information. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0056] Figure 1 It is a flow chart of a multi-modal frequency domain cross-image fusion method provided by an embodiment of the present invention;
[0057] Figure 2 is a schematic diagram of the structure of a multimodal image fusion network provided by an embodiment of the present invention;
[0058] Figure 3 Detailed structural diagram of a convolutional neural network encoder provided by an embodiment of the present invention;
[0059] Figure 4is a schematic diagram of the structure of a cross-modal cross-fusion network provided by an embodiment of the present invention;
[0060] Figure 5 It is a schematic diagram of the hybrid gradient loss structure provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0061] The technical solution of the present invention is described below in conjunction with the accompanying drawings.
[0062] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.
[0063] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same. "of", "corresponding, relevant" and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same.
[0064] In the embodiments of the present invention, sometimes a subscript such as W1 may be written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.
[0065] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0066] The embodiment of the present invention provides a multi-modal frequency domain cross-image fusion method, such as Figure 1 and Figure 2 As shown, the processing flow of the method may include the following steps:
[0067] S1. Construct a multimodal image fusion network; wherein the multimodal image fusion network includes: an encoder network, a cross-modal cross-fusion network (also known as: a high- and low-frequency feature cross-fusion network, such as Figure 2 As shown) and a decoder; the execution steps of the multimodal image fusion network include:
[0068] Use the encoder network to extract different frequency domain features from each input modality image; thus, by designing encoder networks for different frequency domain features, the high and low frequency feature extraction process of different modalities can be realized;
[0069] The cross-modal cross-fusion network is used to achieve the fusion of frequency domain features of different modal images through cross-layer connection and cross-connection, so as to effectively utilize and efficiently transmit high- and low-frequency information of different levels, and improve the transmission efficiency and utilization rate of information;
[0070] Input the fused features into the decoder, decode the fused features, and output the fused image;
[0071] S2. Use multimodal images to train the constructed multimodal image fusion network and output a fused image containing the features of each modality image.
[0072] As a preferred embodiment, in this embodiment, the encoder network includes: a shared feature encoder, a basic conversion network encoder and a detail convolutional neural network encoder;
[0073] Modal images include: infrared images and visible light images;
[0074] The method of extracting different frequency domain features from each input modality image using an encoder network includes:
[0075] A1, cross-modal shallow feature extraction is achieved through shared feature encoder to obtain shallow features of infrared images and visible light images in, Represent the shallow features of infrared image and visible light image respectively;
[0076] In this embodiment, the shared feature encoder includes: 2 transformer layers based on the Restormer block;
[0077] Each shared feature encoder is composed of two stacked layers of transformers based on the recovery transformer block, and the parameters between the two shared feature encoders are independent of each other to better extract shallow features from different modalities of infrared and visible light images;
[0078] The formula of the shared feature encoder is expressed as:
[0079]
[0080] Among them, S(·) represents the process of extracting features by the shared feature encoder, and {I, V} represents the input paired infrared image and visible light image, respectively.
[0081] A2: Use the basic conversion network encoder to further extract features from shallow features to obtain low-frequency basic features in, Represent the low-frequency basic features of infrared images and visible light images respectively;
[0082] In this embodiment, the basic transformation network encoder includes a transformation network module based on a lightweight transformer (Lite-transformer), which is used to extract low-frequency basic features from shallow features;
[0083] Each conversion network module includes: 2 Norm layers, 1 basic Attention layer and 1 MLP layer; Norm is normalization, Attention is attention, and MLP is a multi-layer perceptron;
[0084] The formula of the basic transformation network encoder is expressed as:
[0085] Norm(Attention(Norm(MLP(·))))=B(·)
[0086]
[0087] Among them, B(·) represents the process of extracting features by the basic transformation network encoder.
[0088] A3: Use the detailed convolutional neural network encoder to further extract features from shallow features to obtain high-frequency detail features in, Represent the high-frequency detail features of infrared images and visible light images respectively.
[0089] In this embodiment, the detail convolutional neural network encoder includes three INNs (Invertible Neural Networks) for extracting high-frequency detail features from shallow features; wherein the INN is a reversible neural network;
[0090] like Figure 3 As shown, in each processing of the reversible neural network INN, the transformation process is:
[0091]
[0092] Among them, ⊙ represents the Hadamard product operation, represents the first to cth channels in the kth reversible layer input feature (k = 1, 2, ..., K), K represents the number of reversible layer processing, C represents the number of input feature channels, h×w×c represents the feature scale, CAT(·) is the channel connection operation, I i (i=1,...,3) is an arbitrary mapping function, i represents a different mapping function identifier, and in this embodiment, the bottleneck residual block (BRB) in the mobile network (MobileNetV2) is selected as the mapping function, and after the reversible neural network is repeatedly processed for 3 times, the detail convolutional neural network encoder feature extraction result, i.e., the high-frequency detail feature, is output;
[0093] It should be noted that: Figure 3 A detailed processing flow chart of a reversible neural network INN is given in FIG. After repeating the reversible neural network processing three times, the output result is the detailed convolutional neural network encoder result.
[0094] The formula of the entire detail convolutional neural network encoder is expressed as:
[0095]
[0096] Where D(·) represents the process of extracting features by the detail convolutional neural network encoder.
[0097] In this embodiment, a shared feature encoder, a basic transformation network encoder and a detail convolutional neural network encoder are used to encode the input infrared image and visible light image to obtain the low-frequency basic features and high-frequency detail features of the two modalities. In this way, by designing encoder networks for different frequency domain features, the high and low frequency feature extraction process of different modalities is realized.
[0098] As a preferred embodiment, in this embodiment, if Figure 4 As shown in the figure, the cross-modal cross-fusion network includes high-frequency and low-frequency branches and a channel attention module, which are used to fuse the frequency domain features extracted by the encoder; each high-frequency and low-frequency branch consists of 3 layer feature extraction modules representing different depths;
[0099] The layer feature extraction process of the layer feature extraction module is expressed as:
[0100] LFC(Θ B ') = Θ B ″,LFC(Θ B ′,Θ B ″)=Θ B ″′,
[0101] LFC(Θ D ') = Θ D ″,LFC(Θ D ′,Θ D ″)=Θ D ″′
[0102] in, Represent the low-frequency basic features of infrared images and visible light images respectively, They represent the high-frequency detail features of infrared images and visible light images respectively. LFC(·) stands for layer feature extraction, which is used to extract high- and low-frequency features of different levels. Different stage features from shallow to deep can be generated according to the number of processing times. B ′,ΘB ″,Θ B ″′} represents the characteristics of each stage of low frequency; {Θ D ′,Θ D ″,Θ D ″′} represents the characteristics of each high frequency stage.
[0103] It should be noted that: Figure 4 The fourth layer feature extraction module in represents only the feature extraction result and does not include the processing process.
[0104] As a preferred embodiment, in this embodiment, the use of a cross-modal cross-fusion network to achieve fusion of frequency domain features of different modal images by means of cross-layer connection and cross-connection includes:
[0105] For the obtained frequency domain features of different modality images, the layer feature extraction module is used to extract the features of different stages from shallow to deep;
[0106] In the process of processing stage features, a cross-layer connection method is used to link features at different levels and transfer shallow features to each deep feature. Through this processing method, the cross-modal cross-fusion network can retain the original features while improving the fusion depth of the same-frequency features.
[0107] Based on the requirements of processing different frequency domain feature information, frequency domain feature fusion is completed, and the cross connection method is adopted to extract primary features {F B ,F D} and linked to the cross-layer processing results to obtain the heteromodal same-frequency fusion features {Θ B ,Θ D}, improve the fusion depth between frequency domain features, and realize the mutual transmission between different frequency domain feature information; among them, the process of cross-layer connection and cross-connection is expressed as:
[0108]
[0109] LFC(Θ B ′,Θ B ″,Θ B ″′,F D )=Θ B ,LFC(Θ D ′,Θ D ″,Θ D ″′,F B )=Θ D
[0110] Among them, F(·) base and F(·) detailRespectively represent the primary basic features and primary detail feature extraction process; in this way, the adaptability of the cross-modal cross-fusion network to cross-frequency domain feature fusion can be improved, and the fusion effect can be improved; after cross-layer connection and cross-connection, the cross-modal cross-fusion network will produce heteromodal same-frequency fusion features with strong adaptability, original features and deep fusion {Θ B ,Θ D};
[0111] The heteromodal and same-frequency fusion features are sent to the channel attention module for fusion to obtain the final fusion feature map.
[0112] In this embodiment, a cross-modal cross-fusion network is built to achieve an efficient feature fusion process. The unique cross-layer connection method brings a more complete modal information integration effect to the fused image, and the cross-connection design provides higher robustness for the fusion features.
[0113] As a preferred embodiment, in this embodiment, a mixed gradient loss is added to the commonly used detail loss (MSEloss) and pixel loss (SSIMloss) during the training process. The mixed gradient loss is composed of a surface layer and a gradient layer, ensuring that the multimodal image fusion network training process can learn more complete fusion gradient information.
[0114] As a preferred embodiment, in this embodiment, if Figure 5 As shown in FIG. 1 , the mixed gradient loss is used to preserve the visible light image information as much as possible and reduce information loss. The visible light image is converted into the Lab color space. Since the human eye is more sensitive to changes in brightness and the brightness component directly affects the contrast and structural details of the image, the L brightness channel is extracted as the visible light processing channel. The visible light image L brightness channel is mixed with the infrared image channel to produce a surface layer loss result. At the same time, the Sobel operator is used to extract the gradient of the visible light image L brightness channel and the infrared image channel, and the extracted gradient map is subjected to gradient mixing processing MixGrad(·) to generate the gradient layer loss result. The obtained surface layer loss result and gradient layer loss result are summed according to the weights to obtain the final mixed gradient loss Loss Mixed_Grad (Right now: Figure 2 L Mixed_Grad ), expressed as:
[0115]
[0116] Among them, {V L , I, F} represent the L brightness channel of the visible light image, the infrared image and the fused image respectively. They represent the L brightness channel gradient map of the visible light image, the infrared image gradient map, and the fused image gradient map, respectively. λ is the sum weight of the surface layer loss and the gradient layer loss, and λ=10 is taken here. Finally, the loss function is expressed as:
[0117] Loss = MSEloss + SSIMloss + Loss Mixed_Grad
[0118] Among them, MSEloss is detail loss, which is obtained by weighting the MSE values calculated by the infrared image, visible light image and fused image respectively; SSIMloss is pixel loss, which is obtained by weighting the inverse of the SSIM values calculated by the infrared image, visible light image and fused image respectively.
[0119] In this embodiment, the hybrid gradient loss maximizes the use of the information of the visible light image and the infrared image in the loss function calculation by retaining the light and dark details of the visible light image to the greatest extent, thereby providing better guidance for the training process of the entire multimodal image fusion network, helping the multimodal image fusion network to better utilize the modal information while bringing clearer details to the fusion results.
[0120] In this embodiment, the multimodal image fusion network extracts the low-frequency basic features and high-frequency detail features of the two modal images by fusing visible light images and infrared images, and realizes the fusion process of the two features between modalities through an efficient cross-modal cross-fusion network. The fused image is output through the decoder, and a mixed gradient loss is added during the training process to help network training, which can improve the multimodal image fusion effect, coordinate the infrared image information and the visible light image brightness information, and finally output a fused image with clearer details and richer information.
[0121] In this embodiment, the trained multimodal image fusion network can be used to efficiently extract the features of each modality, obtain the high-frequency detail features and low-frequency basic features of the infrared image and the visible light image, and realize effective fusion of the features, and finally output a fused image with clearer details and richer information including infrared features and visible light features.
[0122] In this embodiment, in order to verify the effectiveness of the multimodal frequency domain cross image fusion method provided by the embodiment of the present invention, the MSRS data set is used to evaluate its performance. The test environment and parameters are set as follows: the number of training iterations is 100 times, the training batch size is 16, the training image size is 320×240, the Adam optimizer is used to optimize the training, the StepLR scheduler is used to optimize the learning rate, and the training environment is 2 RTX2080Ti; the evaluation indicators and results are described as follows:
[0123] The evaluation results of the multimodal frequency domain cross-image fusion method (hereinafter referred to as the method of the present invention) provided in the embodiment of the present invention are compared with other methods in Table 1; in order to reflect the advantages of the method of the present invention in terms of fusion results, multiple indicators such as EN (Entropy), SD (Standard Deviation), SSIM (Structural Similarity), AG (Average Gradient), and PSNR (Peak Signal-to-Noise Ratio) are selected for comparison;
[0124] In this embodiment, entropy (EN) is an evaluation index based on information theory, representing the amount of information in the image, and is mathematically defined as follows:
[0125]
[0126] Among them, L represents the number of gray levels, p l Represents the normalized histogram of the corresponding gray levels in the fused image. The larger the entropy, the more information is contained in the fused image and the better the performance of the fusion method.
[0127] In this embodiment, the standard deviation (SD) indicator is based on a statistical concept that reflects the distribution and contrast of the fused image. The standard deviation is mathematically defined as follows:
[0128]
[0129] Where M represents the length of the processed image, N represents the width of the processed image, F(i,j) represents the image pixel value at position (ij), and μ represents the average value of the fused image. Due to the sensitivity of the human visual system to contrast, high-contrast areas always attract human attention. Therefore, high-contrast fused images tend to produce larger standard deviations, which means that the fused image can achieve good visual effects.
[0130] In this embodiment, the structural similarity index measure (SSIM) is used to simulate image loss and distortion. The index is mainly composed of three parts: correlation loss and brightness and contrast distortion. The product of these three components is the evaluation result of the fused image, which is defined as follows:
[0131]
[0132] Among them, SSIM X,F represents the structural similarity between the source image X and the fused image F; x and f represent the image patches of the source image and the fused image in the sliding window, respectively; σ xf represents the covariance of the source image and the fused image; σ x and σ frepresents standard deviation (SD); μ x and μ f Represent the average values of the source image and the fused image respectively; C1, C2 and C3 are parameters used to stabilize the algorithm; when C1 = C2 = C3 = 0, SSIM is simplified to a universal image quality index. Therefore, the structural similarity between all source images and fused images can be written as follows:
[0133] SSIM=SSIM A,F +SSIM B,F
[0134] Among them, SSIM A,F and SSIM B,F represent the structural similarity between infrared / visible and fused images, respectively.
[0135] In this embodiment, the average gradient (AG) index quantifies the gradient information of the fused image and represents its details and texture. The average gradient index is defined as follows:
[0136]
[0137] in, They represent the horizontal and vertical gradients of the pixel at position (i, j) respectively; the larger the average gradient metric, the more gradient information the fused image contains, and the better the performance of the fusion algorithm.
[0138] In this embodiment, the peak signal-to-noise ratio (PSNR) indicator is the ratio of the peak power to the noise power in the fused image, and thus reflects the distortion in the fusion process. The peak signal-to-noise ratio indicator is defined as follows:
[0139]
[0140] Among them, MSE AF and MSE BF Respectively represent the difference between the fused image and the infrared and visible light images, where A(i,j) represents the infrared image pixel value, B(i,j) represents the visible light image pixel value; r represents the peak value of the fused image. The larger the peak signal-to-noise ratio, the closer the fused image is to the source image, and the smaller the distortion produced by the fusion method.
[0141] Table 1 Comparison of the proposed method with other methods in MSRS dataset
[0142]
[0143] According to the results in Table 1, it can be seen that by comparing with other methods, the method of the present invention has achieved satisfactory results in all indicators, which shows that the infrared visible light fusion results obtained by the method of the present invention have excellent image effects. The method of the present invention has greatly improved the EN value and AG value compared with other methods, indicating that the fused image obtained by the method of the present invention has richer information and better gradient results; the excellent results of the method of the present invention in the SD value also show that the obtained fused image has better visual effects.
[0144] In order to verify the significance of each part of the method described in this embodiment, an ablation experiment was also performed in this embodiment.
[0145] In this embodiment, the effectiveness of the cross-modal cross-fusion network and the mixed gradient loss is verified by performing ablation experiments. Table 2 shows the fusion result indicators of the summed fusion network (Ori+Add) without the cross-modal cross-fusion network and the mixed gradient loss, the original cross-modal cross-fusion network (Ori+Cross), and the original cross-modal cross-fusion network with the mixed gradient loss added (Ori+Cross+G). Here, SF (Spatial Frequency), SSIM (Structural Similarity) structural similarity index measurement, AG (Average Gradient) average gradient, and PSNR (Peak Signal-to-Noise Ratio) peak signal-to-noise ratio are selected for comparison:
[0146] In this embodiment, spatial frequency (SF) is an image quality indicator based on gradients, namely horizontal gradient and vertical gradient, also referred to as spatial row frequency (RF) and column frequency (CF). The spatial frequency indicator can effectively measure the gradient distribution of an image, thereby revealing the details and texture of the image. This indicator is defined as follows:
[0147]
[0148] According to the human visual system, the fused image with large SF is sensitive to human perception with rich edges and textures.
[0149] Table 2 Ablation experiments of the proposed method in MSRS dataset
[0150]
[0151] According to the results in Table 2, it can be seen that: by comparing with other ablation experimental results, the cross-modal cross-fusion network and mixed gradient loss given by the method of the present invention both improve the network fusion effect; after the network adopts the cross-modal cross-fusion network and mixed gradient loss, the fusion result has more gradient information, clearer details and less distortion effect.
[0152] The multimodal frequency domain cross-image fusion method described in the embodiment of the present invention has at least the following advantages:
[0153] 1) The present invention is a multi-modal frequency domain cross-image fusion method, which extracts the high and low frequency features of infrared images and visible light images respectively, and fuses them to produce a fused image with richer information;
[0154] 2) An efficient cross-modal cross-fusion network is designed to obtain high- and low-frequency feature results at different levels by calculating, and then combine cross-layer connection and cross-connection to complete the feature fusion process, thereby improving the information transmission efficiency and fusion effect;
[0155] 3) A loss function calculation method designed to improve the multimodal image fusion effect and improve the detail and texture performance of the fused image;
[0156] 4) The present invention verifies this method on the MSRS dataset and can obtain better results than the most advanced methods.
[0157] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.
[0158] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.
[0159] In the present invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can be represented by: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.
[0160] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0161] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0162] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0163] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0164] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0165] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0166] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0167] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A multimodal frequency domain cross-image fusion method, characterized in that: The method comprises: S1. Construct a multimodal image fusion network; wherein the multimodal image fusion network includes: an encoder network, a cross-modal cross-fusion network and a decoder; the execution steps of the multimodal image fusion network include: Use the encoder network to extract different frequency domain features from each input modality image; A cross-modal cross-fusion network is used to achieve the fusion of frequency domain features of different modal images through cross-layer connection and cross-connection. Input the fused features into the decoder, decode the fused features, and output the fused image; S2. Use multimodal images to train the constructed multimodal image fusion network and output a fused image containing the features of each modality image.
2. The multimodal frequency domain cross image fusion method according to claim 1, characterized in that: The encoder network includes: a shared feature encoder, a basic conversion network encoder, and a detail convolutional neural network encoder; Modal images include: infrared images and visible light images; The method of extracting different frequency domain features from each input modality image using an encoder network includes: Cross-modal shallow feature extraction is achieved through a shared feature encoder to obtain shallow features of infrared images and visible light images in, Represent the shallow features of infrared image and visible light image respectively; The basic conversion network encoder is used to further extract the shallow features to obtain low-frequency basic features. in, Represent the low-frequency basic features of infrared images and visible light images respectively; The detailed convolutional neural network encoder is used to further extract the shallow features to obtain high-frequency detail features. in, Represent the high-frequency detail features of infrared images and visible light images respectively.
3. The multimodal frequency domain cross image fusion method according to claim 2, characterized in that: The shared feature encoder includes: 2 transformer layers based on recovery transformer blocks; Each shared feature encoder is composed of two stacked transformer layers based on the recovery transformer block, and the parameters between the two shared feature encoders are independent of each other; The formula of the shared feature encoder is expressed as: Among them, S(·) represents the process of extracting features by the shared feature encoder, and {I, V} represents the input paired infrared image and visible light image, respectively.
4. The multimodal frequency domain cross image fusion method according to claim 3, characterized in that: The base conversion network encoder includes a conversion network module based on a lightweight converter, which is used to extract low-frequency basic features from shallow features; Each conversion network module includes: 2 Norm layers, 1 basic Attention layer and 1 MLP layer; Norm is normalization, Attention is attention, and MLP is a multi-layer perceptron; The formula of the basic transformation network encoder is expressed as: Norm(Attention(Norm(MLP(·))))=B(·) Among them, B(·) represents the process of extracting features by the basic transformation network encoder.
5. The multimodal frequency domain cross-image fusion method according to claim 4, characterized in that: The detail convolutional neural network encoder includes three INNs, which are used to extract high-frequency detail features from shallow features; wherein the INN is a reversible neural network; In each reversible neural network processing, the transformation process is: Among them, ⊙ represents the Hadamard product operation, represents the 1st to cth channels in the kth reversible layer input feature, k = 1, 2, ..., K, K represents the number of reversible layer processing, C represents the number of input feature channels, CAT(·) is the channel connection operation, I i is an arbitrary mapping function, i=1,...,3, i represents different mapping function identifiers. After the reversible neural network is repeatedly processed for 3 times, the detail convolutional neural network encoder feature extraction result, i.e., the high-frequency detail feature, is output; The formula of the entire detail convolutional neural network encoder is expressed as: Where D(·) represents the process of extracting features by the detail convolutional neural network encoder.
6. The multimodal frequency domain cross-image fusion method according to claim 1, characterized in that: The cross-modal cross-fusion network consists of high-frequency and low-frequency branches and a channel attention module, which are used to fuse the frequency domain features extracted by the encoder; each high-frequency and low-frequency branch consists of three layer feature extraction modules representing different depths; The layer feature extraction process of the layer feature extraction module is expressed as: LFC(W) B ′)=Θ B ″,LFC(Θ B ′,Θ B ″)=Θ B ″′, LFC(W) D ′)=Θ D ″,LFC(Θ D ′,Θ D ″)=Θ D ″′ in, Represent the low-frequency basic features of infrared images and visible light images respectively, They represent the high-frequency detail features of infrared images and visible light images respectively. LFC(·) stands for layer feature extraction, which is used to extract high- and low-frequency features of different levels. Different stage features from shallow to deep can be generated according to the number of processing times. B ′,Θ B ″,Θ B ″′} represents the characteristics of each stage of low frequency; {Θ D ′,Θ D ″,Θ D ″′} represents the characteristics of each high frequency stage.
7. The multimodal frequency domain cross image fusion method according to claim 6, characterized in that: The method of using a cross-modal cross-fusion network to achieve fusion of frequency domain features of different modal images by means of cross-layer connection and cross-connection includes: For the obtained frequency domain features of different modality images, the layer feature extraction module is used to extract the features of different stages from shallow to deep; In the process of processing stage features, cross-layer connection is used to link features at different levels and transfer shallow features to each deep feature. Based on the requirements of processing different frequency domain feature information, frequency domain feature fusion is completed, and the cross connection method is adopted to extract primary features {F B ,F D } and linked to the cross-layer processing results to obtain the heteromodal same-frequency fusion features {Θ B ,Θ D }; The process of cross-layer connection and cross-connection is expressed as: LFC(W) B ′,Θ B ″,Θ B ″′,F D )=Θ B ,LFC(Θ D ′,Θ D ″,Θ D ″′,F B )=Θ D Among them, F(·) base and F(·) detail Respectively represent the primary basic feature and primary detail feature extraction processes; The heteromodal and same-frequency fusion features are sent to the channel attention module for fusion to obtain the final fusion feature map.
8. The multimodal frequency domain cross image fusion method according to claim 1, characterized in that: A mixed gradient loss is added in the training process, where the mixed gradient loss is composed of a surface layer and a gradient layer, to ensure that the multimodal image fusion network training process can learn more complete fusion gradient information.
9. The multimodal frequency domain cross image fusion method according to claim 8, characterized in that: The mixed gradient loss converts the visible light image into Lab color space, extracts the L brightness channel as the visible light processing channel, and mixes the visible light image L brightness channel with the infrared image channel to generate a surface layer loss result. At the same time, the Sobel operator is used to extract the gradient of the visible light image L brightness channel and the infrared image channel, and the extracted gradient map is subjected to gradient mixing processing MixGrad(·) to generate the gradient layer loss result. The obtained surface layer loss result and gradient layer loss result are summed according to the weights to obtain the final mixed gradient loss Loss Mixed_Grad , expressed as: Among them, {V L , I, F} represent the L brightness channel of the visible light image, the infrared image and the fused image respectively. They represent the L brightness channel gradient map of the visible light image, the infrared image gradient map and the fused image gradient map respectively, and λ is the sum weight of the surface layer loss and the gradient layer loss.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Visible light image and infrared image fusion method based on Laplace pyramid
CN112184606A
Multi-modal medical image fusion method based on multi-scale codec
CN116757982A
Feature decomposition-based infrared image and visible light image fusion method
CN118134780A
Multi-modal medical image fusion method
CN118608396A
Object-level infrared-and-visible-light image fusion method based on fully convolutional neural network
WO2024174488A1
Cited By
Multi-modal brain tumor segmentation method based on frequency domain channel attention
CN120580433A
A multi-modal brain tumor segmentation method based on frequency domain channel attention
CN120580433B