Multi-modal fusion method for SAR and visible light image feature enhancement

The features of SAR and visible light images are extracted through a multimodal shared feature encoder and an improved dual-branch independent encoder, and the images are reconstructed through the decoder, solving the problems of high complexity and unstable quality in the prior art, and achieving high resolution and detailed accuracy of image fusion effect.

CN120032216APending Publication Date: 2025-05-23ZHEJIANG UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510094235.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing SAR and visible image fusion methods have high computational complexity and unstable fusion quality when processing large-scale complex battlefield data, making it difficult to meet the actual combat needs.

Method used

Using a multimodal fusion method, the shallow and high-frequency detail features of SAR and visible images are extracted through a multimodal shared feature encoder and improved dual-branch independent encoder, and the image is reconstructed by the decoder, and the loss is calculated to train the model.

Benefits of technology

Through deep learning models, the multimodal images are effectively fused, the resolution and detail accuracy of the fusion images are improved, and the real-time application in complex environments is adapted to the urgent needs of the remote sensing field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032216A_ABST
    Figure CN120032216A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal fusion method for SAR and visible light image feature enhancement, and the method comprises the steps: inputting an original SAR image and a visible light image into a multi-modal shared feature encoder, extracting cross-modal shallow features corresponding to the original SAR image and the visible light image, inputting the cross-modal shallow features into corresponding branches of an improved double-branch independent encoder, obtaining corresponding single-modal low-frequency basic features, and carrying out the recognition of the single-modal low-frequency basic features. Extracting corresponding high-frequency detail features, respectively combining the corresponding low-frequency basic features and the high-frequency detail features, inputting the combined features into a decoder, reconstructing an image, respectively calculating loss, and training an encoder and the decoder; and inputting the paired SAR image and visible light image, extracting the current basic features and high-frequency detail features, inputting the extracted current basic features and high-frequency detail features into a feature fusion layer, inputting the fused basic features and high-frequency detail features into a decoder, and training the fused basic features and high-frequency detail features to generate a fused image. The method improves the resolution and detail precision of the fused image, is suitable for real-time application in a complex environment, is suitable for remote sensing application scenes, and is efficient and reliable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of general image data processing or generation, and in particular to a multimodal fusion method for SAR and visible light image feature enhancement in the field of image processing and deep learning. Background Art

[0002] In modern military operations, obtaining high-precision, high-resolution battlefield situation awareness information is crucial for command decision-making. Remote sensing technology, as a key means of battlefield monitoring and target identification, can provide real-time, multi-dimensional image information support for the military in complex environments. Synthetic Aperture Radar (SAR) and visible light image fusion technology are particularly important in the military field, especially in all-weather and all-day combat environments, showing unique advantages.

[0003] SAR technology generates high-resolution radar images by receiving electromagnetic wave signals reflected by targets. Its advantages include penetrating clouds, rain, fog and dark environments, so that stable images can be formed even in bad weather. This technology is widely used in battlefield monitoring, reconnaissance and target positioning. However, SAR images are insufficient in resolving detailed texture information of ground targets and are difficult to provide sufficient visual details, especially in complex terrain or urban environments, and cannot accurately identify small objects or camouflaged targets.

[0004] In contrast, visible light remote sensing images, with their rich texture and color information, can clearly show the shape, material and detailed structure of the target. However, due to weather and lighting conditions, the quality of visible light images drops sharply at night or in bad weather, and cannot provide continuous situational awareness support for the military.

[0005] In order to solve the limitations of a single sensor, the fusion technology of SAR and visible light images has become a research hotspot in the current military remote sensing field. By fusing the two images, we can not only utilize the all-weather monitoring capability of SAR images, but also combine the rich detail information of visible light images to generate a high-quality comprehensive image, providing more comprehensive intelligence support for command decision-making. However, the current image fusion methods have problems such as high computational complexity and unstable fusion quality when processing large-scale and complex battlefield data, which makes it difficult to meet the needs of actual combat. Summary of the invention

[0006] In order to solve the above technical problems, the present invention provides a multimodal fusion method for SAR and visible light image feature enhancement.

[0007] The technical solution adopted by the present invention is a multimodal fusion method for SAR and visible light image feature enhancement, the method comprising the following steps:

[0008] S1 inputs the original SAR image and visible light image into the multimodal shared feature encoder to extract the shallow features corresponding to the SAR and visible light images respectively. and The multimodal shared feature encoder here is mainly used to extract shallow features including but not limited to similar features of the two modalities, such as similar background information;

[0009] S2 and Input the corresponding branch of the improved dual-branch independent encoder to obtain the corresponding low-frequency basic features and

[0010] S3 and Input the corresponding branch of the improved dual-branch independent encoder to obtain the corresponding high-frequency detail features and

[0011] S4 and and After being combined separately, they are input into the decoder. The decoder consists of multiple RestormerBlocks to reconstruct the image and The losses of the corresponding SAR images and visible light images are calculated respectively, and the multimodal shared feature encoder, improved dual-branch independent encoder and decoder are trained;

[0012] S5 is based on the trained encoder and decoder, inputs the paired SAR image and visible light image, and extracts the current and Then input feature fusion layer;

[0013] S6 will merge the features and Input the decoder, generate a fused image, calculate the loss between the fused image and the original SAR image and visible light image, and train the feature fusion layer;

[0014] S7 inputs the original SAR image and the visible light image into the trained encoder, decoder and feature fusion layer to calculate the fused image F.

[0015] Preferably, in S2, the improved dual-branch independent encoder comprises an improved SAR image independent encoder and an improved visible light image independent encoder corresponding to the SAR image and the visible light image respectively;

[0016] The improved SAR image independent encoder and the improved visible light image independent encoder respectively include a low-frequency basic feature encoder and a high-frequency detail feature encoder that are arranged in coordination.

[0017] Preferably, the low-frequency basic feature encoder of the improved SAR image independent encoder comprises a multi-frequency progressive channel self-attention module, a PoolMLP layer and a DropPath layer arranged in sequence, and the low-frequency basic feature encoder of the improved SAR image independent encoder is input Add the output of the DropPath layer and output The PoolMLP layer contains two 1x1 convolutional layers and an activation function;

[0018] A normalization layer is provided at both the input and output ends of the multi-frequency progressive channel self-attention module.

[0019] Preferably, the multi-frequency progressive channel self-attention module includes a multi-frequency channel attention module and a channel self-attention module arranged in sequence.

[0020] Preferably, the high-frequency detail feature encoder of the improved SAR image independent encoder comprises a plurality of inverse residual modules arranged in sequence;

[0021] The SAR image tensor is divided from the original channel to form two input tensors with C / 2 channels. and Input the first inverted residual module and multiply or sum the output of each inverted residual module except the first one, and output a tensor with a corresponding channel number of C / 2. Add the output of the first inverted residual module and output the corresponding tensor And input each inverted residual module except the first inverted residual module, and Concatenate on the channels to form the tensor output φ ′ sar .

[0022] Preferably, the low-frequency basic feature encoder of the improved visible light image independent encoder comprises a pooling layer, a PoolMLP layer and a DropPath layer arranged in sequence, and the shallow feature of the low-frequency basic feature encoder of the improved visible light image independent encoder is input. Add the output of the DropPath layer and output The PoolMLP layer contains two 1x1 convolutional layers and an activation function;

[0023] A normalization layer is provided at both the input and output ends of the pooling layer.

[0024] Preferably, the high-frequency detail feature encoder of the improved visible light image independent encoder comprises a plurality of SE modules (Squeeze-and-Excitation Blocks) arranged in sequence;

[0025] The visible light image tensor is divided from the original channel to form two input tensors with C / 2 channels. and Input the first SE module and multiply or sum the output of each SE module except the first SE module, and output a tensor with the corresponding channel number C / 2 Add the output of the first SE module and output the corresponding tensor And enter each SE module except the first SE module, and Concatenate on the channels to form the tensor output φ ′ vis .

[0026] In the present invention, the low-frequency basic features include but are not limited to the overall structure and background information, and the high-frequency detail features include but are not limited to the detail information of the edge texture mode.

[0027] Preferably, in S4, the loss And SAR image loss L sar , visible light image loss L vis and the eigendecomposition loss L decomp association.

[0028] Preferably, in S5, and Input the basic feature fusion layer and and Input detail fusion layer, output fused basic features φ BF and detail features φ DF .

[0029] Preferably, the total loss of the fused image is calculated and trained; the total loss Associated with a-Huber loss, gradient loss, and decomposition loss.

[0030] The present invention provides a multimodal fusion method for SAR and visible light image feature enhancement, which comprises the following steps: inputting the original SAR image and the visible light image into a multimodal shared feature encoder, extracting cross-modal shallow features corresponding to the SAR and visible light images respectively, inputting the features into the corresponding branches of an improved dual-branch independent encoder, and obtaining corresponding single-modal low-frequency basic features; then extracting corresponding high-frequency detail features from the corresponding branches of the improved dual-branch independent encoder, respectively combining the corresponding low-frequency basic features and high-frequency detail features and inputting them into a decoder, reconstructing the image and respectively calculating the loss, and training the encoder and decoder; then based on the trained encoder and decoder, inputting paired SAR images and visible light images, extracting the current basic features and high-frequency detail features and then inputting them into a feature fusion layer, inputting the fused basic features and high-frequency detail features into a decoder, training and using them to generate a fused image.

[0031] The beneficial effect of the present invention is that the multimodal images are effectively fused through a deep learning model, which can not only improve the resolution and detail accuracy of the fused image, but also adapt to real-time applications in complex environments, solve the urgent needs of the current remote sensing field, and is suitable for remote sensing application scenarios such as object recognition and disaster monitoring, and is efficient and reliable. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is a flow chart of the method of the present invention;

[0033] Figure 2 It is a structural diagram of the improved SAR image independent encoder of the present invention;

[0034] Figure 3 A structural diagram of an improved visible light image independent encoder according to the present invention;

[0035] Figure 4 This is a structural diagram of the multi-frequency progressive channel self-attention module in the present invention;

[0036] Figure 5 It is a schematic diagram of an embodiment of the present invention. DETAILED DESCRIPTION

[0037] The present invention is further described in detail below in conjunction with embodiments, but the protection scope of the present invention is not limited thereto.

[0038] The present invention relates to a multimodal fusion method for SAR and visible light image feature enhancement, the method comprising the following steps:

[0039] (1) Input the original SAR image and visible light image into the multimodal shared feature encoder to extract the shallow features corresponding to the SAR and visible light images respectively. and

[0040] (2) and Input the corresponding branch of the improved dual-branch independent encoder to obtain the corresponding low-frequency basic features and

[0041] (3) and Input the corresponding branch of the improved dual-branch independent encoder to obtain the corresponding high-frequency detail features and

[0042] (4) and and After combining them separately, input them into the decoder to reconstruct the image and The losses of the corresponding SAR images and visible light images are calculated respectively, and the multimodal shared feature encoder, improved dual-branch independent encoder and decoder are trained;

[0043] (5) Based on the trained encoder and decoder, input the paired SAR image and visible light image to extract the current and Then input feature fusion layer;

[0044] (6) The fused features and Input decoder, generate fused image, calculate the loss of fused image and original SAR image and visible light image, train feature fusion layer

[0045] (7) The original SAR image and the visible light image are input into the trained encoder, decoder and feature fusion layer to calculate the fused image F.

[0046] The steps are described below in conjunction with specific embodiments.

[0047] (1) Input the original SAR image and visible light image into the multimodal shared feature encoder to extract the shallow features corresponding to the SAR and visible light images respectively. and

[0048] Here is the original SAR image I sar With visible light image I vis are paired images of size H×W×1 and H×W×3, respectively. The image pairs are input into a multi-modal shared encoder, which in this embodiment is a Restormer module, which can extract shallow features. Output shallow features of SAR and visible light images and

[0049] The features here are common shallow features, which contain the global shallow basic modal common information of the image (such as background, large-area object contours, etc.).

[0050] (2) and Input the corresponding branch of the improved dual-branch independent encoder to obtain the corresponding low-frequency basic features and

[0051] The improved dual-branch independent encoder comprises an improved SAR image independent encoder and an improved visible light image independent encoder corresponding to the SAR image and the visible light image respectively;

[0052] The improved SAR image independent encoder and the improved visible light image independent encoder respectively include a low-frequency basic feature encoder and a high-frequency detail feature encoder that are arranged in coordination.

[0053] (3) and Input the corresponding branch of the improved dual-branch independent encoder to obtain the corresponding high-frequency detail features and

[0054] The improved SAR image independent encoder and the improved visible light image independent encoder are described below respectively.

[0055] (2-1) Improved SAR image independent encoder

[0056] The low-frequency basic feature encoder of the improved SAR image independent encoder comprises a multi-frequency progressive channel self-attention module, a PoolMLP layer and a DropPath layer arranged in sequence, and the low-frequency basic feature encoder of the improved SAR image independent encoder is input. Add the output of the DropPath layer and output The PoolMLP layer contains two 1x1 convolutional layers and an activation function;

[0057] A normalization layer is provided at both the input and output ends of the multi-frequency progressive channel self-attention module;

[0058] The multi-frequency progressive channel self-attention module includes a multi-frequency channel attention module and a channel self-attention module arranged in sequence.

[0059] In the present invention, for shallow features First, normalize the data using the normalization layer, perform the multi-frequency progressive channel self-attention (MFPA) operation using the multi-frequency progressive channel self-attention module, and then normalize again to obtain the updated features. satisfy,

[0060]

[0061] Among them, LN is layer normalization and MFPA is the multi-frequency progressive channel self-attention module.

[0062] Here, MFPA(X)=PCSA(MFCA(X)), which corresponds to the sequentially set multi-frequency progressive channel self-attention (MFCA) module and progressive channel self-attention (PCSA) module; the multi-frequency progressive channel self-attention module will extract low-frequency features through discrete cosine transform (DCT), and by explicitly modeling low-frequency information, fully retain the global background information of the image, combine global statistical features (average, maximum and minimum values), dynamically adjust channel weights, strengthen important modal information, and suppress redundant features and noise influences.

[0063] The input feature map is C s is the number of channels (feature depth), H s ,W s is the height and width of the feature map, (u k ,v k ) is a frequency index, defining a specific frequency component; is the basis function of the two-dimensional DCT:

[0064]

[0065] Finally, the basic features are extracted through the DropPath operation. DropPath is a random depth operation.

[0066] The high-frequency detail feature encoder of the improved SAR image independent encoder comprises a plurality of inverse residual modules arranged in sequence;

[0067] The SAR image tensor is divided from the original channel to form two input tensors with C / 2 channels. and Input the first inverted residual module and multiply or sum the output of each inverted residual module except the first one, and output a tensor with a corresponding channel number of C / 2. Add the output of the first inverted residual module and output the corresponding tensor And input each inverted residual module except the first inverted residual module, and Concatenate on the channels to form the tensor output φ ′ sar .

[0068] In the present invention, after extracting the shallow features, the system uses the SAR high-frequency detail feature encoder to perform deep-level refinement on the image features and extract the SAR branch low-frequency basic features. Reduce computational complexity while highlighting local texture and structural features in SAR images;

[0069] Specifically, multiple inverted residual blocks (IRBs) are used to extract high-frequency detail features of SAR images. k represents the kth IRB layer, k=0,1,...,N-1, and finally we get

[0070] (2-2) Improved visible light image independent encoder

[0071] The low-frequency basic feature encoder of the improved visible light image independent encoder includes a pooling layer, a PoolMLP layer and a DropPath layer arranged in sequence, and the shallow feature of the low-frequency basic feature encoder of the improved visible light image independent encoder is input. Add the output of the DropPath layer and output The PoolMLP layer contains two 1x1 convolutional layers and an activation function;

[0072] A normalization layer is provided at both the input and output ends of the pooling layer.

[0073] In the present invention, for shallow features First, normalize the data using the normalization layer, perform pooling, and then normalize again to obtain updated features. satisfy

[0074]

[0075] Among them, LN is the layer normalization operation, and Pooling is the pooling operation;

[0076] Then enter Perform PoolMLP operation and DropPath to further optimize and obtain low-frequency basic features satisfy

[0077] The high-frequency detail feature encoder of the improved visible light image independent encoder comprises a plurality of SE modules arranged in sequence;

[0078] The visible light image tensor is divided from the original channel C to form two input tensors with the number of channels C / 2. and Input the first SE module and multiply or sum the output of each SE module except the first SE module, and output a tensor with the corresponding channel number C / 2 Add the output of the first SE module and output the corresponding tensor And enter each SE module except the first SE module, and Concatenate on the channels to form the tensor output φ ′ vis .

[0079] In the present invention, after extracting the shallow features, the system uses the visible light high-frequency detail feature encoder to perform deep-level refinement on the image features and extract the low-frequency basic features of the visible light branches. Effectively capture the dependencies between channels and highlight the channel features that contribute more to the key information in the image;

[0080] Specifically, multiple SE modules (Squeeze-and-Excitation Block) are used to extract the deep detail features of visible light. k represents the kth SEB layer, k=0,1,...,N-1, and finally we get

[0081] (4) and and After being combined, they are input into the decoder. The decoder consists of multiple RestormerBlocks to reconstruct the image. and The losses of the corresponding SAR images and visible light images are calculated respectively, and the multimodal shared feature encoder, improved dual-branch independent encoder and decoder are trained;

[0082] loss And SAR image loss L sar , visible light image loss L vis and the eigendecomposition loss L decomp association.

[0083] In the present invention, in image reconstruction, the basic features (low frequency) and the depth detail features (high frequency) of the SAR image and the visible light image are combined and connected along the channel dimension to obtain and These features are input into the decoder D for image reconstruction to obtain the original SAR image reconstruction result And the original visible light image reconstruction result

[0084] In the present invention, in the loss calculation, the original SAR image I is input sar , reconstructed SAR image Original visible light image I vis , reconstructed visible light image And the feature decomposition related information, SAR image loss L sar satisfy:

[0085]

[0086] Among them, L aa<Dber is the a-Huber loss function, L SSIc is the loss based on the structural similarity index measurement, L rbd is the relative grayscale difference ratio loss, μ and ρ are adjustment parameters used to balance the weights of different loss items in (infrared image loss).

[0087] Visible light image loss L vis Calculation and L sar Similar, satisfying

[0088]

[0089] Eigendecomposition loss L decomp satisfy,

[0090]

[0091] Among them, CC represents the correlation coefficient calculation, express and The correlation coefficient between is the SAR feature tensor, is the visible light feature tensor.

[0092] The total loss satisfies,

[0093]

[0094] Among them, α 1 and α 2 is the adjustment parameter.

[0095] (5) Based on the trained encoder and decoder, input the paired SAR image and visible light image to extract the current and Then input feature fusion layer;

[0096] Will and Input the base feature fusion layer (Base Fuse Layer), and The input detail fusion layer (Detail Fuse Layer), the basic feature fusion layer (Base Fuse Layer) and the detail fusion layer (Detail Fuse Layer) have the same structure as the visible light low-frequency basic encoder, including a pooling layer, a PoolMLP layer, a DropPath layer and two normalization layers set in sequence for feature fusion;

[0097] Output the basic features after fusion and detail features satisfy

[0098]

[0099] (6) The fused features and Input the decoder, generate a fused image, calculate the loss between the fused image and the original SAR image and visible light image, and train the feature fusion layer;

[0100] Calculate the total loss of the fused image and train; the total loss Associated with a-Huber loss, gradient loss and decomposition loss,

[0101]

[0102] Among them, α 3 and α i is the adjustment parameter;

[0103] Gradient loss L hrad satisfy,

[0104]

[0105] Among them, I m To fuse the images, is the fused image gradient, and is the SAR and visible light image gradient;

[0106] Eigendecomposition loss L decomp The same as stage 1;

[0107] a-Huber loss satisfies,

[0108]

[0109] Where δ is the threshold of a-Huber loss.

[0110] In the present invention, in the final loss function training, the fused image F and the original SAR image I are input sar and the original visible light image I vis And feature decomposition related information, calculate the total loss of the fused image A-Huber loss, gradient loss and decomposition loss are included to ensure that the fused image has high quality in both details and global structures.

[0111] (7) The original SAR image and the visible light image are input into the trained encoder, decoder and feature fusion layer to calculate the fused image F.

[0112] The present invention also relates to a computer-readable storage medium in its application, on which is stored a multimodal fusion program for SAR and visible light image feature enhancement, which implements the above-mentioned multimodal fusion method for SAR and visible light image feature enhancement when executed by a processor.

[0113] The present invention also relates to a computer device in its application, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the multimodal fusion method for SAR and visible light image feature enhancement is implemented.

[0114] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0115] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0116] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction means that implements the functions specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.

[0117] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.

[0118] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.

[0119] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A multimodal fusion method for SAR and visible light image feature enhancement, characterized in that: The method comprises the following steps: S1 inputs the original SAR image and visible light image into the multimodal shared feature encoder to extract the shallow features corresponding to the SAR and visible light images respectively. and S2 and Input the corresponding branch of the improved dual-branch independent encoder to obtain the corresponding low-frequency basic features and S3 and Input the corresponding branch of the improved dual-branch independent encoder to obtain the corresponding high-frequency detail features and S4 and and After combining them separately, input them into the decoder to reconstruct the image and The losses of the corresponding SAR images and visible light images are calculated respectively, and the multimodal shared feature encoder, improved dual-branch independent encoder and decoder are trained; S5 is based on the trained encoder and decoder, inputs the paired SAR image and visible light image, and extracts the current Then input feature fusion layer; S6 will merge the features and Input the decoder, generate a fused image, calculate the loss between the fused image and the original SAR image and the visible light image, and train the feature fusion layer; S7 inputs the original SAR image and the visible light image into the trained encoder, decoder and feature fusion layer, and calculates the fused image F.

2. The multimodal fusion method for SAR and visible light image feature enhancement according to claim 1, characterized in that: In S2, the improved dual-branch independent encoder includes an improved SAR image independent encoder and an improved visible light image independent encoder corresponding to the SAR image and the visible light image respectively; The improved SAR image independent encoder and the improved visible light image independent encoder respectively include a low-frequency basic feature encoder and a high-frequency detail feature encoder that are arranged in coordination.

3. The multimodal fusion method for SAR and visible light image feature enhancement according to claim 2, characterized in that: The low-frequency basic feature encoder of the improved SAR image independent encoder comprises a multi-frequency progressive channel self-attention module, a PoolMLP layer and a DropPath layer arranged in sequence, and the low-frequency basic feature encoder of the improved SAR image independent encoder is input. Add the output of the DropPath layer and output The PoolMLP layer contains two 1x1 convolutional layers and an activation function; a normalization layer is set at the input and output ends of the multi-frequency progressive channel self-attention module.

4. The multimodal fusion method for SAR and visible light image feature enhancement according to claim 3, characterized in that: The multi-frequency progressive channel self-attention module includes a multi-frequency channel attention module and a channel self-attention module arranged in sequence.

5. The multimodal fusion method for SAR and visible light image feature enhancement according to claim 2, characterized in that: The high-frequency detail feature encoder of the improved SAR image independent encoder comprises a plurality of inverse residual modules arranged in sequence; The SAR image tensor is divided from the original channel to form two input tensors with C / 2 channels. and Input the first inverted residual module and multiply or sum the output of each inverted residual module except the first one, and output a tensor with a corresponding channel number of C / 2. Add the output of the first inverted residual module and output the corresponding tensor And input each inverted residual module except the first inverted residual module, and Concatenate on the channels to form the tensor output φ ′ sar .

6. The multimodal fusion method for SAR and visible light image feature enhancement according to claim 2, characterized in that: The low-frequency basic feature encoder of the improved visible light image independent encoder includes a pooling layer, a PoolMLP layer and a DropPath layer arranged in sequence, and the shallow feature of the low-frequency basic feature encoder of the improved visible light image independent encoder is input. Add the output of the DropPath layer and output The PoolMLP layer contains two 1x1 convolutional layers and an activation function; A normalization layer is provided at both the input and output ends of the pooling layer.

7. The multimodal fusion method for SAR and visible light image feature enhancement according to claim 2, characterized in that: The high-frequency detail feature encoder of the improved visible light image independent encoder comprises a plurality of SE modules arranged in sequence; The visible light image tensor is divided from the original channel to form two input tensors with C / 2 channels. and Input the first SE module and multiply or sum the output of each SE module except the first SE module, and output a tensor with the corresponding channel number C / 2 Add the output of the first SE module and output the corresponding tensor And enter each SE module except the first SE module, and Concatenate on the channels to form the tensor output φ ′ vis .

8. The multimodal fusion method for SAR and visible light image feature enhancement according to claim 1, characterized in that: In S4, the loss And SAR image loss L sar , visible light image loss L vis and the eigendecomposition loss L decomp association.

9. The multimodal fusion method for SAR and visible light image feature enhancement according to claim 1, characterized in that: In S5, and Input the basic feature fusion layer and and Input detail fusion layer, output fused basic features φ BF and detail features φ DF .

10. The multimodal fusion method for SAR and visible light image feature enhancement according to claim 1, characterized in that: Calculate the total loss of the fused image and train; the total loss Associated with a-Huber loss, gradient loss, and decomposition loss.

Citation Information

Cited By

  • Multi-modal image defogging method and device, electronic equipment and storage medium

    CN121032837A

  • Visible light and infrared image multi-modal fusion method and system based on transition state

    CN122023977A