Multi-scale large kernel convolution optical and SAR image fusion method and device

By employing a multi-scale, large-kernel convolution method for optical and SAR image fusion, the problems of differences in imaging mechanisms and noise effects in optical and SAR image fusion are solved, achieving high-quality image fusion and enhancing the application effect of remote sensing data.

CN120876246APending Publication Date: 2025-10-31WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510902253.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing optical and SAR image fusion methods, due to differences in imaging mechanisms and the influence of noise, cannot fully leverage their complementary advantages, resulting in poor fused image quality. In particular, they have significant limitations in spectral distortion, coarse texture details, and local blurring, hindering the promotion of high-precision applications.

Method used

A multi-scale large-kernel convolutional optical and SAR image fusion method is adopted. By constructing a multi-scale large-kernel convolutional encoder, fusion layer and decoder, the shared and private feature extraction of optical and SAR images is realized. In combination with spatial attention mechanism to suppress noise, a cross-domain multi-scale shallow feature extraction module and a SAR image detail feature large-kernel extraction module are designed and trained in two stages to generate high-quality fused images.

Benefits of technology

It improves the quality of optical and SAR image fusion, enhances the model's efficiency in capturing long-range semantic dependencies from multi-source remote sensing data, suppresses speckle noise in SAR images, and generates fused images with rich details and high spectral fidelity, making it suitable for applications such as dynamic environmental monitoring and disaster impact assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876246A_ABST
    Figure CN120876246A_ABST
Patent Text Reader

Abstract

The invention provides a multi-scale large kernel convolution optical and SAR image fusion method and device, and the method comprises the steps: constructing a fusion network which comprises a multi-scale large kernel convolution encoder, a fusion layer and a decoder; the multi-scale large-kernel convolution encoder comprises a cross-domain multi-scale shallow feature extraction module and an optical and SAR image private encoder branch; a cross-domain multi-scale shallow feature extraction module extracts shared shallow features, outputs the shared shallow features to access optical and SAR image branches respectively, and extracts optical basic and detail features and SAR basic and detail features; wherein the SAR detail features adopt a multi-scale large-kernel integral solution to realize global context perception, and speckle noise is suppressed in combination with a space attention mechanism; two-stage training is implemented, in the first stage, the optical image and the SAR image are reconstructed, and in the second stage, the fused image is generated. According to the method, the fusion image with rich details is generated by utilizing the complementary characteristics between the optical image and the SAR image, and the multi-dimensional quality improvement of the fusion result in spectral fidelity, spatial detail retention and noise suppression level is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of remote sensing image processing technology, and in particular relates to a multi-scale large kernel convolution optical and SAR image fusion technology solution. Background Technology

[0002] In the field of remote sensing image processing, image fusion technology integrates multi-source observation data from different sensors, effectively leveraging the complementary advantages of various sensors in terms of imaging mechanisms, spatial resolution, and information perception capabilities, while suppressing redundant information, thereby generating information-rich fused images. Currently, Earth observation satellites are widely equipped with multi-source imaging devices such as optical sensors and Synthetic Aperture Radar (SAR) to acquire different types of remote sensing data products. Optical remote sensing images, as a passive imaging product, rely on natural illumination to observe ground reflections, possessing a large instantaneous field of view and capable of rapidly covering large areas. However, limited by the energy receiving mechanism, while maintaining good spectral resolution, its spatial resolution is relatively low. In contrast, SAR imaging systems are active microwave remote sensing devices with significant advantages in all-weather, all-day operation. By actively emitting microwave signals and receiving reflected echoes from ground objects, they can stably acquire information under complex weather conditions. Moreover, radar antennas can receive microwave energy reflected by objects. The intensity of this reflected energy is affected by the surface roughness, humidity, and dielectric properties of an object, reflecting its structural characteristics and helping to distinguish different surface objects, thus generating SAR images rich in spatial information. By fully utilizing the spectral advantages of optical imagery and the all-weather characteristics of SAR imagery, we can gain a more comprehensive understanding of Earth's resources and environment, providing more accurate and reliable data support for various applications.

[0003] However, improving image quality through the fusion of optical and SAR images remains challenging because their imaging mechanisms differ, and SAR images contain speckle noise, resulting in poor quality fused images and making it difficult to fully leverage the complementary advantages between different modalities. Currently, there are two main categories of methods for fusion of optical and SAR images: traditional methods and deep learning methods. Traditional methods mainly include component substitution, multi-scale decomposition model methods, and model-based methods. These methods are highly dependent on image correlation, and the fused images may lead to spectral distortion. Multi-scale decomposition methods typically decompose multispectral and synthetic aperture radar images into high-frequency and low-frequency components and design corresponding fusion rules; however, designing different fusion rules for images has always been a research challenge. Model-based fusion methods are divided into variational and sparse representation-based methods. These methods are severely affected by SAR random noise, requiring prior knowledge support, complex model construction, and high operational costs. Deep learning in the field of optical SAR image fusion mainly includes mainstream technical paradigms such as convolutional neural networks (CNNs) and generative adversarial networks (GANs). CNN-based methods can accurately extract features from SAR and optical imagery, but their small receptive field makes it difficult to fully capture the complementarity of optical and SAR images when fusing them. While GAN-based methods improve the interpretability of SAR imagery, they still have shortcomings in handling special features and textures, and the generated scenes vary greatly, with the realism of the images needing improvement.

[0004] In summary, both traditional and deep learning-based methods have improved the quality of optical remote sensing images to some extent using SAR imagery. However, due to the inherent differences between optical and SAR images in imaging geometry, ground object radiometric characteristics, and texture representation, current image fusion results still have significant limitations, such as spectral distortion, coarse texture details, and local blurring. These shortcomings pose challenges to improving the quality of fused images and hinder the further promotion of the technology in high-precision application scenarios. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing technologies by using deeper multimodal sharing and private feature extraction methods, more efficient encoding and decoding optimization methods, and unique processing mechanisms for the differences between optical and SAR. This is expected to further improve the quality and practicality of optical and SAR fused images, providing a high-quality data foundation for diversified applications such as environmental dynamic monitoring and disaster impact assessment.

[0006] The technical solution of this invention is a multi-scale large kernel convolution optical and SAR image fusion method, which constructs a fusion network including a multi-scale large kernel convolution encoder, a fusion layer and a decoder; The multi-scale large-kernel convolutional encoder includes a cross-domain multi-scale shallow feature extraction module, an optical image proprietary encoder branch, and a SAR image proprietary encoder branch. The cross-domain multi-scale shallow feature extraction module extracts shared shallow features between optical and SAR images, and the outputs are respectively connected to the optical image proprietary encoder branch and the SAR image proprietary encoder branch. The optical image proprietary encoder branch extracts optical basic features and optical detail features, and the SAR image proprietary encoder branch extracts SAR basic features and SAR detail features. The SAR detail features are obtained by using multi-scale large-kernel convolution decomposition to achieve global context awareness, and combined with a spatial attention mechanism to suppress speckle noise. The fusion layer is used to fuse the basic features of optical and SAR images, and to fuse the detailed features of optical images and SAR images respectively. The decoder is used to decode the fusion features to generate a fused image; The training process involves two stages. The first stage reconstructs optical and SAR images based on a multi-scale large kernel convolutional encoder and decoder. The second stage generates fused images sequentially through a multi-scale large kernel convolutional encoder, a fusion layer, and a decoder.

[0007] Moreover, the optical image proprietary encoder branch extracts basic optical features through a lightweight converter and extracts detailed optical features through a reversible neural network. The SAR image private encoder branch extracts basic SAR features through a lightweight transformer and sets up a SAR image detail feature large kernel extraction module to extract SAR detail features.

[0008] Moreover, the SAR image detail feature large kernel extraction module uses deep convolution kernels of different scales to capture multi-level contextual information, splices multi-scale features and extracts spatial global relationships through global pooling, and uses spatial attention mechanism to generate feature selection masks and weighted fusion output.

[0009] Furthermore, in the fusion layer, a lightweight converter is used to fuse optical and SAR basic features, a reversible neural network is used to fuse optical detail features, and a SAR image detail feature large kernel extraction module is used to fuse SAR detail features.

[0010] Moreover, the first-stage loss includes SAR image reconstruction loss, optical image reconstruction loss, feature decomposition loss, and spectral angle loss; the second-stage loss includes fused image intensity similarity loss, gradient consistency loss, and feature decomposition loss.

[0011] Furthermore, the cross-domain multi-scale shallow feature extraction module includes a channel feature extraction attention module and a spatial asymmetric convolution module. The channel feature extraction attention module generates a channel attention map through pointwise convolution and multi-scale pooling operations, and dynamically calibrates the channel weights of cross-modal features. The spatial asymmetric convolution module captures the long-range spatial dependencies of multimodal images through cross-scale asymmetric depth convolution.

[0012] Moreover, the decoder adopts the same structure as the cross-domain multi-scale shallow feature extraction module.

[0013] On the other hand, the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the optical and SAR image fusion method of multi-scale large kernel convolution as described above.

[0014] On the other hand, the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the optical and SAR image fusion method of multi-scale large kernel convolution as described above.

[0015] On the other hand, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the optical and SAR image fusion method of multi-scale large kernel convolution as described above.

[0016] Compared with the prior art, the present invention has the following advantages: 1) This invention proposes a multi-scale large kernel convolution optical and SAR image fusion system and method, which is a two-branch cross-modal feature fusion method composed of a multi-scale large kernel convolution encoder, a basic and detail feature fusion layer and a decoder. It promotes the sharing and deep extraction of private features of optical and SAR images, and gradually generates fused images with rich cross-modal complementary features.

[0017] 2) The cross-domain multi-scale shallow feature extraction module designed in this invention achieves pixel-by-pixel comprehensive joint extraction of shared shallow features of optical and SAR images through channel feature extraction and spatial asymmetric convolution, thereby enhancing the model's efficiency in capturing long-range semantic dependencies in multi-source remote sensing data.

[0018] 3) The SAR image detail feature large kernel extraction module designed in this invention simulates global context awareness by decomposing large kernel convolution, which enhances the model's ability to pay attention to relevant context regions in SAR image detail features, and introduces a spatial attention mechanism to suppress SAR image speckle noise and filter out noise redundancy that is unrelated to context information. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0020] Figure 1 This is an overall framework diagram of a multi-scale large-kernel convolution optical and SAR image fusion system provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the cross-domain multi-scale shallow feature extraction module provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the SAR image detail feature large kernel extraction module provided in an embodiment of the present invention; Figure 4 This is a sample image of the image fusion result provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0022] This invention provides a multi-scale large-kernel convolutional method for fusion of optical and SAR images. It outputs a high-quality fused image through a multi-scale large-kernel convolutional encoder, a basic and detail feature fusion layer, and a decoder. In the fusion network constructed in this invention, a cross-domain multi-scale shallow feature extraction module is set up. This module utilizes channel feature extraction and spatial asymmetric convolution to achieve pixel-by-pixel comprehensive joint extraction of shared shallow features from optical and SAR images, enhancing the model's efficiency in capturing long-range semantic dependencies in multi-source remote sensing data. Furthermore, a large-kernel SAR image detail feature extraction module is constructed. This module simulates global context awareness by decomposing large-kernel convolution and introduces a spatial attention mechanism to suppress speckle noise in SAR images. This invention utilizes complementary features between optical and SAR images to generate a detail-rich fused image, promoting multi-dimensional quality improvement in spectral fidelity, spatial detail preservation, and noise suppression levels.

[0023] like Figure 1 As shown in the embodiment, an optical and SAR image fusion method with multi-scale large kernel convolution is provided, wherein the proposed fusion network structure specifically includes: I. Encoder Section This invention constructs a multi-scale large-kernel convolutional encoder, including a cross-domain multi-scale shallow feature extraction module, an optical image proprietary encoder branch, and a SAR image proprietary encoder branch. The output of the cross-domain multi-scale shallow feature extraction module is respectively connected to the optical image proprietary encoder branch and the SAR image proprietary encoder branch.

[0024] The optical image private encoder branch includes an optical basic feature extraction module (preferably LiteTransformer) and an optical image detail feature extraction module (preferably INN). The SAR image private encoder branch includes a SAR image basic feature extraction module (preferably LiteTransformer) and a SAR detail feature big kernel extraction module (proposed in this invention, abbreviated as LKDG).

[0025] The encoder portion of this embodiment is implemented as follows: 1) A global joint extraction and noise suppression of shared features between optical and SAR images is achieved by using a cross-domain multi-scale shallow feature extraction module.

[0026] See Figure 2 The cross-domain, multi-scale shallow feature extraction module designed in this embodiment of the invention comprises a channel feature extraction attention module and a spatial asymmetric convolution module arranged sequentially. The channel feature extraction attention module generates a channel attention map through pointwise convolution and multi-scale pooling operations, dynamically calibrating the channel weights of cross-modal features. The spatial asymmetric convolution module captures the long-range spatial dependencies of multimodal images through cross-scale asymmetric depth convolution.

[0027] The preferred implementation method provided in the embodiments specifically includes the following steps: 1.1) Input features Perform pointwise convolution and The operation initially extracts cross-modal features and obtains tensors. The formula is expressed as:

[0028] In the formula, This represents a 1×1 pointwise convolution. The activation function operation is represented by R, the real number field is represented by B, the batch size is represented by C, the number of channels is represented by H and W, which represent the height and width of the feature map, respectively.

[0029] 1.2) The tensor obtained in step 1.1 Adaptive average pooling (AAP), adaptive max pooling (AAP), and global median pooling (GMP) are used to generate multi-scale channel descriptions, resulting in channel descriptors. , , The formula is expressed as:

[0030] In the formula, It is median pooling. It is max pooling. It is average pooling. The first characteristic tensor in the characteristic tensor The first sample Two-dimensional spatial characteristics of each channel Indicates the first The index of each sample in the batch Indicates the number of channels.

[0031] 1.3) Generate a channel weight map after compressing and activating the channel descriptors: Compress the feature channel number obtained in step 1.2 to the original dimension through the first pointwise convolution. ( (These are the dimensionality reduction coefficients), which are then batch normalized to eliminate inter-modal distribution bias, and then... Activation filters common features; subsequently, a second pointwise convolution restores the feature dimension to its original scale, and a batch normalization layer is added to stabilize the feature distribution. Specifically, a... The function adjusts the three-channel pooling to construct an attention distribution adapted to the characteristics of multi-source data. Finally, element-wise summation of the attention maps generates a robust channel attention map. The formula is expressed as:

[0032] in, This represents the input optical and SAR cross-modal image features. express Activation function Multilayer perceptron; 1.4) The result obtained in step 1.3 By performing channel-by-channel multiplication, the learned attention weight matrix is ​​applied to the initial feature map, achieving dynamic calibration of features from different channels and obtaining channel-weighted features. The specific formula is as follows:

[0033] in, This indicates element-wise multiplication.

[0034] 1.5) The result obtained in step 1.4 Perform 5×5 depthwise convolution Extracting basic features The formula is expressed as:

[0035] in, express Depth-separable convolutions.

[0036] 1.6) The result obtained in step 1.5 Parallel use of multi-scale asymmetric depthwise convolution kernels ( Multi-level features are extracted using the formulas (e.g., 1×11, 11×1, 1×7, 7×1, 1×21, 21×1), and then fused through pointwise convolution to obtain the final result. The formula is expressed as:

[0037] in, Indicates kernel is Depth-separable convolution, .

[0038] 1.7) The result obtained in step 1.6 Construct a spatial weight matrix and combine it with the channel weighted features obtained in step 1.4. Multiply to output the final feature. The formula is expressed as:

[0039] in, express Pointwise convolution, This indicates element-wise multiplication.

[0040] 2) Considering that optical and SAR images each have their own unique modal characteristics, we then perform precise extraction of their basic and detailed features respectively.

[0041] The preferred implementation method provided in the embodiments specifically includes the following parts: 2.1) The result obtained in step 1.7 The optical and SAR image basic feature extraction module, which is constructed using a lightweight transformer (preferably a Lite Transformer module), does not share parameters between the optical and SAR images, ensuring that the basic features unique to each mode are fully extracted.

[0042] For a detailed implementation of the Lite Transformer module, please refer to Zhao, Z., Bai, H., Zhang, J., Zhang, Y., Xu, S., Lin, Z., ... & Van Gool, L. (2023). Cddfuse: Correlation-driven dual-branch feature decomposition for multi-modality image fusion. In Proceedings of the IEEE / CVF conference on computer vision and pattern recognition (pp. 5906-5916). For ease of implementation reference, the following detailed implementation instructions are provided: By capturing long-range dependency features through spatial attention and compressing the embedding dimension to reduce the number of parameters, its flattened feedforward network structure is formulated as follows:

[0043] in, For layer normalization, For multi-head space self-attention, For depthwise separable convolution, The Gaussian Error Linear Unit (Gaussian Error Linear Unit) is a smooth, nonlinear activation function. For input channel dimension To compressed dimensions The linear transformation weight matrix; To compress dimensions To the output channel dimension The linear transformation weight matrix, These are the basic features extracted.

[0044] 2.2) The result obtained in step 1.7 The optical image detail feature extraction module, constructed using a reversible neural network (INN), extracts the detailed features of the optical image and employs an affine coupling layer to achieve lossless feature transmission.

[0045] For a detailed implementation of the reversible neural network module, please refer to Dinh, L., Sohl-Dickstein, J., & Bengio, S. (2017, February). Density estimation using Real NVP. International Conference on Learning Representations . In each reversible layer, It can be set to any mapping without affecting the lossless information transmission in the reversible layer. To calculate the trade-off between cost and feature extraction capability, a bottleneck residual block (BRB) is used in MobileNetV2 as... The formula is:

[0046] in, It represents the Hadamah accumulation. This refers to the number of channels used for affine transformation when splitting the channels. This represents the total number of channels in the feature tensor along the channel dimension. Represents the first to the second input features of the k-th layer. One channel, This is a channel connection operation. It is any mapping function.

[0047] 2.3) The result obtained in step 1.7 The SAR detail feature big kernel extraction (LKDG) module is used to extract SAR image detail features.

[0048] This invention proposes to simulate global context awareness by using large kernel convolution decomposition and introduce a spatial attention mechanism to suppress speckle noise in SAR images. The SAR detail feature large kernel extraction module uses deep convolution kernels of different scales to capture multi-level context information, splices multi-scale features and extracts spatial global relationships through global pooling; and uses the spatial attention mechanism to generate feature selection masks and weighted fusion output.

[0049] like Figure 3 As shown, the implementation process of the SAR image detail feature large kernel extraction module further designed in this embodiment of the invention specifically includes the following steps: 2.3.1) The result obtained in step 1.7 The inputs are fed into 7×7 and 11×11 deep convolutional networks with large kernel convolution decomposition units, respectively. Each deep convolutional layer employs a different convolution rate and receptive field expansion strategy to ensure that contextual information at different scales can be captured. This is represented as:

[0050] in, It represents the kernel size. Depth convolution, It represents the kernel size. Depth convolution.

[0051] 2.3.2) The result obtained in step 2.3.1 Through a Pointwise convolution Further processing, maintaining consistency in channel dimensions, is represented as follows:

[0052] in, represent Pointwise convolution.

[0053] 2.3.3) The result obtained in step 2.3.2 From and Characteristics of these two different receptive field ranges To facilitate subsequent global operations, the formula is as follows:

[0054] in, It is a feature splicing operation.

[0055] 2.3.4) The features obtained from concatenating the channel dimensions in step 2.3.3. Perform global average pooling and global max pooling To efficiently extract global relationships in SAR shadow space. The formula is:

[0056] in, and They are and Symbols describing the spatial features of subsequent SAR images.

[0057] 2.3.5) By allowing the result obtained in step 2.3.4 and Information exchange between them, connecting different spatial pooling features, and using convolutional layers Pooling features are converted into spatial attention maps. Then, for each SAR image's spatial attention map, a sigmoid activation function is used to obtain a spatial selection mask for each image patch with decomposed depthwise convolutional kernels. Specifically:

[0058] in, This represents the Sigmoid function. The representative size is convolution, This indicates a splicing operation at the channel dimension.

[0059] 2.3.6) Select the mask according to its corresponding space. The features in the decomposed depthwise convolution kernel sequence are weighted and then passed through a 1×1 pointwise convolution. Fusion to obtain attention features Specifically:

[0060] in, represent Pointwise convolution, This indicates element-wise multiplication.

[0061] 2.3.7) The attention features obtained in step 2.3.6 and the result of step 1.7 Perform element-wise multiplication and output SAR image detail features Specifically:

[0062] II. Basic and Detailed Feature Fusion Layer This invention proposes to fuse basic and detailed features separately through a fusion layer (used only in training phase 2).

[0063] This invention considers that the inductive bias for fusion of basic and detail features should be similar to that for extraction of basic and detail features in the encoder. Therefore, it employs a Lite Transformer module for basic feature fusion and an INN module for optical image detail feature fusion. The module performs detailed feature fusion of SAR images.

[0064]

[0065] in, These are the features extracted by the encoder. These are the fundamental features of optical images. As a basic feature of SAR images, For optical image detail features. This refers to the detailed features of SAR images.

[0066] III. Decoder Section The decoder is used to decode the fused features to generate a fused image. The decoder structure is consistent with the design of the cross-domain multi-scale shallow feature extraction module, that is, it is used as the basic unit of the decoder. For specific implementation, please refer to part 1) of the aforementioned encoder.

[0067] Based on the above-mentioned fusion network structure, the network is trained, and then the trained network is used to extract the fusion results.

[0068] To improve the quality of the fused image, this invention proposes a two-stage training process: training stage 1 reconstructs the input optical and SAR images, and training stage 2 generates the fused image.

[0069] In the fusion network, after passing through a multi-scale large-kernel convolutional encoder, the decoder concatenates the decomposed features along the channel dimension before decoding. In training phase one, no fusion layer is needed; the multi-scale large-kernel convolutional encoder and decoder directly output the reconstructed optical and SAR feature maps. In training phase two, the multi-scale large-kernel convolutional encoder adds the basic features of the optical and SAR images, passes through a fusion layer, and finally outputs the fused image through the decoder. The specific formula is as follows:

[0070]

[0071] in, These are the fundamental features of optical images. As a basic feature of SAR images, For optical image detail features. For SAR image detail features. and These represent the SAR and optical images reconstructed by the decoder, respectively. For the decoder module, the input consists of basic and detailed features stitched together in the channel dimension. For the final high-quality image output.

[0072] This invention further proposes to constrain the input optical and SAR images through a two-stage loss function. The first training stage reconstructs the input optical and SAR images using a composite loss function that includes spectral fidelity loss and feature decomposition loss. The second training stage generates the fused image using a composite loss function that includes gradient consistency loss and intensity loss.

[0073] Specifically, training phase 1 ensures the quality of the reconstructed image while focusing on suppressing spectral distortion of the optical modes. This phase employs a composite loss function design, including SAR image reconstruction loss, optical image reconstruction loss, feature decomposition loss, and spectral angle loss. The specific formulas are as follows:

[0074] in, This is the loss function for the first stage. and It is the loss in generating reconstructed images from SAR and optical images. It is spectral angle loss. It is the basic and detailed feature decomposition loss, which mainly ensures that no information in the image is lost during encoding and decoding. The specific formula is as follows:

[0075] in, It is a structural similarity index. as well as These are tuning parameters. It is the correlation coefficient operator. is the base of the natural logarithm. It is the loss term balance coefficient. It is the loss of intensity consistency between the original SAR image and the reconstructed SAR image. It is the loss of intensity consistency between the original optical image and the reconstructed optical image. It is the correlation between basic features and optical detail features. It is the correlation between basic features and the detailed features of SAR itself. These are the basic features extracted from SAR images. and These represent the detailed features extracted from SAR and optical images, respectively. The preferred embodiment uses... Set it to 1.01 to ensure that this item is always positive.

[0076] The definition is as follows:

[0077] in, Represents the optical image vector. This represents the reconstructed optical image vector. It is an inverse cosine function, which converts the dot product into the angle between the spectra. The L2 norm of a vector is also called the Euclidean norm.

[0078] In training phase 2, detailed fused images are generated. The loss functions are defined as follows: image intensity similarity loss, gradient consistency loss, and feature decomposition loss.

[0079] in, and These represent the height and width of the image (or feature map), respectively. This represents the summation after calculating the absolute value of each element. Norm, This indicates taking the absolute value of each element. The fused image representing the network output. This represents the Sobel gradient operator. This represents the loss due to the reconstruction constraint. and These are tuning parameters. This is the loss function for the second stage. This represents the intensity similarity loss between the fused image and the source image. This represents the gradient preservation loss. It takes the maximum value at corresponding positions of two input images (or their gradients) pixel by pixel. and These represent the SAR and optical images after the first stage of reconstruction, respectively.

[0080] To demonstrate the technical effects of this invention, the experiments in this embodiment were conducted in an NVIDIA RTX 3090 hardware environment and a Python software environment. During the preprocessing stage, training samples were randomly cropped into 128×128 pixel patches. The number of training epochs was set to 120, with 40 for the first stage and 80 for the second stage. The Adam optimizer was used for training, with an initial learning rate of 10⁻⁴, decreasing by 0.5 every 20 epochs.

[0081] The dataset used in this embodiment is the hunandata dataset released by Wuhan University. This dataset is a multimodal remote sensing image dataset for Hunan Province, a region with complex and diverse terrain, including water bodies, hills, mountains, and continuous valleys. It contains 500 image pairs. The optical images are from the Sentinel-2 dataset, preprocessed to remove cloud and shadow interference, while the SAR images are from the Sentinel-1 dataset. All images were resampled to a 10m resolution, and a 256×256 pixel block was selected for the experiment.

[0082] This invention was compared with four existing methods published in authoritative journals to improve image fusion effects: [Chun-Lin, L. (2010). A tutorial of the wavelet transform. , 21(22), 2.] (Comparison Method 1), [J.-Y. Zhu, T. Park, P. Isola, and AA Efros, "Unpaired image-to-image translation using cycle-consistent adversarial networks." pp. 2223-2232.] (Comparison Method 2), [H. Xu, J. Ma, J. Jiang, X. Guo, HJIT o. PA Ling, and M. Intelligence, “U2Fusion: A unified unsupervised image fusion network,”vol. 44, no. 1, pp. 502-518. 2020.] (Comparison Method 3), [Zhao, Z., Bai, H., Zhang, J., Zhang, Y., Xu, S., Lin, Z., ...&Van Gool, L. (2023). Cddfuse: Correlation-driven dual-branch feature decomposition for multi-modality image fusion. InProceedings of the IEEE / CVF conference on computer vision and patternrecognition (pp. 5906-5916).] (Compare Method 4).

[0083] The corresponding fusion results are as follows Figure 4 As shown. Figure 4 From left to right, each column represents: Sentinel-2 optical image, Sentinel-1 SAR image, results of comparison method 1, results of comparison method 2, results of comparison method 3, results of comparison method 4, and results of the method of this invention. The comparison results show that, compared to other methods, the multi-scale large-kernel convolution optical and SAR image fusion system designed in this invention achieves the fusion of complementary features between optical and SAR images, highlighting the visual perception of the main contour information of the SAR image, and eliminating the interference of inherent speckle noise in the SAR image, ensuring the generation of a fused image rich in detail.

[0084] The corresponding fusion evaluation results are shown in Table 1. The comparison results show that, compared to other methods, the multi-scale large-kernel convolution optical and SAR image fusion system designed in this invention performs best in terms of MI, VIF, and SSIM indices, with VIF, MI, and SSIM being superior to the second-best results of 0.3, 0.1, and 0.26, respectively. This observation clearly demonstrates that the fused image contains rich information, extracts a large amount of information from the source image, and exhibits a high degree of structural similarity.

[0085] Table 1

[0086] Figure 5 A schematic diagram of the physical structure of an electronic device is provided. This device may include a processor, a communications interface, memory, and a communication bus, wherein the processor, communications interface, and memory communicate with each other via the communication bus. The processor can invoke logical instructions from the memory to execute a method for a multi-scale large-kernel convolution optical and SAR image fusion system.

[0087] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0088] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program that can be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer is able to execute the methods provided above for optical and SAR image fusion systems for multi-scale large kernel convolution.

[0089] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the methods provided above for optical and SAR image fusion systems with multi-scale large kernel convolution.

[0090] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0091] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0092] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-scale large-kernel convolution method for optical and SAR image fusion, characterized in that, Construct a fusion network, including a multi-scale large-kernel convolutional encoder, fusion layers, and a decoder; The multi-scale large-kernel convolutional encoder includes a cross-domain multi-scale shallow feature extraction module, an optical image proprietary encoder branch, and a SAR image proprietary encoder branch. The cross-domain multi-scale shallow feature extraction module extracts shared shallow features between optical and SAR images, and the outputs are respectively connected to the optical image proprietary encoder branch and the SAR image proprietary encoder branch. The optical image proprietary encoder branch extracts optical basic features and optical detail features, and the SAR image proprietary encoder branch extracts SAR basic features and SAR detail features. The SAR detail features are obtained by using multi-scale large-kernel convolution decomposition to achieve global context awareness, and combined with a spatial attention mechanism to suppress speckle noise. The fusion layer is used to fuse the basic features of optical and SAR images, and to fuse the detailed features of optical images and SAR images respectively. The decoder is used to decode the fusion features to generate a fused image; The training process involves two stages. The first stage reconstructs optical and SAR images based on a multi-scale large kernel convolutional encoder and decoder. The second stage generates fused images sequentially through a multi-scale large kernel convolutional encoder, a fusion layer, and a decoder.

2. The multi-scale large-kernel convolution optical and SAR image fusion method according to claim 1, characterized in that: The proprietary optical image encoder branch extracts basic optical features through a lightweight converter and extracts detailed optical features through a reversible neural network. The SAR image private encoder branch extracts basic SAR features through a lightweight transformer and sets up a SAR image detail feature large kernel extraction module to extract SAR detail features.

3. The optical and SAR image fusion method based on multi-scale large kernel convolution according to claim 2, characterized in that: The SAR image detail feature large kernel extraction module uses deep convolution kernels of different scales to capture multi-level contextual information, splices multi-scale features and extracts spatial global relationships through global pooling, and uses spatial attention mechanism to generate feature selection masks and weighted fusion output.

4. The optical and SAR image fusion method based on multi-scale large kernel convolution according to claim 2, characterized in that: In the fusion layer, a lightweight converter is used to fuse optical and SAR basic features, a reversible neural network is used to fuse optical detail features, and a SAR image detail feature large kernel extraction module is used to fuse SAR detail features.

5. The optical and SAR image fusion method based on multi-scale large kernel convolution according to claim 1, characterized in that: The first-stage loss includes SAR image reconstruction loss, optical image reconstruction loss, feature decomposition loss, and spectral angle loss; the second-stage loss includes fused image intensity similarity loss, gradient consistency loss, and feature decomposition loss.

6. The optical and SAR image fusion method based on multi-scale large kernel convolution according to claim 1, characterized in that: The cross-domain multi-scale shallow feature extraction module includes a channel feature extraction attention module and a spatial asymmetric convolution module. The channel feature extraction attention module generates a channel attention map through pointwise convolution and multi-scale pooling operations, and dynamically calibrates the channel weights of cross-modal features. The spatial asymmetric convolution module captures the long-range spatial dependencies of multimodal images through cross-scale asymmetric depth convolution.

7. The optical and SAR image fusion method based on multi-scale large kernel convolution according to claim 1, characterized in that: The decoder adopts the same structure as the cross-domain multi-scale shallow feature extraction module.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, it implements the optical and SAR image fusion method of multi-scale large kernel convolution as described in any one of claims 1 to 7.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the optical and SAR image fusion method with multi-scale large kernel convolution as described in any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that: When the computer program is executed by the processor, it implements the optical and SAR image fusion method with multi-scale large kernel convolution as described in any one of claims 1 to 7.