Multispectral and SAR data fusion method, medium and system

Through deep learning fusion network and attention mechanism, the problem of distortion of spectral information in the fusion of low spatial resolution of multispectral images and SAR data is solved, and efficient fusion of multi-source remote sensing data is achieved, improving spatial resolution and maintaining spectral integrity.

CN120355592AActive Publication Date: 2025-07-22BEIJING GEOWAY INFORMATION TECH

Patent Information

Application Number
CN202510837573.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-07-22
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

The prior art is difficult to maintain the spectral information integrity of SAR data while improving the spatial resolution of multispectral images, especially in complex scenarios, and it is difficult to achieve balance between spatial information enhancement and spectral fidelity.

Method used

Deep learning method is adopted to build a deep fusion network, use tile training samples and attention mechanisms to adaptively learn the complementary characteristics of multi-spectral and SAR data, and combine the supervised learning strategy of degradation-reconstruction to achieve spatial resolution improvement and spectral information preservation.

Benefits of technology

It effectively solves the problem of spectral information distortion during the fusion of low spatial resolution of multi-spectral images and SAR data, realizes the complementary fusion of multi-source remote sensing data, and provides a high-quality data foundation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355592A_ABST
    Figure CN120355592A_ABST
Patent Text Reader

Abstract

The invention provides a multispectral and SAR (Synthetic Aperture Radar) data fusion method, medium and system, and belongs to the technical field of electrical digital data processing. A training set of a low-resolution SAR image and a degraded multispectral image is constructed, and an original multispectral image is used as a supervision label; and designing a deep network architecture comprising a feature extraction module, a feature fusion module and an image reconstruction module. And optimizing network parameters by adopting a tile type sample training strategy, inputting the multi-spectral image and the high-resolution SAR image after spatial registration into a trained network in a prediction stage, and outputting a fusion image with improved spatial resolution and complete spectral information. According to the method, adaptive fusion of heterogeneous data features is realized by using an attention mechanism, the spectral integrity is maintained through a quality degradation-reconstruction supervised learning strategy, and the technical problems of low spatial resolution of a multispectral image and spectral information distortion in an SAR data fusion process are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of electronic digital data processing, and in particular, relates to a method, a medium and a system for fusing multi-spectral and SAR data. Background Art

[0002] Remote sensing data fusion is a key technology in the field of earth observation, especially the fusion of multispectral images and synthetic aperture radar (SAR) data. Traditional fusion methods mainly include component replacement-based methods (such as IHS, PCA transformation), multi-resolution analysis-based methods (such as wavelet transform, pyramid decomposition) and statistical learning-based methods (such as sparse representation, dictionary learning). These methods are widely used in different application scenarios, such as land cover classification, change detection and environmental monitoring.

[0003] However, traditional fusion methods have obvious limitations. Although the method based on component replacement can improve the spatial resolution, it often leads to serious spectral distortion; the method based on multi-resolution analysis is difficult to effectively utilize the texture and structure information of SAR images when processing multi-source data; the method based on statistical learning is computationally complex, requires a lot of prior knowledge and is difficult to adjust parameters. In addition, when there is a large resolution difference between multispectral and SAR images, traditional methods cannot effectively maintain spectral integrity.

[0004] Therefore, how to improve the spatial resolution of multispectral images while making full use of the high-resolution structural information of SAR images and maintaining the spectral integrity of multispectral data has become a core technical problem that needs to be solved in this field. Especially in complex scenes with rich ground details, it is difficult for existing technologies to achieve a balance between spatial information enhancement and spectral fidelity. In other words, there are technical problems in existing technologies such as low spatial resolution of multispectral images and distortion of spectral information during SAR data fusion. Summary of the invention

[0005] In view of this, the present invention provides a method, medium and system for fusing multispectral and SAR data, which can solve the technical problems of low spatial resolution of multispectral images and distortion of spectral information in the SAR data fusion process in the prior art.

[0006] The present invention is implemented as follows: A method for fusing multi-spectral and SAR data provided by a first aspect of the present invention includes: When constructing a deep learning training set, downsampling the SAR image to the same spatial resolution as the multi-spectral image to obtain a low-resolution SAR image; performing a downsampling operation on the multi-spectral image first and then an upsampling operation to obtain a degraded multi-spectral image; using the original multi-spectral image as a model training label; cutting the low-resolution SAR image and the degraded multi-spectral image into fixed-size tile training samples; designing a deep fusion network including a feature extraction module, a feature fusion module, and an image reconstruction module; training the deep fusion network using the tile training samples; when predicting, processing the multi-spectral image to be fused and the high-resolution SAR image to have the same spatial range and inputting them into the trained deep fusion network; and outputting a fused multi-spectral image through the deep fusion network.

[0007] Among them, the downsampling operation specifically reduces the spatial resolution of the high-resolution image through an interpolation algorithm; the upsampling operation specifically increases the spatial resolution of the low-resolution image through an interpolation algorithm.

[0008] Among them, the tile training sample specifically cuts a large-size remote sensing image after annotation into fixed-size pixel blocks to form a training sample, and the fixed size is 256×256 pixels.

[0009] Among them, the fused multi-spectral image is specifically a multi-spectral image with improved spatial resolution and intact spectral information.

[0010] Among them, the feature extraction module specifically extracts multi-scale feature representations from the multi-spectral image and the SAR image using a convolutional neural network; the feature fusion module specifically effectively combines multi-source features through an attention mechanism or other fusion algorithms to retain their respective advantageous information; the image reconstruction module specifically restores the fused features to a high-resolution multi-spectral image output through deconvolution or upsampling convolution.

[0011] Among them, in the step of processing the multi-spectral image to be fused and the high-resolution SAR image to have the same spatial range, specifically, the multi-spectral image and the SAR image are adjusted to the same geographic coordinate system and spatial range through geometric transformation.

[0012] Among them, in the step of designing a deep fusion network including a feature extraction module, a feature fusion module, and an image reconstruction module, it also includes an optimization step for the fusion quality assessment equation set, specifically, dynamically optimizing the feature fusion module during the network training stage by constructing a multi-objective fusion quality assessment equation set; the fusion quality assessment equation set includes a structure consistency equation, a spectral fidelity equation, a texture enhancement equation, a global correlation equation, and a fusion determination equation.

[0013] Among them, in the step of processing the multi-spectral image and the high-resolution SAR image to be fused into the same spatial range, a multi-level dynamic cropping optimization step is further included. Specifically, the multi-spectral image to be fused and the high-resolution SAR image are subjected to multi-level dynamic cropping processing, and the cropping tile size is dynamically adjusted according to the complexity of the image content.

[0014] The second aspect of the present invention provides a computer-readable storage medium, in which program instructions are stored. When the program instructions run on a computer, they are used to execute the above method for fusing multi-spectral and SAR data.

[0015] The third aspect of the present invention provides a system for fusing multi-spectral and SAR data, including the above computer-readable storage medium. The system can be any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is set inside the system, and a microprocessor for executing the program instructions stored in the computer-readable storage medium is set inside the system.

[0016] The present invention realizes the efficient fusion of multi-source remote sensing data by designing a specific deep learning model architecture and training strategy. In the training stage, through image degradation simulation and supervised learning of the original multi-spectral image, a mapping relationship from low-quality input to high-quality output is established, effectively solving the problems of information expression and integration in multi-source data fusion. The method of the present invention has significant advantages compared with traditional technologies: First, through the deep feature extraction and attention mechanism fusion module, it can adaptively learn the complementary features of multi-spectral and SAR data, avoiding the limitations of artificial feature design in traditional methods; Second, by adopting the supervised learning strategy of degradation-reconstruction, the network can effectively learn the mapping from the degraded image to the original high-quality image, so that in the prediction stage, the high-resolution multi-spectral image can be assisted by SAR data for reconstruction; Finally, the entire fusion process realizes the improvement of spatial resolution on the premise of maintaining the integrity of spectral information. Therefore, the present invention effectively solves the technical problems of low spatial resolution of multi-spectral images and spectral information distortion in the process of SAR data fusion, realizes the efficient fusion of the advantages of multi-source remote sensing data, and provides a high-quality data basis for various remote sensing applications. Description of the Drawings

[0017] Figure 1 It is a flowchart of the method of the present invention.

[0018] Figure 2 It is a graph of the change of the loss function in the training process of the deep fusion network in Example 2. Detailed Embodiments

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0020] As Figure 1 shown, it is a flowchart of a method for multi-spectral and SAR data fusion provided by the first aspect of the present invention. The method includes the following steps: S01. When constructing a deep learning training set, downsample the SAR image to the same spatial resolution as the multi-spectral image to obtain a low-resolution SAR image; S02. Perform downsampling on the multi-spectral image first and then perform upsampling to reduce the quality of the multi-spectral image and obtain a degraded multi-spectral image; S03. Use the original multi-spectral image as a model training label to supervise the network to learn high-quality multi-spectral image reconstruction; S04. Crop the low-resolution SAR image and the degraded multi-spectral image into fixed-size tile training samples according to the same geographical range; S05. Design a deep fusion network architecture, including a feature extraction module, a feature fusion module, and an image reconstruction module; S06. Use the tile training samples to train the deep fusion network and optimize the network parameters until the verification accuracy reaches a preset threshold; S07. In the prediction stage, process the multi-spectral image to be fused and the high-resolution SAR image into the same spatial range; S08. Input the multi-spectral image to be fused and the high-resolution SAR image within the same spatial range into the trained deep fusion network at the same time; S09. Obtain a fused multi-spectral image with improved spatial resolution and intact spectral information through the output of the deep fusion network.

[0021] Among them, the downsampling operation specifically refers to reducing the spatial resolution of a high-resolution image through an interpolation algorithm. Common algorithms include nearest neighbor interpolation, bilinear interpolation, or bicubic interpolation.

[0022] Among them, the upsampling operation specifically refers to increasing the spatial resolution of a low-resolution image through an interpolation algorithm, but without increasing the actual amount of information, resulting in a reduction in image quality.

[0023] Among them, the tile training sample specifically refers to cropping a large-size remote sensing image into fixed-size blocks such as 256×256 pixels, which is convenient for batch processing by a deep learning network.

[0024] Among them, the feature extraction module specifically refers to using a convolutional neural network to extract multi-scale feature representations from multi-spectral images and SAR images.

[0025] Among them, the feature fusion module specifically refers to effectively combining multi-source features through an attention mechanism or other fusion algorithms, and retaining their respective advantageous information.

[0026] Among them, the image reconstruction module specifically refers to restoring the fused features into a high-resolution multispectral image output through deconvolution or upsampling convolution.

[0027] Among them, spatial registration specifically refers to adjusting the multispectral image and the SAR image to the same geographic coordinate system and spatial range through geometric transformation.

[0028] In step S05, it also includes the optimization step of the fusion quality evaluation equation set. Specifically, a multi-objective fusion quality evaluation equation set is constructed to dynamically optimize the feature fusion module during the network training stage, thereby improving the fusion effect and generalization ability.

[0029] The fusion quality evaluation equation set includes a structure consistency equation, a spectral fidelity equation, a texture enhancement equation, a global correlation equation, and a fusion determination equation, which evaluate the fusion quality from different dimensions. During the network training process, the parameter update of the feature fusion module no longer depends only on a single loss function, but is comprehensively guided by this fusion quality evaluation equation set. The specific implementation method is as follows: First, extract multiple feature representations from the fused multispectral image, including the gradient map of the fused multispectral image, the mean value of each band of the fused multispectral image, the local entropy value of the fused multispectral image, and the feature vector of the fused multispectral image; then compare and calculate these features with the corresponding features of the original multispectral image and the high-resolution SAR image to obtain the structure consistency index, the spectral fidelity index, the texture enhancement index, and the global correlation index respectively; then integrate these indexes into a fusion quality determination value through the fusion determination equation. In the backpropagation stage, the system calculates the gradient according to the fusion quality determination value and updates the parameters of the feature fusion module. Among them, the weight coefficient of the fusion determination equation can also be dynamically adjusted to adapt to different scenario requirements. This multi-objective optimization mechanism avoids the problem of quality degradation in other aspects that may be caused by optimizing a single index by balancing multiple factors such as structure information preservation, spectral information integrity, texture detail enhancement, and global correlation. For example, only focusing on structure consistency may lead to spectral distortion; only paying attention to spectral fidelity may result in the loss of spatial details. The entire optimization process is carried out in an iterative manner. As the training progresses, the system will gradually find the optimal fusion parameters that balance each index, so as to obtain a fused multispectral image that performs well in multiple quality dimensions. This method greatly improves the ability of the deep fusion network to handle complex scenarios, and is particularly suitable for the fusion task of heterogeneous data such as multispectral images and SAR images, where the deep fusion network needs to simultaneously process the complementary information of different source data and improve the overall information content while maintaining the characteristics of the original data.

[0030] The structural consistency equation is used to evaluate the degree of preservation of structural information between the fused multispectral image and the high-resolution SAR image. The inputs include the gradient map of the fused multispectral image, the gradient map of the high-resolution SAR image, the local window size, the edge intensity threshold, and the structural weight parameter, and the output is the structural consistency index; The spectral fidelity equation is used to evaluate the degree of preservation of spectral information between the fused multispectral image and the original multispectral image. The inputs include the mean values of each band of the fused multispectral image, the mean values of each band of the original multispectral image, the band weight vector, the spectral angle distance threshold, and the color distortion tolerance, and the output is the spectral fidelity index; The texture enhancement equation is used to quantify the degree of enhancement of texture details introduced from the high-resolution SAR image during the fusion process. The inputs include the local entropy value of the fused multispectral image, the local entropy value of the original multispectral image, the local entropy value of the high-resolution SAR image, the texture complexity parameter, and the scale factor, and the output is the texture enhancement index; The global correlation equation is used to evaluate the global correlation degree between the fused multispectral image and the original multispectral image and the high-resolution SAR image. The inputs include the feature vectors of the fused multispectral image, the original multispectral image, and the high-resolution SAR image, the feature dimension weight, and the correlation threshold, and the output is the global correlation index; The fusion decision equation is used to calculate the final fusion quality decision value by integrating the structural consistency index, the spectral fidelity index, the texture enhancement index, and the global correlation index. The fusion quality decision value is used as the parameter optimization index of the feature fusion module in network training. The higher the fusion quality decision value, the better the fusion quality.

[0031] In step S07, a multi-level dynamic cropping optimization step is further included. Specifically, the multi-spectral image to be fused and the high-resolution SAR image are subjected to multi-level dynamic cropping processing, and the size of the cropped tiles is dynamically adjusted according to the complexity of the image content to avoid regional segmentation problems caused by fixed-size cropping.

[0032] The multi-level dynamic cropping optimization step first identifies the edge-dense regions and texture-rich regions in the multi-spectral image to be fused and the high-resolution SAR image through an edge detection algorithm, then determines the cropping strategy according to the fusion quality decision value, uses a smaller tile size for the edge-dense regions and the texture-rich regions to retain detailed information, uses a larger tile size for the homogeneous regions to maintain spatial coherence, and finally seamlessly stitches the processed tiles back into the complete fused multi-spectral image, effectively reducing tile boundary artifacts and improving the overall fusion effect.

[0033] Among them, the edge-dense region specifically refers to the region in the image where the gradient value is relatively high and the density of edge feature points exceeds a preset threshold.

[0034] Among them, the texture-rich region specifically refers to the region in the image where the local texture complexity calculated by the gray-level co-occurrence matrix exceeds a preset threshold.

[0035] Among them, the homogeneous region specifically refers to the region in the image where the gray-level change is gentle and the local standard deviation is lower than a preset threshold.

[0036] Among them, the tile boundary artifact specifically refers to the brightness jump or texture discontinuity phenomenon caused by inconsistent processing at the joint of adjacent tiles.

[0037] The specific implementation manners of the above steps are described in detail below.

[0038] The specific implementation manner of step S01 is to perform downsampling on the high-resolution SAR image so that its spatial resolution is consistent with that of the multispectral image. This step first obtains the spatial resolution information of the original SAR image and the multispectral image, calculates the resolution ratio factor between the two, and then uses the bicubic interpolation algorithm to perform downsampling on the SAR image. The bicubic interpolation algorithm generates the target pixel value by weighted summation of 16 adjacent pixels of the original image, and has better edge-preserving ability compared with the nearest neighbor interpolation and bilinear interpolation, effectively reducing block artifacts. The downsampling ratio is usually an integer multiple, for example, from 3-meter resolution to 10-meter resolution, and the ratio factor is 10 / 3. This step aims to create a multi-source data pair with the same spatial resolution, provide paired training samples for the subsequent deep learning model, and ensure that the network can learn the mapping relationship between multi-sensor data.

[0039] The specific implementation manner of step S02 is to perform quality reduction processing on the multispectral image to simulate low-quality data in actual applications. This step first performs downsampling on the original multispectral image, uses a Gaussian low-pass filter for smoothing processing, and then uses the bilinear interpolation algorithm to reduce the resolution to 1 / 2 to 1 / 4 of the original resolution. The downsampling ratio can be adjusted according to the actual application scenario, and usually takes a value of 2 or 4. Then, perform upsampling on the multispectral image with reduced resolution, and use bilinear interpolation to restore it to the original spatial resolution without increasing the actual information content. This degradation process simulates the blurring and detail loss of the multispectral image in actual applications, enabling the network to learn how to recover high-quality images from low-quality inputs, and improving the robustness and generalization ability of the model.

[0040] The specific implementation of step S03 is to set the original high-quality multispectral image as the training label of the deep learning network. This step first preprocesses the original multispectral image, including radiometric correction, atmospheric correction, and normalization, adjusts the pixel value range to the interval of 0-1, and eliminates sensor noise and atmospheric effects. Then, the processed original multispectral image is used as the network training label to calculate the loss function and guide the update of network parameters. Through the supervised learning mechanism, the network learns the ability to reconstruct high-quality multispectral images from degraded multispectral images and low-resolution SAR images, realizing effective data fusion and image quality improvement. The loss function adopts a weighted combination of mean square error and structural similarity index, and the weight ratio is usually set to 0.7:0.3 to balance pixel-level accuracy and structural information preservation.

[0041] The specific implementation of step S04 is to perform spatial registration and cropping operations on the low-resolution SAR image and the degraded multispectral image to generate deep learning training samples. First, image spatial registration is performed based on the geographic information system to ensure that the two data have the same geographic coordinate system and spatial range, and the registration accuracy should be controlled within 0.5 pixels. Then, the registered images are cropped into tiles of a fixed size using the sliding window method. The tile size is usually set to 256×256 or 512×512 pixels, and the overlapping rate of adjacent tiles is set to 25%-50% to increase the number of training samples and avoid edge effects. At the same time, considering the information content of each tile, tiles with cloud coverage exceeding 10% or containing all invalid values are excluded. After cropping, data augmentation operations, including random rotation, flipping, and brightness adjustment, are performed on each tile to further expand the training set and improve the generalization ability of the model.

[0042] The specific implementation of step S05 is to design a deep fusion network architecture, including a feature extraction module, a feature fusion module, and an image reconstruction module. The feature extraction module adopts an improved residual network structure to extract multi-scale features from the multispectral image and the SAR image respectively. The network depth is set to 10-15 layers, the size of each convolutional kernel is 3×3, and the number of channels increases layer by layer from 64 to 256. The feature fusion module adopts a dual attention network combining channel attention mechanism and spatial attention mechanism to perform adaptive weight assignment and fusion on the features of different source data. The attention module contains 2-3 fully connected layers, and the activation function adopts a parametric rectified linear unit. The image reconstruction module consists of 3-5 deconvolution layers, gradually restoring the spatial resolution and reconstructing the complete multispectral image. The last layer uses convolution to adjust the number of channels to be consistent with the number of target multispectral bands. The overall network adopts an encoder-fuser-decoder architecture, and residual connections are added between modules to alleviate the problem of gradient disappearance and promote feature transfer.

[0043] Furthermore, the specific implementation method of the fusion quality evaluation equation optimization step in step S05 is to construct a multi-objective quality evaluation system and dynamically optimize the feature fusion module. The structural consistency equation evaluates the degree of structural information preservation by calculating the correlation coefficient between the fused multi-spectral image and the gradient map of the high-resolution SAR image. The local window size is set to 7×7 or 9×9 pixels, the edge intensity threshold is the first 30% percentile of the gradient amplitude, and the structural weight parameter is set to 0.8 - 0.9. The spectral fidelity equation calculates the similarity between the fused multi-spectral image and the mean values of each band of the original multi-spectral image based on the spectral angle mapping algorithm. The band weight vector is determined according to the information entropy of each band. The spectral angle distance threshold is set to 0.05 radians, and the color distortion tolerance is 0.1. The texture enhancement equation quantifies the texture details introduced from the SAR image using the local entropy value difference. The local entropy calculation window size is 5×5 pixels, the texture complexity parameter is set to 0.7, and the scale factor is set to 1.5 - 3 according to the ratio of the SAR to multi-spectral resolution. The global correlation equation evaluates the overall correlation between the fusion result and the input data based on the canonical correlation analysis method. The feature dimension weight is inversely proportional to the contribution rate of the retained feature variance in the principal component analysis. The correlation threshold is set to 0.6. The fusion decision equation combines the four indices in a weighted summation form, and the weight coefficients are dynamically adjusted by the Bayesian optimization algorithm. The fusion quality decision value ranges from 0 to 1, and the higher the value, the better the fusion quality. The threshold is usually set to 0.75.

[0044] The specific implementation of step S06 is to train the deep fusion network using tile training samples and optimize the network parameters. First, the training sample set is divided into a training set, a validation set, and a test set in a ratio of 7:2:1. Then, the Adam optimization algorithm is used for network training. The initial learning rate is set to 0.001, and the learning rate decays to 0.8 times the original value every 20 epochs. The batch size is set to 8 - 32 according to the hardware conditions, and the total number of training epochs is 100 - 300. After each epoch, the validation set is used to evaluate the model performance, and metrics such as peak signal-to-noise ratio, structural similarity index, and spectral angle mapping error are calculated. When the validation accuracy improvement does not exceed 0.1% for 10 consecutive epochs or reaches the preset threshold of 0.85, the training is stopped early to prevent overfitting. At the same time, the gradient clipping technique is used to limit the maximum value of the gradient norm to 10 to avoid unstable training. Finally, the test set is used for the final performance evaluation to ensure that the model has good generalization ability.

[0045] The specific implementation of step S07 is to preprocess the multi-spectral image to be fused and the high-resolution SAR image during the prediction stage to make their spatial ranges consistent. First, the geographic coordinate information and projection parameters of the two images are obtained separately, and then the multi-spectral image is reprojected to the SAR image coordinate system based on the georegistration technology. The affine transformation algorithm is used in the reprojection process, and the control points are selected based on the maximum mutual information criterion, and the registration error is controlled within 0.3 pixels. After reprojection, image cropping operations are used to ensure that the two data have exactly the same spatial range and coordinate reference system. Finally, the consistency of the raster size, number of rows and columns, and geographic range is checked to ensure that the spatial correspondence is accurate. This step provides input data for spatial position matching in subsequent fusion processing and is a prerequisite for accurate fusion.

[0046] Furthermore, the specific implementation method of the multi-level dynamic cropping optimization step in step S07 is to adaptively adjust the cropping strategy according to the characteristics of the image content. First, the Canny edge detection algorithm is used to identify edge-dense regions, and the high and low thresholds are set to the 80% and 40% percentiles of the gradient magnitude histogram, respectively. Texture-rich regions are identified by the entropy value and contrast of the gray-level co-occurrence matrix. The local entropy value threshold is set to 5.0, and the contrast threshold is set to 0.6. Homogeneous regions are identified by the local standard deviation, and the standard deviation threshold is set to 15% of the standard deviation of the original image. Then, the cropping tile size is dynamically determined according to the region type. Small-sized tiles of 128×128 or 192×192 pixels are used for edge-dense regions and texture-rich regions, and large-sized tiles of 384×384 or 512×512 pixels are used for homogeneous regions. After processing each tile, the Poisson image editing algorithm is used for seamless stitching, and the transition zone width is set to 16-32 pixels. By solving the Poisson equation, the gradient field in the transition region is kept continuous, effectively reducing tile boundary artifacts. This step balances the processing accuracy and efficiency by specifically processing regions of different complexities and improves the overall fusion effect.

[0047] The specific implementation of step S08 is to input the multi-spectral image to be fused and the high-resolution SAR image with consistent spatial ranges into the trained deep fusion network at the same time. First, the input data is preprocessed by normalization, and the pixel values are scaled to the range of 0-1. The multi-spectral image uses min-max normalization, and the SAR image is normalized after logarithmic transformation to reduce the difference in sample distributions. Then, a batch processing data structure is constructed, and each batch contains multiple paired image tiles. The batch size is usually set to 1-4, depending on the hardware memory capacity. During the forward propagation of the network, the multi-spectral image and the SAR image pass through their respective feature extraction branches, and then the complementary information is integrated in the feature fusion module. Finally, the fusion result is output through the image reconstruction module. The entire processing process is in an end-to-end manner without manual intervention. The processing speed is related to the network complexity and hardware performance, and the typical processing time is 0.5-2 minutes per square kilometer.

[0048] The specific implementation of step S09 is to obtain a fused multi-spectral image with improved spatial resolution and intact spectral information through the output of the deep fusion network. First, the network output is de-normalized to restore the pixel values from the range of 0 to 1 to the original dynamic range. Then, quality assessment is performed on the output results, including calculating the spatial resolution gain coefficient, spectral distortion degree, edge preservation index, and noise suppression level. The evaluation thresholds are set to 1.5, 0.08, 0.8, and 0.75 respectively. For local abnormal regions, a post-processing strategy is adopted for optimization, including median filtering to eliminate isolated noise points, edge enhancement algorithms to improve contour clarity, and stripe removal processing based on variational models. Finally, the fused multi-spectral image is output, whose spatial resolution is consistent with that of the high-resolution SAR image, the spectral information is highly consistent with the original multi-spectral image, the spectral angle distance error is controlled within 0.05 radians, the peak signal-to-noise ratio is increased by 3 to 6 dB, and the structural similarity index is increased by 0.05 to 0.15.

[0049] Furthermore, the detailed structure of the deep fusion network is as follows: The overall network adopts a dual-stream encoder-fuser-single-stream decoder architecture. The multi-spectral stream encoder contains 5 residual blocks, each of which contains 2 layers of 3×3 convolutional layers, batch normalization layers, and LeakyReLU activation functions. The number of channels doubles from 64 in each layer until it reaches 512; the SAR stream encoder has a similar structure, but the first layer contains a logarithmic transformation layer to handle the speckle noise characteristics of SAR images. The feature fusion module consists of two parts: a cross-modal feature interaction unit and a multi-scale feature aggregation unit. The cross-modal feature interaction unit uses a channel attention mechanism to calculate the importance weights of feature channels. The attention module contains a global average pooling layer, two fully connected layers, and a Sigmoid activation function; the multi-scale feature aggregation unit extracts and fuses features from different scales through a spatial pyramid pooling structure. The pooling scales are 1×1, 2×2, 4×4, and 8×8 respectively, and then the channel dimension is adjusted through 1×1 convolution. The image reconstruction module contains 4 transposed convolutional layers to gradually restore the spatial resolution, and the number of channels is halved in each layer; at the same time, multi-level skip connections are introduced to splice the features of the corresponding layers of the encoder and the decoder features to promote detail restoration. A global residual connection is adopted at the end of the network to add the upsampled input multi-spectral image to the network output to enhance the network stability. The loss function consists of pixel-level L1 loss, structural similarity loss, spectral angle distance loss, and adversarial loss, with a weight ratio of 0.5:0.2:0.2:0.1, comprehensively optimizing spatial details and spectral fidelity.

[0050] In the present invention, "performing a downsampling operation on the multispectral image first and then an upsampling operation" and "using the original multispectral image as the model training label" constitute a degradation-reconstruction supervised learning strategy. Its technical principle lies in creating a self-supervised learning framework that enables the network to learn the mapping relationship from low-quality input to high-quality output. The downsampling-upsampling operation simulates the quality degradation process of the multispectral image in practical applications by reducing and then increasing the spatial resolution. The resulting degraded multispectral image retains the basic spectral features of the original image but loses spatial details. Using the original non-degraded multispectral image as the training label provides the network with a supervision signal of "ideal output". This design enables the network to learn how to utilize the high-frequency structure information in the SAR image to compensate for the lost spatial details in the multispectral image during the training process while maintaining the integrity of the spectral information. From the perspective of technical effects, this strategy solves the contradiction that traditional fusion methods cannot balance spatial enhancement and spectral fidelity. First, through training with degraded samples, the network learns to distinguish which information can be obtained from the SAR data and which needs to be retained from the multispectral data. Second, the supervision signal ensures that the spectral information is not wrongly replaced during the reconstruction process. Third, this training paradigm of simulating degradation-reconstruction enhances the generalization ability of the model, enabling it to handle input data of different quality levels. Finally, this self-supervised mechanism reduces the dependence on large-scale labeled data and improves the practicality of the method. This optimized design essentially constructs an end-to-end learning framework from degraded multi-source data to high-quality fused data, achieving the optimal balance between spatial and spectral information.

[0051] Furthermore, the technical principle of the optimization step of the fusion quality assessment equation system in S05 lies in constructing a multi-dimensional and multi-objective assessment system. Through the synergistic effect of four equations, namely structural consistency, spectral fidelity, texture enhancement, and global correlation, comprehensive constraints and dynamic optimization of the fusion process are achieved. This method breaks through the limitations of traditional single loss functions, enabling the feature fusion module to balance and adjust in different quality dimensions, and avoiding the problem of performance degradation caused by optimizing a single index. This multi-objective optimization mechanism can dynamically adjust the weights of various evaluation indices according to different scenario characteristics during the backpropagation process, providing more accurate gradient guidance for the feature fusion module, thereby achieving more intelligent parameter updates. Its technical effects are reflected in: the fusion result simultaneously meets multiple requirements such as structure preservation, spectral integrity, texture enhancement, and global consistency; the network can adaptively adjust the fusion strategy according to different regional and scenario characteristics; the model has stronger generalization ability and can adapt to different types of multi-spectral and SAR data combinations; significantly reduces the generation of false information during the fusion process. The technical principle of the multi-level dynamic cropping optimization step in S07 is to break through the limitations of traditional fixed-size tile processing. According to the complexity of the image content, the cropping strategy is intelligently adjusted. For edge-dense areas and texture-rich areas, a finer cropping size is used to retain details, while for homogeneous areas, larger tiles are used to maintain coherence. This content-adaptive processing method synergizes with the aforementioned fusion quality assessment equation system, and the technical effects are manifested as: effectively avoiding the problem of important ground objects or texture features being segmented at tile boundaries; reducing the stitching artifacts at tile boundaries; improving the fusion quality of areas with different complexities; enhancing the system's ability to process large-scale remote sensing images; ensuring the spatial continuity and integrity of the fusion result.

[0052] The second aspect of the present invention provides a computer-readable storage medium, in which program instructions are stored. When the program instructions run on a computer, they are used to execute the above-mentioned method for fusing multi-spectral and SAR data.

[0053] The third aspect of the present invention provides a system for fusing multi-spectral and SAR data, including the above-mentioned computer-readable storage medium. The system can be any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is set inside the system, and a microprocessor for executing the program instructions stored in the computer-readable storage medium is set inside the system.

[0054] Specifically, the principle of the present invention is as follows: The technical principle of the present invention is based on the feature extraction, fusion, and reconstruction capabilities of deep learning, and realizes the complementary advantages of multi-source remote sensing data through a cleverly designed network architecture and training strategy. The core principle lies in solving the differences in physical imaging mechanisms, radiation characteristics, and spatial resolutions between multi-spectral and SAR heterogeneous data, and establishing an effective information mapping and fusion mechanism.

[0055] In the feature extraction stage, the present invention extracts multi-scale feature representations from multi-spectral and SAR data respectively. The multi-spectral branch mainly learns the correlation information and basic spatial structure in the spectral dimension; the SAR branch focuses on extracting structural information such as high-frequency textures and edges. Through the hierarchical feature extraction ability of the convolutional neural network, the system can automatically learn the key features in the two types of data, avoiding the limitations of artificial feature engineering in traditional methods.

[0056] The feature fusion module is the core of the present invention, and an attention mechanism is used to achieve the adaptive fusion of two heterogeneous features. The attention mechanism can dynamically adjust the fusion strategy according to the contributions of different regions and different scale features by learning the feature importance weights, which not only retains the spectral information of the multi-spectral data but also incorporates the high-resolution structural information of the SAR data. This adaptive fusion method is superior to the simple weighted fusion or replacement operations in traditional methods.

[0057] The image reconstruction module converts the fused features into high-resolution multi-spectral outputs and restores the spatial details through deconvolution or upsampling convolution operations. In particular, the present invention adopts a degradation-reconstruction supervised learning strategy, enabling the network to learn the mapping relationship from degraded images to original high-quality images. This strategy effectively avoids the spectral distortion problem that may be caused by direct fusion. Using the original multi-spectral image as a supervision signal, the system can maintain the integrity of spectral information during the reconstruction process, and at the same time use the high-resolution structural information provided by the SAR data to enhance the expression of spatial details.

[0058] In addition, the network training adopts a tiled sample construction strategy, which effectively handles the characteristics of high-dimensional and large-size remote sensing data, improving the model training efficiency and generalization ability. In the prediction stage, spatial registration is used to ensure the consistency of multi-source data, providing a basic guarantee for fusion.

[0059] The following provides a specific Embodiment 1 of the present invention, and the specific implementation of each step in this Embodiment 1 is described in detail as follows.

[0060] The specific implementation of step S01 is to perform downsampling on the high-resolution SAR image to make its spatial resolution consistent with that of the multi-spectral image. First, obtain the resolution of the original SAR image and the resolution of the multi-spectral image , and calculate the resolution scale factor , and then the bicubic interpolation algorithm is used to downsample the SAR image. The bicubic interpolation algorithm is based on the cubic convolution kernel function, which is specifically expressed as follows: ; In the formula, is the normalized distance from the central pixel, and its value range is .

[0061] By performing two-dimensional convolution on the original SAR image through the convolution kernel function, the downsampled low-resolution SAR image can be obtained , and the calculation formula is: ; In the formula, is the original SAR image; is the downsampled low-resolution SAR image; is the target pixel coordinate; is the original pixel coordinate; is the convolution kernel function. This downsampling process realizes the conversion from the high-resolution SAR image to the data with the same resolution as the multispectral image, providing paired training samples for the subsequent deep learning model.

[0062] The specific implementation of step S02 is to perform quality reduction processing on the multispectral image to simulate low-quality data in actual applications. First, perform downsampling on the original multispectral image, and after smoothing with a Gaussian low-pass filter, use the bilinear interpolation algorithm to reduce the resolution. The two-dimensional convolution kernel of the Gaussian low-pass filter is defined as: ; In the formula, is the position of the pixel relative to the center of the convolution kernel; is the standard deviation of the Gaussian function, which controls the smoothing degree, and its value range is 1 to 3.

[0063] The mathematical expression of the downsampling process is: ; In the formula, is the original multispectral image; is the downsampled multispectral image; is the target pixel coordinate; is the original pixel coordinate; represents the convolution operation.

[0064] Then, perform upsampling on the multispectral image with reduced resolution, and use bilinear interpolation to restore it to the original spatial resolution. The bilinear interpolation calculation formula is: ; In the formula, is the degraded multi - spectral image; are the floating - point coordinates of the target pixel; and respectively represent and the integer parts of; and are the weight coefficients. This degradation process simulates the blurring and detail loss of multi - spectral images in practical applications.

[0065] The specific implementation of step S03 is to set the original high - quality multi - spectral image as the training label of the deep - learning network. First, pre - process the original multi - spectral image, including radiometric correction, atmospheric correction, and normalization. Radiometric correction uses the linear stretching method: ; In the formula, is the multi - spectral image after radiometric correction; is the original multi - spectral image; are the pixel coordinates; is the band index; and are respectively the minimum and maximum values of the original data.

[0066] Atmospheric correction uses the dark - target subtraction method: ; In the formula, is the multi - spectral image after atmospheric correction; is the digital quantity of the dark target in the th band, obtained by statistically averaging the 1% pixels in the darkest area of each band.

[0067] Normalization adjusts the pixel value range to the 0 - 1 interval: ; In the formula, is the multi - spectral image after normalization, used as the network training label; and respectively represent the minimum and maximum values of the th band.

[0068] The specific implementation of step S04 is to perform spatial registration and cropping operations on the low - resolution SAR image and the degraded multi - spectral image to generate deep - learning training samples. First, perform image spatial registration based on the geographic information system to ensure that the two data have the same geographic coordinate system and spatial range. Spatial registration uses the affine transformation model: ; In the formula, is the coordinate in the source coordinate system; is the coordinate in the target coordinate system; ~ are the six parameters of the affine transformation matrix, which are calculated based on control point pairs by the least squares method. The solution of the transformation matrix uses the normal equation: ; In the formula, is the coefficient matrix composed of control point coordinates; is the transformation parameter vector; is the target coordinate vector. The control points are selected based on feature matching, and the matching accuracy is controlled within 0.5 pixels.

[0069] Then, the registered image is cropped into fixed-size tiles using the sliding window method. The tile position calculation formula is: ; In the formula, represents the tile pixel set with the position index of ; is the sliding step, usually taking 50% - 75% of the tile size; and are the width and height of the tile respectively, usually set to 256×256 or 512×512 pixels.

[0070] The specific implementation of step S05 is to design a deep fusion network architecture, including a feature extraction module, a feature fusion module, and an image reconstruction module. The feature extraction module uses an improved residual network structure to extract multi-scale features from the multi-spectral image and the SAR image respectively. The mathematical expression of the residual block is: ; In the formula, and represent the feature maps of the th layer and the th layer respectively; represents the residual mapping function; represents the network parameters of the th layer.

[0071] The residual mapping function is defined as a combination of convolution, batch normalization, and activation function: ; In the formula, represents the convolution operation; represents the batch normalization operation; represents the activation function, using the parametric rectified linear unit (PReLU): , where is a learnable parameter with an initial value set to 0.2.

[0072] The feature fusion module adopts a dual attention network that combines channel attention mechanism and spatial attention mechanism. The calculation formula for channel attention is: ; In the formula, is the input feature map; and represent global average pooling and global max pooling operations respectively; represents a multi-layer perceptron composed of two fully connected layers; represents the Sigmoid activation function; is the channel attention weight, with the size being the same as the number of channels of the feature map.

[0073] The calculation formula for spatial attention is: ; In the formula, and represent the average pooling and max pooling results along the channel dimension respectively; represents the concatenation operation on the channel dimension; represents convolutional layer; is the spatial attention weight, with the size being the same as the spatial dimension of the feature map.

[0074] The final feature fusion expression is: ; In the formula, and represent the multi-spectral feature and SAR feature respectively; represents element-wise multiplication; is the fused feature.

[0075] In the optimization steps of the fusion quality assessment equation set, the mathematical expression of the structure consistency equation is: ; In the formula, is the structure consistency index, with the value range [0, 1]. The larger the value, the higher the degree of structure preservation; and represent the gradient maps of the fused multi-spectral image and high-resolution SAR image respectively; is the weight coefficient at position , which is determined by the edge intensity: ; In the formula, is the edge intensity threshold, with the value being the first 30% percentile of the gradient amplitude; is the weight of the non-edge area, and its value range is 0.1 to 0.3.

[0076] The mathematical expression of the spectral fidelity equation is: ; In the formula, is the spectral fidelity index, and its value range is [0, 1]. The larger the value, the higher the degree of spectral preservation; and respectively represent the means of each band of the fused multispectral image and the original multispectral image; is the band index; is the total number of bands; is the weight of the band, which is calculated according to the information entropy: ; In the formula, is the information entropy of the band: , where is the gray value probability distribution.

[0077] The mathematical expression of the texture enhancement equation is: ; In the formula, is the texture enhancement index, and its value range is [0, 1]. The larger the value, the higher the degree of texture enhancement; , and respectively represent the local entropy maps of the fused multispectral image, the original multispectral image, and the high-resolution SAR image; is a small constant to prevent division by zero, and its value is 0.001; is the scale factor, and its value is 1.5 to 3, which is determined according to the ratio of the SAR and multispectral resolutions. The local entropy calculation formula is: ; In the formula, is the normalized gray value of the pixel in the local window.

[0078] The mathematical expression of the global correlation equation is: ; In the formula, is the global correlation index, and its value range is [0, 1]. The larger the value, the higher the degree of overall correlation; represents the canonical correlation coefficient; , and respectively represent the feature vectors of the fused multi-spectral image, the original multi-spectral image, and the high-resolution SAR image; and are weight coefficients, and , usually taking , . The formula for calculating the canonical correlation coefficient is: ; In the formula, represents and 's covariance; and respectively represent and 's variances.

[0079] The mathematical expression of the fusion decision equation is: ; In the formula, is the fusion quality decision value, with a value range of [0, 1]. The higher the value, the better the fusion quality; , , and are weight coefficients, and , with initial values set to 0.3, 0.3, 0.2, and 0.2 respectively, and dynamically adjusted through the Bayesian optimization algorithm: ; In the formula, and respectively represent the weight vectors of the th and th iterations; is the learning rate, with a value range of 0.01 - 0.05; represents the gradient of the fusion quality decision value with respect to the weight vector.

[0080] The specific implementation of step S06 is to train the deep fusion network using tile training samples and optimize the network parameters. The Adam optimization algorithm is used for network training, and its update rules are: ; ; ; ; ; In the formula, represents the current gradient; and respectively represent the first-order momentum and the second-order momentum; and is the momentum after deviation correction; is the updated network parameter; is the learning rate, and the initial value is set to 0.001; and are the momentum decay coefficients, with values of 0.9 and 0.999 respectively; is a small constant to prevent division by zero, with a value of .

[0081] The learning rate is dynamically adjusted using the cosine annealing strategy: ; In the formula, is the learning rate of the th cycle; and are the minimum and maximum learning rates respectively, with values of 0.0001 and 0.001; is the total number of cycles, with a value ranging from 100 to 300.

[0082] The loss function adopts a weighted combination of mean squared error and structural similarity index: ; In the formula, is the mean squared error loss; is the structural similarity loss; is the spectral angle mapping loss; , and are the weight coefficients, with values of 0.5, 0.3 and 0.2 respectively.

[0083] The specific calculation formulas for each loss term are: ; ; ; In the formula, and respectively represent the network prediction output and the true label, and respectively represent the th network prediction output and the th true label; and respectively represent the means of and ; and respectively represent the standard deviations; represents the covariance; and To prevent small constants from dividing by zero; and respectively represent the spectral vectors of the predicted output and the true label at position .

[0084] The specific implementation of step S07 is to preprocess the multi-spectral image to be fused and the high-resolution SAR image during the prediction stage to make their spatial ranges consistent. First, the geographic coordinate information and projection parameters of the two images are obtained respectively, and then the multi-spectral image is reprojected to the SAR image coordinate system based on the georegistration technology. The reprojection adopts the affine transformation model, the control points are selected based on the maximum mutual information criterion, and the registration error is controlled within 0.3 pixels. The formula for mutual information is: ; In the formula, represents the mutual information between images and ; represents the joint probability distribution; and respectively represent the marginal probability distributions. During the registration process, the optimal control point pairs are selected by maximizing the mutual information value.

[0085] The specific implementation method of the multi-level dynamic cropping optimization step is to adaptively adjust the cropping strategy according to the characteristics of the image content. First, the Canny edge detection algorithm is used to identify the edge-dense regions: ; ; In the formula, represents the gradient magnitude at position ; and respectively represent the horizontal and vertical gradients; represents the gradient direction.

[0086] The high and low thresholds of the Canny algorithm are set to the 80th and 40th percentiles of the gradient magnitude histogram respectively. The criterion for determining the edge-dense region is: ; In the formula, represents the determination result of the edge-dense region; represents the Canny edge detection result; represents the statistical window size, with a value of 16 or 32; represents the density threshold, with a value of 0.3 - 0.5.

[0087] The texture-rich regions are identified by the entropy value and contrast of the gray-level co-occurrence matrix: ; ; In the formula, and respectively represent the entropy value and contrast of the gray-level co-occurrence matrix; represents from the gray value to the gray value joint probability; represents the number of gray levels, usually taking values of 16 or 32.

[0088] The determination criterion for the texture-rich area is: ; In the formula, represents the determination result of the texture-rich area; represents the entropy value threshold, taking a value of 5.0; represents the contrast threshold, taking a value of 0.6.

[0089] The homogeneous area is identified by the local standard deviation: ; In the formula, represents the determination result of the homogeneous area; represents the local standard deviation; represents the standard deviation threshold, taking a value of 15% of the standard deviation of the original image.

[0090] Finally, the size of the cropped tile is dynamically determined according to the region type: ; In the formula, represents the position tile size at; , and respectively represent the small, medium, and large tile sizes, with values of 128×128, 256×256, and 512×512 pixels respectively.

[0091] The specific implementation of step S08 is to input the multi-spectral image to be fused and the high-resolution SAR image within the same spatial range into the trained deep fusion network at the same time. First, the input data is preprocessed by normalization, and the pixel values are scaled to the range of 0 to 1: ; In the formula, represents the normalized multi-spectral image; represents the original multi-spectral image; represents the pixel position; represents the band index; and respectively represent the The minimum and maximum values of the band.

[0092] The SAR image is first subjected to logarithmic transformation to reduce the impact of speckle noise, and then normalized: ; ; In the formula, represents the SAR image after logarithmic transformation; represents the normalized SAR image.

[0093] Then a batch data structure is constructed, with each batch containing multiple paired image tiles: ; In the formula, represents the data of one batch; represents the batch size, with a value ranging from 1 to 4, depending on the hardware memory capacity.

[0094] The mathematical expression of the forward propagation process of the network is: ; ; ; ; In the formula, and respectively represent the multispectral encoder and the SAR encoder; and respectively represent the extracted multispectral features and SAR features; represents the feature fusion module; represents the fused features; represents the decoder; represents the fused output result.

[0095] The specific implementation of step S09 is to obtain a fused multispectral image with improved spatial resolution and intact spectral information through the output of the deep fusion network. The network output is first subjected to inverse normalization to restore the pixel values from the 0 - 1 range to the original dynamic range: ; In the formula, represents the fused multispectral image after inverse normalization; represents the normalized fused result output by the network.

[0096] Then the output result is evaluated for quality, including calculating the spatial resolution gain coefficient, spectral distortion degree, edge preservation index, and noise suppression degree: ; In the formula, represents the spatial resolution gain coefficient. A value greater than 1 indicates an improvement in spatial resolution, and the ideal value is the square of the ratio of SAR to multispectral resolution; and represent the gradient maps of the fusion result and the original multispectral image, respectively.

[0097] The formula for calculating the spectral distortion degree is: ; In the formula, represents the spectral distortion degree. The smaller the value, the higher the degree of spectral preservation; represents the original multispectral image after interpolation upsampling; represents the number of bands; represents the image size.

[0098] The formula for calculating the edge preservation index is: ; In the formula, represents the edge preservation index, with a value range of [0, 1]. The larger the value, the higher the degree of edge preservation; and represent the edge maps of the fusion result and the high-resolution SAR image, respectively; represents the edge overlap operation.

[0099] The formula for calculating the noise suppression degree is: ; In the formula, represents the noise suppression degree, with a value range of [0, 1]. The larger the value, the better the noise suppression effect; and represent the variances of the fusion result and the high-resolution SAR image in the homogeneous region, respectively.

[0100] The thresholds of the evaluation indicators are set as: , , , . If the evaluation indicators do not meet the requirements, the model needs to be retrained or the parameters adjusted.

[0101] For local abnormal regions, a post-processing strategy is adopted for optimization. First, the abnormal regions are detected through local standard deviation: ; In the formula, represents the abnormal region marking map; and represent the local standard deviation and the mean value, respectively; Denote the anomaly detection threshold, with a value range of 3 to 5.

[0102] For the detected anomaly regions, median filtering is applied for processing: ; In the formula, denotes the fused multi - spectral image after post - processing; denotes the median filtering operation; denotes the filtering window radius, with a value of 1 or 2.

[0103] Finally, the edge enhancement algorithm is implemented based on unsharp masking: ; In the formula, denotes the finally output fused multi - spectral image; denotes the fused image after Gaussian filtering; denotes the edge enhancement factor, with a value range of 0.5 to 1.5.

[0104] The overall architecture of the deep fusion network adopts a dual - stream encoder - fusion - single - stream decoder structure. The multi - spectral stream encoder and the SAR stream encoder are respectively responsible for extracting features from the multi - spectral image and the SAR image. The encoder consists of multiple residual blocks, and each residual block is expressed as: ; In the formula, and respectively denote the features of the th layer and the th layer; denotes the non - linear mapping function of the th layer. The non - linear mapping function is specifically implemented as: ; In the formula, and respectively denote the convolutional operations of the first layer and the second layer; denotes the batch normalization operation; denotes the activation function, using the parametric rectified linear unit (PReLU).

[0105] The feature fusion module adopts a dual - attention network combining channel attention mechanism and spatial attention mechanism, and its mathematical expression is: ; ; ; ; ; In the formula, and respectively represent the multi - spectral feature and the SAR feature; and respectively represent the channel attention weight and the spatial attention weight; and respectively represent the weighted multi - spectral feature and the SAR feature; represents the fused feature.

[0106] Decoder Consists of multiple transposed convolutional layers, which are used to convert the fused feature into a high - resolution multi - spectral image. The transposed convolution operation is expressed as: ; In the formula, represents the transposed convolution operation. At the same time, multi - level skip connections are introduced to splice the features of the corresponding layers of the encoder and the decoder features: ; In the formula, represents the feature of the th layer of the decoder; represents the feature of the corresponding layer of the encoder; represents the feature splicing operation; represents the total number of layers of the network. At the end of the network, a global residual connection is also adopted to add the up - sampled input multi - spectral image and the network output: ; In the formula, represents the up - sampling operation.

[0107] The loss function consists of pixel - level L1 loss, structural similarity loss, spectral angle distance loss and adversarial loss: ; In the formula, represents the L1 loss; represents the structural similarity loss; represents the spectral angle distance loss; represents the adversarial loss; , , and are weight coefficients, and their values are 0.5, 0.2, 0.2 and 0.1 respectively.

[0108] The specific calculation formulas of each loss term are: ; ; ; ; In the formula, represents the fused result output by the network; represents the true label; and respectively represent the spectral vectors at position ; represents the discriminator network, which is used to judge the consistency between the fused result and the high-resolution SAR image.

[0109] In summary, the multi-spectral and SAR data fusion method effectively combines the improvement of spatial resolution and the preservation of spectral information through deep learning technology. In the training stage, the low-resolution SAR image and the degraded multi-spectral image are used as inputs, and the original multi-spectral image is used as the label to learn the mapping relationship between multi-sensor data. In the prediction stage, the multi-spectral image to be fused and the high-resolution SAR image are input into the trained deep fusion network, and a fused multi-spectral image with improved spatial resolution and intact spectral information is output. Through the optimization of the fusion quality evaluation equation set and the multi-level dynamic cropping optimization, the fusion effect and generalization ability are further improved, which is applicable to the remote sensing image fusion tasks in various complex scenarios.

[0110] To better understand and implement the present invention, the following provides Embodiment 2 of a specific application scenario of the present invention: To solve the problems of insufficient spatial resolution of multi-spectral data and single spectral information of SAR data in agricultural remote sensing monitoring, researchers carried out multi-spectral and SAR data fusion experiments in a certain farmland irrigation area. This area contains various crop types such as rice, rapeseed, and wheat, with an area of approximately 25 square kilometers. The experiment used GF6 multi-spectral data and GF3 SAR data. The GF6 multi-spectral data has a spatial resolution of 8m and includes four bands: blue (450 - 520nm), green (520 - 590nm), red (630 - 690nm), and near-infrared (760 - 900nm). The GF3 SAR data is in the C band with a spatial resolution of 3m. The purpose of the fusion is to generate a fused data product with both the high spatial resolution (3m) of GF3 and the complete spectral information of GF6 to support the evaluation of farmland irrigation status and crop growth monitoring.

[0111] Step S01: Perform downsampling processing on the GF3 SAR data. First, measure the spatial resolutions of the original GF3 SAR data and the GF6 multi-spectral data, and calculate the resolution scale factor α = 8 / 3 ≈ 2.67. Then, use the bicubic interpolation algorithm to perform downsampling processing on the GF3 SAR image to generate a low-resolution SAR image (8m) with the same spatial resolution as the GF6 multi-spectral image. The downsampling uses a 3×3 pixel convolution window and uses the convolution kernel: ; The downsampling results are shown in Table 1: Table 1 Comparison Table of SAR Data Before and After Downsampling

[0112] Step S02: Perform quality reduction processing on the GF6 multispectral image. First, use a Gaussian low-pass filter to smooth the original multispectral image, set the Gaussian kernel parameter σ to 1.5, and then use bilinear interpolation to reduce the resolution to 1 / 3 of the original resolution (i.e., 24m). Subsequently, use bilinear interpolation to upsample the low-resolution multispectral image back to the original resolution (8m) to form a degraded multispectral image. This process simulates the blurring and detail loss of multispectral images in actual applications. The processing results are shown in Table 2: Table 2 Effect Table of Multispectral Data Degradation Processing

[0113] Step S03: Set the original GF6 multispectral image as the training label of the deep learning network. Perform radiometric correction and atmospheric correction on the original multispectral image to eliminate sensor noise and atmospheric effects, and then normalize it to the range of 0-1. The typical pixel value distribution of the original multispectral image is shown in Table 3: Table 3 Typical Ground Object Reflectance Table of the Original Multispectral Image

[0114] Step S04: Crop the low-resolution SAR image and the degraded multispectral image into training samples according to the same geographical range. First, perform spatial registration, adopt an affine transformation model, and control the registration accuracy within 0.35 pixels. Then, use the sliding window method to crop the registered image into fixed-size tiles, set the tile size to 256×256 pixels, and the sliding step to 192 pixels (the overlapping rate of adjacent tiles is 25%). Eliminate the tiles with cloud coverage exceeding 10% and invalid values exceeding 5%. The number of finally generated training samples is shown in Table 4: Table 4 Statistical Table of the Number of Training Samples

[0115] Step S05: Design the deep fusion network architecture. The network includes a feature extraction module, a feature fusion module, and an image reconstruction module. The feature extraction module adopts an improved residual network structure, the network depth is 12 layers, the convolution kernel size of each layer is 3×3, and the number of channels increases layer by layer from 64 to 256. The feature fusion module adopts a dual attention network combining channel attention mechanism and spatial attention mechanism. The network parameter settings are shown in Table 5: Table 5 Network Parameter Setting Table of the Deep Fusion Network

[0116] The fusion quality assessment equations include a structural consistency equation, a spectral fidelity equation, a texture enhancement equation, a global correlation equation, and a fusion determination equation. The edge intensity threshold of the structural consistency equation is set to the top 30% percentile of the gradient magnitude, and the weight β of the non-edge region is set to 0.2. The band weights of the spectral fidelity equation are calculated based on information entropy, and the weight of the near-infrared band is at most 0.38. The scale factor γ of the texture enhancement equation is set to 2.5. The weight coefficients λ1 and λ2 of the global correlation equation are set to 0.6 and 0.4 respectively. The initial weights of the fusion determination equation are set to w1 = 0.3, w2 = 0.3, w3 = 0.2, w4 = 0.2, and the learning rate η is set to 0.02.

[0117] Step S06: Train the deep fusion network using tile training samples. The Adam optimization algorithm is used for network training. The initial learning rate is set to 0.001 and dynamically adjusted using the cosine annealing strategy, with the minimum learning rate being 0.0001. The batch size is set to 16, and the number of training epochs is 200. The loss function uses a weighted combination of mean squared error, structural similarity index, and spectral angle mapping, with a weight ratio of 0.5:0.3:0.2. The changes in the loss function during the training process are shown in Table 6: Table 6 Changes in the loss function during network training

[0118] As Figure 2 shown, it shows the changing trends of the L1 loss, SSIM loss, SAM loss, and total loss during the network training process with respect to the number of training epochs. It clearly shows how the loss values rapidly decrease in the initial stage of training, then gradually stabilize, and the training is prematurely stopped at the 183rd epoch due to reaching the preset accuracy threshold. After the validation accuracy reaches the preset threshold of 0.82, the training is prematurely stopped at the 183rd epoch to prevent overfitting. The evaluation metrics on the test set are shown in Table 7: Table 7 Test set evaluation metrics table

[0119] Step S07: Data preprocessing in the prediction stage. The GF6 multispectral image to be fused and the high-resolution GF3 SAR image are processed to have the same spatial range. First, the two types of data are spatially registered based on the geographic information system. The affine transformation model is adopted, and the control points are selected based on the maximum mutual information criterion, with the registration accuracy controlled within 0.25 pixels. Then, multi-level dynamic cropping processing is performed. The Canny edge detection algorithm is used to identify edge-dense regions, and the high and low thresholds are set to the 78% and 38% percentiles of the gradient magnitude histogram, respectively. Texture-rich regions are identified through the entropy value of the gray-level co-occurrence matrix, and the threshold is set to 5.2. Homogeneous regions are identified through the local standard deviation, and the threshold is set to 13% of the standard deviation of the original image. The dynamic cropping distribution is shown in Table 8: Table 8 Statistical table of multi-level dynamic cropping

[0120] Step S08: The GF6 multispectral image to be fused and the high-resolution GF3 SAR image are simultaneously input into the trained deep fusion network. The multispectral image is normalized using the maximum and minimum values, and the SAR image is normalized after logarithmic transformation. A batch data structure is constructed, and the batch size is set to 2. The entire processing process is in an end-to-end manner, with a total processing time of 32 minutes and an average processing time of 1.28 minutes per square kilometer.

[0121] Step S09: Obtain the fusion result and conduct quality assessment. The network output is first de-normalized to restore the pixel values from the 0-1 interval to the original dynamic range. Then, quality assessment indicators are calculated: the spatial resolution gain coefficient is 2.56, the spectral distortion degree is 0.062, the edge preservation index is 0.863, and the noise suppression degree is 0.826, all of which meet the assessment threshold requirements. For local abnormal regions, 3×3 median filtering is applied for post-processing, and the abnormal detection threshold is set to 4.2. Edge enhancement is achieved using unsharp masking, and the edge enhancement factor is set to 1.2. The final fusion result is shown in Table 9: Table 9 Comparison table of indicators before and after fusion

[0122] Compared with traditional multi-spectral and SAR data fusion methods, the deep learning method adopted in this embodiment has significant advantages. Traditional methods mainly include methods based on the transform domain (such as high-pass filter injection, wavelet transform, principal component analysis, etc.) and methods based on statistical models (such as Bayesian fusion, sparse representation, etc.). These methods usually have problems such as serious spectral distortion, difficulty in simultaneously maintaining spatial and spectral quality, and difficulty in adaptively adjusting algorithm parameters. For example, the wavelet transform method is prone to generating ringing artifacts in the processing of high-frequency information such as farmland boundaries, and the extraction accuracy of general farmland boundaries can only reach about 85%; the principal component analysis method is prone to causing spectral distortion, and the accuracy of NDVI usually drops by 8% - 12%. In contrast, the deep fusion network of the present invention dynamically optimizes the feature fusion module through a multi-objective fusion quality evaluation equation set, achieving a balance in multiple aspects such as maintaining structural information, spectral information integrity, enhancing texture details, and global correlation. At the same time, the multi-level dynamic cropping strategy adaptively adjusts the processing method according to the characteristics of the image content, effectively reducing the edge artifact problem caused by traditional fixed window processing. The experimental results in the farmland irrigation area of Jingzhou, Hubei show that compared with traditional methods, while maintaining a high spatial resolution (3m), the accuracy drop of NDVI in this method is controlled within 2.2%, the extraction accuracy of farmland boundaries is increased to 92.8%, and the recognition rate of small farmlands (area < 0.5 hectares) is increased by 33.0%, providing more accurate data support for precision agriculture management and farmland irrigation monitoring.

[0123] It should be noted that the detailed explanations of the variables involved in the present invention are shown in Tables 10, 11, 12, and 13 below.

[0124] Table 10 Variable Explanation Table (First Part)

[0125] Table 11 Variable Explanation Table (Second Part)

[0126] Table 12 Variable Explanation Table (Third Part)

[0127] Table 13 Variable Explanation Table (Fourth Part)

[0128] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.

Claims

1. A method for fusing multi-spectral and SAR data, characterized in that, Including: When constructing a deep learning training set, downsample the SAR image to the same spatial resolution as the multispectral image to obtain a low-resolution SAR image; first perform a downsampling operation on the multispectral image and then an upsampling operation to obtain a degraded multispectral image; use the original multispectral image as the model training label; crop the low-resolution SAR image and the degraded multispectral image into fixed-size tile training samples; design a deep fusion network including a feature extraction module, a feature fusion module, and an image reconstruction module; use the tile training samples to train the deep fusion network; during prediction, process the multispectral image to be fused and the high-resolution SAR image to have the same spatial range and input them into the trained deep fusion network; output the fused multispectral image through the deep fusion network.

2. The method according to claim 1, wherein The downsampling operation specifically reduces the spatial resolution of the high-resolution image through an interpolation algorithm; the upsampling operation specifically increases the spatial resolution of the low-resolution image through an interpolation algorithm.

3. The method according to claim 2, wherein The tile training samples are specifically formed by labeling a large-size remote sensing image and cropping it into fixed-size pixel blocks to form training samples.

4. The method according to claim 3, characterized in that, The fused multispectral image is specifically a multispectral image with improved spatial resolution and intact spectral information.

5. The method according to claim 4, wherein The feature extraction module specifically uses a convolutional neural network to extract multi-scale feature representations from the multispectral image and the SAR image; the feature fusion module specifically effectively combines multi-source features through an attention mechanism or other fusion algorithms, retaining their respective advantageous information; the image reconstruction module specifically restores the fused features to a high-resolution multispectral image output through deconvolution or upsampling convolution.

6. The method according to claim 5, characterized in that In the step of processing the multispectral image to be fused and the high-resolution SAR image to have the same spatial range, specifically, the multispectral image and the SAR image are adjusted to the same geographic coordinate system and spatial range through geometric transformation.

7. The method according to claim 6, characterized in that In the step of designing a deep fusion network including a feature extraction module, a feature fusion module, and an image reconstruction module, it further includes a step of optimizing the fusion quality evaluation equation set, specifically, dynamically optimizing the feature fusion module during the network training stage by constructing a multi-objective fusion quality evaluation equation set.

8. The method according to claim 7, wherein In the step of processing the multispectral image to be fused and the high-resolution SAR image to have the same spatial range, it further includes a step of multi-level dynamic cropping optimization, specifically, performing multi-level dynamic cropping processing on the multispectral image to be fused and the high-resolution SAR image, and dynamically adjusting the cropped tile size according to the complexity of the image content.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions, and when the program instructions run on a computer, they are used to execute the method for fusing multispectral and SAR data according to any one of claims 1-8.

10. A system for fusing multi-spectral and SAR data, characterized in that, Including the computer-readable storage medium according to claim 9, the system is any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is disposed within the system, and a microprocessor for executing the program instructions stored within the computer-readable storage medium is disposed within the system.

Citation Information

Patent Citations

  • Homogeneous / heterogeneous remote sensing image high-fidelity generalized space-spectrum fusion method

    CN110533600A

  • Spectral image fusion super-resolution method and system and electronic equipment

    CN116630159A

  • Synthetic aperture radar image and multispectral image fusion method

    CN116863283A

  • High-resolution SAR and visible light image fusion method based on unsupervised deep learning

    CN118230103A

  • Semi-supervised spectral image classification method based on auxiliary task

    CN119600349A

Cited By

  • Multi-spatial resolution remote sensing simulation and error influence stripping method

    CN120911142A

  • Multi-spectral-band spectral image structure-based segmentation and optimization method and system

    CN121600006A