Degradation sensing panchromatic sharpening method and system based on three-stage progressive fusion

Through the three-stage progressive fusion method and the dual-path feature mutual enhancement network, the remote sensing image fusion problem in noisy and blurred scenes is solved, high-quality spectral-spatial information fusion is achieved, and the robustness and clarity of remote sensing images are improved.

CN120807308APending Publication Date: 2025-10-17WUHAN UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510821086.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing panchromatic sharpening methods suffer from performance degradation in noisy and blurred scenes and lack effective deblurring and fusion consistency, resulting in blurred image details and spectral distortion, which affects the robustness and quality of remote sensing images.

Method used

A three-stage progressive fusion method is adopted, through the coarse fusion-deblurring-fine fusion process, combined with a dual-path feature mutual enhancement network, using the PLKBlock module and the SMFA module for feature optimization, and an end-to-end loss function is designed to achieve a dynamic balance of spectral and spatial information.

Benefits of technology

It significantly improves the fusion robustness and image quality in noise interference and blur degradation scenarios, improves the spatial resolution and spectral fidelity, and outperforms existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807308A_ABST
    Figure CN120807308A_ABST
Patent Text Reader

Abstract

The invention provides a panchromatic sharpening joint optimization method and system based on three-stage progressive fusion. The method comprises three stages of coarse fusion, deblurring detail enhancement and fine fusion. The method comprises the following steps: firstly, performing feature extraction and preliminary fusion on a low-resolution multispectral image and a high-resolution panchromatic image through a dual-path mutual enhancement network; secondly, a multi-layer stacked deblurring module is introduced to perform deep enhancement on fusion features, and the details and definition of the image are effectively improved in combination with partial large kernel convolution, channel mixing and an element-level attention mechanism; and finally, realizing fine optimization of spectrum consistency and a space structure, and outputting a fused image with high spatial resolution and high spectral fidelity. According to the method, the stability and reconstruction quality of panchromatic sharpening under a complex degradation condition are effectively improved through multi-source remote sensing image collaborative modeling and progressive feature fusion, and the method has good practical value and popularization prospect and is superior to a current panchromatic sharpening method based on deep learning.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of remote sensing image processing, and particularly relates to a multispectral (MS) and panchromatic (PAN) image fusion method and system, which is particularly suitable for a panchromatic sharpening (Pansharpening) task of a remote sensing image with noise and blur degradation. BACKGROUND

[0002] The fusion technology of multispectral images (Multispectral Image, MS) and panchromatic images (Panchromatic Image, PAN), commonly known as panchromatic sharpening (Pansharpening), plays an important role in the field of remote sensing image processing. MS images have spectral information of multiple bands and can provide rich ground feature reflection characteristics, showing strong discrimination ability in land cover classification, vegetation monitoring, water body identification and other applications. However, due to the limitations of sensor resolution and imaging methods, MS images often have low spatial resolution, resulting in insufficient edge definition, texture structure and detail level expression. In contrast, although PAN images only contain one band, their high spatial resolution can clearly show the spatial form and structural details of surface objects, and have obvious advantages in spatial information expression, but lack multi-dimensional spectral information.

[0003] Therefore, how to retain the spectral characteristics of MS images while introducing the spatial structure information of PAN images to generate a fusion image with high spectral resolution and high spatial resolution has become one of the core challenges in remote sensing image processing. An effective fusion method can not only significantly improve the visual quality and interpretability of remote sensing images, but also provide more reliable and accurate input data for fine ground object classification, change detection, target recognition and other high-level remote sensing analysis tasks. With the development of remote sensing imaging technology and artificial intelligence methods, more and more fusion methods based on transform domain, optimization model and deep learning have been proposed, which not only improve the quality of the fusion image, but also face technical problems such as spectral distortion, spatial artifacts and insufficient generalization ability. Therefore, designing a fusion strategy that balances spectral preservation and spatial enhancement is one of the key directions to promote the progress of remote sensing image analysis technology.

[0004] At present, existing panchromatic sharpening technologies can be divided into four categories: component substitution (Component Substitution, CS) method, multiresolution analysis (Multiresolution Analysis, MRA) method, model-based method, and deep learning (Deep Learning, DL) based method.

[0005] Among them, the component substitution method is to project the multispectral image (MS) to multiple image components, and then replace the component containing spatial information with the panchromatic image (PAN) to realize the spatial enhancement of the multispectral image. Representative examples of such methods include the intensity-hue-saturation (IHS) method, the nonlinear IHS (NIHS) method, the band-dependent spatial detail (BDSD) method, and the Gram-Schmidt adaptive (GSA) method. Such methods can effectively enhance the spatial information of the image, but due to the failure to completely decouple the spatial and spectral information, the problem of spectral distortion often occurs.

[0006] The multi-resolution analysis method realizes image fusion by decomposing the spatial details in the panchromatic image and injecting them into the multispectral image. Some studies have also proposed a hybrid method that combines component substitution with multi-resolution analysis to take advantage of both methods and compensate for the limitations of a single method.

[0007] In recent years, model-based image fusion methods have received attention. Such methods usually construct an optimization objective function by constructing prior constraints on spatial consistency and spectral consistency, thereby realizing image fusion. For example, existing literature has proposed a variational panchromatic sharpening method based on gradient sparse representation, and a variational method based on local gradient constraints to consider the gradient difference between the panchromatic image and the multispectral image in local regions and bands. In addition, some methods separate the cartoon component and texture component in the image, and introduce structure and texture similarity constraints to improve the fusion performance. Such methods achieve a good balance between spatial and spectral fidelity, but usually have high computational complexity and are highly dependent on hand-designed parameters.

[0008] With the development of deep learning technology, deep learning-based panchromatic sharpening methods have been widely studied. Such methods automatically extract and learn image features through multiple layers of nonlinear structures, which can effectively model the complex relationship between panchromatic images and multispectral images. Early methods such as PNN (Pansharpening Neural Network) use shallow convolutional networks to realize image fusion, but due to the simple structure, there are limitations in handling complex nonlinear relationships. To improve the network expression ability, subsequent studies have proposed a deep residual panchromatic sharpening network (DRPNN) and introduced a multi-scale residual structure to enhance feature extraction capability. At the same time, some studies have constructed a dual-stream pyramid network to process the panchromatic image and the multispectral image respectively, and realize layer-by-layer spatial detail injection. Although such methods improve the fusion performance, they introduce a large number of convolution operations, resulting in an increase in network parameter quantity.

[0009] In terms of fusing physical models and deep networks, existing methods such as FusionNet simulate the detail injection mechanism in traditional CS and MRA based on residual structure, enhancing the model's interpretability. Another type of method such as Gradient Projection Network (GPPNN) explicitly maps the model-based optimization process into network modules through an iterative structure, combining the advantages of model-driven and data-driven. Further research has proposed a network architecture combining Laplacian pyramid structure and cross-scale loss function to strengthen multi-scale spatial feature expression.

[0010] In addition, existing methods have begun to explore the introduction of Transformer structure into the panchromatic sharpening task. For example, Panformer enhances fusion performance by modeling long-range dependencies between images. Other methods such as HyperDSNet combine multi-detail extraction and spectral attention mechanisms to improve the ability to preserve details and spectral information using deep and shallow fusion structures. At the same time, some research has combined variational models with deep neural networks to construct the implicit function operator in the spatial fidelity constraint through deep networks, thereby improving fusion performance and non-linear expression ability.

[0011] In recent years, deep learning-based image fusion methods have been widely studied in the field of panchromatic sharpening. These methods are typically based on neural network structures that automatically extract spatial structure features and spectral representation features from data, improving image spatial resolution while preserving spectral consistency as much as possible. Therefore, deep learning methods achieve a certain balance between spatial detail enhancement and spectral information preservation.

[0012] However, most existing deep learning-based panchromatic sharpening methods generally assume that the input images are under ideal imaging conditions and do not explicitly consider image degradation factors that may be introduced during remote sensing image acquisition, such as noise and blur effects caused by atmospheric disturbances, platform jitter, or unstable sensor response. In practical applications, these factors are common in remote sensing imaging and can cause image details to be blurred or edges to be unclear, affecting the performance of subsequent fusion models and reducing their robustness.

[0013] To address these issues, some technical solutions attempt to introduce an image deblurring preprocessing module before image fusion to weaken the interference of noise and blur on the fusion results. However, this two-stage processing scheme typically has the following two problems: First, the image deblurring process and the subsequent fusion task are independent of each other, lacking consistency. When the deblurring degree is insufficient, the residual blur areas in the image will interfere with the effective extraction of spatial details, thereby reducing the quality of the fused image. When the deblurring degree is too high, it may introduce structural artifacts or cause spectral distortion, affecting the overall authenticity and consistency of the fused image.

[0014] Secondly, due to the different imaging characteristics of PAN and MS images, the specific forms of blur degradation they suffer during imaging may differ. If independent deblurring is performed on both, the lack of collaborative constraint mechanisms may lead to inconsistencies in the degree of detail enhancement between the two types of images, exacerbating the mismatch between spatial and spectral features and affecting the spectral-spatial consistency of the fusion results.

[0015] How to effectively and appropriately remove the blur in PAN and MS images, control the degree of deblurring, and effectively extract spatial and spectral information from PAN and MS images and effectively fuse them through deep learning is the key to achieving high-quality remote sensing images. SUMMARY

[0016] The present application proposes an end-to-end joint optimization framework to address the performance degradation of existing panchromatic sharpening methods in noisy and blurred scenarios. The deblurring and multispectral-panchromatic image fusion tasks are unified in a three-stage strategy of coarse fusion-deblurring-fine fusion, achieving a dynamic balance between spatial detail recovery and spectral fidelity. The core innovation of this method is the construction of a dual-path feature mutual enhancement network, which effectively solves the artifact diffusion and spectral distortion problems of existing methods under complex degradation conditions through a progressive optimization mechanism guided by cross-modal attention.

[0017] The technical scheme adopted by the present application is: a panchromatic sharpening joint optimization method based on three-stage progressive fusion, which realizes high-quality fusion of multispectral images and panchromatic images through an innovative three-stage processing flow. First, in the data preprocessing and coarse fusion stage, the system performs multi-modal feature extraction and preliminary fusion on the input low-resolution multispectral images (LR-MS) and high-resolution panchromatic images (PAN). Specifically, first, a coarse fusion network (PNN) is used to realize cross-modal feature interaction to generate coarse fusion feature maps with rich spectral-spatial information. Secondly, in the adaptive deblurring and detail recovery stage, the system optimizes the coarse fusion features through a cascaded PLKBlock module. This stage innovatively uses a partial large kernel convolution (PLKConv2d) strategy to expand the receptive field while maintaining computational efficiency, effectively capturing long-range spatial dependencies. In combination with the local texture enhancement capability of the channel mixer (DCCM) and the dynamic feature selection mechanism of the element-level attention (EA), the image detail recovery quality is significantly improved, especially in handling motion blur and Gaussian noise degradation. Finally, in the spectral-spatial fine-grained fusion optimization stage, the system introduces a multi-scale feature pyramid structure to refine the features through a spectral multi-scale fusion module (SMFA). This module uses a multi-branch parallel processing architecture, combining the multi-scale feature extraction capability of the atrous convolution and the dynamic gating feature selection mechanism, to optimize the spatial details while maintaining spectral properties. A loss function suitable for this network is designed, and finally trained on the satellite GF2 dataset. The entire processing flow realizes the coordinated optimization of each stage through end-to-end training, and finally outputs a fusion image with high spatial resolution and accurate spectral properties. The method comprises the following steps: Step 1, obtaining a multispectral image group; Step 2, obtaining a panchromatic image group; Step 3, constructing a three-stage fusion network for panchromatic sharpening, including a coarse fusion module, a feature enhancement module, and a fine fusion module; The coarse fusion module is used to coarsely fuse the panchromatic image and the multispectral image, the feature enhancement module is a multi-layer stacked partial large convolution block PLKBlock, which is used to extract local texture information and long-range dependencies, and the fine fusion module is used to optimize the spectral-spatial features of the features extracted by the multi-layer PLKBlock, and finally realizes high-resolution reconstruction through feature repeated upsampling and convolution layers; Step 4, training the three-stage fusion network combined with the loss function, and realizing the target panchromatic sharpening using the trained network.

[0018] Further, in step 1, a low-resolution multispectral image of a target scene is collected, containing multiple spectral band information; the multispectral image is preprocessed, including radiation correction, geometric correction and registration; spectral features of the multispectral image are extracted, including spectral reflectivity, spectral angle and spectral gradient features; the above features are combined to form a multispectral image feature map group, which is used as the input of the coarse fusion network; In step 2, a high-resolution panchromatic image of the same target scene is collected; the panchromatic image is preprocessed, including denoising, enhancement and edge extraction; spatial features of the panchromatic image are extracted, including edge intensity, texture features and structure tensor; multi-scale features of the panchromatic image are calculated, including Gaussian pyramid and wavelet transform coefficients; the above features are combined to form a panchromatic image feature map group, which is used as the input of the coarse-grained fusion network. Further, the coarse fusion module includes two symmetrical input branches, which process the multispectral image and the panchromatic image respectively, and the preliminary feature after fusion is represented as: ; wherein PNN represents the coarse fusion module, represents the up-sampled multispectral image, represents the panchromatic image; the coarse fusion module includes: an up-sampling module, which up-samples the low-resolution multispectral image to a predetermined high-resolution scale; a first convolutional layer, which extracts features from the up-sampled multispectral image and the panchromatic image respectively to obtain corresponding multispectral feature maps and panchromatic feature maps; a fusion convolutional module, which splices the extracted multispectral feature maps and panchromatic feature maps in channel dimension or spatial dimension; a joint convolutional layer, which further fuses and transforms the spliced features to output joint feature maps.

[0019] Further, each PLKBlock includes: a dual-channel convolutional mixing module DCCM, a large kernel convolution operation PLKConv2d on part of the channels, an element-level attention module EA and a one-dimensional convolution refining module and group normalization operation.

[0020] Further, the structure of the DCCM module includes: a channel expansion convolutional layer, which expands the input feature dimension from C to 2C; a nonlinear activation unit, which maps the features using a Mish function; a channel compression convolutional layer, which compresses the feature dimension from 2C to the original C dimension; a residual connection, which is used to enhance the channel dimension while maintaining the stability of the features; The structure of the PLKConv module includes: a channel division module, which divides the input channels into the first pdim channels and the remaining channels; a large kernel convolution operation, which only applies a 7x7 or 11x11 deep convolution operation to the first pdim channels to expand the receptive field; a channel splicing module, which combines the processed first pdim channels with the unprocessed channels in the channel dimension; a convolutional fusion layer, which normalizes and activates the combined features to generate output features; The EA attention module comprises: a spatial convolution layer, which performs 3*3 or 5*5 convolution on input features to generate an attention map; an activation function layer, which applies a Sigmoid activation function to the convolution result to generate normalized attention weights; and an element-level multiplication layer, which element-wise multiplies the input feature map and the attention weights to achieve significant enhancement of different spatial regions.

[0021] Further, the PLKConv2d divides the input channels, applies convolution operation on only part of the channels to reduce the calculation overhead, and adopts different forward strategies in the training and inference stages respectively, and the mathematical expression is: Training stage

[0022] Inference stage:

[0023] Wherein X represents the output of the partial convolution module, represents the channel subset selected for large convolution operation, is the reserved channel.

[0024] Further, the fine fusion module comprises: a channel splitting unit, which splits the input feature map into two sub-features X and Y along the channel direction; a spatial perception sub-module, which applies down-sampling pooling, variance estimation and spatial enhancement operations to the X branch; a long-range modeling sub-module, which applies a deep multi-layer perception structure to the Y branch for modeling long-range dependencies; and a feature fusion unit, which performs weighted fusion on the outputs of the two sub-branches and applies convolution operation to generate output features; wherein the spatial perception sub-module specifically comprises: average pooling and maximum pooling operations, which compress the X feature map to obtain global semantic information; a variance perception layer, which estimates the variability of local features and generates a modulation factor; an optional deep convolution layer or SE attention structure, which is used to improve the spatial perception ability; and an up-sampling operation, which restores the enhanced features to the original spatial scale; the long-range modeling sub-module comprises: a flattening operation, which expands the spatial dimension into a sequence form; a multi-layer perception structure, which performs nonlinear transformation on each position and enhances feature correlation; position encoding and residual connection, which are used to preserve spatial information and feature consistency.

[0025] Further, a hybrid loss function is designed, including spectral loss, spatial loss and perceptual loss; the spectral loss calculates the spectral difference between the fusion result and the reference image using mean square error; the spatial loss evaluates the spatial detail preservation using structural similarity index; and the perceptual loss uses high-level features extracted by a pre-trained VGG network for constraint.

[0026] Further, the network is trained using the ADAM optimizer adaptive optimization algorithm, and is verified on multiple datasets.

[0027] The application also provides a panchromatic sharpening joint optimization system based on three-stage progressive fusion, comprising the following units: A multispectral image group acquisition unit is configured to acquire multispectral prior image group input: collect low-resolution multispectral images (LR-MS) of a target scene, and perform preprocessing such as radiation correction, geometric correction and image registration; extract multidimensional spectral features such as spectral reflectance, spectral angle and spectral gradient to form a multispectral image feature group as input of the coarse fusion network.

[0028] A panchromatic image group acquisition unit is configured to acquire panchromatic prior image group input: collect high-resolution panchromatic images (PAN) of a target scene, and perform preprocessing such as denoising, enhancement and edge extraction; extract spatial features such as edge intensity, texture and structure tensor, and form a multi-level panchromatic image feature group through multi-scale modeling (such as Gaussian pyramid and wavelet decomposition) as input of the coarse fusion network.

[0029] A joint network construction unit is configured to construct a three-stage fusion network for panchromatic sharpening, comprising a coarse fusion module, a feature enhancement module and a fine fusion module. The coarse fusion module is configured to perform coarse-grained level fusion of the panchromatic image and the multispectral image, the feature enhancement module is a multi-layer stacked partial large convolution block (PLKBlock) configured to extract local texture information and long-range dependency, and the fine fusion module is configured to perform spectral-spatial feature optimization on the features extracted by the multi-layer PLKBlock, and finally realize high-resolution reconstruction through feature repeated up-sampling and convolution layers. A reconstruction unit is configured to train the three-stage progressive fusion network in combination with a loss function, and to realize panchromatic sharpening of the multispectral image and the panchromatic image using the trained network.

[0030] Compared with the prior art, the application has the following advantages and beneficial effects: The application provides a three-stage progressive fusion-based panchromatic sharpening joint optimization method, which innovatively unifies multispectral-panchromatic image fusion and image deblurring in an end-to-end optimization framework, significantly improving the fusion robustness and image quality in noisy and blurred scenes.

[0031] Firstly, the PNN network is used to complement cross-modal features, effectively improving the coupling expression ability of spectral and spatial information. Compared with traditional methods that separately sharpen or fuse images, the method completes feature-level alignment in the coarse fusion stage, reducing information redundancy and mismatching problems.

[0032] Secondly, the application designs a cascaded PLKBlock module, which combines a large receptive field partial large kernel convolution (PLKConv2d) and a channel mixer (DCCM) to jointly model long-range and local texture dependencies of an image, and introduces an element-level attention mechanism (EA) to realize dynamic weight distribution, effectively suppress the diffusion of blur area artifacts while enhancing key structural details. Experimental results show that the design still maintains stable image reconstruction performance under complex degradation conditions such as Gaussian blur, motion blur and noise interference.

[0033] Thirdly, in the fine-grained fusion stage, the application constructs a spectral-spatial multi-scale fusion module (SMFA) to realize dual protection of spectral consistency and spatial definition through multi-branch hollow convolution and dynamic gating mechanism. The SMFA module further optimizes the performance of the fused image in edge transition, texture continuity and local color fidelity, and significantly outperforms traditional sharpening methods based on pyramid or single-scale fusion.

[0034] In addition, the method designs a joint loss function suitable for multi-task target, and trains and verifies on real satellite remote sensing data (such as GF2). Experimental results show that the application is superior to existing deep learning-based panchromatic sharpening methods in terms of spatial resolution enhancement, spectral fidelity, structure restoration and other evaluation indicators. BRIEF DESCRIPTION OF DRAWINGS

[0035] Figure 1 is a comparison chart of simulation fusion results of different deep learning methods on QB satellite, from left to right are panchromatic image, HOU method fusion result, XU method fusion result, the method result, and the true value chart.

[0036] Figure 2 is a coarse fusion network structure diagram.

[0037] Figure 3 is a feature fusion module structure diagram.

[0038] Figure 4 is a three-stage fusion network structure diagram, PNN: coarse fusion network, PLKblock: deep PLK feature extraction module, SMFA: fine fusion module.

[0039] Figure 5 is a comparison chart of simulation fusion results of different deep learning methods on GF2 satellite, from left to right are panchromatic image, HOU method fusion result, XU method fusion result, the method result, and the true value chart.

[0040] Figure 6 is a comparison chart of real fusion results of different deep learning methods on GF2 satellite, from left to right are HOU method fusion result, XU method fusion result, the method result, and panchromatic image. DETAILED DESCRIPTION

[0041] In order to facilitate ordinary technicians in this field to understand and implement the present invention, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0042] This invention mainly aims at the application needs of high-precision remote sensing image fusion and reconstruction, and proposes a new three-stage panchromatic sharpening joint optimization method based on deep learning. The dual-path mutual enhancement feature extraction structure is used to decompose and represent the input low-resolution multispectral image (LR-MS) and high-resolution panchromatic image (PAN), and the spectral and spatial features are extracted respectively. By introducing the PNN module, the deep interaction of multimodal features is realized, and the coupling expression ability of spectral and structural information is effectively improved. On this basis, the PLKBlock module is further used to deblur and restore details of the coarse fusion features, and the large kernel convolution and channel mixing mechanism are combined to enhance the network receptive field and local detail modeling capabilities. Finally, the multi-scale fusion features are finely optimized through the SMFA module, which significantly improves the spatial clarity and spectral fidelity of the fused image, and realizes the accurate reconstruction of high-quality remote sensing images. The multispectral image sharpening results after fusing the panchromatic image are as follows. Figure 1 The method specifically comprises the following steps: Step 1: Obtain a multispectral image set: Collect a low-resolution multispectral image (LR-MS) of the target scene, containing information from multiple spectral bands; preprocess the multispectral image, including radiometric correction, geometric correction, and registration; extract the spectral features of the multispectral image, including spectral reflectance, spectral angle, and spectral gradient features; combine these features to form a multispectral image feature set, which serves as the input to the first branch of the coarse fusion network; Step 2: Obtain a panchromatic image set: Collect high-resolution panchromatic images (PAN) of the same target scene; perform preprocessing on the panchromatic image, including denoising, enhancement, and edge extraction; extract spatial features of the panchromatic image, including edge intensity, texture features, and structure tensor; calculate multi-scale features of the panchromatic image, including Gaussian pyramid and wavelet transform coefficients; combine the above features to form a panchromatic image feature set, which serves as the input to the second branch of the coarse fusion network; Step 3: construct a three-stage fusion network for full color sharpening. Figure 2The coarse fusion network shown realizes the double-branch feature interaction of multispectral and panchromatic images, then performs multi-scale feature extraction through cascaded PLKBlock modules, each PLKBlock includes a channel mixer (DCCM) and a partial large kernel convolution (PLKConv2d); the network introduces an element-level attention mechanism (EA) to realize adaptive feature fusion, and uses an SMFA module to optimize spectral-spatial features; finally, high-resolution reconstruction is realized through feature repeated upsampling and convolution layers (pixelshuffle operation), while the spectral characteristics are maintained; Step 4, training the three-stage fusion network combined with the loss function, and realizing the panchromatic sharpening of the target using the trained network; the three-stage fusion network is trained combined with the loss function: a hybrid loss function is designed, including spectral loss, spatial loss and perceptual loss; the spectral loss calculates the spectral difference between the fusion result and the reference image using mean square error; the spatial loss evaluates the spatial detail preservation using structural similarity index (SSIM); the perceptual loss is constrained using high-level features extracted by a pre-trained VGG network; the network is trained using the ADAM optimizer adaptive optimization algorithm, and is verified on datasets such as GF2, WorldView-3 and QuickBird.

[0043] Further, a panchromatic sharpening joint optimization neural network based on three-stage progressive fusion, the overall structure of which is composed of a coarse fusion network, a feature enhancement module and a fine fusion module, is used to process multi-source image input and realize high-resolution image reconstruction and fusion.

[0044] Further, the coarse fusion network adopts a spectral-spatial complementary fusion mechanism and is realized based on a convolutional neural network, and is used to coarsely fuse a panchromatic image and a multispectral image. The coarse-grained fusion module includes two symmetrical input branches, which respectively process the multispectral image and the panchromatic image, and the fused preliminary feature is represented as: . Wherein represents the up-sampled multispectral image, represents the panchromatic image. As shown in the accompanying Figure 2 , the coarse fusion network includes: an up-sampling module, which up-samples a low-resolution multispectral image MS to a predetermined high-resolution scale for subsequent input of the enhancement module; a first convolutional layer, which respectively extracts features from the up-sampled multispectral image and the PAN image to obtain corresponding multispectral feature maps and panchromatic feature maps; a fusion convolutional module, which splices the extracted multispectral feature maps and panchromatic feature maps in the channel dimension or the spatial dimension; a joint convolutional layer, which further fuses and transforms the spliced features to output a joint feature map; Furthermore, the output features of the above coarse fusion network are input into the feature enhancement module as a multi-layer stacked partial large convolution block (PLKBlock), which is used to extract local texture information and long-range dependencies. Each PLKBlock consists of the following four parts: 1) Doubled Convolutional Channel Mixer (DCCM) is used to enhance the interaction between channels. Figure 3 As shown in the figure, the structure of the DCCM module includes: a channel expansion convolution layer, which expands the input feature dimension from C to 2C; a nonlinear activation unit, which uses the Mish function for feature mapping; a channel compression convolution layer, which compresses the feature dimension from 2C to the original C dimension; and a residual connection, which is used to enhance the channel dimension while maintaining feature stability.

[0045] 2) The large convolution kernel operation PLKConv2d of some channels is used to apply convolution with a large receptive field on some channels to enhance the long-range feature modeling capability. Figure 3 As shown in the figure, the structure of the PLKConv module includes: a channel division module, which divides the input channels into the first pdim channels and the remaining channels; a large kernel convolution operation, which applies a deep convolution operation such as 7×7 or 11×11 only to the first pdim channels to expand the receptive field; a channel splicing module, which merges the processed first pdim channels with the unprocessed channels in the channel dimension; a convolution fusion layer, which normalizes and activates the merged features to generate output features.

[0046] 3) Element-wise Attention (EA) module, used to enhance feature response areas. The EA attention module includes: a spatial convolution layer, which performs 3×3 or 5×5 convolution on the input features to generate attention maps; an activation function layer, which applies a sigmoid activation function to the convolution results to generate normalized attention weights; and an element-wise multiplication layer, which multiplies the input feature map by the attention weights element by element to enhance the saliency of different spatial regions.

[0047] 4) One-dimensional convolution refinement module and group normalization operation are used to fuse and standardize features.

[0048] Furthermore, the output of the PLKBlock is in the form of residual connection:

[0049] in Indicates the Feature representation of the layer.

[0050] Further, the PLKConv2d module divides the input channels, applies convolution operation only on part of the channels to reduce computational overhead, and adopts different forward strategies in training and inference stages respectively. The mathematical expression is: Training stage

[0051] Inference stage:

[0052] where X represents the output of the partial convolution module, represents the subset of channels selected for large convolution operation, is the reserved channel.

[0053] Further, the EA module generates element-level attention weight map using 3x3 convolution plus Sigmoid activation, and applies it to the input feature map for element-wise weighting: . represents the input feature of the EA module, is the output feature of the EA module, represents the sigmod activation function.

[0054] Further, the spectral multi-scale feature aggregation (SMFA) module uses multi-scale feature aggregation to fuse the features extracted by multiple PLKBlock layers, improving the ability to express details, which is expressed as:

[0055] The SMFA module includes: a channel splitting unit that splits the input feature map into two sub-features X and Y along the channel direction; a spatial perception sub-module that applies down-sampling pooling, variance estimation and spatial enhancement operations to the X branch; a long-range modeling sub-module that applies a deep multi-layer perceptron (DMlp) structure to the Y branch to model long-range dependencies; and a feature fusion unit that weights and fuses the outputs of the two sub-branches and applies convolution operation to generate output features. The spatial perception sub-module specifically includes: average pooling and maximum pooling operations to compress the X feature map to obtain global semantic information; a variance perception layer to estimate the variability of local features and generate a modulation factor; an optional deep convolution layer or SE attention structure to enhance spatial perception ability; and an up-sampling operation to restore the enhanced features to the original spatial scale. The long-range modeling sub-module includes: a flattening operation to expand the spatial dimension into a sequence form; a multi-layer perceptron structure to perform non-linear transformation on each position and enhance feature correlation; position encoding and residual connection to preserve spatial information and feature consistency.

[0056] Further, as shown in Figure 4 the entire network structure is based on multi-source image input (including PAN image, low-resolution MS image, up-sampled MS image), through coarse-grained fusion, deep PLK feature extraction, attention correction (i.e. EA attention module), SMFA fine fusion, etc. mechanisms, finally output high-quality reconstructed image: ; Among them, is the panchromatic image, is the high-pass filtered result image of the panchromatic image pan, is the up-sampled multispectral image, is the low-resolution image, f is the three-stage progressive fusion network proposed above, is the result output by the fusion network.

[0057] Based on the above steps, the fusion results of multispectral images and panchromatic images are obtained. In order to compare with other methods, we use HOU(HOU J, CAO Q, RAN R, et al. Bidomain Modeling Paradigm for Pansharpening[C / OL] / / Proceedings of the 31st ACM International Conference on Multimedia. Ottawa ON Canada: ACM, 2023: 347-357[2025-01-16].), XU(XU S, ZHANG J, ZHAO Z, et al. Deep Gradient Projection Networks for Pan-sharpening[C / OL] / / 2021 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, TN, USA: IEEE, 2021: 1366-1375[2024-10-28].) method, and our method on GF2 satellite data. The results are shown in FIGS. 8-10. Figure 5 -Appendix Figure 6 as shown.

[0058] To quantitatively evaluate the quality of the fused image, we introduce the peak signal-to-noise ratio (PSNR), structural similarity (SSIM), spectral angle mapping error (SAM), root mean square error (RMSE), correlation coefficient (CC), relative dimensionless global error (ERGAS), and universal image quality index (UIQC) as evaluation indicators. The larger the values of PSNR, SSIM, and UIQC, the higher the image fusion quality and the more complete the preservation of structure and texture information. The smaller the values of SAM, RMSE, and ERGAS, the smaller the spectral distortion and the closer the image to the true high-resolution image. CC represents spectral consistency, and the closer the value to 1, the more sufficient the preserved spectral information. According to the experimental results, the fused image performs well in multiple indicators, verifying the effectiveness of the method in the task of multispectral image and panchromatic image fusion. The quantitative comparison results on the object-level dataset are as follows: Table 1 Quantitative analysis of different fusion methods on the GF2 dataset

[0059] The quantitative indicator results show that the method proposed in the present application is superior to existing methods in multiple mainstream quantitative indicators, showing stronger spatial detail recovery capability and spectral information preservation capability. In typical remote sensing image fusion tasks, whether in terms of image structural similarity, spectral fidelity, or error control and correlation, the method shows high reconstruction quality. Especially in the face of complex scenes with blur and noise interference, the method can effectively suppress artifact diffusion, improve the clarity and authenticity of the fused image, and has better robustness and generalization ability.

[0060] On the other hand, the embodiment of the present application provides a panchromatic sharpening joint optimization system based on three-stage progressive fusion, which includes the following units: A multispectral image group acquisition unit is used to acquire multispectral prior image group input: collect low-resolution multispectral images (LR-MS) of the target scene, and perform radiation correction, geometric correction, image registration, and other preprocessing; extract spectral reflectance, spectral angle, and spectral gradient, and other multi-dimensional spectral features to form a multispectral image feature group as input for the coarse fusion network.

[0061] A panchromatic image group acquisition unit is used to acquire panchromatic prior image group input: collect high-resolution panchromatic images (PAN) of the target scene, and perform denoising, enhancement, edge extraction, and other preprocessing; extract edge intensity, texture, and structure tensor, and other spatial features, and form a multi-level panchromatic image feature group through multi-scale modeling (such as Gaussian pyramid, wavelet decomposition) as input for the coarse fusion network.

[0062] The joint network construction unit is configured to construct a three-stage fusion network for panchromatic sharpening, including a coarse fusion module, a feature enhancement module, and a fine fusion module; The coarse fusion module is configured to perform coarse-grained level fusion on the panchromatic image and the multispectral image, the feature enhancement module is a multi-layer stacked partial large convolution block (PLKBlock) configured to extract local texture information and long-range dependency, and the fine fusion module is configured to perform spectral-spatial feature optimization on the features extracted by the multi-layer PLKBlock, and finally realize high-resolution reconstruction through feature repeated upsampling and convolution layers. The reconstruction unit is configured to train the three-stage progressive fusion network based on a loss function, and realize panchromatic sharpening of the multispectral image and the panchromatic image by using the trained network.

[0063] The specific implementation manners of each unit are the same as those of each step, and the embodiments of the present application are not described herein.

[0064] It should be understood that the parts not described in detail in the specification are all prior art.

[0065] It should be understood that the above description of the embodiments is relatively detailed, and therefore should not be considered as limiting the scope of patent protection of the present application. Those skilled in the art can make substitutions or modifications without departing from the scope of protection of the claims of the present application, and all fall within the scope of protection of the present application. The scope of protection of the present application should be subject to the appended claims.

Claims

1. A degradation-aware pan-sharpening method based on three-stage progressive fusion, characterized in that: The steps include: Step 1: Obtain a multispectral image group; Step 2, obtaining a full-color image group; Step 3: Construct a three-stage fusion network for pan-sharpening, including a coarse fusion module, a feature enhancement module, and a fine fusion module; The coarse fusion module is used to perform coarse-grained fusion of the panchromatic image and the multispectral image. The feature enhancement module is a multi-layer stack of partially large convolution blocks PLKBlock, which is used to extract local texture information and long-range dependencies. The fine fusion module is used to optimize the spectral-spatial features of the features extracted by the multi-layer PLKBlock, and finally achieve high-resolution reconstruction through repeated feature upsampling and convolution layers. Step 4: Combine the loss function to train the three-stage fusion network, and use the trained network to achieve full color sharpening of the target.

2. The degradation-aware pan-sharpening method based on three-stage progressive fusion according to claim 1, characterized in that: In step 1, a low-resolution multispectral image of the target scene is collected, which contains information of multiple spectral bands; the multispectral image is preprocessed, including radiometric correction, geometric correction, and registration; the spectral features of the multispectral image are extracted, including spectral reflectance, spectral angle, and spectral gradient features; the above features are combined to form a multispectral image feature map group as the input of the coarse fusion network; In step 2, a high-resolution panchromatic image of the same target scene is collected; the panchromatic image is preprocessed, including denoising, enhancement, and edge extraction; the spatial features of the panchromatic image are extracted, including edge intensity, texture features, and structure tensor; the multi-scale features of the panchromatic image are calculated, including Gaussian pyramid and wavelet transform coefficients; the above features are combined to form a panchromatic image feature map group as the input of the coarse-grained fusion network.

3. The degradation-aware pan-sharpening method based on three-stage progressive fusion according to claim 1, characterized in that: The coarse fusion module includes two symmetrical input branches, which process multispectral images and panchromatic images respectively. The preliminary features after fusion are expressed as: ; Among them, PNN represents the coarse fusion module, represents the upsampled multispectral image, Represents a full-color image; the coarse fusion module includes: an upsampling module, which upsamples the low-resolution multispectral image to a predetermined high-resolution scale; the first convolution layer, which extracts features from the upsampled multispectral image and the full-color image respectively to obtain corresponding multispectral feature maps and full-color feature maps; the fusion convolution module, which splices the extracted multispectral feature maps and full-color feature maps in the channel dimension or spatial dimension; the joint convolution layer, which further fuses and transforms the spliced ​​features and outputs a joint feature map.

4. The degradation-aware pan-sharpening method based on three-stage progressive fusion according to claim 1, wherein: Each PLKBlock includes: a dual-channel convolutional mixing module DCCM, a large convolution kernel operation PLKConv2d for some channels, an element-level attention module EA, and a one-dimensional convolution refinement module with group normalization operations.

5. The degradation-aware pan-sharpening method based on three-stage progressive fusion according to claim 4, characterized in that: The DCCM module consists of a channel expansion convolution layer that expands the input feature dimension from C to 2C; a nonlinear activation unit that uses the Mish function for feature mapping; a channel compression convolution layer that compresses the feature dimension from 2C to the original C dimension; and a residual connection that is used to enhance the channel dimension while maintaining feature stability. The structure of the PLKConv2d module includes: a channel partitioning module that divides the input channels into the first pdim channels and the remaining channels; a large kernel convolution operation that applies a 7×7 or 11×11 depthwise convolution operation only to the first pdim channels to expand the receptive field; a channel splicing module that merges the processed first pdim channels with the unprocessed channels in the channel dimension; and a convolutional fusion layer that normalizes and activates the merged features to generate output features. The EA attention module includes: a spatial convolution layer that performs 3×3 or 5×5 convolution on the input features to generate an attention map; an activation function layer that applies a Sigmoid activation function to the convolution result to generate normalized attention weights; and an element-wise multiplication layer that multiplies the input feature map by the attention weights element by element to achieve saliency enhancement of different spatial regions.

6. The degradation-aware pan-sharpening method based on three-stage progressive fusion according to claim 4, characterized in that: PLKConv2d splits the input channels and applies convolution operations only on some channels to reduce computational overhead. It also uses different forward strategies in the training and inference stages. Its mathematical expression is: Training phase Reasoning stage: Where X represents the output of the partial convolution module, represents the subset of channels selected for large convolution operations, To reserve the channel.

7. The degradation-aware pan-sharpening method based on three-stage progressive fusion according to claim 1, characterized in that: The fine fusion module includes: a channel splitting unit that splits the input feature map into two sub-features X and Y along the channel direction; a spatial perception sub-module that applies downsampling pooling, variance estimation, and spatial enhancement operations to the X branch; a long-range modeling sub-module that applies a deep multi-layer perceptron structure to the Y branch to model long-range dependencies; a feature fusion unit that performs weighted fusion on the outputs of the two sub-branches and applies a convolution operation to generate output features; the spatial perception sub-module specifically includes: average pooling and maximum pooling operations to compress the X feature map to obtain global semantic information; a variance perception layer to estimate local feature variability and generate modulation factors; an optional deep convolutional layer or SE attention structure to improve spatial perception capabilities; an upsampling operation to restore the enhanced features to the original spatial scale; and the long-range modeling sub-module includes: a flattening operation to expand the spatial dimensions into a sequence form. The multi-layer perceptron structure performs nonlinear transformation on each position and enhances feature correlation; position encoding and residual connection are used to preserve spatial information and feature consistency.

8. The degradation-aware pan-sharpening method based on three-stage progressive fusion according to claim 1, characterized in that: A hybrid loss function is designed, including spectral loss, spatial loss and perceptual loss; the spectral loss uses the mean square error to calculate the spectral difference between the fusion result and the reference image; the spatial loss uses the structural similarity index to evaluate the preservation of spatial details; the perceptual loss is constrained by the high-level features extracted by the pre-trained VGG network.

9. The degradation-aware pan-sharpening method based on three-stage progressive fusion according to claim 1, characterized in that: The network was trained using the ADAM optimizer adaptive optimization algorithm and validated on multiple datasets.

10. A degradation-aware pan-sharpening system based on three-stage progressive fusion, characterized in that: Includes the following units: A multispectral image group acquisition unit, used for acquiring a multispectral image group; A full-color image group acquisition unit, configured to acquire a full-color image group; A joint network construction unit for constructing a three-stage fusion network for pan-sharpening, including a coarse fusion module, a feature enhancement module, and a fine fusion module; The coarse fusion module is used to perform coarse-grained fusion of the panchromatic image and the multispectral image. The feature enhancement module is a multi-layer stack of partially large convolution blocks PLKBlock, which is used to extract local texture information and long-range dependencies. The fine fusion module is used to optimize the spectral-spatial features of the features extracted by the multi-layer PLKBlock, and finally achieve high-resolution reconstruction through repeated feature upsampling and convolution layers. The reconstruction unit is used to train the three-stage fusion network in combination with the loss function, and use the trained network to achieve full color sharpening of the target.

Citation Information

Cited By

  • Defect detection framework based on differential convolution attention and staged feature reconstruction

    CN121074051A

  • Vegetation classification method based on space-time multi-modal deep learning

    CN121600411A

  • Image feature optimization fusion method, panchromatic sharpening method and product

    CN121904538A