Multispectral image panchromatic sharpening method based on plug-and-play gradient feature guidance fusion
By using a plug-and-play gradient feature-guided fusion method, the problems of insufficient adaptability and gradient information utilization in multispectral image pancolor sharpening methods are solved, achieving efficient high-frequency detail capture and cross-modal structure alignment, thus improving the performance of multispectral image pancolor sharpening.
Patent Information
- Application Number
- CN202511103160.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-08-07
AI Technical Summary
Existing multispectral image panchromatic sharpening methods rely on fixed design patterns, making it difficult to adapt to different datasets or tasks. They also fail to make sufficient use of gradient information and lack flexible modular frameworks, resulting in inadequacies in capturing high-frequency details and cross-modal structure alignment.
A plug-and-play gradient feature-guided fusion method is adopted. By extracting gradient features, guiding fusion and optimizing loss, combined with attention mechanism and learnable weights, multispectral image panchromatic sharpening is achieved, thereby improving reconstruction performance.
It effectively captures high-frequency details, achieves cross-modal structure alignment, improves pan-color sharpening performance, and generates high-quality, high-resolution multispectral images.
Smart Images

Figure CN120976059A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of multi-spectral image panchromatic sharpening, and more specifically relates to a multi-spectral image panchromatic sharpening (Pansharpening) method based on plug-and-play gradient feature guided fusion. BACKGROUND
[0002] Multi-spectral images (MSI) play an important role in earth science applications such as land cover classification, vegetation monitoring, and water body identification due to their rich spectral information. However, due to the limitations of satellite sensor performance and on-board processing capabilities, the acquired multi-spectral images usually have low spatial resolution, making it difficult to capture fine-grained information of complex surfaces. In contrast, panchromatic images (PAN) have higher spatial resolution and can capture structural details more clearly, but their spectral information is relatively limited. In order to combine the advantages of both, researchers have conducted extensive research to develop multi-spectral remote sensing image panchromatic sharpening technology, which fuses the high spatial resolution of panchromatic images with the rich spectral information of multi-spectral images to generate high-resolution multi-spectral images.
[0003] The research focus of existing deep learning-based multi-spectral image panchromatic sharpening methods mainly includes two categories. One class usually adopts a dual-branch architecture to enhance the feature expression ability of specific modal, focusing on the effective reconstruction of spectral and spatial information. These methods achieve decoupling and accurate recovery of spectral fidelity and spatial details by designing separate processing paths for panchromatic images (PAN) and multi-spectral images (MSI). The other class focuses on efficient fusion strategies for MSI and PAN, achieving cross-modal interaction and feature aggregation by introducing attention mechanisms and Mamba structures. In addition, convolution-based optimization techniques are often used as effective baseline solutions to improve the expression ability of structural priors and spatial alignment accuracy.
[0004] However, traditional multispectral image panchromatic sharpening methods usually rely on fixed design patterns, which makes them difficult to adapt to different datasets or tasks. To solve this problem, some recent studies introduce Plug-and-Play (PnP) priors into the multispectral image panchromatic sharpening process, aiming to combine model-based regularization with learned priors. Tao et al. proposed a PnP-based residual detail injection framework, and Shu et al. designed a dual-domain attention mechanism combining Transformers and CNNs. In addition, there are also studies that apply PnP-based variational frameworks to jointly handle image misalignment and fusion problems, further verifying the potential of PnP priors in improving the generalization and adaptability of multispectral image panchromatic sharpening tasks. Recently, a representative work is the adaptive dual-weighted mechanism (ADWM) proposed by Huang et al., which models the redundancy and heterogeneity between and within feature layers using covariance matrices, thereby reducing information overlap and enhancing feature discriminability. Although ADWM is not explicitly designed as a PnP block, it functions similarly and serves as an effective enhancement component integrated into the sharpening process.
[0005] The use of structural priors in existing PnP-based multispectral image panchromatic sharpening methods is still limited, especially the structural guidance based on gradient information has not been fully explored. In addition, few current multispectral image panchromatic sharpening methods provide a modular and flexible framework to seamlessly integrate gradient cues in different panchromatic sharpening architectures.
[0006] The existing multispectral image panchromatic sharpening methods have the following disadvantages: 1. They rely on fixed design patterns, which makes them difficult to adapt to different datasets or tasks; 2. The use of structural priors is still limited, especially the structural guidance based on gradient information has not been fully explored; 3. There is currently a lack of a flexible and modular framework that can seamlessly integrate gradient priors into existing architectures, which is deficient in capturing high-frequency details and achieving cross-modal structural guidance alignment. SUMMARY
[0007] The purpose of the present application is to overcome the shortcomings of the prior art and provide a multispectral image panchromatic sharpening method based on plug-and-play gradient feature guided fusion to fully exploit gradient information, capture high-frequency details, and achieve cross-modal structural guidance alignment, flexibly and efficiently realize multispectral image panchromatic sharpening, and improve panchromatic sharpening performance.
[0008] To achieve the above invention purposes, the multispectral image panchromatic sharpening method based on plug-and-play gradient feature guided fusion comprises the following steps:
[0009] (1) Gradient feature extraction
[0010] Based on low-resolution multispectral images MSI Panchromatic Image I PAN Extract gradient-guided features;
[0011] (2) Gradient-guided fusion
[0012] 2.1) Normalization and convolution mapping are used to obtain the query, key, and value.
[0013] The gradient-guided feature GRAD is normalized, and then passed through a GRAD convolution mapping to obtain the value V. GRAD Key K GRAD For panchromatic image I PAN After normalization, the values are then passed through a PAN convolution mapping to obtain the value V. PAN Key K PAN For multispectral images After normalization, and then through an MSI convolution mapping, the value V is obtained. MSI Key K MSI And query Q MSI ;
[0014] 2.2) Calculate attention
[0015] Value V GRAD Key K GRAD And query Q MSI Attention is obtained after passing through the attention module. GRAD Value V PAN Key K PAN And query Q MSI Attention is obtained after passing through the attention module. PAN Value V MSI Key K MSI And query Q MSI Attention is obtained after passing through the attention module. MSI ;
[0016] 2.3) Generate MSI features
[0017] Attention GRAD Attention PA Attention MSI The MSI features F are modulated using learnable weights α, β, and γ, respectively, and then aggregated through weighted summation and refined by a convolutional layer. MSI :
[0018] F MSI =Conv(Attention) GRAD⊙α+Attention PAN ⊙β+Attention MSI ⊙γ)
[0019] (3), generating high-resolution multispectral image
[0020] MSI features F MSI are input into the baseline model, and the multispectral image I pan image I PAN is gradient-guided fusion, and the fusion output is mapped to the corresponding target dimension through a convolution layer to obtain a high-resolution multispectral image
[0021] (4), loss joint optimization
[0022] The total loss function is calculated
[0023] (5), training multispectral image panchromatic sharpening system
[0024] Based on steps (1), (2), (3), the multispectral image panchromatic sharpening system is constructed, and the total loss function is used The network parameters of the multispectral image panchromatic sharpening system are updated using the gradient descent method until the set conditions are reached.
[0025] (6), multispectral image panchromatic sharpening
[0026] The multispectral image I MSI , the panchromatic image I PAN is sent into the multispectral image panchromatic sharpening system, and is processed according to steps (1), (2), (3) to obtain a high-resolution multispectral image
[0027] The purpose of the present application is achieved.
[0028] The present application is based on a plug-and-play gradient feature guided fusion multi-spectral image panchromatic sharpening method which extracts and integrates gradient guided features (GRAD) to obtain excellent recovery results in terms of structural fidelity, detail preservation and spectral consistency. The gradient feature guided fusion is designed as a highly flexible plug-and-play method which can seamlessly enhance existing architectures. Specifically, the present application explicitly models GRAD by integrating the fine spatial details of high-resolution PAN and the intrinsic spectral information of MSI. At the same time, the attention mechanism and the learnable weighting and residual connection are adopted to realize the selective aggregation and adaptive fusion of image information, thereby significantly improving the MSI quality. In addition, the GRAD is systematically analyzed from the structure, detail and spectrum, and the results are optimized through a multi-objective loss function. In this way, the gradient information is fully tapped to capture high-frequency details, realize cross-modal structure guided alignment, and flexibly and efficiently realize multi-spectral image panchromatic sharpening, thereby improving the performance of panchromatic sharpening. BRIEF DESCRIPTION OF DRAWINGS
[0029] Figure 1 is a structural schematic diagram of a multi-spectral image panchromatic sharpening system designed according to the present application based on a plug-and-play gradient feature guided fusion multi-spectral image panchromatic sharpening method.
[0030] Figure 2 is a flowchart of one specific embodiment of the present application based on a plug-and-play gradient feature guided fusion multi-spectral image panchromatic sharpening method.
[0031] Figure 3 is a schematic diagram of gradient guided fusion principle.
[0032] Figure 4 is a comparison chart of the average channel difference between the high-resolution multi-spectral image obtained by panchromatic sharpening and the reference image. DETAILED DESCRIPTION
[0033] The specific embodiments of the present application will be described below in conjunction with the accompanying drawings, so that those skilled in the art can better understand the present application. It should be particularly noted that in the following description, when the detailed description of known functions and designs may obscure the main content of the present application, these descriptions will be omitted here.
[0034] Figure 1 is a structural schematic diagram of a multi-spectral image panchromatic sharpening system designed according to the present application based on a plug-and-play gradient feature guided fusion multi-spectral image panchromatic sharpening method.
[0035] In order to solve the deficiencies of existing panchromatic sharpening methods in capturing high-frequency details and realizing cross-modal structure guided alignment, the present application constructs a multi-spectral image panchromatic sharpening system which is a flexible and efficient plug-and-play module, such as Figure 1As shown, the present application is developed around three key issues: how to effectively extract gradient information (GRAD), how to introduce it into the fusion process, and how to optimize its fusion mode to improve the reconstruction performance. Corresponding to this are the gradient feature extraction module, the gradient guided fusion module, and the loss joint optimization module (dashed line connection), and the processing process of each module is as shown in Figure 2
[0036] Figure 2 is a specific embodiment flowchart of the present application based on plug-and-play gradient feature guided fusion multi-spectral image panchromatic sharpening method.
[0037] In this embodiment, as shown in Figure 1 , 2 the present application based on plug-and-play gradient feature guided fusion multi-spectral image panchromatic sharpening method includes the following steps:
[0038] Step S1: Gradient feature extraction
[0039] Step S1.1: Shallow gradient feature extraction
[0040] In order to show the modeling of the fusion gradient feature, the shallow layer is extracted from the original gradient feature of MSI and PAN, and the structure prior contained in the panchromatic image PAN(I PAN ) and the multi-spectral image MSI(I MSI ) is modeled by an explicit modeling, and the shallow layer gradient feature is first extracted from the two modalities. These initial features highlight local edge and texture information, providing a basis for subsequent deeper abstract learning. Specifically:
[0041] The low-resolution multi-spectral image I MSI is upsampled to obtain a multi-spectral image with the same size as the panchromatic image I PAN Then the shallow layer gradient feature is extracted:
[0042]
[0043] wherein, denotes a gradient operator, such as a Sobel operator, G MSI is the shallow layer gradient feature of the multi-spectral image, and G PAN is the shallow layer gradient feature of the panchromatic image.
[0044] Step S1.2: Deep gradient feature extraction
[0045] In order to capture high-frequency representations with more semantic and discriminative power, we further input G MSI and G PAN into respective encoders f MSI (·) and f PAN Processing in (·) :
[0046] Φ MSI = f MSI (G MSI ), Φ PAN = f PAN (G PAN )
[0047] where f MSI (·) denotes the encoder of multispectral images, f PAN (·) denotes the encoder of panchromatic images, Φ MSI is the deep gradient feature of multispectral images, and Φ PAN is the deep gradient feature of panchromatic images, which represent higher-level gradient semantic information, such as complex texture primitives or robust boundary structures, and are less sensitive to noise and bias.
[0048] Step S1.3: Gradient-guided feature extraction
[0049] After obtaining the deep gradient features Φ MSI and Φ PAN , they are first concatenated to form a joint representation, and the feature interaction between them is modeled through a fusion network . Subsequently, a channel attention module is introduced to further enhance the fusion result, which can adaptively reweight the channel features according to their importance. The final enhanced gradient features are obtained through a residual connection, and this residual enhancement mechanism ensures that the fused gradient features not only maintain structural consistency with the PAN image but also maintain spectral information consistency with the MSI image. Specifically:
[0050]
[0051] where GRAD is the gradient-guided feature, concat represents channel concatenation, is the fusion network, is the channel attention module.
[0052] Step S2: Gradient-guided fusion
[0053] To better apply the gradient feature information of the fusion representation to the entire pansharpening process, the invention designs an attention module to inject spatial details into the MSI from the PAN and GRAD features, and to enhance them using the spatial context of the MSI itself, achieving selective aggregation and adaptive fusion of spatial information. This part mainly contains three paths, as shown in Figure 3 . Specifically, gradient-guided fusion includes the following steps:
[0054] Step S2.1: Normalization processing, convolution mapping to obtain query, key and value
[0055] The gradient guided feature GRAD is normalized and then respectively mapped by a GRAD convolution to obtain value V GRAD , key K GRAD , the panchromatic image I PAN is normalized and then respectively mapped by a PAN convolution to obtain value V PAN , key K PAN , the multispectral image I MSI is normalized and then respectively mapped by a MSI convolution to obtain value V MSI , key K MSI and query Q MSI .
[0056] Step S2.2: Calculate attention
[0057] Value V GRAD , key K GRAD and query Q MSI are input into the attention module to obtain attention Attention GRAD , value V PAN , key K PAN and query Q MSI are input into the attention module to obtain attention Attention PAN , value V MSI , key K MSI and query Q MSI are input into the attention module to obtain attention Attention MSI .
[0058] In the present application, the GRAD guided texture refinement selectively injects fine-grained edge and texture information from GRAD into MSI representation to enrich its high-frequency details; the PAN guided spatial detail injection adaptively selects and weights spatial detail information according to the corresponding PAN feature map row for each spatial position in the MSI feature map, thereby enhancing spatial fidelity; the spectral-context association models the inter-channel dependency and long-distance spatial context information within MSI, thereby achieving better structure perception enhancement and spectral consistency.
[0059] Step S2.3: Generate MSI feature
[0060] Attention GRAD , Attention PA , Attention MSIThe MSI features F are modulated using learnable weights α, β, and γ, respectively, and then aggregated through weighted summation and refined by a convolutional layer. MSI :
[0061] F MSI =Conv(Attention) GRAD ⊙α+Attention PAN ⊙β+Attention MSI ⊙γ)
[0062] In this invention, the gradient-guided fusion module aims to generate high-quality MSI features that possess both fine spatial detail and maintain spectral accuracy by jointly modeling the spatial information of the PAN, the texture details of the GRAD, and the spectral contextual cues of the MSI. Simultaneously, it effectively reduces spectral distortion and spatial artifacts caused by the uneven correlation between the PAN and the MSI.
[0063] Step S3: Generate high-resolution multispectral images
[0064] MSI feature F MSI Input to the baseline model, at the insertion point, and multispectral image Panchromatic Image I PAN Gradient-guided fusion is performed, and the fused output is mapped to the corresponding target dimension through a convolutional layer to obtain a high-resolution multispectral image.
[0065] Step S4: Joint Loss Optimization
[0066] Step S4.1: Calculate pixel-level reconstruction loss
[0067] Pixel-level reconstruction loss is based on the L1 norm, a widely used loss mechanism in image restoration tasks. Its goal is to minimize the absolute difference between the predicted image and the high-resolution ground truth image, thereby promoting structural similarity and maintaining consistency in the global intensity distribution. Specifically:
[0068]
[0069] in, It is a high-resolution multispectral image. The value at pixel x, It is a high-resolution true multispectral image. The value at pixel x, Ω represents the set of all pixel positions, and N is the total number of pixels.
[0070] Step S4.2: Calculate gradient detail loss
[0071] The pixel-level reconstruction loss based on L1 norm calculation is usually insensitive to high-frequency components, often leading to overly smooth image edges and loss of fine-grained texture details. As a representative of high-frequency information, image gradients can effectively capture local texture, contours, and structural changes. The introduction of gradient detail loss can make up for the shortcomings of L1 norm, explicitly enhancing the reconstruction effect of edges and textures, thereby improving the perceptual clarity and structural perception ability of the output image, especially in areas with rich spatial changes, specifically:
[0072]
[0073] wherein, represents calculating the image gradient.
[0074] Step S4.3: Calculate spectral consistency loss
[0075] For the task of panchromatic sharpening, it is also crucial to maintain consistency between spectral channels while preserving spatial details. The spectral consistency loss is used to constrain the channel relationship between the reconstructed image and the reference image, thereby reducing spectral distortion phenomena such as color shift or band misplacement, improving the realism and practicality of the reconstructed image in remote sensing applications, specifically:
[0076]
[0077] Step S4.4: Calculate total loss function
[0078] The above three loss terms are complementary in function: the pixel-level reconstruction loss is used to strengthen structure alignment, the gradient detail loss is used to enhance detail sharpness, and the spectral consistency loss is used to maintain consistency between bands. Through joint optimization, the total loss function is:
[0079]
[0080] Step S5: Train the multi-spectral image panchromatic sharpening system
[0081] Based on steps S1, S2, and S3, construct the multi-spectral image panchromatic sharpening system, according to the total loss function use gradient descent method to update the network parameters of the multi-spectral image panchromatic sharpening system until the set conditions are met.
[0082] Step S6: Multi-spectral image panchromatic sharpening
[0083] Send the multi-spectral image I MSI and the panchromatic image I PAN into the multi-spectral image panchromatic sharpening system, process according to steps S1, S2, and S3, to obtain a high-resolution multi-spectral image
[0084] Example
[0085] This invention utilizes a plug-and-play gradient feature-guided fusion method for multispectral image panchromatic sharpening. The panchromatic sharpening dataset proposed by Deng et al. was trained and tested on multiple baseline models (PanNet, FusionNet). This dataset includes WV3, QB, and GF2 satellite data. The results were evaluated using PSNR, SSIM, ERGAS, SAM, SCC, and Q2n metrics. The comparison results are shown in Table 1.
[0086]
[0087] Table 1
[0088] Table 1 shows the improvement results of the present invention on multiple baselines in the WV3, QB, and GF2 datasets. The method proposed in this invention is highlighted in bold, showing a significant improvement compared to the baseline models.
[0089] The average channel difference between the generated high-resolution multispectral image and the high-resolution real multispectral image (reference image) is visualized as follows: Figure 4 As shown, the left column of each group displays the baseline results, while the right column displays the results enhanced by the method of this invention; brighter colors indicate greater differences. The method proposed in this invention effectively reduces edge artifacts, improves the fidelity of structural details, and verifies its effectiveness in high-frequency reconstruction.
[0090] Although the illustrative specific embodiments of the present invention have been described above to enable those skilled in the art to understand the invention, it should be understood that the invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the invention as defined and determined by the appended claims, and all inventions utilizing the concept of the present invention are protected.
Claims
1. A multispectral image panchromatic sharpening method based on plug-and-play gradient feature-guided fusion, characterized in that, Includes the following steps: (1) Gradient feature extraction Based on low-resolution multispectral images MSI Panchromatic Image I PAN Extract gradient-guided features (GRAD); (2) Gradient-guided fusion 2.1) Normalization and convolution mapping are used to obtain the query, key, and value. The gradient-guided feature GRAD is normalized, and then passed through a GRAD convolution mapping to obtain the value V. GRAD Key K GRAD For panchromatic image I PAN After normalization, the values are then passed through a PAN convolution mapping to obtain the value V. PAN Key K PAN For multispectral images After normalization, and then through an MSI convolution mapping, the value V is obtained. MSI Key K MSI And query Q MSI ; 2.2) Calculate attention Value V GRAD Key K GRAD And query Q MSI Attention is obtained after passing through the attention module. GRAD Value V PAN Key K PAN And query Q MSI Attention is obtained after passing through the attention module. PAN Value V MSI Key K MsI And query Q MSI Attention is obtained after passing through the attention module. MSI ; 2.3) Generate MSI features Attention GRAD Attention PAN Attention MSI The MSI features F are modulated using learnable weights α, β, and γ, respectively, and then aggregated through weighted summation and refined by a convolutional layer. MSI : F MSI =Conv(Attention GRAD ⊙α+Attention PAN ⊙β+Attention MSI ⊙γ) (3) Generate high-resolution multispectral images MSI feature F MSI Input to the baseline model, at the insertion point, and multispectral image Panchromatic Image I PAN Gradient-guided fusion is performed, and the fused output is mapped to the corresponding target dimension through a convolutional layer to obtain a high-resolution multispectral image. (4) Joint optimization of losses Calculate the total loss function (5) Training a multispectral image panchromatic sharpening system Based on steps (1), (2), and (3), a multispectral image panchromatic sharpening system is constructed, and based on the total loss function... The network parameters of the multispectral image panchromatic sharpening system are updated using the gradient descent method until the set conditions are met. (6) Multispectral image full-color sharpening Multispectral image I MSI Panchromatic Image I PAN The image is fed into a multispectral image panchromatic sharpening system and processed according to steps (1), (2), and (3) to obtain a high-resolution multispectral image.
2. The multispectral image panchromatic sharpening method based on plug-and-play gradient feature-guided fusion according to claim 1, characterized in that, The gradient feature extraction described in step (1) is as follows: 1.1) Shallow gradient feature extraction For low-resolution multispectral images I MSI Upsampling is performed to obtain the image I (similar to the panchromatic image). PAN Multispectral images of the same size Then, shallow gradient features are extracted: in, Let G represent the gradient operator. MSI For shallow gradient features of multispectral images, G PAN This represents the shallow gradient features of a panchromatic image. 1.2) Deep gradient feature extraction F MSI =f MSI (G MSI ),F PAN =f PAN (G PAN ) Among them, f MSI (·) represents the encoder of the multispectral image, f PAN (·) represents the encoder of the panchromatic image, Φ MSI For deep gradient features of multispectral images, Φ PAN This represents the deep gradient features of a panchromatic image; 1.3) Gradient-guided feature extraction Here, GRAD stands for Gradient Guided Features, and concat indicates concatenation by channel. Indicates a converged network. This indicates the channel attention module.
3. The multispectral image panchromatic sharpening method based on plug-and-play gradient feature-guided fusion according to claim 1, characterized in that, The joint optimization of the loss described in step (4) is as follows: 4.1) Calculate pixel-level reconstruction loss in, It is a high-resolution multispectral image. The value at pixel x, It is a high-resolution true multispectral image. The value at pixel x, Ω represents the set of all pixel positions, and N is the total number of pixels; 4.2) Calculate gradient detail loss in, This indicates the calculation of image gradient; 4.3) Calculate the spectral uniformity loss 4.4) Calculate the total loss function
Citation Information
Patent Citations
Multispectral remote sensing image fusion method and device based on residual learning
CN110415199A
Unsupervised panchromatic sharpening method based on dual cycle consistency
CN117152006A
Panchromatic sharpening method for hyperspectral image with extreme resolution ratio
CN117689578A
Panchromatic sharpening method based on multi-resolution panchromatic feature guidance
CN120013808A
Remote sensing image panchromatic sharpening method and system fusing Mama and CNN under detail enhancement guidance
CN120374447A
Cited By
Hyperspectral fusion imaging method based on hierarchical gradient guidance
CN121527632A
Multi-modal correction panchromatic sharpening system and method based on task allocation method
CN121937330A
A multi-modal correction panchromatic sharpening system and method based on a task allocation method
CN121937330B