Panchromatic sharpening method based on cross-modal characteristic decomposition and recombination

By designing a dual-branch feature decomposition module and an interactive feature recombination module in the full-color sharpening method, combined with an image reconstruction module based on a reversible neural network, the information redundancy and spectral distortion problems in the span modal feature extraction and fusion process in the existing technology are solved, and more efficient feature fusion and image quality improvement are achieved.

CN120107112APending Publication Date: 2025-06-06NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510163979.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing full-color sharpening methods have shortcomings in the process of cross-modal feature extraction and fusion, and it is difficult to effectively distinguish the common information shared between modals from the unique detailed information of their respective modals, and there are problems of information redundancy and spectral distortion.

Method used

A dual-branch feature decomposition module is designed to extract low-frequency basic features and high-frequency detail features through the shallow feature extraction module, and the long-distance dependence and global contextual connection between features are enhanced through the interactive feature recombination module. The image reconstruction module based on reversible neural network is used to fusion feature to reduce spectral distortion.

Benefits of technology

The effect of cross-modal feature extraction and fusion is significantly improved, spectral distortion is reduced, spectral distortion caused by improper feature fusion is avoided, and the overall quality and detailed performance of the image is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107112A_ABST
    Figure CN120107112A_ABST
Patent Text Reader

Abstract

The invention discloses a panchromatic sharpening method based on multi-resolution panchromatic feature guidance, which comprises the following steps: firstly, designing a double-branch feature decomposition module for respectively performing feature decomposition on a panchromatic image and a multi-spectral image to extract low-frequency basic features and high-frequency detail features, and enabling the two branches to share the same shallow feature extraction module; correlation constraint is applied to the extracted basic features; then, aiming at the problem of information redundancy caused by direct serial connection and stacking of the extracted features, performing interactive fusion on the features by using an interactive feature recombination module, and enhancing a long-distance dependency relationship and a global context relationship among different modal features; and finally, an INN-based image reconstruction module is provided, and efficient fusion of the basic features and the detail features is realized through multi-level transformation and mapping. According to the method, cross-modal features can be effectively extracted and recombined and fused, spectrum distortion is remarkably reduced, and the spectrum distortion phenomenon caused by improper feature fusion is effectively avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning, and in particular relates to a method for panchromatic sharpening of remote sensing images. Background Art

[0002] Panchromatic sharpening is a key task in multimodal image fusion. It improves the spatial resolution and spectral information of the image by integrating the complementary information of two different modalities, panchromatic and multispectral images. The input features of different modalities are correlated at low frequencies, representing the shared information of the modalities, while the high-frequency features are uncorrelated, representing the unique characteristics of each modality. Specifically, these two images are usually obtained by different satellite sensors from the same observation area, so they have some common features, such as background contours and large-scale environmental features, while also retaining their own unique advantages, such as the rich spectral information of multispectral images and the fine spatial details of panchromatic images. How to effectively fuse the features of these different modalities is the key to improving the performance of panchromatic sharpening.

[0003] Although the existing panchromatic sharpening methods have achieved certain results in their respective research fields and solved some problems of panchromatic sharpening, the relevant algorithms still have certain deficiencies when dealing with cross-modal features. First, the ability to extract cross-modal features is limited, and it is difficult to effectively distinguish the common information shared between modalities from the unique detail information of each modality; second, there is information redundancy between modal features, and the unique advantages of panchromatic images and multispectral images are not fully utilized, resulting in unsatisfactory fusion effects; finally, how to efficiently reorganize features and retain key information in the fusion process to avoid information loss is also an improvement that the existing methods need to make. At present, the panchromatic sharpening framework is mainly implemented in two ways. The first method is to directly connect the panchromatic image and the multispectral image in series and input them into the same network for training. This method has a simpler processing flow, but it is often difficult to effectively distinguish and extract the unique features of each modality because it processes data of different modalities. The second method uses a network with a two-stream structure to extract independent features of the panchromatic image and the multispectral image, and then fuse these features in the later stage of the network. Although this structure can process the respective modal data well, it usually ignores the common characteristics between the modalities, which may lead to the fusion result being less than ideal in terms of spectral or spatial fidelity. Summary of the invention

[0004] In order to overcome the shortcomings of the prior art, the present invention provides a panchromatic sharpening method based on cross-modal feature decomposition and recombination. First, a dual-branch feature decomposition module is designed to perform feature decomposition on the panchromatic image and the multispectral image respectively, so as to extract low-frequency basic features and high-frequency detail features. The two branches share the same shallow feature extraction module to impose correlation constraints on the extracted basic features; then, in order to solve the information redundancy problem caused by the direct series connection and stacking of the extracted features, an interactive feature recombination module is used to interactively fuse the features to enhance the long-distance dependency and global contextual connection between different modal features; finally, an image reconstruction module based on an invertible neural network (INN) is proposed, which realizes the efficient fusion of basic features and detail features through multiple levels of transformation and mapping. The present invention can effectively extract cross-modal features and perform recombination and fusion, significantly reduce spectral distortion, and effectively avoid the spectral distortion caused by improper feature fusion. The technical solution adopted by the present invention to solve its technical problems includes the following steps:

[0005] Step 1: Dataset preparation;

[0006] The data used in the present invention comes from two satellite sensors: Gaofen-2 (GF2) and WorldView-3 (WV3); according to the Wald protocol, the data collected by the two satellites are preprocessed to construct a down-resolution data set; then, a training set, a validation set and a down-resolution test set are divided from the generated panchromatic image blocks and multispectral image blocks; in addition, the test set also includes full-resolution images, which are obtained from the original data through data cutting;

[0007] Step 2: Construct a fusion model based on cross-modal feature decomposition and recombination;

[0008] The fusion model architecture proposed in this invention is mainly divided into three stages: feature decomposition, feature recombination and image reconstruction, which involves three modules, namely, a dual-branch feature decomposition module, an interactive feature recombination module and an INN-based image reconstruction module. The specific construction process is as follows:

[0009] Step 2-1: Construct a dual-branch feature decomposition module;

[0010] The dual-branch feature decomposition module consists of a shallow feature extraction (SFE) module, a detail feature extraction (DFE) module, and a base feature extraction (BFE) module. The SFE module consists of multiple Restomer modules, and the Restomer module can extract modal shallow features through a self-attention mechanism across feature dimensions. The detail feature extraction branch and the base feature extraction branch each have their own shallow feature extraction modules. The shallow feature modules are used to extract features first, and then the DFE module and the BFE module are used to decompose the base features and detail features.

[0011] Step 2-2: Construct an interactive feature recombination module;

[0012] The interactive feature reorganization module realizes effective interaction between features of different modalities through the self-attention mechanism to enhance the information exchange and integration between features. The module firstly M and F P Perform average pooling in the channel dimension to extract global information features and Then, F M With F P Relative global information and Perform cross-concatenation to obtain new feature maps F m (F M and series) and F n (F P and Serial), the feature map F after cross-series m and F n are sent to the self-attention module for processing respectively; finally, the two feature maps processed by self-attention are weighted and merged to generate the final recombined feature F I ;

[0013] Step 2-3: Construct an INN-based image reconstruction module;

[0014] In order to better reconstruct high-resolution multispectral images and enhance the spatial details and spectral quality of images, the present invention proposes an image reconstruction module based on INN; the module consists of multiple invertible modules, and the invertible modules are implemented by INN with affine coupling layers; in order to achieve effective conversion between two branches, a residual multiscale block (RMB) is designed for feature mapping; RMB realizes multi-scale feature extraction through three dilated convolutions with different expansion coefficients, thereby capturing key information in the image through different receptive fields; the outputs of these dilated convolutions are connected through residual connections and finally fused through a 1×1 convolution to obtain the final features;

[0015] First, in each reversible module, the input detail features are processed by additive transformation, and the input basic features are processed by enhanced affine transformation; then the features output by the two branches are connected in series and passed to the next reversible module for processing; finally, the convolution operation is used to restore the features to the required number of spectra of the multispectral image to obtain the final panchromatic sharpening result; through the synergy of basic features and detail features, INN transforms and maps at multiple levels to enhance the spectral details and spatial details of the image, thereby reconstructing a high-resolution multispectral image;

[0016] Step 2-4: Build the overall network structure;

[0017] First, the multispectral image is four times upsampled to achieve the same spatial resolution as the panchromatic image; the two images are then sent to the dual-branch feature decomposition module to obtain their respective low-frequency basic features and high-frequency detail features; then, the obtained basic features P B and M B Input into the interactive feature reconstruction module, interactively learn the basic information extracted from different modalities, and generate the basic information to be reconstructed F B Similarly, the detail feature P D and M D Input to this module, and after the same operation, the detail feature F to be reconstructed is obtained D Finally, the extracted basic features F are reconstructed using an image reconstruction module based on a reversible neural network. B and detail features F D Perform effective fusion and restore the number of spectra required for multispectral images through convolution to obtain high-resolution multispectral images;

[0018] Step 3: Design loss function;

[0019] The loss function of the method of the present invention is composed of a reconstruction loss function, a structural similarity loss function and a feature decomposition loss function; wherein the formula of the reconstruction loss function is as follows:

[0020]

[0021] In the formula, x ref represents the reference image, x represents the generated full-color sharpened image, N represents the total number of pixels in the image, L rec That is, the generated reconstruction loss;

[0022] The formula of the structural similarity loss function is as follows:

[0023] L SSIM =1-SSIM(x ref ,x) (2)

[0024] In the formula, SSIM(·) represents the structural similarity index, which measures the similarity between the reference image and the generated image in local structure;

[0025] The formula of the eigendecomposition loss function is as follows:

[0026]

[0027] In the formula, CC(·) represents the correlation coefficient operator, ε is set to 1.01 to ensure that the loss function is always positive; P D and M D are the detail features of the panchromatic image and the multispectral image, respectively, B and M B They are the basic features of panchromatic and multispectral images respectively;

[0028] Finally, the overall loss function for training the pan-sharpening method based on cross-modal feature decomposition and recombination is defined as follows:

[0029] L=λ rec L rec +λ SSIM L SSIM +λ decomp L decomp (4)

[0030] In the formula, λ rec , SSIM and λ decomp Corresponding to L rec , L SSIM and L decomp The weight coefficient of the loss function;

[0031] Step 4: Train the network model;

[0032] Set the training parameters: batch size, learning rate, weight decay, maximum number of iterations, and optimizer parameters; use the data set preprocessed in step 1 and use the loss function in step 3 for training; after training, the final fusion model based on cross-modal feature decomposition and recombination is obtained;

[0033] Step 5: Test the network model;

[0034] The panchromatic image and multispectral image to be tested are input into the final fusion model based on cross-modal feature decomposition and recombination, and the fused image is output.

[0035] Preferably, the processed data sets are all panchromatic image blocks with a resolution of 256×256 and multispectral image blocks with a resolution of 64×64; wherein the training set, validation set and reduced-resolution test set images are obtained by Wald protocol processing, and the full-resolution test set images are obtained by data segmentation; the test set data volume is 20 groups of reduced-resolution and full-resolution data sets, and the training set and validation set are taken from the data set excluding the reduced-resolution test set, and the data volume ratio is 9:1.

[0036] Preferably, the weight coefficient of the setting loss is: rec is 1, λ SSIM is 1, λ decomp is 2.

[0037] Preferably, the training parameters are set as follows: the Adam optimizer is used for model training, and the batch size is set to 2, and the training cycle is 300; the initial learning rate is set to 0.0001, and then every 50 cycles, the learning rate will automatically decay by 0.5 times, and the learning rate will be gradually reduced during the optimization process to achieve more stable convergence and better model performance.

[0038] The beneficial effects of the present invention are as follows:

[0039] 1. The present invention proposes a dual-branch feature decomposition module to extract detail features and basic features of panchromatic images and multispectral images respectively. In order to enhance the accuracy and interpretability of basic feature extraction, the basic feature extraction branch of multispectral images and the basic feature extraction branch of panchromatic images share the same shallow feature extraction module. This design enhances the correlation between the low-frequency basic features of the two modalities, thereby imposing correlation constraints on the extracted basic features.

[0040] 2. The present invention proposes an interactive feature recombination module, which interactively learns the basic features of multispectral images and panchromatic images and recombines them into new basic features. At the same time, it also interactively learns the detail features of these two modal images and recombines them into new detail features. This processing method ensures that the information extracted from different modalities can be effectively integrated, thereby enhancing the overall quality and detail performance of the image while retaining key spectral and spatial information.

[0041] 3. The present invention proposes an INN-based image reconstruction module. Through its reversibility, INN can accurately capture and process different features of the image at multiple levels, and ensure that the key information of the image will not be lost in the reconstruction process during forward propagation and back propagation. High-quality multispectral images can be obtained, and the spatial details and spectral quality of the image can be enhanced. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 The figure is a flow chart of the method of the present invention.

[0043] Figure 2 It is a structural schematic diagram of the detail feature extraction module in the dual-branch feature decomposition module in the method of the present invention.

[0044] Figure 3 It is a structural schematic diagram of the basic feature extraction module in the dual-branch feature decomposition module in the method of the present invention.

[0045] Figure 4 It is a schematic diagram of the structure of the interactive feature recombination module in the method of the present invention.

[0046] Figure 5 Schematic diagram of the structure of the INN-based image reconstruction module in the method of the present invention.

[0047] Figure 6 It is a schematic diagram of the overall structure of the fusion model based on cross-modal feature decomposition and recombination in the method of the present invention.

[0048] Figure 7 : is a comparison diagram of the down-resolution WV3 test image and the fusion results of all algorithms in the embodiment of the present invention, where Figure 7 (a) is the WV3 multispectral test image, Figure 7 (b) is the WV3 full-color test image. Figure 7 (c) is the reference image, from Figure 7 (d) to Figure 7 (o) The fusion results of BT-H, BDSD-PC, MTF-GLP-FS, TV, PNN, PanNet, FusionNet, GPPNN, PanFormer, DCPNet, PAPS and the method of the present invention on the reduced-resolution WV3 test image are shown in sequence.

[0049] Figure 8 GF2 test image with reduced resolution and the fusion results of all algorithms in the embodiment of the present invention are compared. Figure 8 (a) is the GF2 multispectral test image, Figure 8 (b) is the GF2 full-color test image, Figure 8 (c) is the reference image, Figure 8(d) to Figure 8 (o) The fusion results of BT-H, BDSD-PC, MTF-GLP-FS, TV, PNN, PanNet, FusionNet, GPPNN, PanFormer, DCPNet, PAPS and the method of the present invention on the reduced-resolution GF2 test image are shown in sequence.

[0050] Fig. 9 : is a comparison chart of the full-resolution WV3 test image and the fusion results of all algorithms in the embodiment of the present invention, where Fig. 9 (a) is the WV3 multispectral test image, Fig. 9 (b) is the WV3 full-color test image, from Fig. 9 (c) to Fig. 9 (n) The fusion results of BT-H, BDSD-PC, MTF-GLP-FS, TV, PNN, PanNet, FusionNet, GPPNN, PanFormer, DCPNet, PAPS and the method of the present invention on the full-resolution WV3 test image.

[0051] Fig.10 : is a comparison diagram of the full-resolution GF2 test image and the fusion results of all algorithms in the embodiment of the present invention, where Fig.10 (a) is the GF2 multispectral test image, Fig.10 (b) is the GF2 full-color test image, from Fig.10 (c) to Fig.10 (n) The fusion results of BT-H, BDSD-PC, MTF-GLP-FS, TV, PNN, PanNet, FusionNet, GPPNN, PanFormer, DCPNet, PAPS and the method of the present invention on the full-resolution GF2 test image are shown in sequence. DETAILED DESCRIPTION

[0052] The present invention is further described below in conjunction with embodiments and drawings.

[0053] like Figure 1 As shown, a panchromatic sharpening method based on cross-modal feature decomposition and recombination includes the following steps:

[0054] Step 1: Dataset preparation;

[0055] The data comes from two satellite sensors, Gaofen-2 (GF2) and WorldView-3 (WV3). Specifically, the spatial resolution of the panchromatic image in the GF2 dataset is 0.8 meters, and the resolution of the multispectral image is 3.2 meters, including four bands, covering red, green, blue and near-infrared. The WV3 dataset has higher resolution and spectral information. The spatial resolution of its panchromatic image reaches 0.31 meters, and the resolution of its multispectral image is 1.24 meters. The multispectral image is expanded to 8 bands, covering richer spectral information from visible light to near-infrared.

[0056] Since the real high-resolution multispectral image (i.e., the reference image of the fusion result) does not exist, the present invention constructs a training set and a test set under reduced resolution according to the Wald protocol. According to the Wald protocol, the data collected by the two satellites are preprocessed to generate panchromatic (PAN) image blocks with a resolution of 256×256 and multispectral (MS) image blocks with a resolution of 64×64, and a reduced resolution data set is constructed. Subsequently, 20 groups are randomly selected from the generated PAN and MS image blocks as the reduced resolution test set, and the remaining data are divided into a training set and a validation set at a ratio of 9:1. The training set and the validation set only contain PAN and MS image blocks with reduced resolution. In addition, the test set also includes full-resolution images, which are obtained from the original data through data cutting, specifically consisting of 20 groups of PAN images with a resolution of 256×256 and MS images with a resolution of 64×64.

[0057] Step 2: Construct a fusion model based on cross-modal feature decomposition and recombination;

[0058] The fusion model architecture proposed in this invention is mainly divided into three stages: feature decomposition, feature recombination and image reconstruction, which involves three modules, namely, a dual-branch feature decomposition module, an interactive feature recombination module and an INN-based image reconstruction module. The specific construction process is as follows:

[0059] Step 2-1: Construct a dual-branch feature decomposition module;

[0060] The dual-branch feature decomposition module consists of a shallow feature extraction (SFE) module, a detail feature extraction (DFE) module, and a basic feature extraction module (BFE). Among them, the shallow feature extraction module consists of multiple Restomer modules, and the Restormer module can extract modal shallow features through the self-attention mechanism across feature dimensions. The specific definitions are as follows:

[0061] F S =f SFE (F) (5)

[0062] In the formula, F represents the input image, f SFE(·) is the shallow feature extraction module composed of Restomer, F S Represents the extracted shallow features. The detail feature extraction branch and the basic feature extraction branch have their own shallow feature extraction modules. The shallow feature module is used to extract features first, and then the basic feature extraction module and the detail feature extraction module are used to decompose the basic features and detail features. In order to enhance the accuracy and interpretability of basic feature extraction, the basic feature extraction branch of multispectral images and the basic feature extraction branch of panchromatic images share the same shallow feature extraction module. This design is used to enhance the correlation between the low-frequency basic features of the two modalities, thereby imposing correlation constraints on the extracted basic features and limiting the solution space to enhance the interpretability of the features and the generalization ability of the model.

[0063] The structure of the DFE module is as follows Figure 2 In the DFE module, the size of the input shallow feature is (h, w, c), which is evenly divided in the channel dimension to generate four sub-feature maps and Each sub-feature map is then assigned to a different attention head. In each attention head, adaptive max pooling is used to perform downsampling at different scales to obtain feature maps of multiple resolutions, whose sizes are (h / 8, w / 8, c), (h / 4, w / 4, c), (h / 2, w / 2, c), and (h, w, c). For the convenience of explanation, let the branch with a scaling factor of 8 of adaptive max pooling be the first layer of attention head, the branch with a scaling factor of 4 be the second layer of attention head, the branch with a scaling factor of 2 be the third layer of attention head, and the branch without pooling operation be the fourth layer of attention head. Specifically, except for the fourth layer of attention head, which is not pooled and downsampled, the height and width of each other layer of attention head are halved in turn. Feature extraction starts from the first layer of attention head, and the features enhanced by the attention mechanism are added element by element with the output of the previous layer through upsampling. This process is continued in all layers, and each layer further enhances the features through the self-attention mechanism until the fourth layer. Subsequently, the outputs of all layers are upsampled to the initial size (h, w, c) and concatenated together, and aggregated through a convolutional layer to form the final detail feature map. Figure 2 As shown in the right part. First, the Query (Q), Key (K) and Value (V) are calculated through the fully connected layer. Then, the dot product between Q and K is calculated, and the result is used to weight V to enhance the features. Finally, the enhanced features are processed through a 1×1 convolutional layer to output the final attention feature map. The following formula shows the process of detail feature extraction:

[0064]

[0065] Where Splite(·) represents segmentation along the channel dimension, Attention(·) represents self-attention processing, and MaxPool i (·) represents adaptive max pooling, where i represents the scaling factor of max pooling, UP j (·) represents the upsampling operation, j represents the upsampling scale, and Concat(·) represents the concatenation operation. and They are the feature maps obtained by self-attention in each layer, which are connected in series and fused through convolution to obtain the detail feature F D The DFE module effectively extracts detail features by using a multi-scale attention mechanism and a maximum pooling operation. Multi-scale processing enables the module to capture features at different levels from global to local, while maximum pooling can highlight important texture and edge information and enhance the representation of image details. The self-attention mechanism further enables the module to strengthen key spatial dependencies and accurately capture complex details. In addition, feature fusion between layers ensures the progression of details from coarse to fine, improving the richness of detail features.

[0066] The structure of the BFE module is as follows Figure 3 As shown in the figure, this module first performs shallow feature extraction on the input Perform maximum pooling downsampling to obtain the low-frequency component of the image. Then, it is input into a 3×3 depth convolution to generate non-local structure features. This step helps the module focus on the basic structure of the image rather than the details. In order to optimize the representation of these non-local structures, the input feature F is introduced S The variance of is used as a regulator of the global description to modulate the non-local structural feature F T Through 1×1 convolution, the variance information is integrated into the non-local feature representation. This variance modulation mechanism effectively strengthens the non-local information of the image. Finally, the modulated features are combined with the original input features F S Multiply element by element to extract important low-frequency basic features F B The specific process is as follows:

[0067] F T =DWConv 3×3 (MaxPool 8 (F S )) (12)

[0068]

[0069] In the formula, DWConv 3×3 (·) represents a 3×3 depthwise convolutional layer, σ 2 Indicates F S The variance of represents the characteristics after modulation, represents the element-wise product operation, and φ(·) represents the GELU activation function.

[0070] Step 2-2: Construct an interactive feature recombination module;

[0071] The interactive feature reorganization module realizes effective interaction between features of different modalities through the self-attention mechanism to enhance information exchange and integration between features. The structure is as follows: Figure 4 As shown in Figure 3, its internal self-attention mechanism is the same as that of the DFE module.

[0072] The interactive feature recombination module first reconstructs the two input features F M and F P Perform average pooling in the channel dimension to extract global information features and Then, F M With F P Relative global information and Perform cross-concatenation to obtain new feature maps F m (F M and series) and F n (F P and The feature map F after cross-concatenation m and F n They are sent to the self-attention module for processing, which strengthens the information complementarity between different inputs, increases the relevance of features, and reduces information redundancy. Finally, the two feature maps processed by self-attention are weighted and merged to generate the final recombined feature F I The specific workflow is as follows:

[0073]

[0074] F I =Attention(F m )+Attention(F n ) (19)

[0075] Where Avg(·) represents the average pooling operation in the channel dimension, and Attention(·) represents the self-attention processing.

[0076] Step 2-3: Construct an INN-based image reconstruction module;

[0077] In order to better reconstruct high-resolution multispectral images and enhance the spatial details and spectral quality of the images, the present invention proposes an image reconstruction module based on INN. The module consists of multiple invertible modules, such as Figure 5 (a), the reversible module is implemented using an INN with an affine coupling layer. In order to achieve effective conversion between the two branches, a residual multi-scale block (RMB) is designed for feature mapping. The output of RMB is shown in Figure 5 (b) RMB achieves multi-scale feature extraction through three dilated convolutions with different expansion coefficients, thereby capturing key information in the image through different receptive fields. The outputs of these dilated convolutions are connected through residual connections and finally fused through a 1×1 convolution to obtain the final features. The specific process is as follows:

[0078] Y 1 =Conv d=1 (X) (20)

[0079] Y 2=Conv d=3 (Y 1 ) (twenty one)

[0080] Y 3 =Conv d=5 (Y 2 ) (twenty two)

[0081] Y=Conv 1×1 (Concat(Y 1 , Y 2 , Y 3 )) (twenty three)

[0082] In the formula, Conv d (·) represents the dilated convolution, where d represents the size of the dilation coefficient. 1×1 (·) represents 1×1 convolution, X represents the input features of the RMB module, and Y 1 , Y 2 and Y 3 They are the different scale features extracted by dilated convolution, and Y represents the final output feature.

[0083] In each reversible module, the input detail feature F D Perform additive transformation on the input basic feature F B Perform enhanced affine transformation processing. Then the features of the two branches are connected in series and passed to the next reversible module for processing. Taking the first reversible module as an example, the specific process is as follows:

[0084]

[0085]

[0086] In the formula, exp(·) represents the exponential function in mathematics, Represents a dot product operation. 1 (·),I 2 (·) and I 3 (·) are all functions that use RMB as the mapping. and The features of the two branch outputs are input in series to the next reversible module to further interact with the detail features and basic features. Finally, the convolution operation is used to restore the features to the required number of spectra of the multispectral image to obtain the final panchromatic sharpening result. Through the synergy of basic features and detail features, INN transforms and maps at multiple levels to enhance the spectral details and spatial details of the image, thereby reconstructing a high-resolution multispectral image.

[0087] Step 2-4: Build the overall network structure;

[0088] The overall network structure of the fusion model proposed in this invention is as follows: Figure 6 As shown in Figure 1. First, the multispectral image is upsampled four times to achieve the same spatial resolution as the panchromatic image. The two images are then fed into the dual-branch feature decomposition module to obtain their respective low-frequency basic features and high-frequency detail features, which are defined as follows:

[0089] [P B , P D ] = f DFDM (P) (26)

[0090] [M B , M D ] = f DFDM (M ↑4 ) (27)

[0091] Where P represents the full-color image, M ↑4 Represents the upsampled multispectral image MS ↑4 , f DFDM (·) indicates a dual-branch eigendecomposition module. B and P D They represent the basic features and detail features obtained by extracting and decomposing the full-color image, respectively. B and M D They represent the basic features and detail features obtained from the multispectral image, respectively. Through the dual-branch feature decomposition module, the basic features and detail features of the panchromatic image and the multispectral image can be efficiently extracted and separated, respectively, so as to effectively process the data of two different modalities.

[0092] Then, the obtained basic feature P B and M B Input into the interactive feature reconstruction module, interactively learn the basic information extracted from different modalities, and generate the basic information to be reconstructed F B Similarly, the detail feature P D and M D Input to this module, and after the same operation, the detail feature F to be reconstructed is obtained D This process ensures the effective integration of features during the fusion process, as follows:

[0093] F B =f IFRM (P B , M B ) (28)

[0094] F D =f IFRM (P D , M D ) (29)

[0095] In the formula, f IFRM (·) represents the interactive feature reconstruction module, which solves the information redundancy problem caused by directly stacking or adding features and fully integrates the low-frequency basic information and high-frequency detail information of different modalities.

[0096] Finally, the extracted basic features F are reconstructed using an image reconstruction module based on a reversible neural network. B and detail features F D Effective fusion is performed and the number of spectra required for the multispectral image is restored through convolution to obtain a high-resolution multispectral image. The specific definition is as follows:

[0097]

[0098] In the formula, is the final high-resolution multispectral image, Conv(·) represents the 3×3 convolution used to adjust the channel dimension, and f INN (·) represents the network structure composed of several reversible basic units in the image reconstruction module. INN ensures the reversibility of the conversion process through its structure, greatly reduces the reconstruction error caused by information loss, and makes the fusion result more accurate and reliable.

[0099] Step 3: Design loss function;

[0100] The loss function of the method of the present invention consists of a reconstruction loss function, a structural similarity loss function, and a feature decomposition loss function. The reconstruction loss function measures the difference between the generated image and the reference image through the L1 norm. The L1 norm calculates the absolute value sum of the pixel value difference between the generated image and the real image, which can promote accurate pixel-level restoration. The specific formula is as follows:

[0101]

[0102] In the formula, x ref represents the reference image, x represents the generated full-color sharpened image, N represents the total number of pixels in the image, L rec That is the generated reconstruction loss.

[0103] The structural similarity loss function (SSIM) optimizes image quality by measuring the structural similarity between two images. Compared with the traditional pixel-level error, SSIM pays more attention to the brightness, contrast and structural information of the image. Its formula is as follows:

[0104] L SSIM =1-SSIM(x ref ,x) (2)

[0105] Where SSIM(·) represents the structural similarity index, which measures the similarity between the reference image and the generated image in local structure.

[0106] The feature decomposition loss function is designed to constrain the model's decomposition of detail features and basic features. The basic features of two different modalities, panchromatic images and multispectral images, contain more modal shared information, such as background information and large-scale environmental information, so they are often highly correlated. In contrast, the correlation between detail features of different modalities is low. For example, the detail features of panchromatic images are high-resolution spatial information, and the detail features of multispectral images are rich spectral information. They are both unique features of the modality. Therefore, in order to better separate these two types of features, the following feature decomposition loss function is designed:

[0107]

[0108] In the formula, CC(·) represents the correlation coefficient operator, and ε is set to 1.01 to ensure that the loss function is always positive. D and M D are the detail features of the panchromatic image and the multispectral image, respectively, B and M B are the basic features of panchromatic images and multispectral images respectively. Under the guidance of the feature decomposition loss function, CC(P D , M D ) will gradually decrease, CC(P D , M D) will gradually increase, making L decomp The smaller the value.

[0109] Finally, the overall loss function for training the fusion model based on cross-modal feature decomposition and recombination is defined as follows:

[0110] L=λ rec L rec +λ SSIM L SSIM +λ decomp L decomp (4)

[0111] In the formula, λ rec , SSIM and λ decomp Corresponding to L rec , L SSIM and L decomp The weight coefficient of the loss function.

[0112] Step 4: Train the network model;

[0113] Set the training parameters: batch size, learning rate, weight decay, maximum number of iterations, and optimizer parameters; use the dataset preprocessed in step 1 and train using the loss function in step 3 to obtain the final fusion model based on cross-modal feature decomposition and recombination.

[0114] The Adam optimizer is used for model training, and the batch size is set to 2 and the training cycle is 300. The initial learning rate is set to 0.0001, and then the learning rate will automatically decay by 0.5 times every 50 cycles. The learning rate is gradually reduced during the optimization process to achieve more stable convergence and better model performance.

[0115] Step 5: Test the network model;

[0116] The panchromatic image and multispectral image to be tested are input into the final fusion model based on cross-modal feature decomposition and recombination, and the fused image is output. Specific embodiment:

[0118] 1. Experimental conditions

[0119] The experimental environment is Intel(R) Core(TM) i5-13600KF CPU@5.10GHz, with 32GB of memory, the GPU processor is NVIDIA GEFORCE RTX 3090, the Python version used is 3.9.18, and it is programmed based on the PyTorch deep learning framework.

[0120] 2. Experimental content

[0121] In order to verify the effectiveness of the present invention, the method of the present invention is compared with several representative algorithms in the field of pan-sharpening. These algorithms include traditional algorithms and algorithms based on deep learning. Traditional algorithms include BT-H and BDSD-PC based on component replacement, MTF-GLP-FS based on multi-resolution analysis, and TV based on variational optimization. Algorithms based on deep learning include PNN, PanNet, FusionNet, GPPNN, PanFormer, DCPNet and PAPS. After the network is trained on the reduced-resolution training set, it is tested and compared on the reduced-resolution and full-resolution test sets respectively.

[0122] 3. Evaluation indicators

[0123] For the resolution reduction experiment, the present invention uses multiple quantitative indicators to evaluate the fusion results, including SAM, ERGAS, SCC, Q2n and PSNR, to evaluate the algorithm performance from multiple aspects such as spectral similarity, spatial similarity and comprehensive quality. Among them, SAM evaluates the spectral similarity by measuring the spectral angle between the fusion result and the reference image. The smaller the value, the higher the spectral similarity. ERGAS is used to measure the overall similarity between the fusion result and the reference image. The smaller the value, the closer the fusion result is to the reference image. SCC quantifies the spatial dependence or similarity between the fusion result and the reference image, and evaluates the spatial quality of the image. The higher the value, the better the spatial quality. Q2n (Q4 is used when MS is 4 bands, and Q8 is used when MS is 8 bands) calculates the quality score of multi-band or multi-channel images, evaluates the spectral and spatial fidelity between the fusion result and the reference image, and the higher the value, the higher the spectral and spatial fidelity of the result. PSNR measures the signal-to-noise ratio. The higher the value, the better the reconstruction quality of the image.

[0124] For full-resolution experiments, the spectral distortion factor D is used. λ , spatial distortion coefficient D S and no-reference quality rating QNR as the evaluation index. λ Indicates the degree of spectral distortion between the experimental results and the input low-resolution multispectral image. The smaller the value, the higher the spectral similarity. S It indicates the degree of spatial distortion between the experimental results and the input full-color image. The smaller the value, the higher the spatial similarity. QNR measures the spectral and spatial distortion between the experimental results and the input image. The higher the value, the better the generated image performs in retaining spectral and spatial information.

[0125] 4. Simulation test

[0126] Figure 7The fusion results of all algorithms on the down-resolution WV3 dataset are shown. The box area is enlarged and displayed in the lower right corner of the image to facilitate observation of local details and spectral information. From the fusion results, the traditional algorithms are not as good as the deep learning methods in maintaining spatial and spectral details. For example, the fusion result of BDSD-PC shows slight spectral distortion, and the overall spectral information is inconsistent with the reference image. This may be because the spatial component of the panchromatic image directly used does not match the original multi-spectral information, resulting in distortion of the spectral information. Especially in the enlarged area of ​​the TV algorithm, spectral information that should not exist can be seen on the building. This may be because the TV algorithm accidentally introduced the spectral information of other objects near the edge during the sharpening process, thereby causing spectral distortion. At the same time, although the deep learning algorithm performs well in maintaining spatial details, when processing the spectral details on the road in the enlarged area, except for the method of the present invention, the other algorithms are not effective, which further verifies the excellent performance of the method of the present invention in reconstructing spectral details.

[0127] Table 1 shows the quantitative evaluation results of all algorithms after pan-sharpening on the down-resolution WV3 dataset. Each algorithm is tested using 20 pairs of multispectral and pan-chromatic test images, and the average value of each indicator is recorded. The bold data in the table indicate the optimal results, and the underlined data indicate the suboptimal results. As can be seen from the data in the table, all evaluation indicators of the deep learning algorithm are better than the traditional algorithm, indicating the application prospects of the powerful nonlinear fitting ability of deep learning in the field of pan-sharpening. DCPNet uses a dual-task driving model to make the two tasks of multispectral super-resolution and pan-sharpening mutually promote and constrain each other, and obtains a good fusion effect. It has achieved suboptimal results in all indicators, second only to the method of the present invention. The PanFormer algorithm extracts features of different modalities through a dual-stream structure based on the Transformer, and applies a cross-attention module to merge spectral and spatial information to generate a fusion result. Although the quantitative indicators of PanFormer have achieved good results, they are still not as good as the method of the present invention, which is also based on the Transformer. The method of the present invention has achieved the best results in all indicators, which fully proves its superior performance in retaining high-resolution details of the image and maintaining spectral consistency.

[0128] Table 1 Quantitative comparison of all algorithms on the down-resolution WV3 dataset

[0129]

[0130] Figure 8It is the fusion result of all algorithms on the down-resolution GF2 dataset, showing the full-color sharpening effect of all algorithms on urban houses. In order to facilitate the observation of details, the area marked by the box in the figure is enlarged in the upper right corner of the image. By observing the fusion result diagram, it can be found that there is serious spectral distortion in the traditional methods. For example, the spectral information of the vegetation area is over-enhanced by the BT-H and TV algorithms; while the spectral information of the fusion results of BDSD-PC and MTF-GLP-FS is not sufficiently preserved. From the overall visual effect, the deep learning methods all show good fusion effects. In terms of details, as shown in the enlarged image area, PNN and PanNet are weak in extracting spatial details, and the edges of the houses appear blurred. In addition, PNN, PanNet, GPPNN and DCPNet all failed to fully retain local spectral details. For example, the spectral information of the roof is not obvious in the fusion diagram of these algorithms, especially in the result of GPPNN, the spectral information of the roof is very different from the reference image, showing obvious distortion of spectral information. In comparison, the fused images of PanFormer, PAPS and the method of the present invention show higher consistency with the reference image, especially the method of the present invention, which performs better in terms of detail preservation and accuracy of spectral information.

[0131] Table 2 shows the quantitative evaluation results of all algorithms on the down-resolution GF2 dataset. Each algorithm is tested using 20 pairs of multispectral and panchromatic test images, and the average value of each indicator is recorded. The bold data in the table represent the optimal results, and the underlined data represent the suboptimal results. From the data in the table, it can be seen that all the indicators of the deep learning algorithm perform better than the traditional algorithms. PanFormer achieved suboptimal results on the ERGAS indicator, indicating that the overall error between PanFormer's fused image and the reference image is small and the similarity is high. DCPNet's performance on the GF2 dataset is not as good as that on WV3, indicating the limitations of the dual-task driven model on specific data and the lack of versatility of the model. PAPS achieved good fusion effects on GF2 data, and achieved suboptimal results on SAM, SCC, Q4 and PSNR indicators. The method of the present invention still achieved the best results, especially the SAM indicator and PSNR indicator, which reached 0.7887 and 38.0134 respectively, which is much better than other algorithms.

[0132] Table 2 Quantitative comparison of all algorithms on the reduced-resolution GF2 dataset

[0133]

[0134] In order to comprehensively evaluate the effects of each algorithm under practical application conditions at full resolution, the experiment adopted an analysis method similar to that of the reduced-resolution dataset. Fig. 9This is the fusion result diagram of all algorithms on the full-resolution WV3 dataset. In order to observe the local texture details, the area of ​​interest is marked with a box and enlarged for display. It can be seen that the fusion results of BT-H, BDSD-PC and TV have artifacts and edge blurring at the edges of the details, which may be caused by the distortion of spatial details during the fusion process. MTF-GLP-FS, PNN, GPPNN and DCPNet have different degrees of spectral distortion, among which the spectral distortion of DCPNet is relatively low. The spectral information of buildings in the fused images of the other three algorithms is very different from that of the reference image. In addition, it can be observed from the enlarged area that PanNet, FusionNet and PAPS show over-enhanced light spots in local areas, which not only affects the overall quality of the image, but also leads to the distortion of spatial texture, such as artifacts at the edges. Fig. 9 (n) It can be seen that the comprehensive performance of the method of the present invention is the best. It can more accurately balance the spectral and spatial details and retain the key texture and spectral information.

[0135] Table 3 shows the quantitative evaluation results of all algorithms on the full-resolution WV3 dataset. All data are based on the average test results of 20 pairs of multispectral and panchromatic images. The values ​​highlighted in bold in the table represent the best results, while the underlined values ​​represent the suboptimal results. From the data in the table, we can see that PanFormer has a good performance in terms of spectral distortion coefficient D λ The no-reference quality rating (QNR) achieves suboptimal results, which shows the effectiveness of its bimodal cross-attention mechanism in processing full-resolution images. PAPS only performs well on the spatial distortion coefficient D S The spectral distortion coefficient D λ The relatively poor performance on may indicate that the algorithm has limitations in processing the spectral information of full-resolution images, which may be due to the fact that the algorithm design of its detail enhancement network in maintaining spectral accuracy needs to be optimized. The method of the present invention has the best performance, with a reference-free quality rating QNR of 0.9443, slightly better than PanFormer, and the spectral distortion coefficient D of the method of the present invention is λ and spatial distortion coefficient D S Both achieved the optimal values, reaching 0.0322 and 0.0217 respectively.

[0136] Table 3 Quantitative comparison of all algorithms on the full-resolution WV3 dataset

[0137]

[0138] Fig.10It is the fusion result of all images on the full-resolution GF2 dataset. Some images containing large-scale buildings are selected, and the local areas are marked with boxes and enlarged for display. It can be seen that the vegetation on both sides of the road in the fused images of BT-H and TV has slight spectral distortion, which does not conform to the actual visual sense. BDSD-PC performs poorly in retaining spectral information, especially the spectral information of the land part is obviously missing. From the overall visual effect analysis, the fusion results of the deep learning algorithm have good spectral information. Observing the enlarged part of the image and further observing the local details, it can be seen that the traditional method has different degrees of blurring on both sides of the road, indicating that the traditional method has weak recovery ability for spatial details on the full-resolution dataset. The network structure of PNN and PanNet is relatively simple, and the feature extraction ability is limited, resulting in the clarity of the road in the enlarged area not as good as other deep learning algorithms. The GPPNN and PanFormer algorithms have spatial information distortion, and the road edge information is blurred. This spatial distortion may be because the two algorithms do not process high-frequency information finely enough during feature fusion and detail enhancement, and fail to fully capture the subtle changes on the edge of the road. The method of the present invention has relatively good effects on spectral details and texture features, especially in complex areas such as road edges. The clarity and texture preservation quality make the visual effects of these parts more natural and realistic, demonstrating its ability in high-quality image reconstruction.

[0139] Table 4 shows the quantitative evaluation results of all algorithms on the full-resolution GF2 dataset. All data are based on the average test results of 20 pairs of multispectral and panchromatic images. The bold data in the table represent the best data, and the underlined data represent the suboptimal data. It can be seen from the data that the PAPS algorithm has a good ability to extract spectral information when processing the GF2 dataset. λ Although DCPNet is not good at preserving spectral information, it has good reconstruction ability of spatial detail features. S The results of the proposed method in terms of the GF2 dataset and QNR indicators were second only to those of the proposed method. The proposed method showed excellent panchromatic sharpening ability on the GF2 dataset and achieved the best results in all three indicators. The analysis of the indicators further proved the high efficiency of the proposed method in the field of panchromatic sharpening.

[0140] Table 4 Quantitative comparison of all algorithms on the full-resolution GF2 dataset

[0141]

Claims

1. A panchromatic sharpening method with cross-modal feature decomposition and recombination, characterized in that: The following steps are involved: Step 1: Dataset preparation; The data used in the present invention comes from two satellite sensors: Gaofen-2 (GF2) and WorldView-3 (WV3); according to the Wald protocol, the data collected by the two satellites are preprocessed to construct a down-resolution data set; then, a training set, a validation set and a down-resolution test set are divided from the generated panchromatic image blocks and multispectral image blocks; in addition, the test set also includes full-resolution images, which are obtained from the original data through data cutting; Step 2: Construct a fusion model based on cross-modal feature decomposition and recombination; The fusion model architecture proposed in this invention is mainly divided into three stages: feature decomposition, feature recombination and image reconstruction, which involves three modules, namely, a dual-branch feature decomposition module, an interactive feature recombination module and an INN-based image reconstruction module. The specific construction process is as follows: Step 2-1: Construct a dual-branch feature decomposition module; The dual-branch feature decomposition module consists of a shallow feature extraction (SFE) module, a detail feature extraction (DFE) module, and a base feature extraction (BFE) module. The SFE module consists of multiple Restormer modules, and the Restormer module can extract modal shallow features through a self-attention mechanism across feature dimensions. The detail feature extraction branch and the base feature extraction branch each have their own SFE modules. The SFE module is used to extract features first, and then the DFE module and the BFE module are used to decompose the base features and detail features. Step 2-2: Construct an interactive feature recombination module; The interactive feature reorganization module realizes effective interaction between features of different modalities through the self-attention mechanism to enhance the information exchange and integration between features. The module firstly M and F P Perform average pooling in the channel dimension to extract global information features and Then, F M With F P Relative global information and Perform cross-concatenation to obtain new feature maps F m (F M and series) and F n (F P and Serial), the feature map F after cross-series m and F n are sent to the self-attention module for processing respectively; finally, the two feature maps processed by self-attention are weighted and merged to generate the final recombined feature F I ; Step 2-3: Construct an INN-based image reconstruction module; In order to better reconstruct high-resolution multispectral images and enhance the spatial details and spectral quality of images, the present invention proposes an image reconstruction module based on INN; the module consists of multiple invertible modules, and the invertible modules are implemented by INN with affine coupling layers; in order to achieve effective conversion between two branches, a residual multiscale block (RMB) is designed for feature mapping; RMB realizes multi-scale feature extraction through three dilated convolutions with different expansion coefficients, thereby capturing key information in the image through different receptive fields; the outputs of these dilated convolutions are connected through residual connections and finally fused through a 1×1 convolution to obtain the final features; First, in each reversible module, the input detail features are processed by additive transformation, and the input basic features are processed by enhanced affine transformation; then the features output by the two branches are connected in series and passed to the next reversible module for processing; finally, the convolution operation is used to restore the features to the required number of spectra of the multispectral image to obtain the final panchromatic sharpening result; through the synergy of basic features and detail features, INN transforms and maps at multiple levels to enhance the spectral details and spatial details of the image, thereby reconstructing a high-resolution multispectral image; Step 2-4: Build the overall network structure; First, the multispectral image is four times upsampled to achieve the same spatial resolution as the panchromatic image; the two images are then sent to the dual-branch feature decomposition module to obtain their respective low-frequency basic features and high-frequency detail features; then, the obtained basic features P B and M B Input into the interactive feature reconstruction module, interactively learn the basic information extracted from different modalities, and generate the basic information to be reconstructed F B Similarly, the detail feature P D and M D Input to this module, and after the same operation, the detail feature F to be reconstructed is obtained D Finally, the extracted basic features F are reconstructed using an image reconstruction module based on a reversible neural network. B and detail features F D Perform effective fusion and restore the number of spectra required for multispectral images through convolution to obtain high-resolution multispectral images; Step 3: Design loss function; The loss function of the method of the present invention is composed of a reconstruction loss function, a structural similarity loss function and a feature decomposition loss function; wherein the formula of the reconstruction loss function is as follows: In the formula, x ref represents the reference image, x represents the generated full-color sharpened image, N represents the total number of pixels in the image, L rec That is, the generated reconstruction loss; The formula of the structural similarity loss function is as follows: L SSIM =1-SSIM(x ref ,x) (2) In the formula, SSIM(·) represents the structural similarity index, which measures the similarity between the reference image and the generated image in local structure; The formula of the eigendecomposition loss function is as follows: In the formula, CC(·) represents the correlation coefficient operator, ε is set to 1.01 to ensure that the loss function is always positive; P D and M D are the detail features of the panchromatic image and the multispectral image, respectively, B and M B They are the basic features of panchromatic and multispectral images respectively; Finally, the overall loss function for training the pan-sharpening method based on cross-modal feature decomposition and recombination is defined as follows: L=λ rec L rec +λ SSIM L SSIM +λ decomp L decomp (4) In the formula, λ rec , SSIM and λ decomp Corresponding to L rec , SSIM and L decomp The weight coefficient of the loss function; Step 4: Train the network model; Set the training parameters: batch size, learning rate, weight decay, maximum number of iterations, and optimizer parameters; use the data set preprocessed in step 1 and use the loss function in step 3 for training; after training, the final fusion model based on cross-modal feature decomposition and recombination is obtained; Step 5: Test the network model; The panchromatic image and multispectral image to be tested are input into the final fusion model based on cross-modal feature decomposition and recombination, and the fused image is output.

2. The method for panchromatic sharpening by cross-modal feature decomposition and recombination according to claim 1, characterized in that: The processed data sets are all panchromatic image blocks with a resolution of 256×256 and multispectral image blocks with a resolution of 64×64; among them, the training set, validation set and reduced-resolution test set images are obtained by Wald protocol processing, and the full-resolution test set images are obtained by data segmentation; the test set data volume is 20 groups of reduced resolution and full resolution, and the training set and validation set are taken from the data set excluding the reduced-resolution test set, and the data volume ratio is 9:

1.

3. The method for panchromatic sharpening by cross-modal feature decomposition and recombination according to claim 1, characterized in that: The weight coefficient of the loss in step 3: λ rec is 1, λ SSIM is 1, λ decomp is 2.

4. The method for panchromatic sharpening by cross-modal feature decomposition and recombination according to claim 1, characterized in that: The training parameters are set as follows: the Adam optimizer is used for model training, and the batch size is set to 2 and the training cycle is 300; the initial learning rate is set to 0.0001, and then the learning rate will automatically decay by 0.5 times every 50 cycles, and the learning rate will be gradually reduced during the optimization process to achieve more stable convergence and better model performance.

Citation Information

Cited By

  • High-spatial-resolution hyperspectral image generation method and device, equipment and medium

    CN120746836A

  • Partial discharge mode fusion identification method and system based on neural network

    CN120850229A

  • Panchromatic sharpening method based on high-order state space modeling

    CN120997084A

  • Hyperspectral fusion imaging method based on hierarchical gradient guidance

    CN121527632A

  • A hyperspectral fusion imaging method based on hierarchical gradient guidance

    CN121527632B