Hyperspectral panchromatic sharpening method based on high and low frequency fusion and coordinate-space attention

Through the high-spectral full-color sharpening method of high-spectral full-color sharpening with coordinate-space attention mechanism, the problem of failure to effectively utilize high-spectral full-color sharpening images in the existing technology is solved, and the quality improvement of the hyperspectral full-color sharpening images is achieved.

CN120339634APending Publication Date: 2025-07-18DALIAN NATIONALITIES UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510445485.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing hyperspectral full-color sharpening methods fail to effectively utilize the high-frequency and low-frequency information differences of remote sensing images, limiting the improvement of image fusion quality.

Method used

The high and low frequency information separation module and the coordinate-space attention mechanism are adopted to separate high and low frequency information, feature extraction and fusion, and the differences between high and low frequency information are utilized, and information fusion is combined with the pixel attention module.

Benefits of technology

The quality of hyperspectral full-color sharpened images is significantly improved, the accuracy of spectral characteristics and spatial resolution are maintained, and network performance and fusion quality are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339634A_ABST
    Figure CN120339634A_ABST
Patent Text Reader

Abstract

The invention discloses a hyperspectral panchromatic sharpening method based on high and low frequency fusion and coordinate-space attention, which belongs to the field of hyperspectral panchromatic sharpening, and comprises a hyperspectral panchromatic sharpening network HLFF-CSANet based on high and low frequency information fusion and a coordinate-space attention mechanism, the HLFF-CSANet separates high and low frequency information, and the HLFF-CSANet separates the high and low frequency information; and a coordinate-space attention module is used to extract spectrum and space information, and finally effective fusion is carried out. A large number of experiments carried out on a PaviaCentre data set, a Botswana data set and a Chikusei data set show that a better panchromatic sharpening effect is achieved through the HLFF-CSANet.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of hyperspectral panchromatic sharpening, and particularly relates to a hyperspectral panchromatic sharpening method based on high-low frequency fusion and coordinate-spatial attention. Background Art

[0002] Similar to RGB images, remote sensing images also contain rich high-frequency information and low-frequency information. Both high-frequency information and low-frequency information are important components of remote sensing images and play different roles in the image processing process. High-frequency information mainly includes edges, textures, and fine structures of the image, etc. In hyperspectral panchromatic sharpening, high-frequency information helps to enhance image details and improve spatial resolution. Low-frequency information mainly includes the overall spectral characteristics of the image and large-scale smooth regions. In the process of hyperspectral panchromatic sharpening, low-frequency information provides the overall structure and background information of the image, helps to maintain overall consistency and authenticity, and effectively prevents color distortion during the sharpening process, keeping the spectral characteristics of the image accurate. Recent research shows that super-resolution is significantly correlated with frequency information, and hyperspectral panchromatic sharpening is essentially a super-resolution process. Therefore, making full use of the differences between high-low frequency information and their respective spectral and spatial characteristics can effectively improve the quality of hyperspectral panchromatic sharpened images.

[0003] Most of the current methods do not pay attention to high-low frequency information, ignore the different roles of different frequency information in the image, but process the whole image containing high-low frequency information as the input, thus limiting the potential to further improve the fusion quality. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides a hyperspectral panchromatic sharpening method based on high-low frequency fusion and coordinate-spatial attention, including:

[0005] Obtain a low-resolution hyperspectral image and a panchromatic image;

[0006] Process the low-resolution hyperspectral image and then splice it with the panchromatic image to obtain an initial input image;

[0007] Separate the high-low frequency information of the initial input image through a high-low frequency separation module to obtain high-frequency information and low-frequency information;

[0008] Extract features from the high-frequency information and low-frequency information based on a coordinate-spatial attention module to obtain high-frequency features and low-frequency features;

[0009] Fuse the high-frequency features and the low-frequency features based on an information fusion module to obtain the final high-resolution hyperspectral image.

[0010] Preferably, the process of obtaining the initial input image includes:

[0011] Upsample the low-resolution hyperspectral image until the low-resolution hyperspectral image has the same spatial resolution as the panchromatic image, and then output to obtain the upsampled low-resolution hyperspectral image;

[0012] Stitch the upsampled low-resolution hyperspectral image and the panchromatic image to obtain the initial input image.

[0013] Preferably, the process of separating the high-frequency and low-frequency information of the initial input image by the high-low frequency separation module to obtain high-frequency information and low-frequency information includes:

[0014] The high-low frequency separation module includes an average pooling layer and a dynamic upsampler;

[0015] Downsample the initial input image based on the average pooling layer to obtain low-frequency information;

[0016] Upsample the low-frequency information based on the dynamic upsampler to generate an upsampled low-frequency feature map;

[0017] Subtract the upsampled low-frequency feature map from the original feature map to obtain high-frequency information.

[0018] Preferably, the coordinate-spatial attention module includes a plurality of sequentially stacked Res-CSA blocks, and each Res-CSA block extracts spectral and spatial features by a parallel coordinate attention mechanism and a spatial attention mechanism.

[0019] Preferably, the process of obtaining high-frequency features and low-frequency features includes:

[0020] Perform a convolution operation on the high-frequency information and the low-frequency information to obtain an intermediate feature map;

[0021] Divide the intermediate feature map into two paths. One path obtains a spectral mask through coordinate attention, and the other path obtains a spatial mask through spatial attention;

[0022] Multiply the spectral mask and the intermediate feature map element by element to obtain spectral features, and multiply the spatial mask and the intermediate feature map element by element to obtain spatial features;

[0023] Add the spectral features, the spatial features and the input feature map element by element to obtain the high-frequency features and the low-frequency features.

[0024] Preferably, the coordinate attention includes: two global average pooling layers respectively along the height and width dimensions, a 1x1 convolutional layer for reducing the number of channels from 64 to 64 / r, a ReLU activation layer, two 1x1 convolutional layers for expanding the number of channels from 64 / r to 64, and finally two sigmoid activation layers.

[0025] Preferably, the spatial attention includes: a parallel arrangement of a global average pooling layer and a global max pooling layer, a 1x1 convolutional layer, and a sigmoid activation layer.

[0026] Preferably, the process of fusing the high-frequency features and the low-frequency features based on the information fusion module to obtain the final high-resolution hyperspectral image includes:

[0027] Upsample the low-frequency features, concatenate the upsampled low-frequency features and the high-frequency features, and compress the number of channels through a 1x1 convolutional layer to obtain fused features;

[0028] Construct a pixel attention module based on two convolutional layers and a PA layer, input the fused features into the pixel attention module for calculation to obtain weighted fused features;

[0029] Based on a convolutional layer, adjust the weighted fused features to a feature map with the same number of channels as the input hyperspectral image to obtain an adjusted feature map;

[0030] Add the adjusted feature map and the upsampled low-resolution hyperspectral image pixel by pixel to obtain the final high-resolution hyperspectral image.

[0031] On the other hand, the present invention also provides an electronic device, including a memory, a processor, and a computing program stored in the memory and executable on the processor, and the processor implements the method when executing the computing program.

[0032] On the other hand, the present invention also provides a computer-readable storage medium, which stores a computer program, and the computer program implements the method when executed by a processor.

[0033] Compared with the prior art, the present invention has the following advantages and technical effects:

[0034] The present invention proposes a lightweight hyperspectral pan-sharpening network based on high-low frequency information fusion and coordinate-spatial attention mechanism, called HLFF-CSANet. Compared with existing deep learning-based hyperspectral pan-sharpening methods, the present invention pays attention to the differences between high-frequency and low-frequency information of remote sensing images. Through the proposed high-low frequency information separation module, the high-low frequency information of the image is separated and the spatial and spectral features are learned respectively, maximizing the use of the high-frequency and low-frequency information of the image and effectively improving the network performance. The present invention also proposes a spatial-spectral attention module based on coordinate attention and spatial attention, which can better extract spatial and spectral features. In addition, the present invention uses pixel attention in the information fusion module to better fuse the outputs of the dual-branch network, further improving the quality of fusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0036] Figure 1 Schematic diagram of the HLFF_CSANet structure according to an embodiment of the present invention;

[0037] Figure 2 Structural diagram of the high-low frequency separation module according to an embodiment of the present invention;

[0038] Figure 3 Schematic diagram of the CSA module according to an embodiment of the present invention;

[0039] Figure 4 Schematic diagram of the Res-CSA structure according to an embodiment of the present invention;

[0040] Figure 5 Schematic diagram of the information fusion module according to an embodiment of the present invention;

[0041] Figure 6 Comparison of visual results of different methods according to an embodiment of the present invention on the Pavia_Centre dataset and mean absolute error graph between the reconstructed image and the reference image of the selected image, where (a), PCA, (b), GSA, (c), MTF-GLP-HPM, (d), HySure, (e), HyperPNN2, (f), FusionNet, (g), HyperDSNet, (h), CCC-SSA-UNet, (i), HLFF-CSANet (Ours), (j), reference ground truth;

[0042] Figure 7Visual result comparison of different methods in the embodiments of the present invention on the Botswana dataset and the mean absolute error map between the reconstructed image and the reference image of the selected image, where (a), PCA; (b), GSA; (c), MTF-GLP-HPM; (d), HySure; (e), HyperPNN2; (f), FusionNet; (g), HyperDSNet; (h), CCC-SSA-UNet; (i), HLFF-CSANet (Ours); (j), reference ground truth

[0043] Figure 8 Visual result comparison of different methods in the embodiments of the present invention on the Chikusei dataset and the mean absolute error map between the reconstructed image and the reference image of the selected image, where (a), PCA; (b), GSA; (c), MTF-GLP-HPM; (d), HySure; (e), HyperPNN2; (f), FusionNet; (g), HyperDSNet; (h), CCC-SSA-UNet; (i), HLFF-CSANet (Ours); (j), reference ground truth Detailed implementation manners

[0044] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments

[0045] It should be noted that the steps shown in the flowchart of the drawings may be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here

[0046] Embodiment 1

[0047] In this embodiment, a hyperspectral panchromatic sharpening method based on high-low frequency fusion and coordinate-spatial attention is provided, including:

[0048] Obtain a low-resolution hyperspectral image and a panchromatic image

[0049] Process the low-resolution hyperspectral image and splice it with the panchromatic image to obtain an initial input image

[0050] Separate the high-low frequency information of the initial input image through a high-low frequency separation module to obtain high-frequency information and low-frequency information

[0051] Feature extraction is performed on the high-frequency information and low-frequency information based on the coordinate-spatial attention module to obtain high-frequency features and low-frequency features;

[0052] Based on the information fusion module, the high-frequency features and the low-frequency features are fused to obtain the final high-resolution hyperspectral image.

[0053] The main purpose of HLFF-CSANet is to separate the high- and low-frequency information of remote sensing images, and hand them over to the coordinate-spatial dual attention module for separate processing. Finally, the final image is obtained through the information fusion module. As Figure 1 shown, HLFF-CSANet adopts a dual-branch learning strategy and consists of three parts: the high-low frequency information separation module HLFS, the coordinate-spatial attention module CSA, and the information fusion module IF. HLFF-CSANet takes a low-resolution hyperspectral image (denoted as X) and a panchromatic image (denoted as P) as the initial inputs, and outputs a high-resolution hyperspectral image (denoted as ). Specifically, first, the low-resolution hyperspectral image X is upsampled by bilinear interpolation to obtain an image (denoted as U) with the same spatial resolution as P. Then U and P are concatenated to obtain image O, and then image F is obtained after 3x3 convolution. Then, the proposed HLFS module is applied to separate the high- and low-frequency information of F, and the high-frequency information F H and the low-frequency information F L are output. F H and F L respectively pass through the CSA module to extract their respective spectral and spatial features, and output F H_O and F L_O respectively. Then F L_O and F H_O are input into the IF module together with F for information fusion to obtain image F fuse , and finally the residual image X res of the high-resolution hyperspectral image is obtained after convolution. Finally, the residual image X res is added to image U pixel by pixel to obtain the final fusion result The above process can be described by the following equations:

[0054] U = UpSample(X);

[0055] X res = f HLFF_CSANet (U, P);

[0056]

[0057] where, UpSample(·) represents bilinear interpolation upsampling, f HLFF_CSANet(·,·) represents the proposed method HLFF - CSANet to be introduced in this embodiment.

[0058] In image processing tasks, traditional separation techniques, such as Fourier transform, involve significant computational overhead and are not easily integrated into the network. To minimize the computational cost of separating high - frequency and low - frequency information, this embodiment proposes a new high - low frequency information separation module. In this embodiment, the proposed module will be introduced, and its structural diagram is as Figure 2 shown. This embodiment uses an average pooling layer in this task. Specifically, the average pooling layer downsamples the input feature map (with size B×H×W×C) to obtain the low - frequency feature F L with size B×H / 2×W / 2×C. Then, DySample is used to upsample these features back to the original dimension, i.e., B×H×W×C. Conventional nearest - neighbor upsampling and bilinear interpolation upsampling apply fixed rules to interpolate the low - resolution features, ignoring the semantic information in the feature map. Some dynamic upsamplers perform well, but they have complex structures and large computational amounts. DySample is a lightweight dynamic upsampler. Compared with conventional upsamplers and some dynamic upsamplers, DySample is simpler, faster, and lower - cost, while retaining the effectiveness of dynamic upsampling. After upsampling, the high - frequency feature F L of F is calculated by subtracting F H from the original feature F. This method enables this embodiment to quickly capture the high - frequency and low - frequency information of the image, and the whole process is described by the following equation:

[0059] F L = GAP(F);

[0060] F H = F - DySample(F L );

[0061] where GAP(·) is average pooling and DySample(·) is the DySample upsampling process.

[0062] To further improve the ability of HLFF - CSANet to extract spatio - spectral features in the feature map, this embodiment introduces CSA based on coordinate - spatial attention mechanism. The CSA module adopts N sequentially stacked Res - CSA modules to adaptively emphasize important spectral and spatial feature information. The schematic diagram of the CSA module is as Figure 3 shown, and it can be expressed as follows:

[0063]

[0064] where, F0 represents the input feature map of the CSA module, corresponding to F output by the high - low frequency separation moduleL and F H 。F N represents the output feature map of the CSA module, corresponding to the output F of the CSA module L_O and F H_O 。F k (1 ≤ k ≤ N - 1) represents the intermediate feature map of the CSA module. 1 ≤ k ≤ N represents the functional representation of the k-th Res-CSA block.

[0065] Similar to Res-SSA in CCC-SSA-UNet and DAU in MIRNet, the design principle of the Res-CSA block in this embodiment is also to integrate spectral and spatial feature information by paralleling two different attention mechanisms and incorporate the two attentions into the basic residual module. This integration aims to simultaneously enhance the spatial and spectral feature representations and the stability of network training and accelerate the convergence speed. Different from the above methods that use channel attention to process spectral information, this embodiment adopts the coordinate attention mechanism, which has the advantage of taking into account the position relationship on the basis of channel attention and can better combine channel information and spatial information. In this embodiment, the entire residual coordinate-spatial attention module is named the Res-CSA block. Figure 4 shows the network structure of the Res-CSA block. For the N-th Res-CSA block, its input is the feature map F N-1 。First, F N-1 is passed through a 3x3 convolutional layer to reduce the number of channels to 64, and after a ReLU activation layer and another 3x3 convolutional layer, the feature map F U is extracted, which serves as the input to the attention module. F U is divided into two paths: one path obtains the spectral mask M CA through the coordinate attention module, and this mask is multiplied element-wise with the feature map F U to obtain F CA ; the other path obtains the spatial mask M SA through the spatial attention module, and this mask is multiplied element-wise with the feature map F U to obtain F SA 。Finally, F CA 、F SA and F N-1 are added element-wise to obtain the output F N of the Res-CSA block. This can be expressed by the mathematical formula:

[0066] F U =Conv 3×3 (ReLU(Conv 3×3 (F N-1 )));

[0067] F N = F CA + F SA + F N-1 ;

[0068] where ReLU(·) represents the ReLU layer, Conv 3x3 (·) represents the 3x3 convolutional layer, M CA and M SA represent the spectral mask and the spatial mask, represents the element-wise multiplication operation.

[0069] Specifically, the coordinate attention module for processing spectral information includes two global average pooling layers respectively along the height and width dimensions, a 1x1 convolutional layer for reducing the number of channels from 64 to 64 / r, a ReLU activation layer, two 1x1 convolutional layers for expanding the number of channels from 64 / r to 64, and finally two sigmoid activation layers. Here, r is called the channel reduction ratio, which can be used to reduce the computational complexity of the network model. The core part of the spatial attention module includes the parallel arrangement of the global average pooling and global max pooling layers, a 1x1 convolutional layer, and a sigmoid activation layer. This can be expressed by the mathematical formula:

[0070] M CA_C = ReLU(Conv 1×1 (Concat(X_GAP(F U ), Y_GAP(F U ))))

[0071] M CA_H = σ(Conv 1×1 (Split_H(M CA_C )))

[0072] M CA_W = σ(Conv 1×1 (Split_W(M CA_C )))

[0073]

[0074] where X_GAP(·) represents the global average pooling along the height dimension, Y_GAP(·) represents the global average pooling along the width dimension, Conv 1×1 (·) represents the 1x1 convolutional layer, Concat(·) represents the channel concatenation operation, Split_H(·) and Split_W(·) respectively represent splitting along the height and width, represents the element-wise multiplication, and σ(·) represents the sigmoid activation layer.

[0075] The coordinate attention module filters out the less important spectral information in the feature tensor, enabling the network to adaptively select the key spectral information. The spatial attention module enables the network to pay more attention to the features in the regions closely related to enhancing the spatial details of the hyperspectral image. By combining the coordinate attention and the spatial attention and embedding them into the basic residual module, not only the spatial-spectral feature representation ability of the network is enhanced, but also the stability of network training is improved.

[0076] The main purpose of the Information Fusion (IF) module is to fuse the high- and low-frequency information to reconstruct the high-resolution hyperspectral image. The schematic diagram of IF is as Figure 5 shown.

[0077] In this embodiment, the spectral and spatial information respectively output by the high- and low-frequency branches are used as inputs. First, the output of the low-frequency branch is upsampled to make its resolution the same as that of the output of the high-frequency branch, and then the two outputs are concatenated. At this time, the number of channels will double, so a 1x1 convolution is used to compress the number of channels to obtain F c . F c is then element-wise added to F to obtain F fuse . This process can be described as:

[0078] F c = Conv 1×1 (Concat(F H_O , UpSample(F L_O ));

[0079] F fuse = F + F c ;

[0080] To better fuse the spectral and spatial information of the two branches, inspired by, in this embodiment, a Pixel Attention (PA) module is adopted at the end of IF. This module consists of two convolutional layers and a PA layer between them. The PA layer is only composed of a 1x1 convolution and a Sigmoid function, and then it is used to weight the input features. This method can effectively improve the final performance at a relatively low parameter cost. In this embodiment, the proposed PA module is denoted as f PA (·). F fuse is fed into the PA module, and finally F final is output.

[0081] F final = f PA (F fuse );

[0082] Finally, to match the number of channels of the input hyperspectral image, a convolutional layer is used to generate the final fusion result:

[0083] Xres = Conv 3×3 (F final );

[0084] To improve the accuracy of the fused reconstruction result compared with the actual high - resolution image, various loss functions are usually adopted in the literature, including mean absolute error (MAE) loss, mean squared error (MSE) loss, perceptual loss, spectral loss, and hybrid loss. Perceptual loss can generate more realistic image results. Spectral loss is designed specifically for spectral information and can maintain better spectral consistency. Hybrid loss combines the advantages of multiple loss functions, which can not only preserve spectral information but also take into account spatial details. Although the above - mentioned loss effects are all good, their computational complexity is relatively high, which will increase the difficulty of network training. MSE loss is vulnerable to outliers, which may lead to the wrong allocation of excessive weights to these data points, resulting in overfitting of the model to noise and outlier data points. This may have a negative impact on the performance of the model. In addition, MSE loss is prone to generating non - sparse solutions, that is, most weights are not compressed to zero. This will lead to an increase in computational cost and may further reduce the generalization ability of the model. In contrast, MAE loss is considered more reliable. The main advantage of using MAE loss function in the process of hyperspectral pan - sharpening is that it can resist outliers and ensure a stable convergence trajectory. Therefore, MAE loss is adopted when evaluating the fusion performance of the network.

[0085] In hyperspectral pan - sharpening, MAE loss is used to evaluate the difference between the generated image and the ground truth (GT) image. It measures the error of the model by calculating the absolute difference between the predicted value and the actual value. It can be expressed by the following mathematical formula:

[0086]

[0087] where GT represents the ground truth image, N represents the number of images in the training set, and |·| represents the l1 norm.

[0088] Example 2

[0089] A hyperspectral pan - sharpening method based on high - low frequency fusion and coordinate - spatial attention is provided in this example, including:

[0090] In this example, model experiments are carried out with the following data, and classic models in the field of hyperspectral pan - sharpening are compared to verify the effectiveness of the model.

[0091] In this example, multiple hyperspectral image datasets are applied to the experiment, and these datasets include Pavia_Centre dataset, Botswana dataset, and Chikusei dataset.

[0092] The Pavia_Centre dataset includes hyperspectral images of the center of Pavia, Italy, acquired by the Reflective Optics System Imaging Spectrometer (ROSIS). The dataset consists of 1096×715 pixels and 102 effective spectral bands, with a spectral range from 0.43 μm to 0.86 μm. The spatial resolution of the Pavia_Centre dataset is 1.3 m by 1.3 m per pixel. The Botswana dataset contains images of the Okavango Delta in Botswana taken by the NASA EO-1 Hyperion sensor. It includes 145 spectral bands from 0.4 to 2.5 μm, with a spatial resolution of 30 m and a coverage area of 1476×256 pixels. The Chikusei dataset was acquired by the Headwall Hyperspec-VNIR-C sensor in Chikusei, Ibaraki, Japan. The dataset consists of 128 spectral bands with a wavelength range of 0.36 to 1.02 μm, a spatial resolution of 2.5 m, and an image size of 2517×2335 pixels.

[0093] In the experiments of this embodiment, all three datasets underwent the following preprocessing procedures. To obtain the baseline ground truth (GT) from the datasets, a region was extracted from the upper left corner of the hyperspectral images. Subsequently, this region was divided into several non-overlapping sub-images. These sub-images constituted the Reflective High-Resolution Hyperspectral Image (HR-HSI) dataset, which was used as the baseline GT. The Wald protocol was adopted to generate the corresponding PAN and Low-Resolution Hyperspectral Image (LR-HSI). Among them, the average value of spectral bands 1 to 100 of the HR-HSI was calculated to generate the PAN. To obtain the LR-HSI image, the HR-HSI image was attenuated using an 8x8 Gaussian filter and then downsampled to reduce the spatial size by a factor. Most of the image pairs were selected for training, and the remaining image pairs were used for testing. Table 1 lists the detailed information of the datasets.

[0094] Table 1

[0095]

[0096] To evaluate the effectiveness of the proposed pansharpening method, five different image quality metrics were used. These metrics include ERGAS, Root Mean Square Error (RMSE), Spectral Angle Mapper (SAM), Peak Signal-to-Noise Ratio (PSNR), and Cross-Correlation (CC).

[0097] CC measures the linear correlation between two images. A value close to 1 indicates a high similarity. SAM quantifies spectral similarity by calculating the angle between spectral vectors. The smaller the angle, the higher the similarity. RMSE evaluates the average error between the predicted image and the reference image. A smaller value indicates better quality. ERGAS evaluates the global relative error of multi - spectral or hyperspectral images. A smaller value indicates higher quality. PSNR measures the peak signal - to - noise ratio. The higher the value, the smaller the distortion and the better the image quality.

[0098] In the experiment of this embodiment, the batch size is set to 1, and the Adam optimizer is used for training, where the hyperparameters β1 is set to 0.9 and β2 is set to 0.999. The initial learning rate is set to 0.0004, and the learning rate is halved every 1000 epochs. To optimize the model, this embodiment uses the l1 loss function and conducts a total of 2000 epochs of training. The model of this embodiment is implemented using the PyTorch framework, and the training process is executed on a single GeForce RTX 2080Ti GPU. For the Chikusei dataset, the training duration is about 12 hours; for the Botswana dataset, the training duration is about 1.5 hours; for the Pavia Centre dataset, the training duration is about 1.5 hours.

[0099] To demonstrate the effectiveness, efficiency, and state - of - the - art performance of the proposed HLFF_CSANet, this embodiment conducts comparative experiments on three datasets: the Pavia Centre dataset, the Botswana dataset, and the Chikusei dataset. The method of this embodiment is compared with four traditional pan - sharpening methods and five state - of - the - art deep - learning - based methods. The traditional methods involved in the comparison include PCA, GSA, MTF - GLP HPM, and HySure. The deep - learning - based methods include HyperPNN2, FusionNe, HyperDSNet, and CCC - SSA - UNet. The traditional methods are implemented using the open - source MATLAB toolbox provided by Loncan et al. For the deep - learning methods, this embodiment reproduces the experiments on this embodiment's computer according to the descriptions and parameter settings in the original papers and presents the best results obtained. The following section will show the detailed comparative experiment results of different pan - sharpening methods on the three datasets.

[0100] In this embodiment, HLFF-CSANet was compared with eight other methods on the test set of the Pavia Centre dataset, and the average quantitative results are shown in Table 2. As can be seen from the table, HLFF-CSANet is significantly superior to other methods in various objective metrics. Specifically, HLFF-CSANet is 0.002 higher than the state-of-the-art method in CC, 0.507 higher in PSNR, while reducing by 0.263, 0.0007, and 0.131 in SAM, RMSE, and ERGAS, respectively.

[0101] Table 2

[0102]

[0103]

[0104] In addition to the quantitative comparison results, Figure 6 The visual results of various pan-sharpening methods on the images selected from the test set of the Pavia Centre dataset are shown, as well as the mean absolute error (MAE) graph between the images reconstructed by different methods and the reference image. Here, the 12th image is selected in this embodiment. It can be seen that the images reconstructed by traditional methods are relatively less clear, with a larger mean absolute error and poorer visual quality. In contrast, the deep learning-based methods benefit from the powerful learning ability of the deep neural network, with lower blurriness of the reconstructed images, smaller mean absolute error, and better visual quality. Among these methods, HLFF-CSANet proposed in this embodiment has the highest similarity to the reference image, demonstrating its excellent ability in maintaining spectral fidelity and accurately restoring spatial details.

[0105] Table 3

[0106]

[0107] In this embodiment, HLFF-CSANet was compared with eight other methods on the test set of the Botswana dataset, and the average quantitative results are shown in Table 3. As can be seen from the table, HLFF-CSANet is significantly superior to other methods in various objective metrics. Specifically, HLFF-CSANet is 0.002 higher than the state-of-the-art method in CC, 0.501 higher in PSNR, while reducing by 0.068, 0.0006, and 0.105 in SAM, RMSE, and ERGAS, respectively.

[0108] In addition to the quantitative comparison results, Figure 7It shows the visual results of various pan-sharpening methods on the images selected from the test set of the Botswana dataset, as well as the mean absolute error (MAE) graphs between the images reconstructed by different methods and the reference image. Here, the 16th image is selected in this embodiment. Obviously, the images reconstructed by the traditional methods are more blurred, with a larger mean absolute error and lower visual quality. While the images reconstructed by the deep learning-based method have less blur, smaller mean absolute error and higher visual quality. The HLFF-CSANet method proposed in this embodiment shows the highest similarity to the reference image, which proves the effectiveness and advancement of the method proposed in this embodiment.

[0109] Table 4

[0110]

[0111] In this embodiment, HLFF-CSANet is compared with 8 other methods in the test set of the Chikusei dataset, and the average quantitative results are shown in Table 4. It can be seen from the table that HLFF-CSANet is significantly superior to other methods in various objective metrics. Specifically, HLFF-CSANet is 0.003 higher than the state-of-the-art method in CC, 0.92 higher in PSNR, while reducing by 0.224, 0.0013 and 0.293 in SAM, RMSE and ERGAS respectively.

[0112] In addition to the quantitative comparison results, Figure 8 it also shows the visual results of various pan-sharpening methods on the images selected from the test set of the Chikusei dataset, as well as the mean absolute error (MAE) graphs between the images reconstructed by different methods and the reference image. Here, the 41st image is selected in this embodiment. Obviously, the HLFF-CSANet method proposed in this embodiment has the highest similarity to the reference image, which proves the excellent performance of this method in terms of effectiveness and advancement.

[0113] Table 5 shows the comparison of the computational complexity of different pan-sharpening methods on the Pavia Centre dataset. Two metrics, PSNR and SAM, are used to evaluate the quality of the network-reconstructed images, while the number of parameters (#Params), multiply-accumulate operations (MACs), floating-point operations (FLOPs), and average inference time (Runtime) are used to evaluate the computational complexity of the neural network. As can be seen from Table 5, the HLFF-CSANet of this embodiment achieves leading image reconstruction performance while maintaining small #Params, MACs, and FLOPs. Extensive experiments on publicly available datasets show that HLFF-CSANet has excellent performance compared to other methods. The model of this embodiment effectively utilizes the high and low frequency information of the image during the fusion reconstruction process while maintaining low memory usage.

[0114] Table 5

[0115]

[0116] HLFF-CSANet mainly relies on CSA to extract spatial and spectral features. Therefore, the number N of ResCSABlock is a key factor affecting the network performance. To explore the impact of N on the performance, four groups of comparative experiments were conducted in this embodiment. To adjust the number of parameters within a reasonable range, N was set to 1, 2, 3, and 4 respectively in this embodiment, and other settings remained unchanged. The experiments were conducted on the test set of the Pavia Centre dataset. The results of the objective evaluation metrics are shown in Table 6.

[0117] The results show that when N = 2, the network reaches the highest performance. When N is greater than 2, with the expansion of the network and the increase in computational complexity, the performance does not improve substantially but instead decreases. Considering the machine performance and computational complexity, the number of ResCSABlock was set to 2 in this embodiment.

[0118] Table 6

[0119]

[0120] This embodiment conducted detailed ablation experiments, focusing on determining the effectiveness of the proposed HLFS module, CSA module, and IF module. For the dataset used in the ablation experiments, this embodiment selected the Pavia Centre dataset.

[0121] The main function of the HLFS module is to separate the high and low frequency information and output it to the dual-branch network. In this subsection, to verify the effectiveness of HLFS, this embodiment constructed a variant of the proposed network, removing the HLFS module from the network, that is, directly processing the spliced image of HSI and PAN without high and low frequency separation.

[0122] In this embodiment, experiments were conducted on the Pavia Centre dataset, and the results are shown in Table 7. It can be seen that the dual-branch method that separates and processes high- and low-frequency information respectively has better performance, which strongly proves the effectiveness of the HLFS module.

[0123] Table 7

[0124]

[0125] The CSA module adopts a dual-branch structure to process high-frequency information and low-frequency information respectively, aiming to enhance the spatial-spectral feature representation of high-frequency and low-frequency information. In this subsection, to verify the effectiveness of the CSA module, three network variants were constructed in this embodiment. In the first variant, no processing was performed on the inputs of the dual branches, and they were directly output to the subsequent IF module. In the second variant, only the high-frequency information was processed, and the low-frequency information was directly output to the subsequent IF module. Contrary to the second variant, the third variant only processes the low-frequency information.

[0126] Table 8 shows the experimental results on the Pavia Centre dataset. Obviously, the network structure that processes both high- and low-frequency information has the best performance, and these results prove the effectiveness of the CSA module proposed in this embodiment.

[0127] Table 8

[0128]

[0129] To better fuse the dual-branch information, the PA module was added to the IF module in this embodiment. The experiments in this section mainly verify the effectiveness of the PA module in the IF module. This embodiment was compared with the IF module without the PA module, and the experimental results are shown in Table 9. The results clearly show that the network performance with the PA module is significantly improved.

[0130] Table 9

[0131]

[0132] This embodiment proposes a lightweight hyperspectral pan-sharpening network based on high-low frequency information fusion and coordinate-spatial attention mechanism, called HLFF-CSANet. Compared with existing deep learning-based hyperspectral pan-sharpening methods, this embodiment focuses on the differences between the high-frequency and low-frequency information of remote sensing images. Through the proposed high-low frequency information separation module, the high-low frequency information of the image is separated and the spatial and spectral features are learned respectively, making the most of the high-frequency and low-frequency information of the image and effectively improving the network performance. This embodiment also proposes a spatial-spectral attention module based on coordinate attention and spatial attention, which can better extract spatial and spectral features. In addition, pixel attention is used in the information fusion module in this embodiment to better fuse the outputs of the dual-branch network, further improving the quality of fusion.

[0133] This embodiment compares HLFF-CSANet with several state-of-the-art methods in three different datasets in detail. The results show that HLFF-CSANet has better performance without adding too many parameters. In the ablation experiment, this embodiment discusses in detail the impact of the proposed modules on the network performance, and the results prove that these modules have a significant effect on the improvement of the network performance. The superiority of HLFF-CSANet in practical applications is obvious.

[0134] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A hyperspectral panchromatic sharpening method based on high-low frequency fusion and coordinate-spatial attention, characterized in that, Including: Obtain a low-resolution hyperspectral image and a panchromatic image; Process the low-resolution hyperspectral image and then splice it with the panchromatic image to obtain an initial input image; Separate the high-frequency and low-frequency information of the initial input image through a high-low frequency separation module to obtain high-frequency information and low-frequency information; Extract features from the high-frequency information and low-frequency information based on a coordinate-spatial attention module to obtain high-frequency features and low-frequency features; Fuse the high-frequency features and the low-frequency features based on an information fusion module to obtain a final high-resolution hyperspectral image.

2. The method according to claim 1, wherein The process of obtaining the initial input image includes: Upsample the low-resolution hyperspectral image until the low-resolution hyperspectral image has the same spatial resolution as the panchromatic image and then output it to obtain an upsampled low-resolution hyperspectral image; Splice the upsampled low-resolution hyperspectral image with the panchromatic image to obtain the initial input image.

3. The method according to claim 1, characterized in that The process of separating the high-frequency and low-frequency information of the initial input image through the high-low frequency separation module to obtain high-frequency information and low-frequency information includes: The high-low frequency separation module includes an average pooling layer and a dynamic upsampler; Downsample the initial input image based on the average pooling layer to obtain low-frequency information; Upsample the low-frequency information based on the dynamic upsampler to generate an upsampled low-frequency feature map; Subtract the upsampled low-frequency feature map from the original feature map to obtain high-frequency information.

4. The method according to claim 1, wherein The coordinate-spatial attention module contains multiple sequentially stacked Res-CSA blocks, and each Res-CSA block extracts spectral and spatial features through a parallel coordinate attention mechanism and a spatial attention mechanism.

5. The method according to claim 4, wherein The process of obtaining high-frequency features and low-frequency features includes: Perform a convolution operation on the high-frequency information and low-frequency information to obtain an intermediate feature map; Divide the intermediate feature map into two paths. One path obtains a spectral mask through coordinate attention, and the other path obtains a spatial mask through spatial attention; Multiply the spectral mask and the intermediate feature map element by element to obtain spectral features, and multiply the spatial mask and the intermediate feature map element by element to obtain spatial features; Add the spectral features, the spatial features, and the input feature map element by element to obtain the high-frequency features and low-frequency features.

6. The method according to claim 4, wherein The coordinate attention includes: two global average pooling layers respectively along the height and width dimensions, a 1x1 convolution layer for reducing the number of channels from 64 to 64 / r, a ReLU activation layer, two 1x1 convolution layers for expanding the number of channels from 64 / r to 64, and finally two sigmoid activation layers.

7. The method according to claim 4, characterized in that, The spatial attention includes: a parallel arrangement of a global average pooling layer and a global maximum pooling layer, a 1x1 convolution layer, and a sigmoid activation layer.

8. The method according to claim 2, wherein The process of fusing the high-frequency features and the low-frequency features based on an information fusion module to obtain a final high-resolution hyperspectral image includes: Upsample the low-frequency features, concatenate the upsampled low-frequency features and the high-frequency features, and compress the number of channels through a 1x1 convolutional layer to obtain fused features; Construct a pixel attention module based on two convolutional layers and one PA layer, input the fused features into the pixel attention module for calculation, and obtain the weighted fused features; Based on a convolutional layer, adjust the weighted fused features to a feature map with the same number of channels as the input hyperspectral image to obtain the adjusted feature map; Add the adjusted feature map and the upsampled low-resolution hyperspectral image pixel by pixel to obtain the final high-resolution hyperspectral image.

9. An electronic device, comprising a memory, a processor, and a computing program stored in the memory and executable on the processor, characterized in that When the processor executes the computing program, the method according to any one of claims 1-8 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1-8 is implemented.