A remote sensing image super-resolution fusion method based on frequency domain feature prior

By introducing the frequency domain feature prior and cross-attention module in the super-resolution fusion of remote sensing images, the FU-Net network is constructed, which solves the problems of existing methods in spectral distortion, insufficient flexibility and low computing efficiency, and achieves efficient image reconstruction effect.

CN120279368BActive Publication Date: 2025-09-02TIANJIN POLYTECHNIC UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510759225.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-02
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

The existing remote sensing image super-resolution fusion method has defects in spectral distortion, insufficient flexibility, low computational efficiency, and failure to fully utilize the differences in frequency and spatial domain characteristics, resulting in poor fusion of hyperspectral and multispectral images.

Method used

Using a method based on the feature prior of the frequency domain, by introducing frequency supplementary prior information and cross-attention module, dual feature learning and fusion are carried out in the frequency domain and spatial domain, FU-Net network is built, and multi-scale feature fusion is realized, improving reconstruction quality and adaptability.

Benefits of technology

It significantly improves the reconstruction accuracy and adaptability of super-resolution fusion of remote sensing images, enhances the spatial details and spectral information reconstruction quality of the image, overcomes the shortcomings of traditional methods, and performs outstandingly in detail reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279368B_ABST
    Figure CN120279368B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of remote sensing image technology and specifically discloses a remote sensing image super-resolution fusion method based on frequency domain feature priors, comprising the following steps: S1: acquiring a panchromatic image and a low-resolution hyperspectral image; S2: mapping the panchromatic image and the low-resolution hyperspectral image to the frequency domain, and then extracting a global frequency feature map; S3: inputting the panchromatic image, the low-resolution hyperspectral image, and the global frequency feature map into a FU-Net network, and fusing the panchromatic image and the low-resolution hyperspectral image in combination with the global frequency feature map to obtain fused features at different scales; S4: reconstructing the image using the fused features at different scales to obtain a high-resolution hyperspectral image. The present invention implements dual feature learning and fusion in both the frequency and spatial domains, significantly improving the performance of high-resolution super-resolution image fusion (HISR) tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing images, and in particular to a remote sensing image super-resolution fusion method based on frequency domain feature priors. Background Art

[0002] Hyperspectral imagery (HSI) can provide rich spectral information due to its high resolution in the spectral dimension, supporting fine material classification and identification. However, due to the limitations of sensor technology, HSI usually has a low spatial resolution, making it difficult to meet the demand for high-precision spatial information in some application scenarios. At the same time, multispectral imagery (MSI) has a high spatial resolution, but can only provide information in a small number of spectral bands, and its spectral resolution is far from sufficient. Therefore, a single type of image (HSI or MSI) often cannot meet the needs of both spectral and spatial information. To address this problem, a fusion-based hyperspectral image super-resolution (HISR) technology has been proposed to integrate the complementary information of HSI and MSI to generate high-resolution hyperspectral images (HR-HSI) with both high spectral resolution and high spatial resolution to meet the refinement requirements in practical applications.

[0003] Currently, existing HISR methods mainly include the following categories: extended PAN sharpening method, Bayesian inference-based method and decomposition-based method. Although traditional HISR methods have made some progress in solving the fusion task of hyperspectral and multispectral data, they still have some problems in practical applications:

[0004] 1) Extended PAN sharpening method: Since this type of method directly borrows PAN sharpening technology, it is easy to produce spectral distortion problems during the fusion process, resulting in the fusion result being not accurate in spectral information.

[0005] 2) Methods based on Bayesian inference: These methods rely on assumptions about prior distributions and exhibit poor flexibility under different data structures, making them difficult to adapt to complex and variable hyperspectral data.

[0006] 3) Decomposition-based methods: Matrix decomposition methods face bottlenecks when learning the relationship between space and spectrum. Although tensor decomposition can preserve the three-dimensional structural characteristics, it has high requirements for computing resources, which limits its efficiency in practical applications.

[0007] With the introduction of deep learning, CNN-based methods have improved the performance of HISR by learning local features through convolution operations. However, the convolution receptive field of CNN is limited, making it difficult to capture the long-range dependencies and global characteristics of images. This limitation restricts the performance of CNN methods in handling complex hyperspectral-multispectral fusion tasks. The Transformer architecture has been gradually applied to HISR tasks due to its advantages in capturing long-range dependencies. For example, the spectral-spatial transformer (SST) and other Transformer-based methods have improved fusion performance by exploring the global characteristics of hyperspectral and multispectral data. However, these methods still focus more on learning the long-range dependencies between HSI and MSI, but do not fully consider the differences in the characteristics of the two images in the frequency and spatial domains. Summary of the Invention

[0008] The present invention aims to solve the above problems. To this end, the present invention provides a remote sensing image super-resolution fusion method based on frequency domain feature priors. By introducing frequency-supplemented prior information and an innovative cross-self-attention module, dual feature learning and fusion are achieved in the frequency domain and the spatial domain, significantly improving the performance of the HISR task. The present invention constructs a network architecture for joint modeling of the frequency domain and the spatial domain to fully tap the complementary characteristics of the two types of images. The network can capture long-range dependencies in the frequency domain and the spatial domain through the FU-net structure, and at the same time combine the two characteristics for feature fusion. By gradually integrating multi-scale features, the reconstruction quality of spatial details and spectral information is improved, and the fusion results are further optimized. It not only overcomes the shortcomings of traditional methods in terms of spectral distortion, lack of flexibility, and low computational efficiency, but also effectively improves the reconstruction accuracy and adaptability of deep learning and Transformer methods in HISR tasks.

[0009] The present invention provides a remote sensing image super-resolution fusion method based on frequency domain feature priors, and the technical solution adopted is as follows: comprising the following steps:

[0010] S1: Acquire panchromatic images and low-resolution hyperspectral images;

[0011] S2: Map the panchromatic image and low-resolution hyperspectral image to the frequency domain, and then extract the global frequency feature map;

[0012] S3: The panchromatic image, low-resolution hyperspectral image and global frequency feature map are input into the FU-Net network. Combined with the global frequency feature map, the panchromatic image and low-resolution hyperspectral image are fused to obtain fusion features of different scales.

[0013] S4: Use fusion features of different scales to reconstruct images and obtain high-resolution hyperspectral images.

[0014] Furthermore, in S2, the panchromatic image and the low-resolution hyperspectral image are respectively mapped from the spatial domain to the frequency domain through the Fourier transform module to obtain low-frequency domain features and high-frequency domain features;

[0015] The low-frequency frequency domain features and high-frequency frequency domain features are input into the FRB module to extract the global frequency features and obtain the global frequency feature map.

[0016] Furthermore, the formula for the FRB module is:

[0017]

[0018]

[0019]

[0020]

[0021] in, represents the low-frequency domain characteristics, represents the high-frequency domain characteristics, Indicates splicing, express and The feature map after fusion, represents a 1x1 convolution operation, represents element-wise multiplication, Represents the high-frequency feature map, represents the convolution operation, Represents the feature map of deep feature extraction after multiple convolution kernels, Represents the global frequency feature map.

[0022] Furthermore, in S3, in the FU-Net network, the low-resolution hyperspectral image is scaled and the number of channels is adjusted by downsampling or upsampling at each stage to obtain the output features of the low-resolution hyperspectral image;

[0023] At each stage, the panchromatic image is downsampled or upsampled to obtain the panchromatic image output features that are aligned with the low-resolution hyperspectral image output feature scale and channel level. The panchromatic image output features are fused with the low-resolution hyperspectral image output features through the cross transformation module to obtain the output features of the cross transformation module.

[0024] For the output features of the cross-transformation module with a corresponding global frequency feature map, the output features of the cross-transformation module are multiplied by the global frequency feature map to obtain a fused feature; for the output features of the cross-transformation module without a corresponding global frequency feature map, the output features of the cross-transformation module are directly used as the fused feature; and fused features of different scales are obtained.

[0025] Furthermore, downsampling uses a downsampling module with a 2×2 convolution layer with a stride of 2, and upsampling uses a 2×2 upsampling transposed convolution layer with a stride of 2.

[0026] Furthermore, the working process of the cross-transformation module is as follows:

[0027] The input image features are resized using a reshape operation and then mapped through a linear transformation layer to generate query vectors, key vectors, and value vectors respectively.

[0028] The query vector, key vector and value vector are used to perform cross-attention weighting to obtain the cross-attention weighted PAN features and MSI features;

[0029] The global frequency feature map, the cross-attention weighted PAN feature and the MSI feature are fused, or the cross-attention weighted PAN feature and the MSI feature are fused; fusion features of different scales are obtained.

[0030] Furthermore, in S2, the number of FRB modules is 3, and the input features of some FRB modules are pre-processed by downsampling;

[0031] In S3, the FU-Net network consists of 5 stages;

[0032] The formula for fusion of 5 scale features is expressed as:

[0033]

[0034]

[0035] in, represents the feature fusion of the i-th stage, represents the PAN feature after cross attention weighting in stage i, represents the MSI feature after cross-attention weighting in stage i, Represents the global frequency feature map integrated into the i-th stage, Represents element-wise multiplication.

[0036] Furthermore, the image reconstruction is expressed as:

[0037]

[0038]

[0039]

[0040] in, represents a 3x3 convolution operation, represents the first reconstruction stage characteristics, represents the characteristics of the second reconstruction stage, represents a high-resolution hyperspectral image, Indicates splicing, represents a 2×2 upsampling transposed convolutional layer with a stride of 2, Represents the low-resolution hyperspectral image input to the FU-Net network.

[0041] The above one or more technical solutions in the embodiments of the present invention have at least one of the following technical effects:

[0042] 1. This method, based on frequency domain feature priors and a cross-Transformer, ensures the effective fusion of multi-stage spatial and spectral features, significantly enhancing the model's super-resolution reconstruction performance, particularly in detail reconstruction. Compared to various existing fusion algorithms and deep learning methods on multiple public datasets, it outperforms multiple evaluation metrics, including spectral correlation coefficient and image quality index.

[0043] 2. This invention combines frequency domain prior information to provide a wealth of global and detailed feature complements for high-resolution hyperspectral image reconstruction. By mapping the image from the spatial domain to the frequency domain through Fourier transform, it effectively separates low-frequency and high-frequency components. Low frequencies provide global context, while high frequencies focus on local details, avoiding global information loss caused by local convolution in the spatial domain. The FRB module deeply mines frequency domain information, extracting global frequency features that serve as important prior knowledge for subsequent processing.

[0044] 3. The FU-Net network of the present invention adopts a dual-branch multi-scale feature interaction strategy. The LR-HS branch and the frequency-spectral branch implement downsampling, upsampling, and cross-transformation modules to efficiently integrate spatial and spectral information, enhance feature expression capabilities, and improve image reconstruction quality and accuracy. The cross-transformation module leverages the Transformer self-attention mechanism to deeply mine and fuse the complementary information of the panchromatic and multispectral images, generating richer and more accurate feature representations.

[0045] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0047] Figure 1 It is a flow chart of the method provided by the present invention.

[0048] Figure 2 It is a structural diagram of the FRB module provided by the present invention.

[0049] Figure 3 It is a structural diagram of the cross-conversion module provided by the present invention.

[0050] Figure 4 It is a structural diagram of the image reconstruction module provided by the present invention.

[0051] Figure 5 This is a subjective comparison experiment diagram of the Pavia Center dataset provided by the present invention. DETAILED DESCRIPTION

[0052] To make the purpose, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention. Obviously, the embodiments described are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0053] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the embodiment of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0054] The following combination Figures 1 to 5 The present invention is further described in detail, and a remote sensing image super-resolution fusion method based on frequency domain feature prior is described as follows:

[0055] In this embodiment, Figure 1 As shown in the figure, a remote sensing image super-resolution fusion method based on frequency domain feature prior is provided. The method combines the frequency domain prior information to provide a large number of global and detail features for the high-resolution hyperspectral image (HR-HS) reconstruction task, including the following steps:

[0056] S1: Acquire panchromatic images (PAN) and low-resolution hyperspectral images (LR-HS).

[0057] Since the low-resolution hyperspectral image obtained in this embodiment is relatively small in size, an upsampling step is performed first.

[0058] S2: Map the panchromatic image and low-resolution hyperspectral image to the frequency domain, and then extract the global frequency feature map.

[0059] First, the panchromatic image and the low-resolution hyperspectral image are passed through the Fourier transform module to obtain low-frequency and high-frequency domain features, respectively. Mapping the image from the spatial domain to the frequency domain can effectively separate the low-frequency and high-frequency components of the image. Among them, the low-frequency component represents the overall structure and spectral distribution of the image, providing stable global context information for the network; the high-frequency component focuses on local details such as edges and textures, compensating for the shortcomings of traditional spatial domain convolution in modeling long-range dependencies.

[0060] For multispectral images ,in, Represents the spatial coordinates of the image, and its two-dimensional discrete Fourier transform formula is:

[0061]

[0062] in, The image is The size in the direction, The image is The size in the direction, is the frequency domain coordinate, is the frequency domain representation after transformation, i.e. The basic form of isofrequency domain feature representation.

[0063] The Fourier transform module, based on the Fourier transform principle, extracts global frequency features. These features characterize different image attributes. For example, low-frequency components correspond to global information such as the image's overall structure, main shape, and color distribution, while high-frequency components contain detailed information such as edges and textures. By extracting features in the frequency domain, it effectively avoids the global information loss that can occur when using only local convolution operations in the spatial domain.

[0064] The low-frequency frequency domain features and high-frequency domain features The FRB module performs deep mining on the image information in the frequency domain, extracts the global frequency features contained therein, and obtains the prior global frequency feature map. and The network is linked by splicing (Cat) operations, and then global information is extracted through a series of deep convolution layers. By extracting features in the frequency domain, the problem of global information loss that may be caused by using only local convolution operations in the spatial domain is effectively avoided. The extracted frequency domain information features will serve as important prior knowledge for subsequent processing, such as Figure 2 The specific formula is as follows:

[0065]

[0066]

[0067]

[0068]

[0069] in, represents the low-frequency domain characteristics, represents the high-frequency domain characteristics, Indicates splicing, express and The feature map after fusion, represents a 1x1 convolution operation, represents element-wise multiplication, Represents the high-frequency feature map, represents the convolution operation, Represents the feature map of deep feature extraction after multiple convolution kernels, Represents the global frequency feature map.

[0070] like Figure 1 As shown in Figure 1, this embodiment uses three FRB modules. In order to correspond to the dimensions in FU-Net, the input features of the last two FRB modules are downsampled at different scales. That is, the input features of the first FRB module are the original low-frequency domain features and high-frequency domain features, and the output is The input features of the second FRB module are the low-frequency domain features and high-frequency domain features that have been downsampled (half) once, and the output is The input features of the third FRB module are the low-frequency domain features and high-frequency domain features that have been downsampled twice, and the output is .

[0071] S3: The panchromatic image, low-resolution hyperspectral image, and global frequency feature map are fed into the FU-Net network. Combined with the global frequency feature map, the panchromatic image and low-resolution hyperspectral image are fused to obtain fusion features at different scales. The prior global frequency feature map is fed into the FU-Net network, and the panchromatic image and low-resolution hyperspectral image are deeply fused by combining residual blocks of different scales and convolutional layers.

[0072] In this embodiment, the FU-Net network has a five-layer structure, and processes the panchromatic image and the low-resolution hyperspectral image in five stages. The three global frequency feature maps obtained in step S2 are integrated into the last three stages.

[0073] Specifically, the FU-Net network consists of two branches, namely the LR-HS branch and the frequency domain-spectral branch.

[0074] In the LR-HS branch of the FU-Net network, at each stage, the low-resolution hyperspectral image is scaled and the number of channels adjusted through downsampling or upsampling to obtain the low-resolution hyperspectral image output features. This process aims to refine and enhance the expressive power of the features. According to the conventional structure of the FU-Net network, downsampling is performed in the first two stages, and upsampling is performed in the third and fourth stages. The specific operation can be described by the following formula:

[0075]

[0076]

[0077] in, represents the output features of the low-resolution hyperspectral image at stage i, Represents the low-resolution hyperspectral image input to the FU-Net network. Represents the residual module. Represents the downsampling module of a 2×2 convolutional layer with a stride of 2, which increases the number of channels of the feature map through a layer of depthwise separable convolution. Represents a 2×2 upsampling transposed convolution layer with a stride of 2, which can achieve upsampling and channel reduction at the same time. The processed spatial texture detail features are transferred to the next-scale module for more refined feature fusion.

[0078] In the frequency-spectral branch of the FU-Net network, at each stage, the panchromatic image is downsampled or upsampled to produce panchromatic image output features that are scale- and channel-aligned with the output features of the low-resolution hyperspectral image, ensuring that the spatial dimensions of the panchromatic image output features match those of the low-resolution hyperspectral image. The panchromatic image output features are then fused with the low-resolution hyperspectral image output features via a cross-transformer block (CTB) to produce the output features of the cross-transformer block. Finally, the output features of the cross-transformer block are multiplied with the corresponding global frequency feature map to achieve effective feature fusion and enhance the spatial detail and spectral accuracy of the low-resolution hyperspectral image. In the first few stages (downsampling), where there is no corresponding global frequency feature map, the output features of the cross-transformer block are directly used as the fused features to be output in this step. Specifically, for cross-transformer block output features with corresponding global frequency feature maps, the cross-transformer block output features are multiplied with the global frequency feature map to produce fused features. For cross-transformer block output features without corresponding global frequency feature maps, the cross-transformer block output features are directly used as the fused features. Ultimately, fused features of different scales are obtained. The specific operation can be expressed as:

[0079]

[0080]

[0081] in, represents the output features of the panchromatic image at stage i, Represents the full-color image input to the FU-Net network.

[0082]

[0083]

[0084] in, represents the cross-transformation module, represents the fusion features of the i-th stage, Represents the global frequency feature map integrated into stage i. There is no corresponding global frequency feature map for stages 1 and 2, but there is a corresponding global frequency feature map for stages 3-5.

[0085] The cross-transformation module is a key module for multi-source image data fusion processing in this network algorithm. This module mainly processes the input full-color image and multi-spectral image feature map. The core theory is based on the self-attention mechanism in the Transformer architecture, which can deeply mine and fuse the complementary information of the two types of images to generate richer and more accurate feature representations. Figure 3As shown, the working process of the cross-transformation module can be expressed as:

[0086] First, the input image features are resized using the Reshape operation to adapt them to the subsequent linear transformation operation. Then, they are mapped through the Linear transformation layer to generate the query vector (Query, Q), key vector (Key, K), and value vector (Value, V). The formula is expressed as:

[0087]

[0088]

[0089]

[0090] in, represents the reshape layer, represents the linear transformation layer, represents the query vector of the i-th stage, represents the key vector of stage i, A vector of values ​​representing the i-th stage.

[0091] Then, the cross attention is weighted and the formula is expressed as:

[0092]

[0093] in, represents the activation layer, represents a 1x3x3 convolution operation, represents the PAN feature after cross attention weighting in stage i, represents a 3x3 convolution operation, represents the MSI feature after cross-attention weighting in the i-th stage.

[0094] Finally, feature fusion is performed to obtain fusion features of different scales. The last three stages also need to incorporate the global frequency feature map. The formula is expressed as:

[0095]

[0096] .

[0097] The dual-branch multi-scale feature interaction strategy can ensure the efficient integration of spatial and spectral information, thereby improving the quality and accuracy of the final image reconstruction. When using the FU-Net structure for HISR fusion tasks, feature maps of different scales can obtain different detail features of the original image, thereby enhancing the feature expression capability. At the same time, the introduction of the Transformer structure in the encoder stage can effectively capture global context information and promote the understanding of image semantics. The decoder stage is responsible for mapping the features extracted by the encoder back to the spatial domain to generate a fusion result. At this time, the Fourier transform is no longer applicable, but the Transformer structure can still be used to focus on spatial features and assist in the generation of high-resolution information. This method is based on frequency domain feature priors and cross-Transformers. By ensuring the effective fusion of multi-stage spatial-spectral features, it greatly enhances the super-resolution reconstruction performance of the model, especially in detail reconstruction.

[0098] S4: Use fusion features of different scales to reconstruct images and obtain high-resolution hyperspectral images.

[0099] In the final image reconstruction stage, multi-scale feature fusion is used to progressively combine the original spectral features from the LR-HS branch with the reconstructed features from the frequency-spectral branch. After obtaining the output of the final stage, a 3×3 convolution is used to map the fused features into the final high-resolution hyperspectral image.

[0100] like Figure 4 As shown, this embodiment uses the image reconstruction module to reconstruct the five fusion features of different scales output by the five cross-transformation modules and the low-resolution hyperspectral image to obtain a high-resolution hyperspectral image. It is expressed as:

[0101]

[0102]

[0103]

[0104] in, represents a 3x3 convolution operation, represents the first reconstruction stage characteristics, represents the characteristics of the second reconstruction stage, Represents high-resolution hyperspectral imagery.

[0105] This example compares this method with existing excellent fusion algorithms and deep learning methods on three public datasets. The two fusion algorithms are PCA and GFPCA, and the seven deep learning methods are DARN, HSRNet, HyperPNN, SSFCNN, SSRNet, HyperDSNet, and HyperRefiner. The evaluation metrics used are SCC (spectral correlation coefficient), SAM (spectral angle mapper), ERGAS (relative global error), UIQI (universal image quality index), SSIM (structural similarity), PSNR (peak signal-to-noise ratio), and TestTime (test time, in milliseconds). ↑ indicates higher values, while ↓ indicates lower values. The three public datasets are Pavia Center, Botswana, and Chikusei. The experimental results are shown in Table 1.

[0106] Table 1 Average quantitative index results on three different datasets

[0107]

[0108] This example also conducted a subjective comparison experiment on the Pavia Center dataset. The experimental results are as follows: Figure 5 The experimental results demonstrate the leading performance of this method in the field of HISR technology.

[0109] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A remote sensing image super-resolution fusion method based on frequency domain feature priors, characterized by: The following steps are involved: S1: Acquire panchromatic images and low-resolution hyperspectral images; S2: Map the panchromatic image and low-resolution hyperspectral image to the frequency domain, and then extract the global frequency feature map; In S2, the panchromatic image and the low-resolution hyperspectral image are respectively mapped from the spatial domain to the frequency domain through the Fourier transform module to obtain low-frequency domain features and high-frequency domain features; Input the low-frequency domain features and high-frequency domain features into the FRB module, extract the global frequency features, and obtain the global frequency feature map; The formula of the FRB module is: in, represents the low-frequency domain characteristics, represents the high-frequency domain characteristics, Indicates splicing, express and The feature map after fusion, represents a 1x1 convolution operation, represents element-wise multiplication, Represents the high-frequency feature map, represents the convolution operation, Represents the feature map of deep feature extraction after multiple convolution kernels, Represents the global frequency feature map; S3: Input the panchromatic image, low-resolution hyperspectral image, and global frequency feature map into the FU-Net network; combine the global frequency feature map to fuse the panchromatic image and low-resolution hyperspectral image to obtain fusion features of different scales; In S3, in the FU-Net network, the low-resolution hyperspectral image is scaled and the number of channels is adjusted by downsampling or upsampling at each stage to obtain the low-resolution hyperspectral image output features; At each stage, the panchromatic image is downsampled or upsampled to obtain the panchromatic image output features that are aligned with the low-resolution hyperspectral image output feature scale and channel level; the panchromatic image output features are fused with the low-resolution hyperspectral image output features through the cross transformation module to obtain the output features of the cross transformation module; For the output features of the cross-transformation module with the corresponding global frequency feature map, the output features of the cross-transformation module are multiplied by the global frequency feature map to obtain a fused feature; for the output features of the cross-transformation module without the corresponding global frequency feature map, the output features of the cross-transformation module are directly used as the fused feature; and fused features of different scales are obtained; S4: Use fusion features of different scales to reconstruct images and obtain high-resolution hyperspectral images.

2. The remote sensing image super-resolution fusion method based on frequency domain feature prior according to claim 1, characterized in that: The downsampling module uses a 2×2 convolution layer with a stride of 2, and the upsampling module uses a 2×2 upsampling transposed convolution layer with a stride of 2.

3. The remote sensing image super-resolution fusion method based on frequency domain feature prior according to claim 1, characterized in that: The working process of the cross conversion module is: The input image features are resized using a reshape operation and then mapped through a linear transformation layer to generate query vectors, key vectors, and value vectors respectively. The query vector, key vector and value vector are used to perform cross-attention weighting to obtain the cross-attention weighted PAN features and MSI features; The global frequency feature map, the cross-attention weighted PAN feature and the MSI feature are fused, or the cross-attention weighted PAN feature and the MSI feature are fused; fusion features of different scales are obtained.

4. The remote sensing image super-resolution fusion method based on frequency domain feature prior according to claim 3, characterized in that: In S2, the number of FRB modules is 3, and the input features of some FRB modules are preprocessed by downsampling; In S3, the FU-Net network consists of 5 stages; The formula for feature fusion is expressed as: in, represents the feature fusion of the i-th stage, represents the PAN feature after cross attention weighting in stage i, represents the MSI feature after cross-attention weighting in stage i, Represents the global frequency feature map integrated into the i-th stage, Represents element-wise multiplication.

5. The remote sensing image super-resolution fusion method based on frequency domain feature prior according to claim 4, characterized in that: Image reconstruction is expressed as: in, represents a 3x3 convolution operation, represents the first reconstruction stage characteristics, represents the characteristics of the second reconstruction stage, represents a high-resolution hyperspectral image, Indicates splicing, represents a 2×2 upsampling transposed convolutional layer with a stride of 2, Represents the low-resolution hyperspectral image input to the FU-Net network.

Citation Information

Patent Citations

  • Multispectral and hyperspectral image fusion method based on spatial frequency collaborative network

    CN119314012A

  • High-quality HRHS image generation method based on double-subnetwork structure

    CN119625112A