Remote sensing image super-resolution fusion method based on frequency domain feature prior

By performing feature learning and fusion in the frequency domain and spatial domain, using the FU-Net network of the frequency domain feature prior and cross-attention module, the problems of spectral distortion and low computing efficiency in the super-resolution fusion of remote sensing images are solved, and efficient image reconstruction effect is achieved.

CN120279368AActive Publication Date: 2025-07-08TIANJIN POLYTECHNIC UNIV

Patent Information

Application Number
CN202510759225.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-08
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

Existing remote sensing image super-resolution fusion methods have shortcomings in spectral distortion, insufficient flexibility, low computational efficiency, and difficulty in capturing long-distance dependencies and global characteristics of images, especially the performance of CNN-based methods is limited when dealing with hyperspectral-multi-spectral fusion tasks.

Method used

Using a method based on the prior feature of the frequency domain, by introducing frequency supplementary prior information and cross-attention modules, dual feature learning and fusion are realized in the frequency domain and spatial domain, a network architecture with joint modeling of the frequency domain and spatial domain is constructed, and multi-scale feature fusion is used to improve reconstruction quality and adaptability.

Benefits of technology

It significantly improves the reconstruction accuracy and adaptability of super-resolution fusion of remote sensing images, overcomes the shortcomings of traditional methods, enhances the spatial details and spectral information reconstruction quality of the image, and improves the performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279368A_ABST
    Figure CN120279368A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of remote sensing images, and particularly discloses a frequency domain feature prior-based remote sensing image super-resolution fusion method, which comprises the steps of S1, acquiring a panchromatic image and a low-resolution hyperspectral image; s2, mapping the panchromatic image and the low-resolution hyperspectral image to a frequency domain, and then extracting to obtain a global frequency feature map; s3, inputting the panchromatic image, the low-resolution hyperspectral image and the global frequency feature map into an FU-Net network, and fusing the panchromatic image and the low-resolution hyperspectral image in combination with the global frequency feature map to obtain fusion features of different scales; and S4, performing image reconstruction by using the fusion features of different scales to obtain a high-resolution hyperspectral image. According to the method, dual feature learning and fusion are realized in a frequency domain and a space domain, and the performance of an HISR task is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing images, and particularly to a remote sensing image super-resolution fusion method based on frequency-domain feature prior. Background Art

[0002] Hyperspectral images (HSIs) can provide rich spectral information for fine material classification and recognition due to their high resolution in the spectral dimension. However, due to the limitations of sensor technology, HSIs usually have low spatial resolution, making it difficult to meet the demand for high-precision spatial information in some application scenarios. At the same time, multispectral images (MSIs) have high spatial resolution but can only provide information in a small number of spectral bands, and their spectral resolution is far from sufficient. Therefore, a single type of image (HSI or MSI) often has difficulty in meeting the requirements of both spectral and spatial information. To address this issue, a fusion-based hyperspectral image super-resolution (HISR) technique has been proposed to integrate the complementary information of HSIs and MSIs to generate high-resolution hyperspectral images (HR-HSIs) with both high spectral and spatial resolutions to meet the refined requirements in practical applications.

[0003] Currently, existing HISR methods mainly include the following categories: extended PAN sharpening methods, Bayesian inference-based methods, and decomposition-based methods. Although traditional HISR methods have made some progress in solving the fusion task of hyperspectral and multispectral data, there are still some problems in their practical applications: 1) Extended PAN sharpening methods: Since these methods directly borrow PAN sharpening technology, they are prone to spectral distortion problems during the fusion process, resulting in inaccurate spectral information in the fusion results.

[0004] 2) Bayesian inference-based methods: These methods rely on assumptions about prior distributions and show poor flexibility under different data structures, making it difficult to adapt to complex and variable hyperspectral data.

[0005] 3) Decomposition-based methods: Matrix decomposition methods face bottlenecks in learning the relationship between space and spectrum, while tensor decomposition can retain three-dimensional structural characteristics but has high computational resource requirements, limiting the efficiency of its practical applications.

[0006] With the introduction of deep learning, CNN-based methods learn local features through convolutional operations, improving the performance of HISR. However, the convolutional receptive field of CNN is limited, making it difficult to capture long-range dependencies and global characteristics of images. This limitation restricts the performance of CNN methods in dealing with complex hyperspectral-multispectral fusion tasks. Due to its advantage in capturing long-range dependencies, the Transformer architecture has been gradually applied to HISR tasks. For example, spectral-spatial Transformer (SST) and other Transformer-based methods improve the fusion performance by exploring the global characteristics of hyperspectral and multispectral data. However, these methods still focus more on learning the long-range dependencies between HSI and MSI, but do not fully consider the characteristic differences between the two images in the frequency domain and spatial domain. Summary of the Invention

[0007] The present invention aims to solve the above problems. To this end, the present invention provides a remote sensing image super-resolution fusion method based on frequency-domain feature prior, which realizes dual feature learning and fusion in the frequency domain and spatial domain by introducing prior information of frequency supplementation and an innovative cross self-attention module, significantly improving the performance of the HISR task. The present invention constructs a network architecture for joint modeling of the frequency domain and spatial domain, fully exploiting the complementary characteristics of the two types of images. The network can capture long-range dependencies in the frequency domain and spatial domain through the FU-net structure, and at the same time combine the two characteristics for feature fusion, and improve the reconstruction quality of spatial details and spectral information by gradually integrating multi-scale features, further optimizing the fusion result. It not only overcomes the defects of traditional methods in terms of spectral distortion, lack of flexibility, low computational efficiency, etc., but also effectively improves the reconstruction accuracy and adaptability of deep learning-based and Transformer-based methods in the HISR task.

[0008] The present invention provides a remote sensing image super-resolution fusion method based on frequency-domain feature prior, and the technical solution adopted is as follows: including the following steps: S1: Obtain a panchromatic image and a low-resolution hyperspectral image; S2: Map the panchromatic image and the low-resolution hyperspectral image to the frequency domain, and then extract the global frequency feature map; S3: Input the panchromatic image, the low-resolution hyperspectral image and the global frequency feature map into the FU-Net network, and combine the global frequency feature map to fuse the panchromatic image and the low-resolution hyperspectral image to obtain fusion features of different scales; S4: Use the fusion features of different scales for image reconstruction to obtain a high-resolution hyperspectral image.

[0009] Further, in S2, the panchromatic image and the low-resolution hyperspectral image are respectively passed through a Fourier transform module to be mapped from the spatial domain to the frequency domain, obtaining low-frequency frequency-domain features and high-frequency frequency-domain features; The low-frequency frequency-domain features and the high-frequency frequency-domain features are input into the FRB module to extract global frequency features, obtaining a global frequency feature map.

[0010] Further, the formula of the FRB module is: Among them, represents the low-frequency frequency-domain features, represents the high-frequency frequency-domain features, represents concatenation, represents and the feature map after fusion of represents a 1x1 convolution operation, represents element-wise multiplication, represents the high-frequency feature map, represents a convolution operation, represents the feature map obtained by performing depth feature extraction through multiple convolutional kernels, represents the global frequency feature map.

[0011] Further, in S3, in the FU-Net network, at each stage, the low-resolution hyperspectral image is subjected to scale transformation and adjustment of the number of channels through downsampling or upsampling, obtaining low-resolution hyperspectral image output features; At each stage, the panchromatic image undergoes downsampling or upsampling to obtain panchromatic image output features aligned with the low-resolution hyperspectral image output features in terms of scale and channel layer. The panchromatic image output features and the low-resolution hyperspectral image output features are subjected to feature fusion through a cross-transformation module, obtaining the output features of the cross-transformation module; For the output features of the cross-transformation module with corresponding global frequency feature maps, the output features of the cross-transformation module are multiplied by the global frequency feature maps to obtain fusion features; for the output features of the cross-transformation module without corresponding global frequency feature maps, the output features of the cross-transformation module are directly used as fusion features; different-scale fusion features are obtained.

[0012] Further, the downsampling uses a downsampling module with a 2×2 convolutional layer with a stride of 2, and the upsampling uses a 2×2 transposed convolutional layer with a stride of 2.

[0013] Furthermore, the working process of the cross-transformation module is as follows: The dimensionality of the input image features is adjusted using a reshaping operation, and then mapped through a linear transformation layer to generate query vectors, key vectors, and value vectors respectively; The query vectors, key vectors, and value vectors are used for weighted cross-attention to obtain the PAN features and MSI features after cross-attention weighting; The global frequency feature map, the PAN features and MSI features after cross-attention weighting are subjected to feature fusion, or the PAN features and MSI features after cross-attention weighting are subjected to feature fusion; to obtain fused features of different scales.

[0014] Furthermore, in S2, the number of FRB modules is 3, and the input features of some FRB modules are pre-processed by downsampling; In S3, the FU-Net network includes 5 stages; The formula for feature fusion at 5 scales is expressed as: Among them, represents the feature fusion at the i-th stage, represents the PAN features after cross-attention weighting at the i-th stage, represents the MSI features after cross-attention weighting at the i-th stage, represents the global frequency feature map incorporated at the i-th stage, represents element-wise multiplication.

[0015] Furthermore, the image reconstruction is expressed as: Among them, represents a 3x3 convolution operation, represents the features of the first reconstruction stage, represents the features of the second reconstruction stage, represents the high-resolution hyperspectral image, represents concatenation, represents a 2×2 transposed convolutional layer with a stride of 2 for upsampling, represents the low-resolution hyperspectral image input to the FU-Net network.

[0016] One or more of the above technical solutions in the embodiments of the present invention have at least one of the following technical effects: 1. The present invention is based on frequency-domain feature priors and cross-Transformer, ensuring the effective fusion of multi-stage spatial-spectral features, greatly enhancing the super-resolution reconstruction performance of the model, especially outstanding in detail reconstruction. Compared with a variety of existing fusion algorithms and deep learning methods on multiple public datasets, it performs better in multiple evaluation metrics such as spectral correlation coefficient and image quality index.

[0017] 2. The present invention combines frequency-domain prior information to provide a large number of global and detailed feature supplements for high-resolution hyperspectral image reconstruction. The image is mapped from the spatial domain to the frequency domain through Fourier transform, effectively separating the low-frequency and high-frequency components. The low-frequency provides global context, and the high-frequency focuses on local details, avoiding the loss of global information caused by local convolution in the spatial domain. The FRB module deeply mines the frequency-domain information to extract global frequency features, which serve as important prior knowledge for subsequent processing.

[0018] 3. The FU-Net network of the present invention adopts a dual-branch multi-scale feature interaction strategy. The LR-HS branch and the frequency-domain-spectral branch perform operations such as downsampling, upsampling, and cross-transformation modules to achieve efficient integration of spatial and spectral information, enhance the feature expression ability, and improve the quality and accuracy of image reconstruction. The cross-transformation module relies on the Transformer self-attention mechanism to deeply mine and fuse the complementary information of the panchromatic image and the multispectral image to generate a richer and more accurate feature representation.

[0019] Additional aspects and advantages of the present invention will be given in part in the following description, will become apparent in part from the following description, or will be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the embodiments or the description of the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0021] Figure 1 is the flowchart of the method provided by the present invention.

[0022] Figure 2 is the structural diagram of the FRB module provided by the present invention.

[0023] Figure 3 is the structural diagram of the cross-transformation module provided by the present invention.

[0024] Figure 4 is the structural diagram of the image reconstruction module provided by the present invention.

[0025] Figure 5It is the subjective comparison experiment diagram of the Pavia Center dataset provided by the present invention. Detailed implementation manners

[0026] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, rather than all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention. The following embodiments are used to illustrate the present invention, but cannot be used to limit the scope of the present invention.

[0027] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0028] The following is in conjunction with Figures 1 to 5 A further detailed description of the present invention is made to describe a remote sensing image super-resolution fusion method based on frequency domain feature prior of the present invention: In this embodiment, as Figure 1 shown, a remote sensing image super-resolution fusion method based on frequency domain feature prior is provided, which combines frequency domain prior information to provide a large amount of global and detailed feature supplements for the high-resolution hyperspectral image (HR-HS) reconstruction task, including the following steps: S1: Obtain a panchromatic image (PAN) and a low-resolution hyperspectral image (LR-HS).

[0029] Since the size of the low-resolution hyperspectral image obtained in this embodiment is small, a step of upsampling is first performed.

[0030] S2: Map the panchromatic image and the low-resolution hyperspectral image to the frequency domain, and then extract the global frequency feature map.

[0031] First, the panchromatic image and the low-resolution hyperspectral image are respectively passed through the Fourier transform module to obtain the low-frequency frequency-domain features and the high-frequency frequency-domain features. Mapping the image from the spatial domain to the frequency domain can effectively separate the low-frequency and high-frequency components of the image. Among them, the low-frequency component represents the overall structure and spectral distribution of the image, providing stable global context information for the network; the high-frequency component focuses on local details such as edges and textures, making up for the deficiency of traditional spatial-domain convolution in long-range dependence modeling.

[0032] For the multispectral image , where represents the spatial coordinates of the image, and its two-dimensional discrete Fourier transform formula is: where is the size of the image in the direction, is the size of the image in the direction, is the frequency-domain coordinate, is the transformed frequency-domain representation, that is, the basic form of the frequency-domain features such as in the image.

[0033] The Fourier transform module extracts the global frequency features contained therein based on the Fourier transform principle. These features can characterize different attributes of the image. For example, the low-frequency part corresponds to the global information such as the overall structure, main shape, and color distribution of the image, and the high-frequency part contains the detailed information such as the edges and textures of the image. By extracting features in the frequency domain, the problem of global information loss that may occur when only using local convolution operations in the spatial domain is effectively avoided.

[0034] The low-frequency frequency-domain features and the high-frequency frequency-domain features are input into the FRB module together. The FRB module deeply mines the information of the image in the frequency domain, extracts the global frequency features contained therein, and obtains the prior global frequency feature map. In the FRB module, and are linked through the concatenation (Cat) operation, and then a series of deep convolutional layers are used to extract global information. By extracting features in the frequency domain, the problem of global information loss that may occur when only using local convolution operations in the spatial domain is effectively avoided. The extracted frequency-domain information features will be used as important prior knowledge for subsequent processing. As shown in Figure 2 , the specific formula is as follows: Among them, represents the low-frequency frequency domain feature, represents the high-frequency frequency domain feature, represents concatenation, represents and the feature map after fusion of represents the 1x1 convolution operation, represents element-wise multiplication, represents the high-frequency feature map, represents the convolution operation, represents the feature map obtained by performing deep feature extraction through multiple convolutional kernels, represents the global frequency feature map.

[0035] As Figure 1 shown, in this embodiment, three FRB modules are used. In order to correspond to the dimensions in FU-Net, the input features of the latter two FRB modules are downsampled at different scales. That is, the input features of the first FRB module are the original low-frequency frequency domain feature and high-frequency frequency domain feature, and the output is ; the input features of the second FRB module are the low-frequency frequency domain feature and high-frequency frequency domain feature that have been downsampled once (by half), and the output is ; the input features of the third FRB module are the low-frequency frequency domain feature and high-frequency frequency domain feature that have been downsampled twice, and the output is .

[0036] S3: Input the panchromatic image, low-resolution hyperspectral image, and global frequency feature map into the FU-Net network. Combining the global frequency feature map, fuse the panchromatic image and the low-resolution hyperspectral image to obtain fusion features at different scales. Input the prior global frequency feature map into the FU-Net network, and perform deep fusion on the panchromatic image and the low-resolution hyperspectral image by combining different-scale residual blocks and convolutional layers.

[0037] In this embodiment, the FU-Net network has a 5-layer structure, processes the panchromatic image and the low-resolution hyperspectral image in 5 stages, and the three global frequency feature maps obtained in step S2 are incorporated in the latter 3 stages.

[0038] Specifically, the FU-Net network includes 2 branches, namely the LR-HS branch and the frequency domain-spectral branch.

[0039] In the LR-HS branch of the FU-Net network, at each stage, scale transformation and adjustment of the number of channels are performed on the low-resolution hyperspectral image through downsampling or upsampling to obtain the output features of the low-resolution hyperspectral image. This process aims to refine and strengthen the feature expression ability. According to the conventional structure of the FU-Net network, downsampling is performed in the 1st - 2nd stages, and upsampling is performed in the 3rd - 4th stages. The specific operations can be described by the following formula: Among them, represents the output features of the low-resolution hyperspectral image in the i-th stage, represents the low-resolution hyperspectral image input to the FU-Net network. represents the residual module. represents the downsampling module of a 2×2 convolutional layer with a stride of 2, which increases the number of channels of the feature map through a depthwise separable convolution. represents a 2×2 transposed convolutional layer for upsampling with a stride of 2, which can simultaneously perform upsampling and channel reduction. The processed spatial texture detail features are transmitted to the module at the next scale for finer feature fusion.

[0040] In the frequency domain - spectral branch of the FU-Net network, at each stage, the panchromatic image undergoes downsampling or upsampling to obtain the output features of the panchromatic image that are aligned with the output features of the low-resolution hyperspectral image in terms of scale and channel layer, so that the spatial size of the output features of the panchromatic image matches that of the output features of the low-resolution hyperspectral image. Then, the output features of the panchromatic image and the output features of the low-resolution hyperspectral image are fused through a cross-transformer block (CTB) to obtain the output features of the cross-transformer block. Finally, the output features of the cross-transformer block are multiplied by the corresponding global frequency feature map to achieve effective feature fusion, enhancing the spatial details and spectral accuracy of the low-resolution hyperspectral image. In the previous stages (downsampling), since there is no corresponding global frequency feature map, the output features of the cross-transformer block are directly used as the fused features to be output in this step. That is, for the output features of the cross-transformer block with a corresponding global frequency feature map, the output features of the cross-transformer block are multiplied by the global frequency feature map to obtain the fused features; for the output features of the cross-transformer block without a corresponding global frequency feature map, the output features of the cross-transformer block are directly used as the fused features; finally, fused features of different scales are obtained. The specific operations can be expressed as: Among them, represents the output features of the panchromatic image in the i-th stage, Represents the panchromatic image input to the FU-Net network.

[0041] Among them, Represents the cross-transformation module, Represents the fused feature at the i-th stage, Represents the global frequency feature map incorporated in the i-th stage. In the first and second stages, there is no corresponding global frequency feature map. In the third to fifth stages, there is a corresponding global frequency feature map.

[0042] The cross-transformation module is a key module for multi-source image data fusion processing in this network algorithm. This module mainly processes the input panchromatic image and the multi-spectral image feature map. The core theory relies on the self-attention mechanism in the Transformer architecture, which can deeply mine and fuse the complementary information of the two types of images to generate richer and more accurate feature representations. As Figure 3 shown, the working process of the cross-transformation module can be expressed as: First, use the Reshape operation to adjust the dimension of the input image features to adapt them to subsequent linear transformation operations, and then map them through the Linear transformation layer to generate the query vector (Query, Q), key vector (Key, K), and value vector (Value, V) respectively. The formula is expressed as: Among them, Represents the Reshape layer, Represents the Linear transformation layer, Represents the query vector at the i-th stage, Represents the key vector at the i-th stage, Represents the value vector at the i-th stage.

[0043] Then, perform the weighting of cross-attention. The formula is expressed as: Among them, Represents the activation layer, Represents the 1x3x3 convolution operation, Represents the PAN feature after cross-attention weighting at the i-th stage, Represents the 3x3 convolution operation, Represents the MSI feature after cross-attention weighting at the i-th stage.

[0044] Finally, feature fusion is performed to obtain fused features at different scales. The global frequency feature map also needs to be incorporated in the last three stages. The formula is expressed as: .

[0045] The dual-branch multi-scale feature interaction strategy can ensure the efficient integration of spatial and spectral information, thereby improving the quality and accuracy of the final image reconstruction. When using the FU-Net structure for the HISR fusion task, feature maps at different scales can obtain different detailed features of the original image, thus enhancing the feature expression ability. At the same time, introducing the Transformer structure in the encoder stage can effectively capture global context information and promote the understanding of image semantics. The decoder stage is responsible for mapping the features extracted by the encoder back to the spatial domain to generate the fusion result. At this time, the Fourier transform is no longer applicable, but the Transformer structure can still be used to focus on spatial features and assist in generating high-resolution information. This method is based on the frequency-domain feature prior and cross-Transformer, and by ensuring the effective fusion of multi-stage spatial-spectral features, it greatly enhances the super-resolution reconstruction performance of the model, especially in terms of detail reconstruction.

[0046] S4: Use the fused features at different scales for image reconstruction to obtain a high-resolution hyperspectral image.

[0047] In the final image reconstruction stage, the original spectral features from the LR-HS branch and the reconstructed features from the frequency-domain-spectral branch are fused in a step-by-step manner through multi-scale feature fusion. After obtaining the output of the final stage, the fused features are then mapped into the final high-resolution hyperspectral image through a 3×3 convolution.

[0048] As Figure 4 shown, this embodiment uses the image reconstruction module to perform image reconstruction on the five fused features at different scales output by the five cross-transform modules and the low-resolution hyperspectral image to obtain a high-resolution hyperspectral image. It is expressed as: Among them, represents the 3x3 convolution operation, represents the feature of the first reconstruction stage, represents the feature of the second reconstruction stage, represents the high-resolution hyperspectral image.

[0049] In this embodiment, the present method is compared with existing excellent fusion algorithms and deep learning methods on three public datasets. The two fusion algorithms include PCA and GFPCA, and the seven deep learning methods include DARN, HSRNet, HyperPNN, SSFCNN, SSRNet, HyperDSNet, and HyperRefiner. The evaluation metrics selected are SCC (Spectral Correlation Coefficient), SAM (Spectral Angle Mapper), ERGAS (Relative Global Error), UIQI (Universal Image Quality Index), SSIM (Structural Similarity), PSNR (Peak Signal-to-Noise Ratio), and TestTime (test time in milliseconds). ↑ indicates that the higher the value, the better, and ↓ indicates that the lower the value, the better. The three public datasets are Pavia Center, Botswana, and Chikusei. The experimental results are shown in Table 1.

[0050] Table 1 Results of average quantitative metrics on three different datasets This embodiment also conducted a subjective comparison experiment on the Pavia Center dataset, and the experimental results are as Figure 5 shown. The experimental results prove the leading performance of the present method in the field of HISR technology.

[0051] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A super-resolution fusion method for remote sensing images based on prior frequency domain features, characterized in that It includes the following steps: S1: Obtain a panchromatic image and a low-resolution hyperspectral image; S2: Map the panchromatic image and the low-resolution hyperspectral image to the frequency domain, and then extract the global frequency feature map; S3: Input the panchromatic image, the low-resolution hyperspectral image, and the global frequency feature map into the FU-Net network; Combine the global frequency feature map to fuse the panchromatic image and the low-resolution hyperspectral image to obtain fusion features at different scales; S4: Use the fusion features at different scales for image reconstruction to obtain a high-resolution hyperspectral image.

2. A super-resolution fusion method for remote sensing images based on the prior of frequency domain features as claimed in claim 1, wherein In S2, the panchromatic image and the low-resolution hyperspectral image are respectively mapped from the spatial domain to the frequency domain through the Fourier transform module to obtain the low-frequency frequency domain feature and the high-frequency frequency domain feature; Input the low-frequency frequency domain feature and the high-frequency frequency domain feature into the FRB module to extract the global frequency feature and obtain the global frequency feature map.

3. A super-resolution fusion method for remote sensing images based on prior knowledge of frequency domain features as described in claim 2, characterized in that, The formula of the FRB module is: Among them, represents the low-frequency frequency domain feature, represents the high-frequency frequency domain feature, represents concatenation, represents and the feature map after fusion, represents the 1x1 convolution operation, represents element-wise multiplication, represents the high-frequency feature map, represents the convolution operation, represents the feature map obtained by performing depth feature extraction through multiple convolutional kernels, represents the global frequency feature map.

4. The super-resolution fusion method of a remote sensing image based on the prior of frequency domain features as claimed in claim 1, wherein In S3, in the FU-Net network, perform scale transformation and adjustment of the number of channels on the low-resolution hyperspectral image through downsampling or upsampling at each stage to obtain the output feature of the low-resolution hyperspectral image; At each stage, the panchromatic image undergoes downsampling or upsampling to obtain the output feature of the panchromatic image aligned with the scale and channel level of the output feature of the low-resolution hyperspectral image; Feature fusion is performed on the output feature of the panchromatic image and the output feature of the low-resolution hyperspectral image through the cross-transformation module to obtain the output feature of the cross-transformation module; For the output feature of the cross-transformation module with a corresponding global frequency feature map, multiply the output feature of the cross-transformation module by the global frequency feature map to obtain the fusion feature; For the output feature of the cross-transformation module without a corresponding global frequency feature map, directly use the output feature of the cross-transformation module as the fusion feature; Obtain fusion features at different scales.

5. The super-resolution fusion method of a remote sensing image based on the prior of frequency domain features as described in claim 4, wherein, The downsampling uses a downsampling module with a 2×2 convolutional layer with a stride of 2, and the upsampling uses a 2×2 transposed convolutional layer with a stride of 2.

6. The super-resolution fusion method of remote sensing images based on the prior of frequency domain features according to claim 4, characterized in that The working process of the cross-transformation module is: Use the reshaping operation to adjust the dimension of the input image feature, and then map it through the linear transformation layer to generate the query vector, key vector, and value vector respectively; Use the query vector, key vector, and value vector for cross-attention weighting to obtain the cross-attention weighted PAN feature and MSI feature; Perform feature fusion on the global frequency feature map, the cross-attention weighted PAN feature, and the MSI feature, or perform feature fusion on the cross-attention weighted PAN feature and the MSI feature; Obtain fusion features at different scales.

7. A super-resolution fusion method for remote sensing images based on the prior of frequency domain features as claimed in claim 6, wherein In S2, the number of FRB modules is 3, and the input features of some FRB modules are preprocessed by downsampling; In S3, the FU-Net network includes 5 stages; The formula expression of feature fusion is: Among them, represents the feature fusion in the i-th stage, represents the PAN feature after cross-attention weighting in the i-th stage, represents the MSI feature after cross-attention weighting in the i-th stage, represents the global frequency feature map incorporated in the i-th stage, represents element-wise multiplication.

8. A remote sensing image super-resolution fusion method based on prior knowledge of frequency domain features as claimed in claim 7, wherein, Image reconstruction is expressed as: Among them, represents a 3x3 convolution operation, represents the features in the first reconstruction stage, represents the features in the second reconstruction stage, represents the high-resolution hyperspectral image, represents concatenation, represents a 2×2 transposed convolutional layer with a stride of 2, represents the low-resolution hyperspectral image input into the FU-Net network.

Citation Information

Patent Citations

  • Multispectral and hyperspectral image fusion method based on spatial frequency collaborative network

    CN119314012A

  • High-quality HRHS image generation method based on double-subnetwork structure

    CN119625112A

  • Multi-angle panchromatic and multispectral image progressive fusion method

    CN119942285A

  • Hyperspectral image sharpening method based on dynamic frequency enhancement

    CN120031748A

Cited By

  • Remote sensing image semantic segmentation method fusing frequency domain modeling and lightweight linear attention

    CN121170289A