Remote sensing image sharpening method based on frequency domain hybrid expert and spatial domain difference
By combining frequency domain hybrid expert and spatial domain difference remote sensing image sharpening methods with Fourier transform and hybrid expert structure, the problem of poor sharpening effect of low resolution multispectral images in existing technologies is solved, and efficient image enhancement and feature preservation are achieved.
Patent Information
- Application Number
- CN202511553045.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-29
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-10-29
AI Technical Summary
Existing low-resolution multispectral image sharpening methods cannot effectively utilize the spatial differences between panchromatic images and low-resolution hyperspectral images when generating high spatial resolution multispectral images, resulting in poor sharpening effects. Furthermore, deep learning-based methods have limited representation capabilities when dealing with remote sensing images of multiple ground features.
A remote sensing image sharpening method based on frequency domain hybrid expert and spatial domain difference is adopted. The amplitude and phase components of panchromatic image and low-resolution hyperspectral image are decoupled by Fourier transform. The hybrid expert is used to learn in the frequency domain and fuse spatial difference features to generate high-resolution hyperspectral image.
It effectively improves the spatial resolution of low-resolution multispectral images while preserving the original spectral information, suppressing artifacts and spectral distortion, and the output results are more faithful to the real features of ground objects, thus improving the computational efficiency and generalization ability of the model.
Smart Images

Figure CN121032853A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a kind of based on frequency domain hybrid expert and space difference remote sensing image sharpening method, belong to difference learning and remote sensing image processing technical field. BACKGROUND
[0002] With the development of remote sensing field, remote sensing image plays an important role in many fields, such as map service, precision agriculture and urban development. However, due to hardware limitations, satellite sensors can only capture high-resolution single-channel panchromatic images and low-resolution hyperspectral images, i.e. high spatial resolution multispectral images cannot be captured. Low-resolution multispectral image sharpening is an important but challenging remote sensing image enhancement task, aiming to generate high spatial resolution multispectral images from coupled low spatial resolution multispectral images and panchromatic images.
[0003] Existing low-resolution multispectral image sharpening methods can be roughly divided into two categories: traditional methods and deep learning-based methods. Typical traditional methods include component replacement, multi-resolution analysis, variational optimization, etc. These methods rely on subjective prior assumptions in the sharpening process, although they have a good theoretical explanation, but in actual scenarios, the sharpening quality of the generated images is not good compared to deep learning methods. Deep learning-based methods owe their excellent representation capabilities under the layering paradigm, and they can directly learn dense prior knowledge in the feature space and exhibit strong sharpening performance.
[0004] The low-resolution multispectral image sharpening task aims to enhance the spatial details of low spatial resolution multispectral images under the guidance of panchromatic images, including high-frequency structural information and low-frequency background information of the image. Although existing deep learning-based methods have conducted some research in the frequency domain, however, considering that remote sensing images have diverse ground object categories, previous frequency-based research methods only use a single set of learnable parameters to learn the image components in the frequency domain, which to some extent restricts the model's representation ability when facing multispectral remote sensing images. On the other hand, considering that in the spatial domain, there is a large difference between panchromatic images and low-resolution hyperspectral images, especially that each channel of the low-resolution hyperspectral image contains different semantic information, therefore effectively learning the spatial domain difference between panchromatic images and low-resolution hyperspectral images can guide the model to generate more refined sharpening results. SUMMARY
[0005] The purpose of the present application is to overcome the above-mentioned shortcomings, and to provide a remote sensing image sharpening method based on frequency domain hybrid expert and spatial difference, which realizes superior low-resolution multispectral image sharpening effect.
[0006] The technical solution adopted by the present application is: A remote sensing image sharpening method based on frequency domain hybrid expert and spatial domain difference, comprising the following steps: S1. Obtain a remote sensing image, construct a three-tuple data set including a panchromatic image, a low-resolution hyperspectral image and a ground truth image, preprocess and divide the data set; S2. Fourier transform is used on the panchromatic image and the up-sampled low-resolution hyperspectral image respectively to obtain corresponding amplitude components and phase components to realize frequency domain decoupling, and the amplitude components of the two images are channel spliced and convolved to obtain an amplitude feature map, and the phase components of the two images are channel spliced and convolved to obtain a phase feature map; S3. In the frequency domain, the amplitude feature map is learned by using a hybrid expert amplitude component learning module, and the phase feature map is learned by using a hybrid expert phase component learning module, the hybrid expert amplitude component learning module and the hybrid expert phase component learning module both include an expert library and a calculated hybrid expert weight library, the output feature map of each expert is weighted and fused with the corresponding weight in the hybrid expert weight library to obtain a corresponding component learning feature map; S4. In the spatial domain, the panchromatic image and the up-sampled low-resolution hyperspectral image are subtracted wave by wave to extract the spectral difference information between them, and a final spatial domain difference feature is obtained through convolution operation and ReLU activation function; S5. The amplitude component learning feature map and the phase component learning feature map are inverse Fourier transformed to obtain enhanced representation features of the low-resolution hyperspectral image in the frequency domain, the enhanced representation features in the frequency domain are spliced with the spatial domain difference feature in the channel dimension, and then input into a spatial-frequency domain feature fusion module for fusion; S6. The fused feature and the spatial domain difference feature are added to obtain a network output image, and then the loss with the ground truth is calculated, and finally a high-resolution hyperspectral image is obtained through continuous iteration and optimization; S7. The model is trained using the training set, and the test set is used to test the model in different multispectral scenes to obtain the final model.
[0007] In the above method, the hybrid expert amplitude component learning module of step S3, the expert library contains N experts, each expert is designed as a lightweight depth separable convolution, and the convolution kernel size is 1x1; the input amplitude feature map is processed in parallel by using global average pooling GAP and global maximum pooling GMP, and the results are added, then the aggregated feature vector is input into a fully connected layer FC to obtain, the output dimension is equal to the number of experts N, the output of this layer can be regarded as the activation score of each expert, the score vector is input into the SoftMax function to generate the final expert weight, and these weights jointly constitute the hybrid expert weight library, the output feature map of each expert is multiplied by the corresponding weight value in the hybrid expert weight library, and finally the weighted results are summed to obtain the amplitude component learning feature map.
[0008] The K largest value vectors are selected as input of the SoftMax function to generate the final expert weight.
[0009] The structure of the mixed expert phase component learning module is the same as that of the mixed expert amplitude component learning module, except that the phase feature map is learned to obtain a phase component learning feature map.
[0010] The spatial-frequency domain feature fusion module in step S5 is composed of a 3*3 convolution, a 1*1 convolution and a ReLU activation function, wherein the output channel number of the 1*1 convolution is the input channel number (i.e., the band number) of the low-resolution hyperspectral image.
[0011] Step S6 uses the mean absolute value loss function to supervise the training of the output feature map mapped back to the target dimension in the spatial-frequency domain feature fusion module and the true value high spatial resolution multispectral image, and simultaneously uses the coefficient of variation loss function to prevent the imbalance of the use frequency of the mixed expert.
[0012] Another object of the present application is to provide a remote sensing image sharpening device based on frequency domain mixed experts and spatial domain differences, comprising a memory, a processor and a computer program stored on the memory and executable on the processor, wherein the processor executes the program to realize the remote sensing image sharpening method based on frequency domain mixed experts and spatial domain differences as described above.
[0013] A remote sensing image sharpening storage device has a computer program stored thereon, and the program is executed to realize the steps of the remote sensing image sharpening method based on frequency domain mixed experts and spatial domain differences as described above.
[0014] The present application has the following advantages: (1) In the frequency domain, the Fourier transform, the advanced mixed expert structure and the physical characteristics of the remote sensing image are closely combined, and the mixed expert is used to learn in the decoupled frequency domain component. Compared with the traditional "one-size-fits-all" single network, this method can intelligently call the most suitable expert combination for processing according to the different input ground objects. This makes the model not only enhance the spatial resolution of the image, but also maximize the preservation of the original spectral information, effectively suppress the artifacts and spectral distortion, and the output result is more faithful to the real characteristics of the ground object. The application of the Top-K mechanism not only innovates the model structure, but also significantly improves the computing efficiency and the generalization ability of the model.
[0015] (2) For the spatial domain, the method can fully exploit the semantic difference features of the panchromatic image and the low-resolution hyperspectral image, further realize the complementary of the spatial-frequency domain features, and provide a novel and efficient technical path for the low-resolution hyperspectral image sharpening task.
[0016] (3) The application effectively sharpens low-resolution multispectral remote sensing images by combining mixed experts in decoupled frequency domain component learning and relying on image spatial domain difference learning, and to some extent, makes up for the shortcomings of existing deep learning-based methods in the task of low-resolution multispectral remote sensing image sharpening. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 The method flowchart of the application is shown in the figure. Figure 2 The model architecture diagram of the method of the application is shown in the figure. DETAILED DESCRIPTION
[0018] The application will be further described below in combination with specific embodiments.
[0019] Embodiment 1: A remote sensing image sharpening method based on frequency domain mixed experts and spatial domain differences, comprising the following steps: S1. Obtain a remote sensing image, construct a three-tuple data set including a panchromatic image, a low-resolution hyperspectral image, and a ground truth image, and preprocess and divide the data set: Considering the rigor of the experimental data of the low-resolution hyperspectral image sharpening task, the method of the application uses the PanCollection data set recognized by the field, and the data format strictly follows the Wald protocol. The data set contains remote sensing images captured by World-View3 (8 bands, 11 bits), GaoFen2 (4 bands, 10 bits), and QuickBird (4 bands, 11 bits) satellites. Specifically, the data set divides the data into a training set and a validation set in a ratio of 9:1, and each group of data is composed of three tuples, i.e., a panchromatic image / low-resolution hyperspectral image / ground truth image, with a size of 64x64x1 / 16x16xC / 64x64xC, and C is the number of image bands. The test set data uses 20 three-tuple images, with a size of 256x256x1 / 64x64xC / 256x256xC. We perform intensity upper limit reduction processing on the intensity of all samples in the loaded data set to accelerate the convergence speed of the proposed image enhancement model in the training process. Specifically, for 11-bit remote sensing images, we normalize the data by dividing by 2047, and for 10-bit remote sensing images, we normalize the data by dividing by 1023.
[0020] S2. Fourier transform is used on the panchromatic image and the upsampled low-resolution hyperspectral image to obtain corresponding amplitude components and phase components for frequency domain decoupling, and channel concatenation and convolution operations are performed on the amplitude components of the two images to obtain amplitude feature maps, and channel concatenation and convolution operations are performed on the phase components of the two images to obtain phase feature maps: As Figure 2As shown, since there is a 4-fold difference in spatial resolution between the low spatial resolution multispectral images captured by the satellite sensor and the panchromatic image, the low spatial resolution multispectral images of the training set are first upsampled to 64x64xC by pixel rearrangement, and then the panchromatic image and the upsampled low spatial resolution multispectral image are decoupled in the frequency domain.
[0021] For frequency domain decoupling, we use two-dimensional discrete Fourier transform to convert the panchromatic image and the upsampled multispectral image from the spatial domain to the complex components in the Fourier frequency domain, respectively, and then calculate the corresponding real and imaginary parts according to the complex components, and further calculate the corresponding amplitude components and phase components to achieve frequency domain decoupling. Among them, the component channel number of the panchromatic image is 1, and the component channel number of the multispectral image is the band number C. In order to realize efficient learning of the decoupled frequency domain components by the subsequent hybrid expert, we will subsequently concatenate the amplitude and phase components corresponding to the panchromatic image and the multispectral image, and then map the component features to a high-dimensional space for subsequent learning through convolution operation with a convolution kernel size of 1.
[0022] S3. In the frequency domain, the amplitude feature map is learned by the hybrid expert amplitude component learning module, and the phase feature map is learned by the hybrid expert phase component learning module. The hybrid expert amplitude component learning module and the hybrid expert phase component learning module both include an expert library and a hybrid expert weight library calculated, and the output feature map of each expert is weighted and fused with the corresponding weight in the hybrid expert weight library to obtain the corresponding component learning feature map: In the frequency domain, the same structure of hybrid expert is used to learn the amplitude component and the phase component respectively. In the face of remote sensing images with rich and diverse ground objects, the hybrid expert can adaptively activate its adaptive experts according to the specific amplitude low-frequency background information and phase high-frequency structure information of the input image, aiming to maximize the learning of the most faithful image features.
[0023] There are two sub-modules with the same structure but independent parameters: the hybrid expert amplitude component learning module and the hybrid expert phase component learning module. Taking the hybrid expert amplitude component learning module as an example, its core goal is to adaptively activate the corresponding spectral fidelity expert or content context expert according to the low-frequency background information and overall spectral distribution of the input image, and its detailed processing flow is as follows: The input is the amplitude feature map A obtained by 1x1 convolution after channel concatenation (Concat), which is calculated as follows: , Hybrid expert library: we predefine an expert library containing N experts. Each expert is designed as a lightweight depth separable convolution with a convolution kernel size of 1x1. This design can efficiently perform feature transformation and information extraction between channels while maintaining relatively low computational complexity.
[0024] Mixed expert weight bank: We apply global average pooling GAP and global max pooling GMP to the input amplitude feature map A in parallel, and add the results to obtain a compact vector that can represent the global context and the most salient features. Then we send the aggregated feature vector to a fully connected layer FC to obtain F A The output dimension is equal to the number of experts N. The output of this layer can be regarded as the activation scores of each expert. To improve the efficiency and specialization of the model, we introduce a sparse gating strategy. From the N activation scores, we select the top K values with the largest values, and the remaining N-K scores are set to negative infinity. It ensures that only a few most relevant experts are activated each time, avoids resource waste, and encourages each expert to learn more discriminative features. Finally, we send the Top-K processed score vector to the SoftMax function to generate the final expert weights. At this time, only the selected K experts have weights greater than zero, and the remaining weights are zero. The sum of all weights is 1. These weights together constitute the mixed expert weight bank, whose internal process is as follows: , , , F i A For the i-th score in the activation score, i∈1,2,…N. W A Amplitude mixed expert weight bank, whose dimension is the number of experts N, corresponding to the weight of each expert.
[0025] Weighted fusion: The input amplitude feature map is sent to all N experts in the expert bank in parallel. Then, the output feature map of each expert is multiplied by its corresponding weight value in the mixed expert weight bank, and finally all the weighted results are summed. Mathematically, the process can be represented as: , Where E i A is the i-th expert in the amplitude mixed expert bank, W i A is the i-th weight in the amplitude expert weight bank, A ’ is the amplitude component learning feature map obtained after mixed expert amplitude component learning.
[0026] The mixed expert phase component learning module adopts exactly the same structure, but it is applied to the independent phase feature PThe training is performed on the above, and the goal is to learn how to activate the corresponding texture expert or edge expert according to the structural information of the input image, so as to accurately enhance and reconstruct the high-frequency details, and obtain the phase component learning feature map P obtained after learning the mixed expert amplitude component ’ .
[0027] S4. The full-color image is subtracted from the up-sampled low-resolution hyperspectral image wave by wave, and the spectral difference information between the two is extracted, and the final spatial domain difference feature is obtained through convolution operation and ReLU activation function: For spatial domain difference learning, we subtract the up-sampled low-resolution hyperspectral image wave by wave from the full-color image to explicitly extract the spectral difference information between the two. That is, for each waveband content in the hyperspectral image, the difference feature is calculated with the full-color image of a single waveband. These difference features reveal the semantic information that the low-resolution hyperspectral image needs to pay more attention to in the spatial domain. We use a convolution kernel of size 3 to dig deeper into the difference representation, and then pass it through a ReLU activation function to obtain the final spatial domain difference feature. The process is shown below: , , where the purpose of the duplicate operation is to copy the channel number of the full-color image PAN to the same number as the MS channel number, which is used to realize the subsequent wave subtraction, which is equivalent to Figure 2 the idea of wave subtraction. Diff is the final spatial domain difference feature.
[0028] S5. The amplitude component learning feature map and the phase component learning feature map are inverse Fourier transformed to obtain the enhanced representation features of the low-resolution hyperspectral image in the frequency domain. The enhanced representation features in the frequency domain are concatenated with the spatial domain difference features in the channel dimension, and then input into the spatial-frequency domain feature fusion module for fusion: Before spatial-frequency domain feature fusion, we use the amplitude component learning feature map A ’ and the phase component learning feature map P ’ to obtain the enhanced representation content of the low-resolution hyperspectral image in the frequency domain using two-dimensional discrete inverse Fourier transform. Subsequently, we concatenate it with the difference feature Diff learned in the spatial domain in the channel dimension, and then perform spatial-frequency domain feature fusion. The spatial-frequency domain feature fusion module consists of a 3x3 convolution, a 1x1 convolution, and a ReLU activation function, which aims to realize complementary learning of spatial-frequency domain features. The output channel number of the 1x1 convolution is the input channel number C (i.e. the number of wavebands) of the low-resolution hyperspectral image, which aims to map the features back to the target dimension C.
[0029] S6. The fused features are added to the spatial difference features to obtain a network output image, and then the loss of the network output image and the ground truth is calculated, and through continuous iteration and optimization, a high-resolution hyperspectral image is finally obtained.
[0030] S7. The model is trained using the training set, and the final model is obtained by testing the model in different multi-spectral scenes using the test set: The absolute value loss function is used in the application L 1. The output feature map mapped back to the target dimension C in the spatial-frequency domain feature fusion module and the true value high spatial resolution multispectral image are supervised trained, and the squared coefficient of variation loss function is used L scv to prevent the use frequency of mixed experts from being uneven. For the training setting, the 8-band dataset is trained for about 360 rounds, and the 4-band dataset is trained for about 150 rounds, and the Adam optimizer is used as a whole, with two parameters of 0.9 and 0.999. In addition, the squared coefficient of variation loss function is multiplied by the hyperparameter of 0.3, the initial learning rate is 0.001, and the drop coefficient is multiplied by 0.5 every 100 rounds of training, and the total number of amplitude experts and phase experts is 4, and 2 experts are activated each time. In the model training, the loss value corresponding to the training set is calculated after each round of iteration, and the model after a certain number of iterations is used as the final model. In the model test stage, the model is tested in different multi-spectral scenes using the test set, and high spatial resolution multispectral images can be obtained by adaptive sharpening for different low spatial resolution multispectral images.
[0031] The absolute value loss function calculation formula is: , wherein Output is the output image of the network, and GT is the ground truth image; The squared coefficient of variation loss function calculation formula is: , wherein W A is the amplitude expert weight bank, W P is the phase expert weight bank, and SCV is responsible for calculating the mean along the batch dimension.
[0032] Embodiment 2: The remote sensing image sharpening device based on frequency domain mixed experts and spatial domain differences comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor executes the program to realize the remote sensing image sharpening method based on frequency domain mixed experts and spatial domain differences as described in Embodiment 1.
[0033] A remote sensing image sharpening storage device based on frequency domain hybrid expert and spatial domain difference, which stores a computer program, and when the program is executed, the steps of the remote sensing image sharpening method based on frequency domain hybrid expert and spatial domain difference as described in Embodiment 1 are implemented.
[0034] The above is a further description of the application in combination with the embodiments, and the protection scope of the application is not limited thereto.
Claims
1. A remote sensing image sharpening method based on frequency domain hybridization experts and spatial domain differences, characterized in that, The steps include the following: S1. Acquire remote sensing images, construct triplet data including panchromatic images, low-resolution hyperspectral images, and ground truth images, preprocess and divide the dataset; S2. The amplitude and phase components of the panchromatic image and the upsampled low-resolution hyperspectral image are obtained by Fourier transform to achieve frequency domain decoupling. The amplitude components of the two images are then concatenated and convolved to obtain amplitude feature maps, and the phase components of the two images are then concatenated and convolved to obtain phase feature maps. S3. In the frequency domain, the amplitude feature map is learned using the hybrid expert amplitude component learning module, and the phase feature map is learned using the hybrid expert phase component learning module. Both the hybrid expert amplitude component learning module and the hybrid expert phase component learning module include an expert library and a calculated hybrid expert weight library. The output feature map of each expert is weighted and fused with its corresponding weight in the hybrid expert weight library to obtain the corresponding component learning feature map. S4. Subtract the panchromatic image from the upsampled low-resolution hyperspectral image band by band to extract the spectral difference information between the two, and obtain the final spatial difference features through convolution operation and ReLU activation function. S5. The amplitude component learning feature map and the phase component learning feature map are inverse Fourier transformed to obtain the enhanced representation features of the low-resolution hyperspectral image in the frequency domain. The enhanced representation features in the frequency domain are then concatenated with the spatial difference features in the channel dimension and then input into the spatial-frequency domain feature fusion module for fusion. S6. The fused features are added to the spatial difference features to obtain the network output image. Then, the loss between the network output image and the ground truth is calculated. After continuous iterative optimization, a high-resolution hyperspectral image is finally obtained. S7. Train the model using the training set, and test the model using the test set in different multispectral scenarios to obtain the final model.
2. The remote sensing image sharpening method based on frequency domain hybridization expert and spatial domain difference as described in claim 1, characterized in that, The hybrid expert amplitude component learning module described in step S3 has an expert library containing N experts. Each expert is designed as a lightweight depthwise separable convolution with a kernel size of 1×1. The input amplitude feature map is processed in parallel using Global Average Pooling (GAP) and Global Max Pooling (GMP), and the results are summed. The aggregated feature vector is then fed into a fully connected layer (FC), whose output dimension is equal to the number of experts N. The output of this layer can be regarded as the activation score of each expert. The score vector is fed into the SoftMax function to generate the final expert weights. These weights together constitute the hybrid expert weight library. The output feature map of each expert is multiplied by its corresponding weight value in the hybrid expert weight library, and the weighted results are finally summed to obtain the amplitude component learning feature map.
3. The remote sensing image sharpening method based on frequency domain hybridization expert and spatial domain difference according to claim 2, characterized in that, Step S3 selects the top K score vectors with the largest values and feeds them into the SoftMax function to generate the final expert weights.
4. The remote sensing image sharpening method based on frequency domain hybridization expert and spatial domain difference as described in claim 1, characterized in that, The structure of the hybrid expert phase component learning module is the same as that of the hybrid expert amplitude component learning module, except that it learns the phase feature map to obtain the phase component learning feature map.
5. The remote sensing image sharpening method based on frequency domain hybridization expert and spatial domain difference according to claim 1, characterized in that, The spatial-frequency domain feature fusion module described in step S5 consists of a 3×3 convolution, a 1×1 convolution, and a ReLU activation function, wherein the number of output channels of the 1×1 convolution is the number of input channels of the low-resolution hyperspectral image.
6. A remote sensing image sharpening method based on frequency domain hybridization expert and spatial domain difference according to claim 1, characterized in that, Step S6 uses the mean absolute value loss function to supervise the training of the output feature map mapped back to the target dimension in the spatial-frequency domain feature fusion module and the ground-value high spatial resolution multispectral image, while using the squared coefficient of variation loss function to prevent uneven use of hybrid experts.
7. A remote sensing image sharpening device based on frequency domain hybridization and spatial domain difference, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the remote sensing image sharpening method based on frequency domain mixing experts and spatial domain differences as described in any one of claims 1-6.
8. A remote sensing image sharpening and storage device based on frequency domain hybridization experts and spatial domain differences, wherein a computer program is stored thereon, characterized in that, When the program is executed, the steps of the remote sensing image sharpening method based on frequency domain mixing experts and spatial domain differences as described in any one of claims 1-6 are implemented.
Citation Information
Patent Citations
Multispectral remote sensing image enhancement method based on frequency domain-space double-domain learning
CN118587097A
Low-light image enhancement method based on frequency domain and spatial domain perception
CN118674628A
Remote sensing image space-time fusion method and system based on deep Fourier Transform network
CN118941903A
Remote sensing image arbitrary scale super-resolution method based on dynamic scale frequency domain convolution
CN119559052A
Deep learning remote sensing change detection method based on mixed space-frequency expert
CN120472319A